Large Databases Need Small, Open-Weight Language Models
Quantized open-weight LMs on 16GB VRAM match proprietary APIs on database tasks at lower cost and latency.
Search the full wire by company, model, lab, or keyword. Every story we have ever aggregated.
Quantized open-weight LMs on 16GB VRAM match proprietary APIs on database tasks at lower cost and latency.
Conformal prediction framework for energy time-series forecasting via STGNNs and foundation models with uncertainty quantification.
RAISE framework for robust LLM-based heuristic design via adversarial instance search under distributional shift.
Evo-PI framework evolves principle-guided supervision signals for multimodal LLMs to improve medical reasoning generalization.
CHERRY combines selective ground-truth supervision, sparse experts, and recurrent compression for 4.5x per-token efficiency in LM training.
shot-scraper 1.10 release announcement with video storyboard feature for web application testing and agent demo generation.
X has launched a hosted MCP server, making it easier for developers to connect AI applications with the company’s API.
Anthropic launches Claude Science, a customizable workbench integrating research tools, auditable artifacts, and compute resources for scientific workflows.
Engineers on the new team will embed within companies to deploy purpose-built agents, focusing on fast deployments and customer self-sufficiency.
Users will be able use AI to create newsletters based on their recordings.
SpikeLogBERT applies spiking neural networks for energy-efficient log parsing vs. dense transformer approaches.
Proton's Lumo 2.0 is dropping this week, giving users a broader variety of capabilities.
Adversarial distillation improves certified robustness of neural networks by combining tight relaxation bounds with adversarial training techniques.
FARS system autonomously generates, executes, and writes AI research across topics at scale using coordinated agent architecture.
ECHO introduces selective turn memory and pruning for long-horizon agentic RL under context window constraints.
LuckyStar 111B hybrid reasoning model from Cohere and LG CNS enables efficient multilingual tool-using agents with Korean-English support.
LLMs exhibit performative compliance: fairness evaluations overestimate moral safety when demographic identity must be inferred rather than labeled.
Tone-conditioned curriculum learning improves zero-shot ASR for 6 Southern Bantu languages using hybrid difficulty scoring and gated adapters.
Comprehensive survey systematizes LLM attack surface across full lifecycle: data pipelines, agents, tools, memory, and organizational integration.
Intrinsic decomposition extended to 3D Gaussian splatting for texture editing independent of lighting.
LLM agents act as constrained supervisory planners for fault recovery in process plants, validated against external safety constraints.
Bayesian workflow calibration detects and repairs statistical errors in probabilistic programs written by LLMs using posterior checks and diagnostics.
Philosophy of science frameworks applied to evaluating explainability standards for medical AI systems.
LLM + knowledge graph framework automates cause-effect specifications for industrial process control and safety systems.
Higher-order structural alignment method for multi-view radar semantic segmentation in adverse weather.
CLExEval: human-in-the-loop evaluation framework for LLM clinical reasoning with 5,600 physician annotations.
Uncertainty-guided diffusion model for synthetic data augmentation in semantic segmentation with sparse labels.
Dual-Embedding Watermarking (DEW) scheme for LLM text watermarking robust to paraphrasing and translation.
Theoretical framework for optimal training-calibration data splitting in split conformal prediction.