Bayesian Uncertainty Propagation for Agentic RAG Pipelines: A Proof-of-Concept Study on Multi-Hop Question Answering
Bayesian uncertainty propagation framework for agentic RAG systems identifies failure points in multi-hop reasoning.
Search the full wire by company, model, lab, or keyword. Every story we have ever aggregated.
Bayesian uncertainty propagation framework for agentic RAG systems identifies failure points in multi-hop reasoning.
Svarna open-source corpus workbench integrates 507M words across Modern Greek registers and dialects.
Survey of 200+ VIS4ML papers maps how humans inject knowledge into ML workflows via interactive visualization.
Zero-shot benchmark evaluates Claude, GPT-5.4, and Gemini on fine-grained 13-class emotion classification.
The Builders Stage is returning to TechCrunch Disrupt 2026, bringing together 10,000+ founders, startup operators, and investors for practical conversations. and Q&A on what it takes to build and scale successful companies. Register now to save up to $330.
Controlled CV-versus-DV quantum neural network comparison on wafer-map defect classification for yield.
LeNEPA self-supervised method for time-series representation learning without augmentation.
Aionoscope diagnostic tool debugs latent-state accessibility in frozen time-series representations.
Compares strategies for learning cardiac motion priors in implicit neural representations to accelerate optimization.
Uses diffusion/flow models as geometric maps to enable gradient descent on learned data manifolds.
Studies emotion steering in text-to-speech models via activation steering on speech language models and flow-matching modules.
Meta is developing plans for a cloud infrastructure business, selling access to AI compute power and models. The move would pit it against the big cloud providers like Amazon Web Services, Google Cloud, and Microsoft Azure.
Measures instability in LLM persona-driven outputs for multiple-choice QA across performance, outcome, and question dimensions.
Proposes ILLUME+, a post-hoc explainability framework for cancer drug response prediction using multivariate gene interactions.
Develops post-training pruning methods for Diffusion Transformers that account for their unique weight distribution and skip-connections.
Proposes GMHF framework combining meta-learning with human feedback to guide synthetic data generation for domain adaptation.
Introduces Graph-PRefLexOR, a graph-native reasoning model fine-tuned with GRPO for traceable materials discovery hypothesis generation.
Presents MAGNET, a multi-agent narrative engine with persona-grounded characters, and ATLAS, a verification pipeline for long-form story generation.
Proposes Valdi, latent diffusion world models for fast Model Predictive Control that combines online training with diffusion-based dynamics.
Frontier AI models maintain capability leads on most metrics while smaller models may converge on validation loss as scaling continues.
Mixed-precision LLM quantization shows perplexity-based layer importance poorly predicts reasoning performance; combining task and general data improves post-quantization results.
New benchmark for span-level hallucination detection in RAG systems across code, tool output, structured documents, and natural language.
MultiSynt/MT releases 4.8 trillion tokens of open synthetic parallel pre-training data across 36 European languages via Tower+ and OPUS-MT translation.
DeWorldSG framework generates 3D semantic scene graphs from RGB-D sequences using probabilistic Gaussian distributions and depth-guided filtering.
Flat minima optimization applied to 3D Gaussian Splatting improves generalization on sparse-view novel view synthesis.
Theoretical proof that binary tree mechanism achieves optimal log^(3/2) n error dependence for differentially private continual counting.
Study of ethos and pathos rhetorical appeals in silent readers' interpretations of social media messages using labeled datasets.
SEA architecture enables safe self-modifying agents by freezing base models and gating adaptations through anytime-valid certificates against error budgets.
Production-scale clinical NLP study of inference-time gating with Llama-3.3 70B generator and MMed-Llama-3.1 70B verifier over 167K narratives shows pattern-memory filtering limitations.
The Google Home Speaker is Google’s first smart speaker in years. And it’s pretty! | Photo: Jennifer Patison Tuohy / The Verge Smart speakers have spent the past few years searching for a compelling second act. Beyond music, timers, and controlling your lights, they've struggled to justify taking up space on the kitchen counter. AI promised to change that. Amazon debuted its new hardware powered by a revamped Alexa last fall, and now it's finally Google's turn. The Google Home Speaker is the company's first new smart speaker in six years and its first "built for Gemini." After years of neglec...