Time-Varying Data as Sheaves: an Invitation to Narratives
Mathematical framework for time-varying data using sheaf theory and narratives, applicable across scientific disciplines.
Search the full wire by company, model, lab, or keyword. Every story we have ever aggregated.
Mathematical framework for time-varying data using sheaf theory and narratives, applicable across scientific disciplines.
Training-Free Task Vectors enable LLM behavioral steering without fine-tuning by mapping activation vectors to weight-space edits.
Audit methodology effects rival demographic bias in LLM decision-making: charitable-aid benchmark results fail to replicate across hiring, lending, triage domains.
Human study on chain-of-thought reasoning representations as explanations: evaluates whether formats help users assess LLM outputs across complexity levels.
MIT researcher uses GPT-5.6 Sol with Codex to autonomously run quantum computing experiments, analyze results, and calibrate qubits.
Answer-distribution trajectories track full predictive distributions over LLM reasoning steps, revealing competing hypotheses beyond endpoint accuracy.
Legal and responsibility analysis of AI-assisted code production: examines ownership, producer identity, and quality engineering duties.
Astrophysics study: inverse modeling to simultaneously recover Galactic potential and stellar distribution from Gaia DR3 kinematics.
Diffusion models for constrained discrete tasks (Sudoku, N-queens): compares standard vs. error-correcting sampling strategies.
Knowledge graph-based evaluation framework to assess contextual understanding in LLMs beyond surface-level metrics.
Deposon scattering layer for auditable LLM reasoning paths with conservation constraints and machine epsilon precision.
Physics-informed deep learning with SE-ResNet reduces false VT alarms in ICU using Windkessel hemodynamic model.
Theoretical proof that frozen transformers perform data generation via in-context learning without parameter updates.
Gander: multimodal streaming agent supporting full-duplex video/speech/text interaction with real-time interrupts and proactive feedback.
GraphFAS: distributed system for interpretable graph feature generation in fraud detection via Boruta-based automated selection.
Google Cloud expands its enterprise AI push with Accenture, betting on forward-deployed engineers to drive adoption and overcome deployment bottlenecks.
30B MoE analysis: pretraining loss alone fails to predict post-SFT performance; solution density better predicts downstream quality.
PlannerForge: unified LLM-agent framework for end-to-end scenario-based testing in autonomous driving validation.
Complexity analysis of fitting and learning propositional formulas with arbitrary Boolean function bases.
SQLMorph: query mutation framework for reproducible Text-to-SQL evaluation on complex enterprise schemas.
ONECYL benchmark for graph-based surrogate modeling of CFD on unstructured meshes across flow regimes.
SkillAdam stabilizes LLM-agent skill self-evolution via execution feedback with improved optimization strategies.
Fact-Ablated Evaluation framework audits whether LLMs faithfully use evidence or rely on parametric knowledge in fact-checking.
AuK, open-source multimodal model for unified speech generation and editing via natural-language instructions and audio context.
Robustness audit via NC-MCAR metric reveals LMs destabilize answers when source-attributed cues contradict evidence.
Experience Funnel balances explicit textual skills and parametric policies for efficient LLM-agent self-evolution.
BatchNorm artifacts confound machine unlearning evaluation: single forward pass on retain data reverses forgetting metrics.
Auditable evidence framework improves speech deepfake detection by replacing single scores with explainable decision signals.
Methodological comparison of MAE, Solar, and UniMMQA frameworks tracing evolution of multimodal QA architectures.
Q2D-Web large-scale benchmark for agentic RAG evaluates first-stage retrievers on agent-reformulated queries at production scale.