TraceLab: Characterizing Coding Agent Workloads for LLM Serving
TraceLab releases 4,300 real coding-agent sessions with 350K LLM steps for analyzing serving workloads across agents and models.
Search the full wire by company, model, lab, or keyword. Every story we have ever aggregated.
TraceLab releases 4,300 real coding-agent sessions with 350K LLM steps for analyzing serving workloads across agents and models.
Analyzes convergence of continual classification in homogeneous deep networks via projection theory.
Poller uses LLMs as poetry evaluators by role-playing authors for Chinese poetry understanding tasks.
Linguistic Firewall improves multi-agent routing by using geometric profiling to verify agent competence empirically rather than self-description.
Hybrid quantum-classical workflow for protein-ligand binding energy via generative-ML-assisted quantum selected CI on Fujitsu FX700.
Clover tool tracks student engagement with AI code completion via behavioral metrics and attention checks to measure critical evaluation.
OTF decomposes visual transitions into sparse primitives to learn robust latent action representations in multi-object scenes.
TRACE framework detects emotional entrainment in dyadic speech via temporal modeling; introduces DyadEE dataset with synthetic disruption controls.
Proposes lifelong learning framework for autonomous driving policies to accumulate corrective knowledge from deployment errors without catastrophic forgetting.
Identifies entity binding failures in tool-augmented LLM agents as distinct safety problem: selecting correct tool but acting on wrong entity.
μFlow: one-class deepfake detector trained only on real images generalizes across GAN and diffusion model generators without synthetic supervision.
ITSPACE proximal method optimizes Bures-Wasserstein objective on covariance matrices for domain adaptation and Gaussian embeddings.
TIDAL's new policy will prevent AI-generated music from making money on its service.
Knowledge distillation strategy for visual quantum RL: train classical visual teacher, freeze encoder, distill policy into classical/quantum heads.
RAPS-DA: regime-aware peer specialization framework handles heterogeneous knowledge conflicts in RAG by disentangling reliability levels of retrieved context.
Entropic Learnability Horizon (ELH) framework unifies information theory and topology to explain generalization in overparameterized deep networks.
DeepReinforce releases Ornith-1.0, MIT-licensed open-weights model (9B–397B variants) for agentic coding, built on Gemma 4 and Qwen 3.5, achieving SOTA on coding benchmarks.
Muon optimizer avoids slow saddle-to-saddle dynamics in matrix factorization; learns balanced solutions with uniform convergence rates across modes.
DR-ACI constructs prediction intervals for doubly robust pseudo-outcomes under temporal dependence via adaptive conformal inference.
Random Network Distillation enables lightweight clustering in federated learning without coupling to main training loop, reducing communication overhead.
Post-Hoc Concept Bottleneck Models often learn predictive artifacts rather than semantically meaningful concepts; proposes faithfulness evaluation beyond task accuracy.
CUDA optimization techniques (shared memory, pre-transposed weights, fused kernels) achieve 1.41x speedup on shallow networks on Tesla T4.
Multi-channel Multigrid preconditioner learned via neural networks to solve high-wavenumber Helmholtz equations with phase-space coarsening.
Google AI explains full-stack approach to AI development, covering integrated hardware, software, and model design.
A new proposal would ban the sale of Americans' health and location information to data brokers - including information people reveal to an AI chatbot like ChatGPT or Claude. In the coming weeks, Senator Elizabeth Warren (D-MA) and Representative Mary Gay Scanlon (D-PA) are planning to debut a new version of the Health and Location Data Protection Act that's better suited to the AI era. The former version of the bill, first introduced in June 2022, prohibited data brokers from collecting and selling health and location data. Four years later, it's expanded to ban other companies from selling ...
SIMAX framework generates synthetic clinician-patient dialogues with human-coded annotations to evaluate AI communication coding at scale.
Factorizable Normalizing Flows model parameter-dependent density morphing by factorizing the flow, enabling tractable learning across exponentially many configurations.
AI agents are quickly moving beyond chat. They inspect code, run tests, read documents, search knowledge bases, query internal systems, and operate for hours on... AI agents are quickly moving beyond chat. They inspect code, run tests, read documents, search knowledge bases, query internal systems, and operate for hours on behalf of a user. This unlocks productivity, but can also give agents access to sensitive enterprise data and the ability to complete tasks and take action across business systems, making a secure, governed environment essential. Source
Argues LLMs lack situated perception of physical world causality and agency; identifies this primitive as necessary prerequisite for artificial superintelligence.
COHORT automates enterprise network threat mitigation via multi-agent LLM workflow proposing, implementing, and refining mitigations on emulated topologies.