Bridge Evidence: Static Retrieval Utility Does Not Predict Causal Utility in Multi-Step Agentic Search
ReAct-style agents show static document relevance scores fail to predict utility in multi-step reasoning; gap measured on HotpotQA.
Search the full wire by company, model, lab, or keyword. Every story we have ever aggregated.
ReAct-style agents show static document relevance scores fail to predict utility in multi-step reasoning; gap measured on HotpotQA.
Across 101 enterprises, the infrastructure that feeds AI agents their business context is being built faster than it can be trusted. Retrieval-augmented generation is already the default context source, and provider-native retrieval has quietly overtaken the dedicated vector databases that define the category — yet a majority of enterprises have already watched their agents produce confident, wrong answers traced to missing or inconsistent context. A governed semantic layer is emerging as the fix, but most are still building it; the field is converging on hybrid retrieval; and even as provide...
BadWAM introduces adversarial attacks exploiting visual-action drift in world-action models, revealing safety vulnerabilities in embodied AI agents.
Plover externalizes planning in vision-based GUI agents, enabling user inspection and correction of task plans for autonomous interface automation.
Across 157 enterprises, organizations are granting AI agents more autonomy while trusting the evaluations meant to gate that autonomy less. Half have already shipped an agent that passed their internal evaluations and then failed a customer in production; only one in twenty fully trusts automated evaluation today; and the most-cited weakness is that evaluations do not align with real-world outcomes. Yet two-thirds already allow, or are actively engineering toward, deploying agent changes to production on automated evaluation alone — with no human in the loop. The result is an evaluation gap —...
A video analytics AI agent that can perceive, reason, and act based on massive amounts of video footage must be integrated with existing workflows and... A video analytics AI agent that can perceive, reason, and act based on massive amounts of video footage must be integrated with existing workflows and applications to be useful. These include content management systems, messaging platforms, databases, ticket queue, and escalation paths. This integration is challenging because video systems, enterprise knowledge bases… Source
Agentic AI changes the infrastructure pattern for AI factories. One request can trigger many model calls, tool calls, memory lookups, policy checks, storage... Agentic AI changes the infrastructure pattern for AI factories. One request can trigger many model calls, tool calls, memory lookups, policy checks, storage accesses, and network transfers before a final answer is produced. As more agents run at once and carry context across steps, users, tools, services, and sessions, infrastructure must move, protect, retrieve, and reuse data fast enough to keep… Source
DoorDash is opening a limited beta of dd-cli, a command-line tool that lets developers and AI agents search stores, build carts, and place orders from the terminal, marking another step toward software designed for AI agents instead of just humans.
Multi-agent LLM framework using SFT+DPO enables sustained partisan behavior in political coalition simulation, circumventing RLHF neutrality bias.
OmniaBench: unified benchmark evaluating LLM-based agents across diverse scenarios with explicit state spaces for systematic capability characterization.
LongStraw: GPU-efficient execution stack for million-token RL post-training with GRPO, bridging inference context length vs. post-training gap for agents.
StructureClaw: artifact-centered benchmark for evaluating LLM agents on complete structural engineering workflows with verifiable evidence chains.
Proof-or-Stop: lifecycle control framework for autonomous coding agents using mechanically verifiable evidence gates for transitions between states.
Cars24 deploys OpenAI voice and chat agents handling 1M+ monthly conversation minutes, recovering 12% lost leads via agentic workflows.
Across 101 enterprises, agent orchestration is consolidating onto model-provider platforms — Anthropic’s Claude leads by a wide margin — chosen for the gravity of the underlying model and judged on reliable multi-step execution. But the ambition runs well ahead of the reality: most deployed “agents” are still chatbot wrappers, the control plane enterprises expect is deliberately hybrid to avoid lock-in, and real-time fiscal control over token burn remains the exception. This wave of VentureBeat Pulse Research examines enterprise agent orchestration: which platforms enterprises run on, what dr...
OpenUSD is an open, extensible framework that provides a common scene description language for physical AI. It enables teams to bring CAD data, simulation... OpenUSD is an open, extensible framework that provides a common scene description language for physical AI. It enables teams to bring CAD data, simulation assets, and real-world telemetry into a shared, physically accurate view of the world. Until now, building a USD implementation has typically required adapting a large existing codebase— even for teams that need a specific memory footprint… Source
Study shows agent-optimization gains may not compound over time; proposes Terminal-Bench 2.0 to test continual learning on deployed agents.
TRACE assigns per-turn rewards for long-horizon multi-tool agents to improve credit assignment beyond sparse outcome rewards.
OpenAI is finally releasing some hardware. No, it isn't the mysterious AI-powered device the company is developing with former Apple designer Jony Ive, a project already tangled up in a messy lawsuit. Instead, it's a product designed to be used with its coding platform, Codex. The device, a square-shaped block of buttons called Codex Micro, is a collaboration between the AI company and keyboard maker Work Louder. OpenAI said it is a limited-run collaboration that will give users more ways to monitor and manage their agents. The pad closely resembles Work Louder's Creator Micro 2, and marketin...
DeepStress: stress-testing framework for search agents' robustness to degraded retrieval quality and evidence reliability.
Experience Memory Graph enables one-shot error correction for LLM agents via structured trajectory memory, reducing API costs and improving generalization.
SPyCE co-evolves reusable skills and policies for multimodal agents, distilling visual reasoning trajectories to improve generalization across tasks.
The guy behind TCP/IP is working on a standard for identifying AI agents in the wild.
User permission frameworks for AI agents address prompt injection and unauthorized actions; bridges product-level design with enforcement mechanisms.
Willison/Ronacher reflect on how AI agents may erode institutional knowledge and shared understanding embedded in code review friction.
LLM agents often waste computational budget through redundant context re-reading; proposes task-complexity estimation to optimize execution scope.
PalmClaw framework enables native on-device LLM agents on mobile phones via direct API calls rather than GUI automation.
Coding AI agents are becoming practical operators for long-running machine learning (ML) workflows. They can inspect repositories, set up runtimes, resolve... Coding AI agents are becoming practical operators for long-running machine learning (ML) workflows. They can inspect repositories, set up runtimes, resolve build issues, launch experiments, monitor execution, analyze metrics, and summarize results. For reinforcement learning (RL) research, this matters because meaningful metrics often appear only after the essential experiment infrastructure… Source
What if autonomous coding AI agents could push your vision reasoning models above 90% accuracy with almost no manual effort? When adapting vision reasoning... What if autonomous coding AI agents could push your vision reasoning models above 90% accuracy with almost no manual effort? When adapting vision reasoning models to production video tasks, developers often lose days to data formatting, container setup, training scripts, baseline evaluation, and hyperparameter sweeps before they even know whether post-training improves accuracy. Source