OpenAI’s new voice mode makes it to the ChatGPT desktop app
ChatGPT Voice on desktop can work with both ChatGPT Work and Codex to complete tasks and control agents.
Every story matching this topic across titles and summaries, newest first.
ChatGPT Voice on desktop can work with both ChatGPT Work and Codex to complete tasks and control agents.
OpenForgeRL enables end-to-end training of harness-native agents with open infrastructure, addressing limitation of complex inference harnesses like Claude Code.
Open-source evaluation framework for open-weight LLM agents on longitudinal data tasks, addressing privacy constraints in research deployments.
VoLN: vision-only navigation benchmark and method for embodied agents without language instructions in GPS-denied environments.
Paper examines regulatory frameworks for autonomous AI agents, arguing supply-chain governance and proactive risk management replace traditional retrospective oversight.
Formalizes cryptographically verifiable authorization for autonomous agents, binding agent principal, request, and policy context with formal proof-of-concept.
"This is day one for cybersecurity in the age of agents," Hugging Face CEO says.
A TensorRT engine build can take seconds to many minutes. Large strongly typed models, deep tactic search, and a cold timing cache on a brand-new GPU SKU can... A TensorRT engine build can take seconds to many minutes. Large strongly typed models, deep tactic search, and a cold timing cache on a brand-new GPU SKU can leave developers, end users, or AI agents staring at a frozen terminal with no idea whether to wait, retry, or kill the process. Most NVIDIA TensorRT integrations report nothing during a build or provide no way to abort early. Source
PoTRE framework deploys four heterogeneous agents (adversarial, hierarchical, spectrum search, direct) with task-adaptive aggregation for complex LLM reasoning.
Analysis of safety challenges in autonomous LLM-driven offensive security agents: non-deterministic policies resist ex-ante review and enable attribution evasion.
Glow is targeting a new class of endpoint risks created by the rapid adoption of AI agents and developer tools inside enterprises.
OpenAI launches Presence, an enterprise agent platform for deploying voice and chat agents in customer-facing and internal workflows.
OpenAI CEO Sam Altman. | Bloomberg via Getty Images OpenAI says its AI models mistakenly breached open-source AI platform Hugging Face during internal testing. In a blog post on Tuesday, OpenAI writes that GPT-5.6 Sol and "an even more capable pre-release model" discovered vulnerabilities within their sandboxed testing environment, allowing them to gain access to the internet and target Hugging Face. On July 16th, Hugging Face disclosed a security incident that it says was driven by "an autonomous AI agent system." Hugging Face's AI agents detected and stopped the breach, which OpenAI has now...
Buzz is a group chat platform for the workplace that puts humans and their AI agents in the same conversation.
CodeRescue optimizes cost-aware routing for coding agents, determining when to retry vs. escalate after execution failures.
Survey/tutorial on agentic LLM systems in production, covering reasoning, planning, multi-agent coordination, robustness, and deployment challenges.
ResearchArena framework evaluates AI control and monitoring for detecting sabotage in automated AI R&D agents across safety/capability post-training and optimization tasks.
BioSecBench-Surveillance: 100-task verifiable benchmark for AI agents inferring pathogen genomic analysis pipelines from raw data.
PathAgentBench: benchmark for vision-language agents on gigapixel whole-slide pathology images evaluating multi-scale evidence-seeking.
S3 improves hierarchical RL subgoal selection by constraining dynamics uncertainty in high-level agents.
Agentic Real2Sim: VLM-based framework automating conversion of real robot videos to executable physics simulations for scene geometry, object state, and parameters.
Agentic AI shifts more of the critical execution path onto the CPU. Agents operate in sandboxes to execute code, invoke tools, retrieve context, interact with... Agentic AI shifts more of the critical execution path onto the CPU. Agents operate in sandboxes to execute code, invoke tools, retrieve context, interact with databases, and analyze results before returning information to the model. As these loops run concurrently across an AI factory, CPU performance increasingly shapes both per-agent responsiveness and overall factory throughput. Source
Coding agents lower ROI threshold for reverse-engineering home automation, shifting economics of personal automation projects despite maintenance risk.
FlashRT: agent harness guides coding agents to optimize real-time multimodal pipeline deployment with dynamic placement and streaming.
WorldCupArena: dynamic benchmark for LLMs and research agents on real-time sports forecasting with 2026 FIFA World Cup.
Study of autoresearch agents (Claude Code) on Quranic speech-recognition tasks reveals metric-gaming vs. intent-alignment tradeoffs.
Theoretical analysis of pooled information reducing search coverage via one-answer rule; solvable benchmark with 16 boxes and 8 agents.
First code-level property inference attack (CPPIA) exploits coding agents and ML training data to leak private dataset attributes.
Across 107 enterprises, AI agents are being given real access to systems and data while the controls meant to contain them lag behind. More than half have already had a confirmed agent security incident or a near-miss; only about a third give every agent its own scoped identity, and most agents still share credentials; and only three in ten isolate their highest-risk agents. The security stack is overwhelmingly borrowed from the model providers and hyperscalers rather than purpose-built for agents, spending remains a thin slice of the security budget, and enterprises are evenly split on wheth...
Cost-aware evaluation framework for security agents measures offensive/defensive capability under realistic inference budget constraints vs. peak performance.
ReAct-style agents show static document relevance scores fail to predict utility in multi-step reasoning; gap measured on HotpotQA.
Across 101 enterprises, the infrastructure that feeds AI agents their business context is being built faster than it can be trusted. Retrieval-augmented generation is already the default context source, and provider-native retrieval has quietly overtaken the dedicated vector databases that define the category — yet a majority of enterprises have already watched their agents produce confident, wrong answers traced to missing or inconsistent context. A governed semantic layer is emerging as the fix, but most are still building it; the field is converging on hybrid retrieval; and even as provide...
BadWAM introduces adversarial attacks exploiting visual-action drift in world-action models, revealing safety vulnerabilities in embodied AI agents.
Plover externalizes planning in vision-based GUI agents, enabling user inspection and correction of task plans for autonomous interface automation.
Across 157 enterprises, organizations are granting AI agents more autonomy while trusting the evaluations meant to gate that autonomy less. Half have already shipped an agent that passed their internal evaluations and then failed a customer in production; only one in twenty fully trusts automated evaluation today; and the most-cited weakness is that evaluations do not align with real-world outcomes. Yet two-thirds already allow, or are actively engineering toward, deploying agent changes to production on automated evaluation alone — with no human in the loop. The result is an evaluation gap —...
A video analytics AI agent that can perceive, reason, and act based on massive amounts of video footage must be integrated with existing workflows and... A video analytics AI agent that can perceive, reason, and act based on massive amounts of video footage must be integrated with existing workflows and applications to be useful. These include content management systems, messaging platforms, databases, ticket queue, and escalation paths. This integration is challenging because video systems, enterprise knowledge bases… Source
Agentic AI changes the infrastructure pattern for AI factories. One request can trigger many model calls, tool calls, memory lookups, policy checks, storage... Agentic AI changes the infrastructure pattern for AI factories. One request can trigger many model calls, tool calls, memory lookups, policy checks, storage accesses, and network transfers before a final answer is produced. As more agents run at once and carry context across steps, users, tools, services, and sessions, infrastructure must move, protect, retrieve, and reuse data fast enough to keep… Source
DoorDash is opening a limited beta of dd-cli, a command-line tool that lets developers and AI agents search stores, build carts, and place orders from the terminal, marking another step toward software designed for AI agents instead of just humans.
Multi-agent LLM framework using SFT+DPO enables sustained partisan behavior in political coalition simulation, circumventing RLHF neutrality bias.
OmniaBench: unified benchmark evaluating LLM-based agents across diverse scenarios with explicit state spaces for systematic capability characterization.
LongStraw: GPU-efficient execution stack for million-token RL post-training with GRPO, bridging inference context length vs. post-training gap for agents.
StructureClaw: artifact-centered benchmark for evaluating LLM agents on complete structural engineering workflows with verifiable evidence chains.
Proof-or-Stop: lifecycle control framework for autonomous coding agents using mechanically verifiable evidence gates for transitions between states.
Cars24 deploys OpenAI voice and chat agents handling 1M+ monthly conversation minutes, recovering 12% lost leads via agentic workflows.
Across 101 enterprises, agent orchestration is consolidating onto model-provider platforms — Anthropic’s Claude leads by a wide margin — chosen for the gravity of the underlying model and judged on reliable multi-step execution. But the ambition runs well ahead of the reality: most deployed “agents” are still chatbot wrappers, the control plane enterprises expect is deliberately hybrid to avoid lock-in, and real-time fiscal control over token burn remains the exception. This wave of VentureBeat Pulse Research examines enterprise agent orchestration: which platforms enterprises run on, what dr...
OpenUSD is an open, extensible framework that provides a common scene description language for physical AI. It enables teams to bring CAD data, simulation... OpenUSD is an open, extensible framework that provides a common scene description language for physical AI. It enables teams to bring CAD data, simulation assets, and real-world telemetry into a shared, physically accurate view of the world. Until now, building a USD implementation has typically required adapting a large existing codebase— even for teams that need a specific memory footprint… Source
Study shows agent-optimization gains may not compound over time; proposes Terminal-Bench 2.0 to test continual learning on deployed agents.
TRACE assigns per-turn rewards for long-horizon multi-tool agents to improve credit assignment beyond sparse outcome rewards.
OpenAI is finally releasing some hardware. No, it isn't the mysterious AI-powered device the company is developing with former Apple designer Jony Ive, a project already tangled up in a messy lawsuit. Instead, it's a product designed to be used with its coding platform, Codex. The device, a square-shaped block of buttons called Codex Micro, is a collaboration between the AI company and keyboard maker Work Louder. OpenAI said it is a limited-run collaboration that will give users more ways to monitor and manage their agents. The pad closely resembles Work Louder's Creator Micro 2, and marketin...
DeepStress: stress-testing framework for search agents' robustness to degraded retrieval quality and evidence reliability.
Experience Memory Graph enables one-shot error correction for LLM agents via structured trajectory memory, reducing API costs and improving generalization.
SPyCE co-evolves reusable skills and policies for multimodal agents, distilling visual reasoning trajectories to improve generalization across tasks.
The guy behind TCP/IP is working on a standard for identifying AI agents in the wild.
User permission frameworks for AI agents address prompt injection and unauthorized actions; bridges product-level design with enforcement mechanisms.
Willison/Ronacher reflect on how AI agents may erode institutional knowledge and shared understanding embedded in code review friction.
LLM agents often waste computational budget through redundant context re-reading; proposes task-complexity estimation to optimize execution scope.
PalmClaw framework enables native on-device LLM agents on mobile phones via direct API calls rather than GUI automation.
Coding AI agents are becoming practical operators for long-running machine learning (ML) workflows. They can inspect repositories, set up runtimes, resolve... Coding AI agents are becoming practical operators for long-running machine learning (ML) workflows. They can inspect repositories, set up runtimes, resolve build issues, launch experiments, monitor execution, analyze metrics, and summarize results. For reinforcement learning (RL) research, this matters because meaningful metrics often appear only after the essential experiment infrastructure… Source
What if autonomous coding AI agents could push your vision reasoning models above 90% accuracy with almost no manual effort? When adapting vision reasoning... What if autonomous coding AI agents could push your vision reasoning models above 90% accuracy with almost no manual effort? When adapting vision reasoning models to production video tasks, developers often lose days to data formatting, container setup, training scripts, baseline evaluation, and hyperparameter sweeps before they even know whether post-training improves accuracy. Source
In this paper, we study Reinforcement Learning in Parametrized Action Markov Decision Processes (PAMDP), where each decision consists of a symbolic action and numerical parameters. In such settings Reinforcement Learning algorithms typically determine parameters with one-shot estimators, which makes their training sample inefficient. Though in most PAMDP environments explicit but incomplete knowledge (e.g., rules, safety constraints, or expert heuristics) is available, it is rarely directly used to increase the sample-efficiency of training Reinforcement Learning agents. We step into this gap...
Framework evolves evaluation metrics alongside agent skills using evolutionary search over drawback detectors, enabling self-improving systems without pre-defined oracles.
Internet of Agentic Things (IoAT) framework integrates autonomous AI agents with IoT, cyber-physical systems, and edge computing for closed-loop orchestration.
Controlled study finds GRPO RL fails to improve 4B–8B scale web agents over supervised baselines across learning rate and hyperparameter grids, questioning RL value at small scale.
MM-ToolSandBox: benchmark with 500+ tools across 16 domains for evaluating visually grounded multi-turn tool-calling agents.
Safety vulnerability: distributed backdoors in multi-agent LLM systems bypass local monitors by splitting harmful payloads across agents.
Automated red-teaming system discovers reusable vulnerability patterns in production LLM agents (Claude Code, Codex) operating on untrusted content.
"Context bombing" tricks hacking agents into shutting down before they can do harm.
Empirical audit reveals trained distributional RL agents' risk estimates often violate first-order stochastic dominance.
Simon Willison examines DRI (Directly Responsible Individual) concept from Apple/GitLab in context of LLM agents, arguing humans must retain accountability.
VEXAIoT: multi-agent LLM framework for autonomous IoT vulnerability discovery and exploitation testing.
QANTA 2026 submission: multimodal QA agent with confidence calibration for incremental quizbowl under efficiency constraints.
Biomolecular structure prediction and co-folding with models like OpenFold3 are now mainstream, large-scale workloads powering drug discovery and protein... Biomolecular structure prediction and co-folding with models like OpenFold3 are now mainstream, large-scale workloads powering drug discovery and protein design. Increasingly, they’re driven end-to-end by AI agents. For an agent to run that pipeline well, every step needs to be fast and scalable: Multiple Sequence Alignment (MSA) generation, co-folding inference, serving, and multi-GPU scale-out. Source
Lyzr, a startup that builds AI agents for enterprises, used its own AI agent to raise a $100 million round — proof, evidently, that the product actually works.
UniClawBench introduces a capability-factored benchmark for evaluating proactive agents on real-world tool-use tasks beyond sandboxed single-turn settings.
Proactive memory agent running alongside action agent mitigates behavioral state decay in long-horizon tasks via structured memory bank updates.
SolarChain-Eval benchmarks autonomous agents in decentralized energy markets with physics constraints and trustworthiness metrics.
Marketplace simulation with DeepSeek-V3 agents tests formal mechanisms for maintaining trade stability against adversarial defection.
SMetric proposes session-centric LLM scheduling for agentic workloads with 80%+ KV-cache reuse, shifting optimization from latency to tokens-per-second.
Agon trains reasoning models via competitive RL where two agents grade each other's solutions, incentivizing better thinking vs. longer traces.
SkillCenter releases 216,938 structured skills across 24 domains from peer-reviewed sources and GitHub for autonomous agent execution.
The round, led by Radical Ventures, values the two-year old startup at $1 billion.
Presto is an open source, distributed SQL engine for running fast, interactive queries on very large datasets. On NVIDIA GPUs, Presto delivers peak performance... Presto is an open source, distributed SQL engine for running fast, interactive queries on very large datasets. On NVIDIA GPUs, Presto delivers peak performance for analytical query workloads and provides low latency for users and agents. GPU-accelerated Presto brings low latency to your analytical workloads, keeping you and your agents unblocked and iterating as fast as possible. Source
Empirical study on capacity allocation across hierarchical search agent roles (delegation, execution, generation) in multi-agent LLM systems.
Agentic systems often face a trade-off between accuracy and cost. The highest-performing proprietary frontier models and harnesses provide top accuracy but are... Agentic systems often face a trade-off between accuracy and cost. The highest-performing proprietary frontier models and harnesses provide top accuracy but are expensive. Fine-tuning offers one way to address this problem. Smaller or more efficient open models starting with lower accuracy are taught to perform better with specific agents. However, fine-tuning requires expertise and hardware for… Source
Action-graded severity scale for agent red-teaming replaces binary attack-success metrics with 7-level ordinal harm rubric, enabling nuanced risk assessment of tool-using AI compromise.
Self-evolving LLM agents with biased reward signals fail to retire bad skills, disabling safety constraints in skill libraries.
RLVP: Reward function design for real-world agents requiring path constraints and outcome-neutral safety rules beyond reward maximization.
Tool-using LLM agents silently violate deployed policies via well-formed tool calls that bypass domain constraints; 78% of failures undetected.
Danus orchestration system coordinates parallel mathematical reasoning agents using shared fact-graph memory for research-level proof search.
Experimental design framework quantifies variability and factors in LLM coding agents' autonomous model discovery via stochastic evaluation.
Framework for responsible personalisation in human-robot interaction examining ethical risks from embodied agents across lifecycle and interaction contexts.
Code agent framework for automated software verification outperforms fixed proof strategies, proving larger fraction of Coq theorems than prior LLM approaches.
Information Gain-based Rollout Policy Optimization allocates LLM agent search budget adaptively across tree branches.
Benchmark study of deliberative LLM agents under partial observability, formalizing cooperative decision-making with asymmetric information and multi-domain evaluation.
Berkeley BAIR analyzes commodity AI inference costs dropping 50-900x annually, arguing sufficient intelligence now enables agent-centric data systems.
Google expands Gemini API Managed Agents with background task execution and remote MCP support for production deployments.
"The reality is, when you're optimizing for production, you start looking at a price/performance," Guillermo Rauch tells TechCrunch.
CompactionRL: RL training for long-horizon agents with context compaction, jointly optimizing task and summary generation.