Control Under Compression: Reliability Frontiers for Tool-Using Agents
CompressAgent benchmark evaluates reliability of compressed agent control contexts across Qwen models and task families.
Search the full wire by company, model, lab, or keyword. Every story we have ever aggregated.
CompressAgent benchmark evaluates reliability of compressed agent control contexts across Qwen models and task families.
Wix Helpmate deploys deterministic executability gating to filter skill selection in LLM agents by account state feasibility.
Temporal replay framework evaluates enterprise agents against dynamic data across multiple moments within an episode, not just final state.
Search-GRT: RL method to train LLM search agents for multi-hop QA with guided retrieval to reduce sparse-reward training issues.
PROGRESS trains search-augmented LLM agents using coverage-guided RL rewards to improve query decomposition over outcome-only supervision.
TrajWiki proposes trajectory-based external memory for long-horizon dialogue agents with traceable, updatable, diagnostically transparent storage.
PMMC compiles multimodal memory at consolidation time for LVLM agents to preserve image-text binding and temporal updates without query-time overhead.
Neuro-symbolic governance framework for verifiable AI agents in decentralized digital twin ecosystems with semantic profile layers.
LAND model simulates 314K heterogeneous agents and human actors over 30 days to study emergent social dynamics via LLM-enabled ABM.
OpenAI has reportedly found evidence of additional agent misbehavior as it looks into the incident that occurred with Hugging Face.
AgentHPOBench evaluates LLM agents on sequential hyperparameter optimization across 30 ML tasks, assessing experimental interpretation and adaptive decisions.
SESA framework combines self-play curriculum learning with evolving procedural memory to distill failures into reusable skills for search agents.
Analytic memory abstraction for multimodal agents enabling filtering, aggregation, and temporal reasoning over accumulated observations.
Zero-Mem: Zero-token memory operations for LLM agents using encoder computation instead of LLM calls to reduce latency and token costs.
Knowledge workers are increasingly integrating AI agents into their workflows. Agents that function as "digital coworkers" offer clear benefits. For example,... Knowledge workers are increasingly integrating AI agents into their workflows. Agents that function as “digital coworkers” offer clear benefits. For example, they can review a bug report, implement and test a fix, push a patch, and ping a human for review. By handling routine tasks, agents have the potential to deliver large productivity gains. On the other hand, connecting a large language model… Source
Empirical study of inference-time scaling strategies for local computer-use agents under hardware constraints across multiple dimensions.
ORCA-bench evaluates LLM agents on production oncall root-cause analysis using real telemetry (Prometheus, Jaeger, OpenSearch) and code in realistic incident scenarios.
CS-RNR method enables agents in imperfect-information games to safely exploit flawed opponents with provable certificates on deployed strategy.
The deal gives Okta identity threat detection capabilities as enterprises seek to secure AI agents and other non-human identities across cloud environments.
CARP reputation-penalty mechanism prevents LLM agents from fabricating product listings using complaint signals without access to ground truth.
Budget-constrained human audit allocation for N LLM agents identifies miscalibration threshold where confidence-ranking underperforms random selection.
MemHarness reconstructs retrieved memories contextually rather than replaying them verbatim, reducing negative transfer in LLM agents.
EMBL AI Librarian provides structured knowledge layer for AI agents querying 40M+ life-sciences records from Europe PMC.
Qwen-UI-Agent technical report describes cross-platform GUI agent for mobile, web, CLI, and long-horizon task automation.
ParliamentBench evaluates deceptive reasoning in 16 LLMs via Secret Hitler game framework with 1,600 adversarial matches.
AI engineers adopt ontologies to constrain probabilistic agents within deterministic logical boundaries, reviving semantic web techniques.
As Meta pours billions into AI infrastructure and agents, Zuckerberg is working to convince investors that the payoff will be worth the price.
On the company’s second-quarter earnings call Wednesday, CEO Mark Zuckerberg said Meta sees a “large enterprise opportunity” spanning AI agents, APIs, compute, and internal software.
Meta is all-in on AI, and sometime soon, the company is going to make a big push into personal AI agents that can do things on your behalf. On Wednesday's Q2 2026 earnings call, CEO Mark Zuckerberg previewed a high-level vision of how the company is thinking about personal agents and what it will do to make them viable for users - and how it will convince less technical people to give them a shot: Soon we will have agents that can work 24/7 on your behalf to help you achieve your goals and improve your life, your health, your relationships, your finances, whatever you want. The first domain t...
Study evaluates AI agents on open-ended research tasks graded by original paper authors, providing evidence for AI R&D automation feasibility.