OmegaUse-OfficeVal: Benchmarking LLM Agents on Long-Horizon Office-Suite Tasks with Economic Grounding
OmegaUse-OfficeVal benchmark evaluates LLM agents on 100 long-horizon office-suite tasks with cost grounding and human labor baselines.
Search the full wire by company, model, lab, or keyword. Every story we have ever aggregated.
OmegaUse-OfficeVal benchmark evaluates LLM agents on 100 long-horizon office-suite tasks with cost grounding and human labor baselines.
Cost-aware stopping mechanism (CAM-DF) for LLM agent tool acquisition balancing task coverage against cost, context load, and privacy.
Setoka benchmark evaluates hierarchical user understanding in memory-augmented personalized agents beyond explicit fact retrieval.
AgentSnare uses adaptive deceptive observations to mislead LLM-based penetration testing agents, defeating static artifact recognition.
The startup analyzes calls, messages and CRM data to identify effective sales techniques and turn them into playbooks for AI agents.
TREK benchmark evaluates LLM agents on complex, executable travel itinerary planning with verifiable constraints.
Three-class detection framework distinguishes humans, bots, and AI agents in browser automation traffic; binary classifiers misclassify 39.1% of agents as human.
Two-call self-refinement outperforms five-agent pipeline on Qwen2.5-7B; multi-agent systems suffer error accumulation, dropping GSM8K accuracy to 45% with JSON format.
TSDS framework deploys ReAct agents at edge via convergence probe for reasoning budget and perplexity-based deferral to cloud model.
The deal is Cyera's third acquisition this year.
Desktop-Delta Bench isolates whether computer-use agents understand GUI state transitions caused by actions, beyond end-task success metrics.
UniMem hybrid episodic-parametric memory for LLM agents resolves stability-plasticity dilemma across boundary-agnostic evolving task streams.
MemLens provides value-aware memory management for LLM agents with interactive analytics, treating memory records as first-class objects to reduce redundancy.
OpenAI field report documents how AI coding agents accelerate scientific computing workflows in genomics and adjacent domains.
Google ships Gemini API Managed Agents with Gemini 3.6 Flash and tool-use hooks for production agent deployment.
Messier consolidates 957K evaluation records across 30 benchmarks and 714 agents to enable comparable cross-benchmark agent assessment.
Custom harness engineering distributes security controls for AI coding agents without vendor lock-in, enabling organizational scaling.
RSIBench-Data isolates LLM agents' data-centric research capability in post-training loops, decoupling research from systems optimization.
OpenAI's Akshay Nathan details ChatGPT Work product strategy: Sites, memory, subagents, finance, no-code tools scaling from 0 to 10M users.
HiSkill hierarchical skill graph framework organizes LLM agent trajectories into structured graphs linking high-level skills to atomic operations for long-horizon tasks.
Joint agent-speculator RL aligns next tool-call prediction with deployed agent behavior by unifying speculator and agent in single model, reducing latency.
WorkSurface-Bench evaluates enterprise agents on multi-source knowledge routing across documents, tables, and graphs with 1,151 tasks.
HYSET evaluates tool sets jointly for LLM agents via hyperedge prediction, improving on sequential tool retrieval.
Perplexity has expanded its agentic Personal Computer tool to Windows, allowing computers running the world's most popular OS to be used as a locally run AI system. Like the Mac version that Perplexity launched in April, Personal Computer for Windows operates like a "general-purpose digital worker" that can access local files and apps to perform actions on your behalf, such as creating documents and updating spreadsheets. This launch builds on Personal Computer integrations that Perplexity launched for Microsoft's 365 workspace apps and Teams virtual meeting software in May. Personal Computer...
Multi-turn long-horizon planning for foundation agents via controlled pre/post-training with single/multi-teacher on-policy distillation.
APPA: IFC framework for LLM agents handling mixed-confidentiality data via context branching and prospective access control against injection attacks.
Study of correctness degradation in code repair agents under forced revision loops; 30 HumanEval benchmarks show revision doesn't guarantee reliability.
SIREN applies LLM agents to end-to-end extreme-weather early warning with multi-step decision workflows.
Evaluates fuzz testing methods for safety-critical RL agents across robotics and autonomous systems with standardized metrics.
Robotics learns from bitter lesson scaling; AI agents complete week-long programming tasks; OpenAI discovers accidental vulnerability in reasoning model.