The Archive
Search the full wire by company, model, lab, or keyword. Every story we have ever aggregated.
Prime Intellect raises $130M Series A to help enterprises build their own AI agents
The round, led by Radical Ventures, values the two-year old startup at $1 billion.
Running Low-Latency Analytical Workloads with GPU-Accelerated Presto on NVIDIA GB200 NVL72
Presto is an open source, distributed SQL engine for running fast, interactive queries on very large datasets. On NVIDIA GPUs, Presto delivers peak performance... Presto is an open source, distributed SQL engine for running fast, interactive queries on very large datasets. On NVIDIA GPUs, Presto delivers peak performance for analytical query workloads and provides low latency for users and agents. GPU-accelerated Presto brings low latency to your analytical workloads, keeping you and your agents unblocked and iterating as fast as possible. Source
Think Big, Search Small: Where Capacity Matters in Hierarchical Search Agents?
Empirical study on capacity allocation across hierarchical search agent roles (delegation, execution, generation) in multi-agent LLM systems.
Create a LangChain Deep Agents Harness Profile for NVIDIA Nemotron 3 Ultra to Improve Performance
Agentic systems often face a trade-off between accuracy and cost. The highest-performing proprietary frontier models and harnesses provide top accuracy but are... Agentic systems often face a trade-off between accuracy and cost. The highest-performing proprietary frontier models and harnesses provide top accuracy but are expensive. Fine-tuning offers one way to address this problem. Smaller or more efficient open models starting with lower accuracy are taught to perform better with specific agents. However, fine-tuning requires expertise and hardware for… Source
Beyond Attack-Success Rate: Action-Graded Severity Scale for Tool-Using AI Agents
Action-graded severity scale for agent red-teaming replaces binary attack-success metrics with 7-level ordinal harm rubric, enabling nuanced risk assessment of tool-using AI compromise.
The Blind Curator: How a Biased Judge Silently Disables Skill Retirement in Self-Evolving Agents
Self-evolving LLM agents with biased reward signals fail to retire bad skills, disabling safety constraints in skill libraries.
RLVP: Penalize the Path, Reward the Outcome
RLVP: Reward function design for real-world agents requiring path constraints and outcome-neutral safety rules beyond reward maximization.
Reason Less, Verify More: Deterministic Gates Recover a Silent Policy-Violation Failure Mode in Tool-Using LLM Agents
Tool-using LLM agents silently violate deployed policies via well-formed tool calls that bypass domain constraints; 78% of failures undetected.
Danus: Orchestrating Mathematical Reasoning Agents with Fact-Graph Memory
Danus orchestration system coordinates parallel mathematical reasoning agents using shared fact-graph memory for research-level proof search.
An Experimental Design Approach to Evaluating Agentic AI's Autonomous Model Discovery
Experimental design framework quantifies variability and factors in LLM coding agents' autonomous model discovery via stochastic evaluation.
Responsible Personalisation: The Double-Edged Sword of Personalisation in Human-Robot Interaction
Framework for responsible personalisation in human-robot interaction examining ethical risks from embodied agents across lifecycle and interaction contexts.
Harnessing Code Agents for Automatic Software Verification
Code agent framework for automated software verification outperforms fixed proof strategies, proving larger fraction of Coq theorems than prior LLM approaches.
Information Gain-based Rollout Policy Optimization: An Adaptive Tree-Structured Rollout Approach for Multi-Turn LLM Agents
Information Gain-based Rollout Policy Optimization allocates LLM agent search budget adaptively across tree branches.
LLM Agents for Deliberative Collaboration: A Study on Joint Decision Making Under Partial Observability
Benchmark study of deliberative LLM agents under partial observability, formalizing cooperative decision-making with asymmetric information and multi-domain evaluation.
Intelligence is Free, Now What? <br> Data Systems for, of, and by Agents
Berkeley BAIR analyzes commodity AI inference costs dropping 50-900x annually, arguing sufficient intelligence now enables agent-centric data systems.
Expanding Managed Agents in Gemini API: background tasks, remote MCP and more
Google expands Gemini API Managed Agents with background task execution and remote MCP support for production deployments.
Vercel CEO Guillermo Rauch on the fight to split off models from agents
"The reality is, when you're optimizing for production, you start looking at a price/performance," Guillermo Rauch tells TechCrunch.
CompactionRL: Reinforcement Learning with Context Compaction for Long-Horizon Agents
CompactionRL: RL training for long-horizon agents with context compaction, jointly optimizing task and summary generation.
SovereignPA-Bench: Evaluating User-Owned Personal Agents under Evolving Intent, Platform Mediation, and Consent Constraints
SovereignPA-Bench evaluates user-owned personal agents on privacy, consent, and user sovereignty across evolving intent and platform mediation.
Graph Sparse Sampling: Breaking the Curse of the Horizon in Continuous MDP Planning
Graph Sparse Sampling reduces exponential sampling complexity in continuous MDP planning by sharing futures across multiple agents and lookahead depths.
Multiplayer Interactive World Models with Representation Autoencoders
First multiplayer world model for dynamic environments attributes scene changes to multiple agents' actions, trained on 10K hours of Rocket League gameplay.
MetaSkill-Evolve: Recursive Self-Improvement of LLM Agents via Two-Timescale Meta-Skill Evolution
MetaSkill-Evolve enables recursive self-improvement of LLM agents; agents rewrite and evolve their own improvement procedures from execution traces.
Untrusted Content Masking for Web Agents with Security Guarantees
Defense mechanism for web agents against prompt injection via content masking with formal security guarantees.
When Claws Remember but Do Not Tell: Stealthy Memory Injection in Persistent Personal Agents
Stealth memory injection attack on persistent personal agents: adversary injects poisoned data via untrusted content, remains hidden, affects future behavior.
Latent Programming Horizons in Coding Agents
Residual streams in coding agents linearly encode program properties (parse status, test results, regressions) up to 0.83 AUC.
GameEngineBench: Evaluating Coding Agents on Real C++ Runtime Environments
GameEngineBench evaluates coding agents on C++ implementation tasks in Unreal Engine real-time environments.
SkillOpt-Lite: Better and Faster Agent Self-evolution via One Line of Vibe
SkillOpt-Lite: minimal skill optimization pipeline for autonomous agents via zeroth-order optimization, grounded in Claude.
No Time Like the Present: Agentic Test-Time Training for LLM Agents
Continuous test-time training for LLM agents across multi-turn episodes, balancing adaptation with drift prevention.
Vercel's Andrew Qu on why agents are a new kind of software
Vercel's Andrew Qu discusses eve agent framework, emphasizing skills, sandboxes, and agent-readable web design as architectural primitives.