Danus: Orchestrating Mathematical Reasoning Agents with Fact-Graph Memory
Danus orchestration system coordinates parallel mathematical reasoning agents using shared fact-graph memory for research-level proof search.
Search the full wire by company, model, lab, or keyword. Every story we have ever aggregated.
Danus orchestration system coordinates parallel mathematical reasoning agents using shared fact-graph memory for research-level proof search.
Experimental design framework quantifies variability and factors in LLM coding agents' autonomous model discovery via stochastic evaluation.
Framework for responsible personalisation in human-robot interaction examining ethical risks from embodied agents across lifecycle and interaction contexts.
Code agent framework for automated software verification outperforms fixed proof strategies, proving larger fraction of Coq theorems than prior LLM approaches.
Information Gain-based Rollout Policy Optimization allocates LLM agent search budget adaptively across tree branches.
Benchmark study of deliberative LLM agents under partial observability, formalizing cooperative decision-making with asymmetric information and multi-domain evaluation.
Berkeley BAIR analyzes commodity AI inference costs dropping 50-900x annually, arguing sufficient intelligence now enables agent-centric data systems.
Google expands Gemini API Managed Agents with background task execution and remote MCP support for production deployments.
"The reality is, when you're optimizing for production, you start looking at a price/performance," Guillermo Rauch tells TechCrunch.
CompactionRL: RL training for long-horizon agents with context compaction, jointly optimizing task and summary generation.
SovereignPA-Bench evaluates user-owned personal agents on privacy, consent, and user sovereignty across evolving intent and platform mediation.
Graph Sparse Sampling reduces exponential sampling complexity in continuous MDP planning by sharing futures across multiple agents and lookahead depths.
First multiplayer world model for dynamic environments attributes scene changes to multiple agents' actions, trained on 10K hours of Rocket League gameplay.
MetaSkill-Evolve enables recursive self-improvement of LLM agents; agents rewrite and evolve their own improvement procedures from execution traces.
Defense mechanism for web agents against prompt injection via content masking with formal security guarantees.
Stealth memory injection attack on persistent personal agents: adversary injects poisoned data via untrusted content, remains hidden, affects future behavior.
Residual streams in coding agents linearly encode program properties (parse status, test results, regressions) up to 0.83 AUC.
GameEngineBench evaluates coding agents on C++ implementation tasks in Unreal Engine real-time environments.
SkillOpt-Lite: minimal skill optimization pipeline for autonomous agents via zeroth-order optimization, grounded in Claude.
Continuous test-time training for LLM agents across multi-turn episodes, balancing adaptation with drift prevention.
Vercel's Andrew Qu discusses eve agent framework, emphasizing skills, sandboxes, and agent-readable web design as architectural primitives.
At an internal meeting, the Meta CEO reportedly said that AI development efforts were not moving as quickly as anticipated.
Simon Willison uses DSPy to evaluate and optimize SQL system prompts in Datasette Agent, a tool enabling agents to execute read-only database queries.
Introduces Iterative VibeCoding, a benchmark for studying prompt-injection and code-injection attacks in autonomous AI agents shipping iterative PRs to persistent codebases.
Dual-channel debate framework reveals that LLM agents develop latent objectives and shift public utterances based on social structure, audience, and role without explicit prompting.
CNeVA: controllable simulated agents with interpretable behavior latents for realistic traffic simulation, enabling edge-case testing without real-world risk.
TestEvo-Bench executable benchmark evaluates test-code co-evolution for automated testing agents; mined from real repositories.
EvoPolicyGym benchmark for autonomous policy evolution; evaluates how agents iteratively improve executable policies under fixed budgets.
Geoffrey Litt argues developers must maintain deep code comprehension to effectively collaborate with AI coding agents and avoid cognitive debt.
Constraint-based oversight for coding agents scales cheaper than scaffolding; access control and tooling principles from human teams transfer directly.