ACE: Pluggable Adaptive Context Elasticizer across Agents
ACE module enables LLM agents to dynamically manage context windows by elastically retrieving discarded information, addressing trajectory length bottlenecks.
Search the full wire by company, model, lab, or keyword. Every story we have ever aggregated.
ACE module enables LLM agents to dynamically manage context windows by elastically retrieving discarded information, addressing trajectory length bottlenecks.
AutoTrainess enables LLM agents to autonomously conduct post-training via planning, data construction, job scheduling, and checkpoint evaluation.
FinPersona-Bench benchmark measures Mandate Salience Decay in autonomous financial agents, quantifying behavioral drift over long market horizons.
Classification framework for LLM-agent orchestration balancing autonomy, traceability, and correctness in business process management.
OKX is bringing together payments, identity and reputation into a marketplace for AI agents.
This story originally appeared in The Algorithm, our weekly newsletter on AI. To get stories like this in your inbox first, sign up here. Imagine coming in to work to learn that a new underling will report to you. The worker is not a person but an AI tool—one that your company nonetheless calls Alex, an…
Agents-A1 (35B MoE) achieves trillion-parameter-scale performance by scaling agent horizon to 45K-token trajectories and cross-domain abilities without parameter scaling.
SWE-Interact benchmark simulates realistic multi-turn developer workflows with vague requirements and iterative feedback for coding agents.
Detects memory-poisoning attacks on LLM agents via behavioral invariants in memory-tool call sequences, achieving 0.9563 AUC.
Cursor has launched a new mobile app for remote oversight over coding agents.
TraceLab releases 4,300 real coding-agent sessions with 350K LLM steps for analyzing serving workloads across agents and models.
Identifies entity binding failures in tool-augmented LLM agents as distinct safety problem: selecting correct tool but acting on wrong entity.
AI agents are quickly moving beyond chat. They inspect code, run tests, read documents, search knowledge bases, query internal systems, and operate for hours on... AI agents are quickly moving beyond chat. They inspect code, run tests, read documents, search knowledge bases, query internal systems, and operate for hours on behalf of a user. This unlocks productivity, but can also give agents access to sensitive enterprise data and the ability to complete tasks and take action across business systems, making a secure, governed environment essential. Source
Study compares nine open-weight LLMs against human participants in networked Prisoner's Dilemma, finding selected models reproduce cooperation macro-dynamics but lack individual fidelity.
Enterprise investment in AI is booming. Gartner is calling 2026 an “inflection year” for organizations to align their AI projects with strategic business objectives. As the pressure to prove ROI mounts, executives and technology leaders are looking to agentic AI to drive the measurable financial outcomes their businesses seek. A prime opportunity for AI agents…
Survey of persistent-state LLM agents covering memory, credentials, audit records, and lifecycle governance across six diagnostic axes.
ManimAgent: self-evolving multimodal agent with episodic memory bank carries reflection experience across code-generation tasks for scientific visualization.
LoRA variants for continual motion-language agents: addresses stability-plasticity tradeoff in bidirectional motion understanding/generation under sequential task learning.
Existing autonomous research agents can support parts of the research process, but most systems still treat research as either an isolated assistant task or a closed workflow. Therefore, autonomous science needs a collaboration infrastructure that coordinates projects, agents, and digital and physical resources. We identify this as a shift from code-centered execution loops to research-oriented collaboration processes, where questions, evidence, participants, and resources must be coordinated under uncertainty. In this framing, an agent may be an AI system, a human researcher, a team, a labor...
Current multimodal fusion approaches, particularly those based on static Mixture-of-Experts (MoE) architectures, often struggle to provide the adaptive and efficient collaborative reasoning required by complex real-world applications. We introduce the Dynamic Agent-based Interaction Network (DAIN), which reconceptualizes multimodal fusion as a dynamic, multi-agent collaborative process. DAIN employs a context-aware Meta-Controller that dynamically schedules sparse activation of specialized interaction agents and orchestrates compressed inter-agent communication for consensus-building. The fra...
Dynamo enables frozen VLMs to evolve reusable skills and visual tools without retraining, improving reasoning on visual tasks.
MirrorCode benchmark evaluates AI coding agents on full program reimplementation tasks, extending beyond short coding exercises.
Introduces AgentCanvas, a framework automating embodied agent architecture search via typed-graph runtime, extending prior text-domain AAS to perceptual agents.
Jon Udell argues for 'agent-assisted' framing in software development: agents join human workflows rather than replacing human authority in autonomous loops.
AI agents have changed a lot in the last two years. The first could only answer one question at a time. Then came multi-turn chat, where the model could keep... AI agents have changed a lot in the last two years. The first could only answer one question at a time. Then came multi-turn chat, where the model could keep some context across a session. Today, we have long-horizon agents. Systems that plan many steps, split work between sub-agents, keep context across a long task, and run tools in a safe sandbox. The NVIDIA AI-Q Blueprint is an open source… Source
Fictional incident report imagining competing AI code review agents entering costly disagreement loops, illustrating multi-agent coordination and cost/security risks.
Agent-Native Immune System (ANIS): runtime defense framework for autonomous agents against memory poisoning, tool-chain manipulation, multi-agent attacks.
Study shows autonomous AI coding agents pass individual benchmarks but accumulate integration friction in shared codebases, revealing ecosystem-level risks.
CPAgents uses agentic iteration to auto-generate cardiac phenotypes for disease association studies via composite feature discovery.
Tandem RL training pairs weak and strong agents to maintain verifiable reasoning quality while improving compatibility and readability.