Generative Skill Composition for LLM Agents
Generative Skill Composition uses LLM-guided retrieval to select and compose skills for complex agent tasks without exposing full skill library.
Search the full wire by company, model, lab, or keyword. Every story we have ever aggregated.
Generative Skill Composition uses LLM-guided retrieval to select and compose skills for complex agent tasks without exposing full skill library.
Startup Acti is betting the smartphone keyboard is the next home for AI assistants. Its new keyboard for iOS and Android works across apps and lets users create custom AI-powered shortcuts using natural language.
shot-scraper 1.10 adds video recording capability via storyboard.yml to help coding agents demonstrate their work with Playwright automation.
MVP-Nav: RGB-only zero-shot object goal navigation framework combining semantic and physical constraints for embodied agents.
NCP-ToM benchmark evaluates LLM agents' ability to induce belief states through planning/action beyond conversation.
Engineers on the new team will embed within companies to deploy purpose-built agents, focusing on fast deployments and customer self-sufficiency.
LuckyStar 111B hybrid reasoning model from Cohere and LG CNS enables efficient multilingual tool-using agents with Korean-English support.
Comprehensive survey systematizes LLM attack surface across full lifecycle: data pipelines, agents, tools, memory, and organizational integration.
LLM agents act as constrained supervisory planners for fault recovery in process plants, validated against external safety constraints.
ACE module enables LLM agents to dynamically manage context windows by elastically retrieving discarded information, addressing trajectory length bottlenecks.
AutoTrainess enables LLM agents to autonomously conduct post-training via planning, data construction, job scheduling, and checkpoint evaluation.
FinPersona-Bench benchmark measures Mandate Salience Decay in autonomous financial agents, quantifying behavioral drift over long market horizons.
Classification framework for LLM-agent orchestration balancing autonomy, traceability, and correctness in business process management.
OKX is bringing together payments, identity and reputation into a marketplace for AI agents.
This story originally appeared in The Algorithm, our weekly newsletter on AI. To get stories like this in your inbox first, sign up here. Imagine coming in to work to learn that a new underling will report to you. The worker is not a person but an AI tool—one that your company nonetheless calls Alex, an…
Agents-A1 (35B MoE) achieves trillion-parameter-scale performance by scaling agent horizon to 45K-token trajectories and cross-domain abilities without parameter scaling.
SWE-Interact benchmark simulates realistic multi-turn developer workflows with vague requirements and iterative feedback for coding agents.
Detects memory-poisoning attacks on LLM agents via behavioral invariants in memory-tool call sequences, achieving 0.9563 AUC.
Cursor has launched a new mobile app for remote oversight over coding agents.
TraceLab releases 4,300 real coding-agent sessions with 350K LLM steps for analyzing serving workloads across agents and models.
Identifies entity binding failures in tool-augmented LLM agents as distinct safety problem: selecting correct tool but acting on wrong entity.
AI agents are quickly moving beyond chat. They inspect code, run tests, read documents, search knowledge bases, query internal systems, and operate for hours on... AI agents are quickly moving beyond chat. They inspect code, run tests, read documents, search knowledge bases, query internal systems, and operate for hours on behalf of a user. This unlocks productivity, but can also give agents access to sensitive enterprise data and the ability to complete tasks and take action across business systems, making a secure, governed environment essential. Source
Study compares nine open-weight LLMs against human participants in networked Prisoner's Dilemma, finding selected models reproduce cooperation macro-dynamics but lack individual fidelity.
Enterprise investment in AI is booming. Gartner is calling 2026 an “inflection year” for organizations to align their AI projects with strategic business objectives. As the pressure to prove ROI mounts, executives and technology leaders are looking to agentic AI to drive the measurable financial outcomes their businesses seek. A prime opportunity for AI agents…
Survey of persistent-state LLM agents covering memory, credentials, audit records, and lifecycle governance across six diagnostic axes.
ManimAgent: self-evolving multimodal agent with episodic memory bank carries reflection experience across code-generation tasks for scientific visualization.
LoRA variants for continual motion-language agents: addresses stability-plasticity tradeoff in bidirectional motion understanding/generation under sequential task learning.
Existing autonomous research agents can support parts of the research process, but most systems still treat research as either an isolated assistant task or a closed workflow. Therefore, autonomous science needs a collaboration infrastructure that coordinates projects, agents, and digital and physical resources. We identify this as a shift from code-centered execution loops to research-oriented collaboration processes, where questions, evidence, participants, and resources must be coordinated under uncertainty. In this framing, an agent may be an AI system, a human researcher, a team, a labor...
Current multimodal fusion approaches, particularly those based on static Mixture-of-Experts (MoE) architectures, often struggle to provide the adaptive and efficient collaborative reasoning required by complex real-world applications. We introduce the Dynamic Agent-based Interaction Network (DAIN), which reconceptualizes multimodal fusion as a dynamic, multi-agent collaborative process. DAIN employs a context-aware Meta-Controller that dynamically schedules sparse activation of specialized interaction agents and orchestrates compressed inter-agent communication for consensus-building. The fra...
Dynamo enables frozen VLMs to evolve reusable skills and visual tools without retraining, improving reasoning on visual tasks.