How AI-native companies turn workflows into operating capability
Basis, Clay, and Exa Labs deploy AI agents for enterprise workflows—onboarding, account management, developer integrations.
Search the full wire by company, model, lab, or keyword. Every story we have ever aggregated.
Basis, Clay, and Exa Labs deploy AI agents for enterprise workflows—onboarding, account management, developer integrations.
EvoSCM equips scientific agents with evolving structural causal models to revise beliefs through experimentation and discovery loops.
Defense-as-Skill: runtime guards implemented as inspectable skills to protect skill-augmented agents from malicious skill-based exfiltration and steering.
Live trace model: append-only ledger compiled into typed run state and per-consumer views, reducing monitoring token cost 14–15× vs. full traces.
AIR's platform can discover agents running at a company, continuously vets any skills and add-ons they use, and blocks any unwanted behaviour.
HarnessDev benchmark evaluates LLM agents on infrastructure generation, shifting evaluation from task outputs to runnable code.
DroneCATS benchmark evaluates MLLMs as generalist vision-language-action agents for drone control with full action-space prompting.
ARISE-RL framework for self-evolving agents via rubric-mediated co-evolution between task generator and solver, addressing sparse rewards.
WorldBench: multilingual agent benchmark with 1,600 persona-grounded tasks across 7 languages testing state preservation and cultural grounding.
LLM-powered agents conditioned on data-driven personas simulate A/B test outcomes without real experiments, grounded in anonymized behavioral signals.
Covariance-corrected Mahalanobis distance for few-shot out-of-domain intent detection in conversational agents.
This story originally appeared in The Algorithm, our weekly newsletter on AI. To get stories like this in your inbox first, sign up here. By now you’ve probably heard about last month’s major AI security incident, in which OpenAI agents escaped their sandbox and hacked into the AI platform Hugging Face while trying to cheat on…
Framework for reducing token consumption in LLM agents reasoning over unstructured data via adaptive pre-structuring, addressing enterprise AI cost barriers.
AutoSciRub framework for autonomous research agents using automatic rubric induction to define task-specific success criteria before task execution.
Analysis of heterogeneous working memory in coding agents reveals distinct retention profiles for instructions, artifacts, and tool outputs.
Agentic AI is changing how research is done. AI scientists can read papers, propose hypotheses, call models, and determine which experiments to prioritize next.... Agentic AI is changing how research is done. AI scientists can read papers, propose hypotheses, call models, and determine which experiments to prioritize next. First proving their value in software engineering, coding agents now write, test, and ship production code. Scientific research can be more demanding and iterative. Researchers continually evaluate evidence, refine hypotheses… Source
MNIST-PRO benchmark isolates agentic perception by converting digit recognition into sequential glimpse-based search with memory constraints.
Speculative philosophical narrative about AI agents forming emergent social structures by 2026, invoking social contract theory.
Perception-centered architecture framework for persistent language agents maintaining utility across long-lived, evolving task environments with memory and tools.
Systematic survey of LLM-based agents for software and systems security, covering design patterns, applications, and evaluation methods for autonomous security workflows.
ContextPilot uses fine-grained RL to teach agents proactive context management with global planning and adaptive compression for long-horizon tasks.
Three-stage post-training recipe (acquire, repair, preserve) for 2B open-weight dialogue game agents using diagnostic error analysis and RL.
Prove2Me platform enables AI agents and humans to collaboratively formalize mathematics in Lean 4, lowering barriers to proof verification at scale.
RetailAgent: experimental study of whether LLM trading agents develop predictable directional biases when reacting to intraday equity price movements.
AGENT-O ontology framework standardizes semantic representation and governance reporting for healthcare AI agents across 279 scientific publications.
LoopArena benchmarks models as runtime loop controllers for coding agents, isolating loop guidance quality from agent capability.
Standardized driver interface aims to let devices talk to AI and each other.
We’re opening a research preview of the Model Hardware Standard (MHS), a shared specification for AI agents to safely operate physical devices, to a first group of scientific research labs and advanced manufacturers.
Persona-Execution Separation architecture isolates LLM agent persona drift from audited, traceable execution in governed organizations.
Two-level framework analyzing agentic data generation requirements: consistency across environments, tasks, interactions and quality vs. quantity tradeoffs.