Mark Zuckerberg tells staff that AI agents haven’t progressed as quickly as he’d hoped
At an internal meeting, the Meta CEO reportedly said that AI development efforts were not moving as quickly as anticipated.
Search the full wire by company, model, lab, or keyword. Every story we have ever aggregated.
At an internal meeting, the Meta CEO reportedly said that AI development efforts were not moving as quickly as anticipated.
Simon Willison uses DSPy to evaluate and optimize SQL system prompts in Datasette Agent, a tool enabling agents to execute read-only database queries.
Introduces Iterative VibeCoding, a benchmark for studying prompt-injection and code-injection attacks in autonomous AI agents shipping iterative PRs to persistent codebases.
Dual-channel debate framework reveals that LLM agents develop latent objectives and shift public utterances based on social structure, audience, and role without explicit prompting.
CNeVA: controllable simulated agents with interpretable behavior latents for realistic traffic simulation, enabling edge-case testing without real-world risk.
TestEvo-Bench executable benchmark evaluates test-code co-evolution for automated testing agents; mined from real repositories.
EvoPolicyGym benchmark for autonomous policy evolution; evaluates how agents iteratively improve executable policies under fixed budgets.
Geoffrey Litt argues developers must maintain deep code comprehension to effectively collaborate with AI coding agents and avoid cognitive debt.
Constraint-based oversight for coding agents scales cheaper than scaffolding; access control and tooling principles from human teams transfer directly.
Agent-based patching of LLVM compiler missed optimizations; agents struggle with generalization beyond single cases, compared via benchmark.
Pipeline automating physics research from arXiv corpus to publication via LLM agents with external literature grounding to reduce hallucination.
AgenticSTS introduces bounded-memory architecture for LLM agents with typed retrieval instead of transcript accumulation for long-horizon tasks.
Paper-replication workflow enables coding agents to verify computational claims in scientific ML papers with recorded evidence and automated verification.
ContextNest open specification for verifiable context governance in autonomous AI agents with provenance, version control, and integrity guarantees.
Introspection co-founder explains autoresearch loops, agent recipes, and self-improving systems while arguing humans remain essential to AI software development.
Cursor's Forward Deployed Engineers help enterprises implement AI agents as software factories, per Pauline Brunet.
Audit exposes reliability issues in coding-agent benchmarks (GSO, SWE-Perf, SWE-fficiency): runtime instability, scoring artifacts, selection bias.
Cloudflare is giving AI companies until September 15 to separate web crawlers used for search from those used for AI training and agents, or risk being blocked by default on many publisher sites.
Reinforcement learning (RL) is central to aligning language models, from reinforcement learning with human feedback (RLHF) within AI assistants to newer... Reinforcement learning (RL) is central to aligning language models, from reinforcement learning with human feedback (RLHF) within AI assistants to newer reinforcement learning with verifiable rewards (RLVR) workflows for reasoning and agent tasks. RL is now becoming a practical technique for specialized AI where enterprises need more accurate agents for domain-specific workflows. Source
DiscoPER framework enables open-ended autonomous scientific discovery via iterative meta-reflection and cross-finding synthesis in LLM agents.
Case study on software engineering with frontier AI coding agents reveals shift from implementation scarcity to governance, inspection, and maintainability challenges.
OpenAgent formalizes generalization gaps in LLM tool-use agents across query, action, observation, and domain shifts via controlled sandbox evaluation.
Test-time control framework DART-VLN mitigates memory decay and loop inefficiencies in vision-language navigation agents.
Study shows moderate LLM agent personality expression outperforms extremes on trust and goal-adoption in conversational behavior-change tasks.
SWE-Doctor: LLM-based code agent using multi-faceted bug reproduction tests for runtime diagnosis to improve software patch generation.
SEA architecture enables safe self-modifying agents by freezing base models and gating adaptations through anytime-valid certificates against error budgets.
Anthropic’s Claude Sonnet 5 brings stronger agentic capabilities, lower pricing, and improved safety, positioning the model as a cheaper alternative to Opus, GPT-5.5, and Gemini Pro.
Anthropic releases Claude Sonnet 5, a frontier model optimized for coding, agents, and professional workflows at scale.
QVal proposes efficient evaluation method for dense supervision signals in long-horizon LLM agents without expensive end-to-end training.