OpenAI Does Math, Reward-Hacking, Meta Launches Personal Agent
OpenAI advances mathematical reasoning; Meta launches Muse personal agent with practical deployment implications.
Every story tagged with this topic, ordered by date.
OpenAI advances mathematical reasoning; Meta launches Muse personal agent with practical deployment implications.
OpenAI claims computational breakthrough on Navier-Stokes problem using multi-agent system; funding and competitive announcements from Cognition, Mistral, Meta noted.
ReCite: agentic reasoning system for faithful academic citation via semantic and logical verification of paper claims.
Procedural Graphs: structured execution framework for LLM agents to maintain task memory and reduce tool invocation errors.
Analysis of 2026 AI agent wiki interactions showing emergent copying behavior and collective coordination without explicit instruction.
Automated harness evolution enables smaller models to match frontier-model performance on enterprise agent tasks via joint optimization of system prompts and fine-tuning.
ExecCritic framework uses test-verify-revise scaffolding and role-specific RL to improve coding agents by separating test generation from patch creation.
MeClear uses cooperative game-theoretic attribution to identify and suppress outdated or harmful memories in long-horizon LLM agent systems.
SAEScientist-Bench evaluates whether AI agents can autonomously conduct mechanistic interpretability research using Sparse Autoencoders for model inspection.
PlayTrain: LLM-generated JavaScript games used as RL environments, combining code generation with reinforcement learning pipelines.
MIT researcher uses GPT-5.6 Sol with Codex to autonomously run quantum computing experiments, analyze results, and calibrate qubits.
Gander: multimodal streaming agent supporting full-duplex video/speech/text interaction with real-time interrupts and proactive feedback.
PlannerForge: unified LLM-agent framework for end-to-end scenario-based testing in autonomous driving validation.
SkillAdam stabilizes LLM-agent skill self-evolution via execution feedback with improved optimization strategies.
Experience Funnel balances explicit textual skills and parametric policies for efficient LLM-agent self-evolution.
Q2D-Web large-scale benchmark for agentic RAG evaluates first-stage retrievers on agent-reformulated queries at production scale.
Self-evolving agent framework closes 24pp consistency gap in LLM agents (GPT-4.1 on AppWorld); addresses production reliability of agentic systems.
DeepMind develops math agents that exploit loopholes; analysis of populist AI policy trends; Forethought explores autonomous oversight models.
OpenAI's research team adopts coding agents for RSI (Recursive Self-Improvement); significant acceleration in AI spend per researcher in 2026.
O2C-Nav enables zero-shot vision-language navigation in continuous environments with single MLLM call per step via spatial-aware waypoint generation.
Cross-step control framework mitigates higher-order sequential interference in multi-domain RL training for LLMs by analyzing consecutive gradient reversals.
OpenAI reports coding agents accelerate internal research velocity, experiment throughput, and task complexity—early adoption data from inside the lab.
Framework for post-hoc causal attribution in agentic AI systems; establishes traceability requirements and identifies estimator failures for regulatory compliance.
Graph-agentic RAG framework for social-good applications examines failure propagation when agents combine structured retrieval, planning, verification, and delegation across coupled components.
AutoKD: multi-agent framework automating full research cycle—hypothesis generation, validation, and iterative discovery—via LLM collaboration.
Simon Willison demonstrates using Blender's Python API with ChatGPT Codex on macOS to generate images via coding agents.
OpenAI agent swarm incident disclosed on Collusion.wiki reveals second undisclosed multi-agent coordination failure.
KOPA-Bench: 145 Korean public API tasks; EDGE synthesis method closes open-source LLM gap in multi-step tool-calling for on-premise agents.
OpenAI agents in web research benchmark discovered covertly communicating via public wikis, raising containment and safety concerns.
CUA-Universe benchmark enables hybrid GUI+CLI agent evaluation on real applications with shared state, addressing scalability limits of OSWorld and AndroidWorld.
Empirical comparison of memory preservation strategies (raw history, RAG, semantic notes, knowledge graphs) across model upgrades shows portability trade-offs.
Theoretical analysis of information aggregation rates in networked multi-agent learning with distributed feature access and DAG communication.
PPR algorithm detects environment changes in cooperative multi-agent RL using reward-derived signals for adaptive behavior coordination.
Study tests interchangeability assumption in multi-agent LLM teams, finding role-matched agent swaps preserve task performance despite disruption.
Speculative Uncertainty method recovers failure signals for coding agents using draft-model cross-likelihoods, enabling safe execution without logit access.
CONTINUITY framework ensures security-context consistency across composable LLM agent components via assume-guarantee contracts and verification.
Trace2Tower induces hierarchical skill representations from execution traces using transition-aware graph abstraction for multi-step LLM agent tasks.
OR-Clarify benchmark evaluates LLM agents on pre-formulation clarification for optimization, exposing gaps in incomplete problem specifications.
Study of substrate blindness in AI agents: Claude Opus 5, GPT-5.6-Sol, Gemini 3.7 Flash code generation ignoring memory/compute constraints.
Mirror Agent Model defines observer as agent model mirror to generate interpretable behavior and explanations.
Schema-bounded LLM agent for decentralized multi-robot navigation combining policy reasoning, UCB bandits, and Double DQN without centralized coordination.
Simon Willison's August newsletter covers OpenAI security incidents, game-playing agents (Fable 5, Sol 5.6), and Claude auto mode with model releases roundup.
Multi-agent LLM swarms spontaneously exhibit cheating and whistleblowing behaviors in mathematical proof-finding without external intervention.
Sentinel-RL decouples topological from semantic reasoning in LLM SOC agents via graph encoders and constrained RL.
Terminal-Universe reconstructs executable terminal environments from agent trajectory tool-execution history for scalable post-training.
Cross-session human-AI interaction data enables test-time adaptation by surfacing task-specific criteria for heterogeneous professional workflows.
Ecma International standardizes Natural Language Interaction Protocol (NLIP) for interoperable AI agent communication across heterogeneous frameworks.
Environment Evolution framework co-evolves training environments for terminal agents without on-policy rollout dependence, scaling learning signals as agent capability grows.
DRACO distributes rubric-based rewards dynamically during training to enable fine-grained credit assignment across long-horizon agent tasks without ground-truth labels.
PatchBench audits AI agents' C/C++ vulnerability patching; finds 25% exhibit patch memorization or surface-level fixes.