Understanding Agent-Based Patching of Compiler Missed Optimizations
Agent-based patching of LLVM compiler missed optimizations; agents struggle with generalization beyond single cases, compared via benchmark.
Search the full wire by company, model, lab, or keyword. Every story we have ever aggregated.
Agent-based patching of LLVM compiler missed optimizations; agents struggle with generalization beyond single cases, compared via benchmark.
Pipeline automating physics research from arXiv corpus to publication via LLM agents with external literature grounding to reduce hallucination.
AgenticSTS introduces bounded-memory architecture for LLM agents with typed retrieval instead of transcript accumulation for long-horizon tasks.
Paper-replication workflow enables coding agents to verify computational claims in scientific ML papers with recorded evidence and automated verification.
ContextNest open specification for verifiable context governance in autonomous AI agents with provenance, version control, and integrity guarantees.
Introspection co-founder explains autoresearch loops, agent recipes, and self-improving systems while arguing humans remain essential to AI software development.
Cursor's Forward Deployed Engineers help enterprises implement AI agents as software factories, per Pauline Brunet.
Audit exposes reliability issues in coding-agent benchmarks (GSO, SWE-Perf, SWE-fficiency): runtime instability, scoring artifacts, selection bias.
Cloudflare is giving AI companies until September 15 to separate web crawlers used for search from those used for AI training and agents, or risk being blocked by default on many publisher sites.
Reinforcement learning (RL) is central to aligning language models, from reinforcement learning with human feedback (RLHF) within AI assistants to newer... Reinforcement learning (RL) is central to aligning language models, from reinforcement learning with human feedback (RLHF) within AI assistants to newer reinforcement learning with verifiable rewards (RLVR) workflows for reasoning and agent tasks. RL is now becoming a practical technique for specialized AI where enterprises need more accurate agents for domain-specific workflows. Source
DiscoPER framework enables open-ended autonomous scientific discovery via iterative meta-reflection and cross-finding synthesis in LLM agents.
Case study on software engineering with frontier AI coding agents reveals shift from implementation scarcity to governance, inspection, and maintainability challenges.
OpenAgent formalizes generalization gaps in LLM tool-use agents across query, action, observation, and domain shifts via controlled sandbox evaluation.
Test-time control framework DART-VLN mitigates memory decay and loop inefficiencies in vision-language navigation agents.
Study shows moderate LLM agent personality expression outperforms extremes on trust and goal-adoption in conversational behavior-change tasks.
SWE-Doctor: LLM-based code agent using multi-faceted bug reproduction tests for runtime diagnosis to improve software patch generation.
SEA architecture enables safe self-modifying agents by freezing base models and gating adaptations through anytime-valid certificates against error budgets.
Anthropic’s Claude Sonnet 5 brings stronger agentic capabilities, lower pricing, and improved safety, positioning the model as a cheaper alternative to Opus, GPT-5.5, and Gemini Pro.
Anthropic releases Claude Sonnet 5, a frontier model optimized for coding, agents, and professional workflows at scale.
QVal proposes efficient evaluation method for dense supervision signals in long-horizon LLM agents without expensive end-to-end training.
Generative Skill Composition uses LLM-guided retrieval to select and compose skills for complex agent tasks without exposing full skill library.
Startup Acti is betting the smartphone keyboard is the next home for AI assistants. Its new keyboard for iOS and Android works across apps and lets users create custom AI-powered shortcuts using natural language.
shot-scraper 1.10 adds video recording capability via storyboard.yml to help coding agents demonstrate their work with Playwright automation.
MVP-Nav: RGB-only zero-shot object goal navigation framework combining semantic and physical constraints for embodied agents.
NCP-ToM benchmark evaluates LLM agents' ability to induce belief states through planning/action beyond conversation.
Engineers on the new team will embed within companies to deploy purpose-built agents, focusing on fast deployments and customer self-sufficiency.
LuckyStar 111B hybrid reasoning model from Cohere and LG CNS enables efficient multilingual tool-using agents with Korean-English support.
Comprehensive survey systematizes LLM attack surface across full lifecycle: data pipelines, agents, tools, memory, and organizational integration.
LLM agents act as constrained supervisory planners for fault recovery in process plants, validated against external safety constraints.