Are local models becoming “good enough” faster than expected?
Local LLMs reaching production-grade performance on routine tasks (coding, summarization, agents), driving adoption of hybrid cloud-local workload strategies.
Search the full wire by company, model, lab, or keyword. Every story we have ever aggregated.
Local LLMs reaching production-grade performance on routine tasks (coding, summarization, agents), driving adoption of hybrid cloud-local workload strategies.
Perplexity's Personal Computer brings AI agents to your Mac, and is now open to everyone.
BAMI mitigates precision and ambiguity bias in GUI grounding agents without retraining using masked prediction distribution attribution.
SIRA framework improves retrieval-augmented agents by modeling expert search priors, reducing retrieval rounds and latency for organizational knowledge bases.
RAO trains recursive agents to delegate sub-tasks recursively, enabling divide-and-conquer inference scaling for longer contexts and harder problems.
Framework for parsing and verifying source attribution in LLM research agents; evaluates citation accuracy via AST parsing and reproducible verification.
SkillOS: self-evolving LLM-agent framework learning long-horizon skill curation policies from streaming task interactions.
AI CFD Scientist demonstrates LLM-based agents automating computational fluid dynamics discovery with physics validation loops.
Patch2Vuln reconstructs security vulnerabilities from Linux binary patches using LLM agents without source code access.
CRONA applies multi-agent RL to cross-modal embodied navigation, decomposing monolithic models into modality-specialized agents for flexible deployment.
NeuroAgent automates heterogeneous neuroimaging preprocessing and analysis via LLM-driven agentic framework coordinating modality-specific toolchains.
Save to Spotify is a new command-line tool designed specifically for AI agents like OpenClaw, Claude Code, or OpenAI Codex. If you're the kind of person who collects research on a topic, then feeds it through their AI of choice to create audio summaries and personal podcasts, this lets you save them right alongside the latest episode of The Vergecast and Welcome to Night Vale on Spotify. To set it up, you need to download and install the Save to Spotify CLI from GitHub. Then you just prompt your AI agent as normal, but tack on "and save to Spotify," and it should show up right in your podcast...
Parloa uses OpenAI models to build voice-driven customer service agents with simulation and real-time deployment capabilities for enterprises.
LongSeeker proposes Context-ReAct paradigm for elastic context management in long-horizon search agents, maintaining trajectory at variable detail levels.
Coding agent with executable Python world models, verification, and simplicity-bias refactoring solves 25 public ARC-AGI-3 games without task-specific logic.
Also, rate limits will double for Pro and Max users of tools like Claude Code.
Long-horizon LLM agents depend on intermediate information-gathering turns, yet training feedback is usually observed only at the final answer, because process-level rewards require high-quality human annotation. Existing turn-level shaping methods reward turns that increase the likelihood of a gold answer, but they require answer supervision or stable task-specific verifiers. Conversely, label-free RL methods extract self-signals from output distributions, but mainly at the answer or trajectory level and therefore cannot assign credit to intermediate turns. We propose Self-Induced Outcome Po...
Anthropic launches 10 finance-focused AI agents via Claude Cowork and Managed Agents for KYC screening, pitchbook generation, and month-end close workflows.
SAP plans to buy German AI startup Prior Labs and invest heavily in it. It is also prohibiting customers' agents use to a select few like Nvidia's NemoClaw.
OpenSeeker-v2: SFT on informative trajectories achieves frontier LLM search agent capabilities without full RL pipeline.
MOSAIC-Bench evaluates coding agents' vulnerability to multi-stage attack chains that decompose malicious goals into innocuous sequential tasks, exposing alignment gaps in deployed systems.
Anthropic releases ten Cowork and Claude Code plugins plus Microsoft 365 integrations and MCP app for financial services.
The automotive cockpit is undergoing a fundamental shift from rule-based interfaces to agentic, multimodal AI systems capable of reasoning, planning, and... The automotive cockpit is undergoing a fundamental shift from rule-based interfaces to agentic, multimodal AI systems capable of reasoning, planning, and acting. In most vehicles on the road today, in-vehicle assistants still rely on fixed command-response patterns: interpret a phrase, trigger an action, reset. While effective for well-defined tasks, this approach doesn’t scale to modern… Source
Argues frontier AI failures in open-ended tasks (scientific assistance, agents, personalization) stem from objective ambiguity rather than capability gaps; proposes contextual multi-objective optimization.
Generative AI’s explosive first chapter was defined by humans sending requests and models responding. The agentic chapter is different. Agents don't... Generative AI’s explosive first chapter was defined by humans sending requests and models responding. The agentic chapter is different. Agents don’t follow a pre-determined sequence of actions. They call tools, spawn sub-agents with different tasks and models, retain information in memory, manage their own context window, and decide for themselves when they’re finished. In doing so… Source
ProgramBench: 200-task evaluation showing agents struggle to rebuild large binaries from scratch without cheating vulnerabilities.
The Seattle-based startup's Series A round was led by Glilot Capital, NFX and SignalFire, TechCrunch has exclusively learned.
OpenAI and PwC partner to deploy AI agents for enterprise finance automation, forecasting, and CFO workflow modernization.
FlexSQL agent flexibly explores schemas and data during text-to-SQL generation, enabling recovery from early mistakes.
Opinion piece argues LLM agents require jointly formulated actions and plans with human actors rather than isolated architectures.