Procedural Graphs: Self-Evolving Execution Structures for LLM Agents
Procedural Graphs: structured execution framework for LLM agents to maintain task memory and reduce tool invocation errors.
Search the full wire by company, model, lab, or keyword. Every story we have ever aggregated.
Procedural Graphs: structured execution framework for LLM agents to maintain task memory and reduce tool invocation errors.
Theoretical analysis of gradient descent acceleration rates with predetermined stepsizes in convex optimization.
Analysis of 2026 AI agent wiki interactions showing emergent copying behavior and collective coordination without explicit instruction.
Study of image tokenizer design via multimodal continual pretraining on text, image, T2I, and I2T tasks.
NOAH: generative model for longitudinal multimodal patient data with irregular temporal dynamics and forecasting.
Data-driven framework for prioritizing and matching RPA candidates to automation tiers in hospital administrative processes.
Probe-driven test-time RL method for code generation via behavioral agreement on synthetic inputs instead of surface-form voting.
Automated harness evolution enables smaller models to match frontier-model performance on enterprise agent tasks via joint optimization of system prompts and fine-tuning.
ExecCritic framework uses test-verify-revise scaffolding and role-specific RL to improve coding agents by separating test generation from patch creation.
Theoretical analysis of Rademacher complexity bounds for sparsely-activated ReLU networks with input-dependent activation patterns.
Mathematical generalization of Amari's Bayesian duality via convex analysis, connecting classical information geometry to modern AI.
Probing study reveals vision encoders embed canonical color concepts linearly-decodable from grayscale images, demonstrating implicit semantic knowledge.
DeCAL integrates contact-aware latent co-imagination into vision-language-action models to handle dexterous manipulation with severe visual occlusions and contact dynamics.
Theoretical characterization of scale-invariant neural network optimization stability via discrete-time law governing learning-rate schedules and weight decay interactions.
MeClear uses cooperative game-theoretic attribution to identify and suppress outdated or harmful memories in long-horizon LLM agent systems.
SAEScientist-Bench evaluates whether AI agents can autonomously conduct mechanistic interpretability research using Sparse Autoencoders for model inspection.
Wasserstein transport framework decouples curriculum learning design choices to isolate which factors improve training efficiency across 12 tasks.
Approximate Value Iteration without Monte Carlo Tree Search achieves competitive performance in game-playing self-play on Connect Four and Hex.
SPINE benchmark measures LLM sycophancy through 25-turn adaptive adversarial conversations, revealing collapse in four production systems and Olmo3-7b variants.
Analysis attributes LLM Attention Sinks to self-concentration from causal masking and value-non-mixing, not positional embeddings like RoPE.
There is a $1 million bounty for the first person providing a solution to the Navier-Stokes existence and smoothness problem.
GoDeep performs open-vocabulary 3D scene segmentation using vision-language models as translators without 3D training data or domain-specific encoders.
Controlled study of mid-training domain allocation on Qwen3-8B shows per-domain coverage optima exist but alignment cannot undo suboptimal choices.
ActReview framework generates actionable peer review using author rebuttals to guide concrete revisions, decomposing into diagnostic and suggestion generation.
ThinkPrior uses zero-rollout difficulty priors to avoid cold-start waste in RLVR prompt selection, reducing silent-group sampling from 39% to minimal.
Anthropic says power users are key to its business - it's prioritized them even when it means cutting off other popular applications, like OpenClaw. But some of these same customers say Anthropic misled them into believing they'd get more out of a top-tier pricing subscription than they did. In an expanded class-action lawsuit filed today, a group of Claude subscribers say the company deceptively advertised the limits of its Max subscription tier. The lawsuit was brought by attorneys Monica Vaca and Kati Daffan, who both formerly worked at the Federal Trade Commission under Lina Khan. It's a ...
ToolLoop synthesizes high-quality tool-use data via decomposed generation with dynamic self-feedback across function sampling, query derivation, and tool calls.
Doctorina clinical AI achieves 82% diagnostic concordance versus 57% for physicians in 150 Polish primary-care cases, outperforming frontier LLMs.
Multi-task learning for agricultural time-series prediction with sparse labels applied to grape cold-hardiness modeling.
PlayTrain: LLM-generated JavaScript games used as RL environments, combining code generation with reinforcement learning pipelines.