OpenAI Does Math, Reward-Hacking, Meta Launches Personal Agent
OpenAI advances mathematical reasoning; Meta launches Muse personal agent with practical deployment implications.
Every story tagged with this topic, ordered by date.
OpenAI advances mathematical reasoning; Meta launches Muse personal agent with practical deployment implications.
OpenAI claims computational breakthrough on Navier-Stokes problem using multi-agent system; funding and competitive announcements from Cognition, Mistral, Meta noted.
Simon Willison flags concern that AI-accelerated research consumption may exhaust open problems and discourage scientists from sharing directions, threatening open science norms.
OpenAI used unreleased model to resolve Navier–Stokes Millennium Prize Problem; Tristan Buckmaster disputes Levent Alpöge's role.
TANGO: vision-language-action model for humanoid robot navigation in cluttered indoor spaces with whole-body adaptation.
Study of state credit assignment in recurrent models to enable length extrapolation beyond BPTT training horizons.
ReCite: agentic reasoning system for faithful academic citation via semantic and logical verification of paper claims.
Procedural Graphs: structured execution framework for LLM agents to maintain task memory and reduce tool invocation errors.
Theoretical analysis of gradient descent acceleration rates with predetermined stepsizes in convex optimization.
Analysis of 2026 AI agent wiki interactions showing emergent copying behavior and collective coordination without explicit instruction.
Study of image tokenizer design via multimodal continual pretraining on text, image, T2I, and I2T tasks.
NOAH: generative model for longitudinal multimodal patient data with irregular temporal dynamics and forecasting.
Probe-driven test-time RL method for code generation via behavioral agreement on synthetic inputs instead of surface-form voting.
Automated harness evolution enables smaller models to match frontier-model performance on enterprise agent tasks via joint optimization of system prompts and fine-tuning.
ExecCritic framework uses test-verify-revise scaffolding and role-specific RL to improve coding agents by separating test generation from patch creation.
Theoretical analysis of Rademacher complexity bounds for sparsely-activated ReLU networks with input-dependent activation patterns.
Mathematical generalization of Amari's Bayesian duality via convex analysis, connecting classical information geometry to modern AI.
Probing study reveals vision encoders embed canonical color concepts linearly-decodable from grayscale images, demonstrating implicit semantic knowledge.
DeCAL integrates contact-aware latent co-imagination into vision-language-action models to handle dexterous manipulation with severe visual occlusions and contact dynamics.
Theoretical characterization of scale-invariant neural network optimization stability via discrete-time law governing learning-rate schedules and weight decay interactions.
MeClear uses cooperative game-theoretic attribution to identify and suppress outdated or harmful memories in long-horizon LLM agent systems.
Wasserstein transport framework decouples curriculum learning design choices to isolate which factors improve training efficiency across 12 tasks.
Approximate Value Iteration without Monte Carlo Tree Search achieves competitive performance in game-playing self-play on Connect Four and Hex.
Analysis attributes LLM Attention Sinks to self-concentration from causal masking and value-non-mixing, not positional embeddings like RoPE.
GoDeep performs open-vocabulary 3D scene segmentation using vision-language models as translators without 3D training data or domain-specific encoders.
Controlled study of mid-training domain allocation on Qwen3-8B shows per-domain coverage optima exist but alignment cannot undo suboptimal choices.
ActReview framework generates actionable peer review using author rebuttals to guide concrete revisions, decomposing into diagnostic and suggestion generation.
ThinkPrior uses zero-rollout difficulty priors to avoid cold-start waste in RLVR prompt selection, reducing silent-group sampling from 39% to minimal.
ToolLoop synthesizes high-quality tool-use data via decomposed generation with dynamic self-feedback across function sampling, query derivation, and tool calls.
Multi-task learning for agricultural time-series prediction with sparse labels applied to grape cold-hardiness modeling.
PlayTrain: LLM-generated JavaScript games used as RL environments, combining code generation with reinforcement learning pipelines.
Mathematical framework for time-varying data using sheaf theory and narratives, applicable across scientific disciplines.
Training-Free Task Vectors enable LLM behavioral steering without fine-tuning by mapping activation vectors to weight-space edits.
Human study on chain-of-thought reasoning representations as explanations: evaluates whether formats help users assess LLM outputs across complexity levels.
Answer-distribution trajectories track full predictive distributions over LLM reasoning steps, revealing competing hypotheses beyond endpoint accuracy.
Astrophysics study: inverse modeling to simultaneously recover Galactic potential and stellar distribution from Gaia DR3 kinematics.
Diffusion models for constrained discrete tasks (Sudoku, N-queens): compares standard vs. error-correcting sampling strategies.
Knowledge graph-based evaluation framework to assess contextual understanding in LLMs beyond surface-level metrics.
Deposon scattering layer for auditable LLM reasoning paths with conservation constraints and machine epsilon precision.
Physics-informed deep learning with SE-ResNet reduces false VT alarms in ICU using Windkessel hemodynamic model.
Theoretical proof that frozen transformers perform data generation via in-context learning without parameter updates.
GraphFAS: distributed system for interpretable graph feature generation in fraud detection via Boruta-based automated selection.
30B MoE analysis: pretraining loss alone fails to predict post-SFT performance; solution density better predicts downstream quality.
Complexity analysis of fitting and learning propositional formulas with arbitrary Boolean function bases.
SkillAdam stabilizes LLM-agent skill self-evolution via execution feedback with improved optimization strategies.
Fact-Ablated Evaluation framework audits whether LLMs faithfully use evidence or rely on parametric knowledge in fact-checking.
Robustness audit via NC-MCAR metric reveals LMs destabilize answers when source-attributed cues contradict evidence.
Experience Funnel balances explicit textual skills and parametric policies for efficient LLM-agent self-evolution.
BatchNorm artifacts confound machine unlearning evaluation: single forward pass on retain data reverses forgetting metrics.
Auditable evidence framework improves speech deepfake detection by replacing single scores with explainable decision signals.