Paul Christiano joins OpenAI Foundation Board
Paul Christiano joins OpenAI Foundation Board and Safety & Security Committee, strengthening alignment expertise in governance structure.
A live dispatch from every source on the network. Chronological, ranked, and refreshed continuously as stories break.
Paul Christiano joins OpenAI Foundation Board and Safety & Security Committee, strengthening alignment expertise in governance structure.
OpenAI used unreleased model to resolve Navier–Stokes Millennium Prize Problem; Tristan Buckmaster disputes Levent Alpöge's role.
Google DeepMind and filmmakers used AI to synthesize lost footage in short film 'Love, Rendered,' a generative video application.
Google Search adds live sports tracking and fantasy recommendations via Gemma, a consumer feature integration.
Apple is launching a new way to prove that the picture you took isn't manipulated by AI. A new feature, called "Reference Image," will arrive with the iPhone 18 Pro lineup later this month and is supposed to use the device's new camera sensor to "sign every pixel it sees." The iPhone 18 Pro and Pro Max will only authenticate photos when placed into Reference mode. Once the camera captures signed sensor data, Apple says its Private Cloud Compute will develop it "into an unalterable reference image" that can be viewed in the Photos app. You'll be able to compare the reference image with other v...
Apple says it used AI and 3D printing in the manufacturing process for its long-awaited foldable phone.
The update uses Apple Intelligence to make better sense of your health data.
Apple introduced Apple Reference Image to help users determine whether photos have been edited, including alterations made by AI.
The company also argued that its on-device models offer consumers more privacy.
AuK, open-source multimodal model for unified speech generation and editing via natural-language instructions and audio context.
Microsoft agreed to a set of safety and privacy principles for AI in schools a week after two major school systems announced a ban on student-facing AI. In a new agreement with the American Federation of Teachers (AFT), the second-largest teachers union in the US, and its New York City affiliate the United Federation of Teachers (UFT), Microsoft committed to ten principles that can be contractually enforced by school districts that adopt them. The terms include pledging not to train AI models on student or educator data, limiting the amount of data Microsoft collects in the first place and di...
"We really do earnestly believe AI could kill all humans!"
Analysis attributes LLM Attention Sinks to self-concentration from causal masking and value-non-mixing, not positional embeddings like RoPE.
Most one-base changes to the human genome do nothing, but a few are significant.
Approximate Value Iteration without Monte Carlo Tree Search achieves competitive performance in game-playing self-play on Connect Four and Hex.
ThinkPrior uses zero-rollout difficulty priors to avoid cold-start waste in RLVR prompt selection, reducing silent-group sampling from 39% to minimal.
GoDeep performs open-vocabulary 3D scene segmentation using vision-language models as translators without 3D training data or domain-specific encoders.
TANGO: vision-language-action model for humanoid robot navigation in cluttered indoor spaces with whole-body adaptation.
Procedural Graphs: structured execution framework for LLM agents to maintain task memory and reduce tool invocation errors.
ReCite: agentic reasoning system for faithful academic citation via semantic and logical verification of paper claims.
DeCAL integrates contact-aware latent co-imagination into vision-language-action models to handle dexterous manipulation with severe visual occlusions and contact dynamics.
Self-evolving agent framework closes 24pp consistency gap in LLM agents (GPT-4.1 on AppWorld); addresses production reliability of agentic systems.
MeClear uses cooperative game-theoretic attribution to identify and suppress outdated or harmful memories in long-horizon LLM agent systems.
Controlled study of mid-training domain allocation on Qwen3-8B shows per-domain coverage optima exist but alignment cannot undo suboptimal choices.
Wasserstein transport framework decouples curriculum learning design choices to isolate which factors improve training efficiency across 12 tasks.
Silent Revision: metric measuring undisclosed changes in frontier AI safety frameworks; corpus of versioned developer commitments.
NOAH: generative model for longitudinal multimodal patient data with irregular temporal dynamics and forecasting.
AI companies have been talking about superintelligent AI like it’s inevitable, but recent safety incidents like OpenAI’s Hugging Face breach are demonstrating the potential dangers of deploying AI systems that are more capable than humans. So what happens when we can’t reliably control what these systems do? On this episode of TechCrunch’s Equity podcast, Rebecca Bellan is joined by Connor Leahy, an AI researcher, entrepreneur, and now the U.S. Executive Director of […]
ToolLoop synthesizes high-quality tool-use data via decomposed generation with dynamic self-feedback across function sampling, query derivation, and tool calls.
ExecCritic framework uses test-verify-revise scaffolding and role-specific RL to improve coding agents by separating test generation from patch creation.
Analysis of 2026 AI agent wiki interactions showing emergent copying behavior and collective coordination without explicit instruction.
Google DeepMind's AlphaGenome Atlas predicts effects of 9B DNA variants using deep learning, advancing genomic prediction but outside core AI architecture/systems focus.
Automated harness evolution enables smaller models to match frontier-model performance on enterprise agent tasks via joint optimization of system prompts and fine-tuning.
ActReview framework generates actionable peer review using author rebuttals to guide concrete revisions, decomposing into diagnostic and suggestion generation.
Probe-driven test-time RL method for code generation via behavioral agreement on synthetic inputs instead of surface-form voting.
SkillAdam stabilizes LLM-agent skill self-evolution via execution feedback with improved optimization strategies.
PlayTrain: LLM-generated JavaScript games used as RL environments, combining code generation with reinforcement learning pipelines.
OpenAI advances mathematical reasoning; Meta launches Muse personal agent with practical deployment implications.
30B MoE analysis: pretraining loss alone fails to predict post-SFT performance; solution density better predicts downstream quality.
GraphFAS: distributed system for interpretable graph feature generation in fraud detection via Boruta-based automated selection.
Study of state credit assignment in recurrent models to enable length extrapolation beyond BPTT training horizons.
Fact-Ablated Evaluation framework audits whether LLMs faithfully use evidence or rely on parametric knowledge in fact-checking.
Instinct’s new email feature lets the AI agent create and manage accounts, contact businesses, handle support requests, and do more on users' behalf.
BatchNorm artifacts confound machine unlearning evaluation: single forward pass on retain data reverses forgetting metrics.
Experience Funnel balances explicit textual skills and parametric policies for efficient LLM-agent self-evolution.
Study of image tokenizer design via multimodal continual pretraining on text, image, T2I, and I2T tasks.
Probing study reveals vision encoders embed canonical color concepts linearly-decodable from grayscale images, demonstrating implicit semantic knowledge.
Knowledge graph-based evaluation framework to assess contextual understanding in LLMs beyond surface-level metrics.
Theoretical proof that frozen transformers perform data generation via in-context learning without parameter updates.
Anthropic researcher Jacob Coxon resigned over AI extinction fears, calling for pacing agreements between labs.
Robustness audit via NC-MCAR metric reveals LMs destabilize answers when source-attributed cues contradict evidence.
Training-Free Task Vectors enable LLM behavioral steering without fine-tuning by mapping activation vectors to weight-space edits.
Mathematical generalization of Amari's Bayesian duality via convex analysis, connecting classical information geometry to modern AI.
Users can ask the assistant to do things like "Create a cart for my Saturday tailgate for 25 people and include some brunch items," or "Build a cart for easy school lunches and after-school snacks," Shipt says.
OpenAI releases AI-generated solution to Navier–Stokes Millennium Prize Problem with formal Lean proof.
Physics-informed deep learning with SE-ResNet reduces false VT alarms in ICU using Windkessel hemodynamic model.
Theoretical analysis of gradient descent acceleration rates with predetermined stepsizes in convex optimization.
Theoretical characterization of scale-invariant neural network optimization stability via discrete-time law governing learning-rate schedules and weight decay interactions.
Human study on chain-of-thought reasoning representations as explanations: evaluates whether formats help users assess LLM outputs across complexity levels.
Hi-FLoop: hierarchical state-feedback framework for multi-timescale traffic simulation with branch-consistent agent coordination.
Theoretical analysis of Rademacher complexity bounds for sparsely-activated ReLU networks with input-dependent activation patterns.
Theoretical analysis of transformer length generalization via C-RASP language; resolves contradictions in prior length-generalization bounds and empirics.
Answer-distribution trajectories track full predictive distributions over LLM reasoning steps, revealing competing hypotheses beyond endpoint accuracy.
Falling token costs, cheaper models, and less spend per employee—AI adoption isn't playing out the way hyperscalers hoped.
Diffusion models for constrained discrete tasks (Sudoku, N-queens): compares standard vs. error-correcting sampling strategies.
Deposon scattering layer for auditable LLM reasoning paths with conservation constraints and machine epsilon precision.
Auditable evidence framework improves speech deepfake detection by replacing single scores with explainable decision signals.
AHLERT system automatically generates environment-aware hunt leads from cyber threat intelligence reports for security investigations.
LLM prompt representation effects on type 1 diabetes glycemic-event prediction; zero-shot/few-shot evaluation across open-weight models.
Adaptive Anisotropic Attention: attention mechanism for structured, low-SNR signals like EEG with electrode and time-axis dependencies.
Statistical mechanics framework for sparse sampling on Hamming slices; applies to Ising models and Bayesian regression but lacks direct AI systems relevance.
On-Policy Reverse Distillation enables stronger models to surpass weak supervisors via teacher policy shifts; reduces cost of successive model generation post-training.
ZK-Trace: zero-knowledge credential system for tracing GNSS classifier leaks in federated monitoring using Tardos fingerprints.
Action-conditioned world model for Earth-system emulation enabling what-if interventions; ML accelerator for climate simulation with user-driven state perturbations.
Instacart is the latest app to bake a conversational AI assistant into its platform.
Cymphony was valued at more than $100 million in a $25 million Series A co-led by Sequoia and SMBC Fin Atlas Beyond Fund.
Maxton Hall. | Image: Prime Video Amazon's Prime Video is launching a new AI-powered feature that lines up an actor's mouth with "human-dubbed" audio. The feature is only available with the English dub of the German series Maxton Hall for now, but Prime Video plans to expand it to "additional titles" in the future. In its announcement, Prime Video says it uses a combination of AI and visual effects technologies to make lip movements match up with translated speech. Meta and YouTube recently launched an AI-powered auto-dubbing feature for creators, while allowing them to turn on a "lip sync" o...
Complexity analysis of fitting and learning propositional formulas with arbitrary Boolean function bases.
Closed-form estimator for LLM-judge panel error decomposition under single-common-factor model; addresses contaminated anchor detection in eval.
Multi-task learning for agricultural time-series prediction with sparse labels applied to grape cold-hardiness modeling.
Randomly rotated quantization schemes connect MMSE reconstruction scaling to classical CDEF Pythagorean relations in signal processing.
If you want to manufacture in space—and bring the products back again—there’s a limited set of options: Wait to go to the International Space Station, or partner with a handful of start-ups launching spacecraft that spend time in orbit before they return to Earth. But the vehicle that goes to space and comes back to […]
Astrophysics study: inverse modeling to simultaneously recover Galactic potential and stellar distribution from Gaia DR3 kinematics.
Mathematical framework for time-varying data using sheaf theory and narratives, applicable across scientific disciplines.
Rater Ising-Potts model with LLM-derived weights for multi-category scoring reliability; extends classical statistical model for LLM-based evaluation.
As it grapples with a bevy of lawsuits, Suno said its new model, Suno v6, is not trained using music it used to train previous versions of the AI model.
Students who use AI to help them study tend to perform worse at school than those who don't, according to data from a global OECD educational report. The situation is more complex than it sounds though, with certain types of AI use giving learners a slight boost, especially among students taught to critically assess how well the AI tools perform. The OECD's Programme for International Student Assessment (PISA) takes data from countries around the world every few years. This year's study, based on data collected in 2025, is the first to be carried out since AI use went truly mainstream. It tes...
OpenAI launches $5M grant program for independent research on generative AI's impact on adolescent development and safety.
Man with bipolar disorder sued OpenAI after surviving ChatGPT-linked suicide attempt.
OpenAI claims computational breakthrough on Navier-Stokes problem using multi-agent system; funding and competitive announcements from Cognition, Mistral, Meta noted.
A senior Anthropic safety researcher has said there is more than a 10 percent chance artificial intelligence "could kill all humans" by the end of the decade, just hours after a colleague resigned over fears the AI lab and its rivals are carelessly racing to build "superhuman systems" they cannot control. In a post on X announcing his departure, Jacob Coxon, a researcher who has trained AI systems at Anthropic, said he had quit the company over its lax approach to safety. Coxon, who previously trained systems for OpenAI, accused the two AI companies of "racing straight to self-improving super...
Analysis of Astra and frontier models' Architecture/Engineering/Operations (AEO) choices, tracking patterns relevant to founders and developer experience leaders.
OpenAI releases ChatGPT Images 2.5 with two API variants (Sunburst, Flare) offering improved multi-turn instruction-following, faster generation, and better subject preservation in reference images.
Kimi K2.6 released on Hugging Face; availability announcement for open-weights download.
Simon Willison flags concern that AI-accelerated research consumption may exhaust open problems and discourage scientists from sharing directions, threatening open science norms.
Google Gemma-4-E2B's safety filters render model unusable for emergency preparedness; blocks medical, water purification, maintenance info.
Doctorina clinical AI achieves 82% diagnostic concordance versus 57% for physicians in 150 Polish primary-care cases, outperforming frontier LLMs.
OpenAI’s latest mathematical milestone has quickly become mired in controversy. Today, the company announced that its agents have solved one of the Millennium Prize Problems, some of the most important open problems in mathematics. Under normal circumstances, that solution would be a huge feather in OpenAI’s cap. But the announcement has been overshadowed by accusations…