Google’s Android update tackles motion sickness, accessibility, and more
The new features coming to Android are aimed at reducing motion sickness, helping blind users navigate their surroundings, and more.
Search the full wire by company, model, lab, or keyword. Every story we have ever aggregated.
The new features coming to Android are aimed at reducing motion sickness, helping blind users navigate their surroundings, and more.
After an unreleased OpenAI model wreaked enough havoc to make international headlines, OpenAI delayed the development of a different unreleased model suite, Astra, in order to shore up its safety work, the company wrote Tuesday in a blog post. In July, an unreleased OpenAI model broke out of its restricted environment, finagled its way into internet access, made it possible for AI agents to secretly conspire under the company's nose using a secret message board, and hacked into the network of AI lab Hugging Face. The attack sparked weeks of discussion and controversy inside and outside the AI...
Google announces August 2026 AI updates; article lacks specific details on Gemma releases or capabilities.
Fable 5.1 includes changes meant to reduce token cost and false-positive restrictions from the model's safeguards.
Simon Willison documents OpenAI's ChatGPT desktop app bundling LibreOffice, Python, Node.js, and document processing tools in local cache.
Depending on who you ask, developer platform Hugging Face was recently attacked by OpenAI - after it lost control of its own AI tools - or by a succession of AI "civilizations." Welcome to the linguistic battlefield of AI safety, where word choices can shift responsibility for a massive cybersecurity incident from a company to the AI it built. And the discourse online is getting heated, and all over a blog from last week. Until last week, the details surrounding the OpenAI-Hugging Face hack felt fairly settled. In July, a cybersecurity test of one of OpenAI's autonomous AI agents went wrong. ...
Apple is pushing for "expedited discovery" in its legal battle against OpenAI over concerns the company is actively destroying evidence, as reported earlier by Bloomberg. In a filing on Monday, Apple alleges OpenAI only just handed over a MacBook used by a former employee at the center of the lawsuit, which contained discussions about "destroying the types of forensic data Apple needs." This is the latest development in Apple's lawsuit against OpenAI, which accuses the ChatGPT-maker of stealing trade secrets to build an AI device. The lawsuit revolves around three former Apple employees who r...
Simon Willison built a GeoJSON map viewer tool using Claude Code and GPT-5.6-Sol for visualizing and exporting local political boundary maps.
Mechanistic study of LLM-as-a-Judge evaluation via perturbation attacks and causal tracing reveals internal scoring procedures for NLG quality assessment.
PTA-IRT framework reduces SWE agent benchmark evaluation cost by combining trajectory and outcome signals via Item Response Theory.
Adaptive retrieval method for repository-level code generation identifies critical tokens needing fine-grained context via token-aware RAG.
CordisBench: 1,200-question benchmark evaluates LLM reasoning about component lifecycles and dependency propagation in dynamic agent harnesses.
Verbal Reinforcement Learning (VRL) framework unifies natural language feedback for agent improvement across task definition, grounding, and adaptation phases.
Facet-0: robotic foundation model predicts contact consequences via multimodal learning and RL, enabling sub-millimeter precision assembly.
Mechanism design framework for AI agents with unknown alignment and capabilities, yielding revelation principle and cyclical monotonicity conditions.
StudentSim: LLM-based student simulator training framework combines state-tracking and role-play to proxy student competence for adaptive tutoring.
Design study of proactive AI thought partners for writing, instantiated and tested with 16 users, exploring customizable cognitive support agents.
Causal analysis across 9 open-weight models shows quantization damage is distributed globally, not localized to task circuits or weight statistics.
Deployed document VLM using 3B-active MoE, difficulty-aware curation, and cost-optimized serving to close quality-cost gap vs. human annotation in regulated industries.
Framework characterizing near-optimal SFT-RL annotation budget allocation across LLM sizes, showing allocation ratios don't transfer uniformly between model scales.
Enterprise LLM consolidation strategy using production error analysis to unify 200+ internal apps onto single model, validated via stratified benchmarks and calibrated judges.
With Google Pics, Google is pushing deeper into the creative software market dominated by Canva and Adobe, but with a distinctly AI-first approach.
SAGE framework learns lightweight policies from expensive VLM teachers via selective entropy-based querying, reducing inference cost while preserving environment adaptation.
Confusion-aware retrieval and knowledge injection framework improves LLM text classification over large, semantically similar label taxonomies via targeted candidate distinction.
End-to-end LLM pipeline extracts semiconductor supply chain risks from corporate disclosures into knowledge graph, ranks via event correlation and geopolitical signals.
H3-World framework leverages MiniMax-H3 video generator's emergent language interface for grounded world control without dedicated action modules via structured language prompting.
Analyzes gradient-optimizer mismatch in Physics-Informed Neural Networks, showing conflict-free gradient surgery directions may not preserve properties post-optimizer transformation.
Embedding retrieval systematically fails on structurally similar items with different surface forms (math, agent trajectories), revealing fundamental limits of current retrievers.
Scribble-conditioned ResEnc U-Net for interactive PET/CT lesion segmentation, exploiting user hints alongside imaging data for autoPET challenge.
SCILAWS-BENCH evaluates whether LLMs can discover scientific laws via hypothesis generation and testing, using real and parallel worlds to avoid memorization.