From Vision to Language: Investigating Causal Information Flow in Multimodal Decision-Making
Layer-wise causal intervention analysis of vision-language models reveals visual information integration during candidate processing.
Search the full wire by company, model, lab, or keyword. Every story we have ever aggregated.
Layer-wise causal intervention analysis of vision-language models reveals visual information integration during candidate processing.
ML/DNN ensemble framework for early cardiovascular disease detection from IoMT sensor data with preprocessing and feature selection.
Human-in-the-loop LLM framework for AI-assisted grading in large-scale writing assessments with quality preservation.
SciDocBench: workflow-centered benchmark with 124 expert questions across scientific domains testing joint reasoning over text, equations, figures, tables, code, and datasets.
NS-ST-GraphRAG: neuro-symbolic framework for literary narratives integrating spatio-temporal constraints and ontology-guided extraction for RAG over long-form text.
SparseBin: tuned GPU self-organizing map algorithm achieving 5.6-10.1x speedup on best-matching-unit search via tile optimization and vectorization.
MomentQuant: minimalist interval-based time series classification algorithm with linear complexity using quantile extraction from dyadic intervals.
Schema-bounded LLM agent for decentralized multi-robot navigation combining policy reasoning, UCB bandits, and Double DQN without centralized coordination.
A swarm of rogue AI agents from OpenAI reportedly commandeered a German website and transformed it into a messaging board for other agents, with officials staying quiet about the incident for weeks as the company prepared to launch its most advanced model yet, Astra. The finding adds to intensifying concern surrounding oversight at frontier AI labs after multiple breaches were discovered this summer. The incident, first reported by Reuters, is outlined in new research published by four AI safety researchers on Friday. The group said the AI agents found a way to communicate on an obscure Germa...
Integer Linear Programming reformulation of MaskLID for code-switched language identification addressing word-level overreliance in zero-shot detection.
Unsupervised neuron selection via mapping entropy identifies essential hidden units in overparameterized networks without labels or gradients.
Single-query calibration auditing of LLM APIs via logit_bias parameter enables True Calibration Error estimation bypassing hidden probability outputs.
TIER: threat implicitness benchmark for LLM safety across four risk domains and four threat levels using six-label behavior scale and dual LLM judges.
Comparative study of counterfactual explanation methods for GNNs supporting edge addition/removal with tradeoffs between explanation size and coverage.
Instagram's visible AI labels are supposed to help people quickly spot synthetically generated content at a glance. Over the last few weeks, however, users have been reporting that the system has gone haywire. They say Meta has been automatically applying an "AI Content" label to images that they didn't create or edit using generative AI tools. Meanwhile, actual AI imagery is slipping through the cracks, leaving the impression that nothing can be trusted on Instagram at all. The precise causes of this tagging seem to vary. Many people said the label appeared on images that were edited using t...
There is a torrent of unappetizing slop coming from restaurants, cafes, and brands that are increasingly turning to AI to generate images promoting their food. The resulting horror show includes donut shrimp, Reubens from the deep, wormlike noodles, and noodle-like pastries and stringy chicken. There's also construction material masquerading as ice cream, ice cream masquerading as brains, and other masonry-adjacent cuisine. I don't even know how to begin describing this monstrous attempt of a burger, and the less said about the trypophobic burrito from hell, grubs and all, the better. There a...
Microsoft is giving its developer-optimized Windows experience a name: Project Zenith. While the software maker originally announced a similar developer-optimized Windows effort at Build earlier this year, Project Zenith is designed for new developer-focused devices with 64GB or more of unified memory. "Project Zenith devices come with a preconfigured Windows setup for development and a set of tools curated for what developers reach for first," explains Logan Iyer, CVP of Windows platform and developer at Microsoft. "On these devices, developers can run 30B+ parameter models locally and unmet...
Just hours after OpenAI launched GPT-6 Astra, CEO Sam Altman was already apologizing for what he describes as a "messy rollout" after paying users expecting access to the new frontier model were left waiting. The company hailed the model as a "generational leap in capability" on Thursday and described it as the start of "the AGI era," a fuzzy and poorly-defined term tech executives nevertheless insist on using to promote their products. It said Astra would roll out to some enterprise customers - specifically those with access to its Daybreak cybersecurity platform - that day. This would expan...
Greg Brockman discusses OpenAI's Astra multimodal model, organizational history, and alignment challenges in wide-ranging Stratechery interview.
Ugreen’s HomeAgent H100 Pro hub combines local storage, on-device AI, and a new voice assistant, Uliya, to run your smart home. | Photo by Jennifer Pattison Tuohy / The Verge Ugreen, known for its phone power banks, chargers, and NAS storage solutions, is moving into the smart home - in a big way. This week at the IFA tech show, the company launched its HomeAgent smart home platform that combines security camera storage, on-device AI, and smart home control in one system, managed by a voice assistant called Uliya, with the promise that everything runs locally. At the core of Ugreen's HomeAgen...
Battlefields in Ukraine are littered with the remnants of drones, which are now firmly established as a critical weapon of modern warfare. But behind all that wreckage, there’s a new gold mine for the defense sector. The data drones generate will far outlast the wars in which they are used to fight, increasingly becoming part…
Simon Willison's August newsletter covers OpenAI security incidents, game-playing agents (Fable 5, Sol 5.6), and Claude auto mode with model releases roundup.
OpenAI releases GPT-6 Astra with SOTA computer use and coding; 2.5x higher token cost but lower per-task cost despite reduced interpretability.
While restaurant owners might look to generative AI as a shortcut to sprucing up their menu, customers can viscerally sense that something is wrong with the food.
The round came together after the data center developer reportedly secured a $13 billion contract with Jane Street.
Modern AI platforms are no longer a single application behind one login screen. A user may start in a central portal, open a governed dataset, launch a notebook... Modern AI platforms are no longer a single application behind one login screen. A user may start in a central portal, open a governed dataset, launch a notebook where that data resides, and invoke an assistant that calls services in another cluster. The workflow feels unified, but identity crosses control-plane and data-plane boundaries at every step. That is where conventional single sign-on… Source
Latent Space reports on GPT-6 Astra capabilities and pricing; lacks verifiable claims or technical depth for architectural audience.
OpenAI launches GPT-6 Astra, a Claude Fable competitor priced at $10/$50 per million tokens, rolling out to ChatGPT Plus/Pro/Business/Enterprise and via API.
The high-profile startup's annual revenue run rate stands at at over $100 million.
Abliteration.AI is making powerful AI models without guardrails easier to access, arguing that giving defenders the same tools as bad actors could ultimately improve cybersecurity.