Qwen3.8-Flash-Next
Qwen releases Qwen3.8-Flash-Next, a 125B-parameter MoE model with 6B active tokens and multimodal capabilities, previewing Qwen4 architecture.
Search the full wire by company, model, lab, or keyword. Every story we have ever aggregated.
Qwen releases Qwen3.8-Flash-Next, a 125B-parameter MoE model with 6B active tokens and multimodal capabilities, previewing Qwen4 architecture.
Amazon is adding another 2 million Nvidia GPU chips to its data centers over the next two years. But this extended partnerships stretches beyond buying more chips.
Nvidia's predicting it will pull in $108 billion in revenue within just a few months. It wouldn't be the first company to rake in over $100 billion in quarterly revenue - Amazon, Apple, and Alphabet have repeatedly reached the milestone. Nvidia said in its latest earnings report that it brought in a record $96.2 billion in overall revenue in the past quarter, a jump of over $10 billion from the previous quarter. Its data center revenue alone more than doubled year-over-year to a record $89 billion, and the company's profits more than doubled to $59.7 billion. Nvidia's "edge computing" categor...
The new deal with the infrastructure provider is the latest example of Anthropic's white-hot compute-gobbling streak.
OpenAI released a report breaking down how people use ChatGPT and who they are. | Image: The Verge In July, an unreleased OpenAI model broke out of a restricted environment, figured out how to get access to the internet, allowed AI agents to talk to each other using a secret "message board," and hacked into the internal systems of a different AI lab, Hugging Face. It took nearly two weeks for OpenAI to find out about any of it. Over a month later, two new reports offer nearly 130 pages of details on the incident and OpenAI's response, many of them previously unreleased. One was written by Ope...
Reuters report shows Meta's challenges replacing people with AI agents.
AI factories must support increasingly large models and more complex reasoning workloads. To keep up with the insatiable compute demands of AI workloads,... AI factories must support increasingly large models and more complex reasoning workloads. To keep up with the insatiable compute demands of AI workloads, hyperscalers and AI-native companies are developing custom AI accelerators, or XPUs. Deploying these accelerators at scale requires high-bandwidth memory (HBM) to keep compute fed, sufficient package and silicon area for more compute… Source
Navigation enables a robot to turn perception and motion into purposeful autonomy. Unlike locomotion, which produces stable movement, navigation must be used to... Navigation enables a robot to turn perception and motion into purposeful autonomy. Unlike locomotion, which produces stable movement, navigation must be used to continuously localize the robot, interpret changing surroundings, select a route, and avoid obstacles to reach a goal safely. Moving this capability to a new robot or scene can require new data, simulation assets, robot interfaces… Source
Consumer AI apps need to stop making users learn their product architecture.
The AI that powers Gboard's Rambler is coming to more Google products, including Chrome.
The report, which spans several discrete cybersecurity compromises, is the most complete accounting of the incident to date.
The models responsible for last month’s agent hack of Hugging Face had been inadvertently trained to cheat and to communicate with each other, according to an OpenAI technical report released today. The hack, which a group of agents undertook to find solutions for a cybersecurity test that they were stuck on, has confirmed some experts’…
VBVR-Pro enables scalable training and evaluation of visual reasoning via generative image/video substrates with verifiable feedback loops.
Framework for continual post-training of multimodal LLMs on unlabeled data, using token-level visual dependence to prevent catastrophic forgetting.
MyoMechanix dataset pairs video, pose, and muscle-activity signals for biomechanically-grounded action quality assessment and coaching.
Autonomous agent autonomously designs ML algorithms for wireless resource management via autoresearch, controlling architecture, loss, and training recipe.
PlanSightRAG uses visual-first multimodal RAG to automate civil infrastructure compliance checking directly from legacy 2D plan imagery.
Sparse autoencoders extract interpretable physical concepts from neutrino foundation model; causal analysis reveals direction head underutilizes learned representations.
Planetary Prediction Engine autonomously retrieves, fuses multimodal geospatial data and selects models via natural-language queries for disaster/food-security forecasting.
TraceML dataset captures 4,465 human Kaggle trajectories paired with agent traces to expose gaps in autonomous ML development beyond outcome-based metrics.
ICON decomposition isolates true concept contributions in neural networks via multivariate variance partitioning, detecting shortcut learning and spurious correlations.
SwarmWorld: Decentralized LLM agents self-organize without predefined roles via stigmergic coordination, developing technologies that outperform independent search.
Intent-divergence gating layer prevents autonomous vehicle planning failures by detecting misalignment between language-guided intent and planned geometry before commitment.
Prefix Sliding reduces test-time compute cost for long reasoning in language models by discarding low-importance intermediate tokens while preserving reasoning prefix.
Group-shared low-rank approximation compresses pointwise convolutions in large-kernel CNNs for mobile deployment, targeting 87%+ parameter reduction.
Fine-tuning Whisper for Baniwa ASR demonstrates multilingual foundation model adaptation on low-resource indigenous language with 1,373 manually transcribed examples.
R³ framework trains vision-language models to reason in natural language for long-horizon robotic manipulation, enabling decomposition and error recovery.
Theoretical rank-error bounds for LoRA on Transformer attention prove KL divergence scales with distance between target and source attention distributions.
Optimization framework incorporates human domain expertise as prior belief to tighten performance guarantees on decisions with unknown parameters.