Less Experts, Faster Decoding: Cost-Aware Speculative Decoding for Mixture-of-Experts
Cost-aware speculative decoding for Mixture-of-Experts LLMs optimizes expert activation patterns to reduce inference cost beyond token acceptance alone.
Search the full wire by company, model, lab, or keyword. Every story we have ever aggregated.
Cost-aware speculative decoding for Mixture-of-Experts LLMs optimizes expert activation patterns to reduce inference cost beyond token acceptance alone.
CARE-PPO combines PPO fine-tuning with loss prediction to jointly train LLMs for accurate numerical estimates and calibrated confidence signals in quantitative prediction tasks.
SeRIn proposes architectural separation of modality-specific refinement from cross-modal fusion in multimodal LMs for sentiment analysis.
Method for panoptic symbol spotting in CAD floor plans using text-aided multimodal analysis for industrial digitalization.
Internet of Agentic Things (IoAT) framework integrates autonomous AI agents with IoT, cyber-physical systems, and edge computing for closed-loop orchestration.
Demis Hassabis, during a panel session at the World Economic Forum in Davos, Switzerland. | Image: Bloomberg via Getty Images Demis Hassabis thinks the world needs an AI watchdog with the power to hit the brakes if frontier models become too dangerous. Writing in a blog post, the Google DeepMind CEO and cofounder said the US should lead the initiative, arguing that the country is the best place to set global standards "given its economic and technical standing." The organization, which could resemble existing regulators like the Financial Industry Regulatory Authority, would be made up of lea...
Jetson-PI enables real-time onboard VLA model deployment on low-power Jetson devices via foresight-aligned asynchronous inference for robot control.
EG-VAR uses Lean 4 formal verification with tool-calling to ground LLM reasoning in attested evidence and kernel-checked inference chains, eliminating hallucination.
Establishes rigorous first-principles framework for identifying extractable memorization in LLMs using probability-based matching to distinguish training sequence reproduction from predictability.
AdaPCLA framework for generating longitudinal EHR data with improved tail event modeling via adaptive prior calibration for privacy-preserving rare subpopulation representation.
Controlled study finds GRPO RL fails to improve 4B–8B scale web agents over supervised baselines across learning rate and hyperparameter grids, questioning RL value at small scale.
Theoretical framework proposing atomic compression and compositional reuse as foundations for scalable intelligence across cognitive, biological, and computational systems.
OpenAI guidance on ROI metrics for enterprise AI agent deployments: measuring work-per-dollar, efficiency gains, and workflow scaling.
OpenAI integrates Codex capabilities into ChatGPT, shifting toward a multi-purpose platform rather than pure chat interface.
New hyperscale data centers can't set up shop in New York for up to a year now that Governor Kathy Hochul (D) has signed the nation's first statewide moratorium. But a bill passed by the state legislature that could restrict even more developments still awaits her signature. The order blocks new environmental permits for data centers over 50 megawatts, which the governor's office says will give the state time to come up with the regulations needed to protect residents from rising energy prices and environmental impact. That's higher than the 20 megawatt threshold approved by state lawmakers. ...
They're rolling up their sleeves again, seemingly out of fear of missing AI's defining moment and, presumably, the irresistible allure of making even more money -- potentially a lot more.
Codex usage grew 10x to 7M users in 6 months; article questions whether it has outpaced Claude Code amid sparse adoption metrics.
Simon Willison shares a caching technique for uvx Python tools in GitHub Actions using UV_EXCLUDE_NEWER environment variable.
Uber Chief Product Officer Sachin Kansal walks TechCrunch through the company's financial-services ambitions, its increasingly complicated relationship with Waymo, its new AV Labs data operation, and how AI is starting to show up in ways riders and drivers will actually notice.
Cohere releases Tiny Aya Expedition, a multilingual model supporting 70+ languages for on-device and educational AI applications.
Singapore-based video generation startup PixVerse closed a Series C extension on the strength of 15 million monthly active users, it said.
The company is raising at least $75 million, led by Robot, with significant participation from USV and other prominent investors.
Peter Gostev built DOOMQL, a Doom-like game engine using SQLite queries for all mechanics and rendering, implemented in Python.
Simon Willison documents productivity spike in Datasette project correlating with Opus 4.8 and GPT-5.5 releases via GitHub commit frequency analysis.
In a surprising blog post on Monday, Microsoft CEO is warning enterprises of the dangers of using proprietary models like Anthropic's and OpenAI's.
Siri AI in iOS 27. iOS 27 escaped the developer world today with the launch of the first public beta. I've been testing the new operating system since early June, looking for quirks and seeing if it can live up to the hype Apple promised in the keynote. This year's iOS upgrades are what one might call a Snow Leopard update. That means it's light on new features and instead focused on fixing things that were broken and speeding up processes across the OS. App launches, Photos search results, and AirDrop transfers should all be faster. The Messages app now supports in-line replies and end-to-en...
OpenAI accused of conspiring with former Apple employees to steal trade secrets.
Useful quantum computers will require fault tolerant logical operations. Researchers are actively exploring many different quantum error correction (QEC) codes... Useful quantum computers will require fault tolerant logical operations. Researchers are actively exploring many different quantum error correction (QEC) codes to enable this, improving the Logical Error Rates (LER) of Quantum Processing Units (QPUs). While it is well understood how to run logical operations with surface codes (which belong to the topological code family) via lattice surgery… Source
Apple’s trade secrets lawsuit against OpenAI contains allegations that range from employees joking about unauthorized access to Apple’s systems to claims that job candidates were asked to bring Apple hardware to interviews. Here are the complaint’s most eye-catching claims.
This story originally appeared in The Algorithm, our weekly newsletter on AI. To get stories like this in your inbox first, sign up here. Anthropic—currently the world’s most valuable AI company, with a nearly $1 trillion valuation—has a reputation for publishing strange and heady research. It’s looking into whether AI models can feel pain, for example,…