A computational approach to maximum likelihood thresholds for colored Gaussian graphical models
Computational methods for maximum likelihood thresholds in colored Gaussian graphical models under high-dimensional constraints.
Search the full wire by company, model, lab, or keyword. Every story we have ever aggregated.
Computational methods for maximum likelihood thresholds in colored Gaussian graphical models under high-dimensional constraints.
MultiGhostBench: 928-book multilingual benchmark for detecting LLM-generated text across six languages and distribution shifts.
Theoretical analysis of SGD dynamics via percolation modeling reveals discrete phase transitions in neural network subnetwork merging.
AgentScope: neuro-symbolic framework for diagnosing LLM agent failures via behavioral abstractions and structured analysis.
NE-R1: adaptive retrieval-augmented NER with on-demand mechanism for long-tail and domain-specific entity recognition.
Foundational ontology framework for formalizing contradictions in dialogue-based human-robot and human-agent interactions.
Safe-Stop: learned framework for humanoid robot emergency stopping via stoppability estimation and reach-avoid reasoning.
Algorithm for stable matching maximizing Nash social welfare in classical marriage problem with fairness objectives.
Multimodal framework for subcellularly resolved cell embeddings using RNA, protein sequences, and structural information.
SonicCaps: 15M-caption audio dataset with fine-grained descriptions from Qwen3-Omni for improved audio-language retrieval.
AGI Maze Prediction Datasets benchmark evaluates Transformer world modeling through procedurally generated grid-world transition and state prediction tasks.
SALA framework improves in-context learning for complex reasoning by learning task-specific logical alignment instead of matching fixed demonstration patterns.
ORB-SVM hybrid model applies machine learning to brain tumor detection from MRI scans with reduced parameter overhead versus deep learning.
PGM-attack demonstrates poisoning vulnerability in PGM-index learned indexes through adversarial key insertion to inflate segment counts.
Representational Empowerment scores candidate model elements by capacity expansion for future modeling and planning across environments.
Counter-GEO-Bench evaluates defenses against generative engine optimization attacks that inject misinformation into LLM-synthesized search answers.
DiffIE uses conditional discrete diffusion for multi-output open information extraction, replacing autoregressive generation with parallel trajectory sampling.
Systems survey of GUI agents analyzes observation, memory, action, and runtime efficiency across web, mobile, and desktop environments.
Methods for estimating Bayes-optimal BER and AUC extend prior work on optimal accuracy to imbalanced and noisy classification settings.
Critique refinement and deployment scaffolds make alignment evaluations harder to distinguish from real deployments by adding inference-time compute refinement.
Anthropic releases Claude Fable/Mythos 5.1 with SOTA performance, 75% cache cost reduction, and 70% increased output token throughput.
Paint.NET developer credits Claude AI for reverse-engineering Direct2D wrapper to run on WINE; demonstrates LLM utility for systems programming.
Cohere publishes guide on small language model deployment, enterprise adoption patterns, and SLM strategy considerations.
Claude Fable 5.1 achieves 52.6% on Terminal-Bench-Science 0.1; Willison reports improved coding/creative task performance vs. prior versions.
Anthropic details enterprise safety practices and customer collaboration on frontier model safeguards.
Google has reportedly been reaching out to a number of Hollywood's biggest studios, hoping to strike licensing agreements that would allow it to train its AI models on copyrighted material in exchange for massive piles of cash. In theory, these deals would be a win-win: a huge financial boon to the studios that would also give Google a significant advantage in its race to outpace other AI companies. But while there's little downside for Google, these deals pose a distinct set of risks to each of the studios that might be thinking about going all-in on AI. Getting into bed with Google this way...
AI model-training startup AfterQuery has reportedly raised a round that valued it at $3.2 billion, just five months after announcing its $30 million Series A at a $300 million valuation in April.
Anthropic says its newest AI models, Fable 5.1 and Mythos 5.1, address criticisms from customers about price, data retention, and overzealous safeguards. The company claims Claude Fable 5.1 offers stronger performance than Fable 5, but costs around 25 percent less typically and up to 45 percent less for complex agentic tasks, thanks to reduced pricing on cached data that was already processed and stored. Along with the announcement, a slew of early impressions popped up, including from Every CEO Dan Shipper, who claims, "It's the strongest coding model we've used, but now it's fast, token-eff...
OpenAI previewed the precautions it is taking as it prepares to release Astra, its newest, cyber-critical LLM.