Language Identification via Compositional Data Analysis: A Linear-Time Classifier Based on Log-Ratio Geometry
Linear-time language ID via compositional data analysis: CLR-transformed character/bigram frequencies on simplex geometry.
Search the full wire by company, model, lab, or keyword. Every story we have ever aggregated.
Linear-time language ID via compositional data analysis: CLR-transformed character/bigram frequencies on simplex geometry.
Across 107 enterprises, AI infrastructure spending is accelerating well ahead of the ability to see or steer its economics. Most organizations run their AI on a familiar base of hyperscalers and model-provider APIs, yet the next dollar is aimed at specialized compute almost none of them use today; a majority intend to switch or add providers within the year, many within a quarter. Buying decisions turn on integration and total cost of ownership rather than headline token price — which is fortunate, because most enterprises cannot yet see their unit economics clearly: GPUs sit at half utilizat...
In-place tokenizer expansion for LLMs: reallocate vocabulary post-training to reduce latency/energy for underrepresented languages.
Data-driven algorithms learn optimal block-replacement interval for machines from operational failure histories under censoring.
PRISM: hidden-state analysis shows content danger and physical danger are separable LLM signals; proposes single-layer filter for embodied-agent safety.
NeuronSoup proposes asynchronous, shared-neuron temporal graph architecture replacing backpropagation with delay-mediated signal propagation.
Symbal detects systematic misalignments in MLLM-generated image captions by identifying recurring errors tied to specific visual features.
VEXMLM extends XLM-R with language-specific tokenizers for Ge'ez-script languages (Amharic, Tigrinya) to reduce OOV rates in low-resource African NLP.
Theoretical analysis of bias delocalization in Hamiltonian Monte Carlo and underdamped Langevin samplers, extending prior work on MCMC convergence.
Across 101 enterprises, the infrastructure that feeds AI agents their business context is being built faster than it can be trusted. Retrieval-augmented generation is already the default context source, and provider-native retrieval has quietly overtaken the dedicated vector databases that define the category — yet a majority of enterprises have already watched their agents produce confident, wrong answers traced to missing or inconsistent context. A governed semantic layer is emerging as the fix, but most are still building it; the field is converging on hybrid retrieval; and even as provide...
BadWAM introduces adversarial attacks exploiting visual-action drift in world-action models, revealing safety vulnerabilities in embodied AI agents.
MM-IssueLoc benchmark isolates visual evidence impact in multimodal software repository issue localization across 23 languages and 652 instances.
Expert-in-the-loop annotation framework for depression diagnosis using LLMs, grounding labels in DSM-5-TR criteria for explainable mental health AI.
Two-stage MDP formulation for reinforcement learning in masked diffusion language models, accounting for both token placement and position unmasking.
Novel decomposition framework for uncertainty quantification deriving epistemic/aleatoric measures from subjective risk and proper losses.
Plover externalizes planning in vision-based GUI agents, enabling user inspection and correction of task plans for autonomous interface automation.
Critical analysis of item response theory reliability for AI benchmarks, highlighting regime mismatches between IRT assumptions and benchmark data distributions.
Across 157 enterprises, organizations are granting AI agents more autonomy while trusting the evaluations meant to gate that autonomy less. Half have already shipped an agent that passed their internal evaluations and then failed a customer in production; only one in twenty fully trusts automated evaluation today; and the most-cited weakness is that evaluations do not align with real-world outcomes. Yet two-thirds already allow, or are actively engineering toward, deploying agent changes to production on automated evaluation alone — with no human in the loop. The result is an evaluation gap —...
Hybrid neural-physics ODE framework with RTS smoothing for state/parameter estimation in dynamical systems with unknown components.
T²MLR fuses cached middle-layer representations across decoding steps to enable persistent reasoning in transformers with minimal inference overhead.
Benchmark evaluates six MLLMs on scientific visualization literacy using 49 standardized assessment items across 8 visualization techniques.
Study investigates linear representations of grammaticality in neural language models beyond probability-based measures.
MedFailBench is clinician-built open-source benchmark categorizing medical AI failures by severity and safety gate type with 44 synthetic cases.
Opinion piece examines AI's transition from tool to autonomous research participant via industrialization of science, citing DOE Genesis Mission.
Research investigates scaling laws for Behavior Foundation Models in humanoid robot control using large-scale behavioral data.
On-Policy Delta Distillation introduces delta-signal reward for RL post-training, measuring difference between teacher and base models.
A video analytics AI agent that can perceive, reason, and act based on massive amounts of video footage must be integrated with existing workflows and... A video analytics AI agent that can perceive, reason, and act based on massive amounts of video footage must be integrated with existing workflows and applications to be useful. These include content management systems, messaging platforms, databases, ticket queue, and escalation paths. This integration is challenging because video systems, enterprise knowledge bases… Source
Google integrates third-party app connections into Search's AI Mode for secure data access and interaction.
With this new update, Google is expanding AI Mode beyond answering questions and into completing tasks across the apps they use regularly.