Data Driven Block Replacement Scheduling
Data-driven algorithms learn optimal block-replacement interval for machines from operational failure histories under censoring.
Search the full wire by company, model, lab, or keyword. Every story we have ever aggregated.
Data-driven algorithms learn optimal block-replacement interval for machines from operational failure histories under censoring.
PRISM: hidden-state analysis shows content danger and physical danger are separable LLM signals; proposes single-layer filter for embodied-agent safety.
NeuronSoup proposes asynchronous, shared-neuron temporal graph architecture replacing backpropagation with delay-mediated signal propagation.
Symbal detects systematic misalignments in MLLM-generated image captions by identifying recurring errors tied to specific visual features.
VEXMLM extends XLM-R with language-specific tokenizers for Ge'ez-script languages (Amharic, Tigrinya) to reduce OOV rates in low-resource African NLP.
Theoretical analysis of bias delocalization in Hamiltonian Monte Carlo and underdamped Langevin samplers, extending prior work on MCMC convergence.
Across 101 enterprises, the infrastructure that feeds AI agents their business context is being built faster than it can be trusted. Retrieval-augmented generation is already the default context source, and provider-native retrieval has quietly overtaken the dedicated vector databases that define the category — yet a majority of enterprises have already watched their agents produce confident, wrong answers traced to missing or inconsistent context. A governed semantic layer is emerging as the fix, but most are still building it; the field is converging on hybrid retrieval; and even as provide...
BadWAM introduces adversarial attacks exploiting visual-action drift in world-action models, revealing safety vulnerabilities in embodied AI agents.
MM-IssueLoc benchmark isolates visual evidence impact in multimodal software repository issue localization across 23 languages and 652 instances.
Expert-in-the-loop annotation framework for depression diagnosis using LLMs, grounding labels in DSM-5-TR criteria for explainable mental health AI.
Two-stage MDP formulation for reinforcement learning in masked diffusion language models, accounting for both token placement and position unmasking.
Novel decomposition framework for uncertainty quantification deriving epistemic/aleatoric measures from subjective risk and proper losses.
Plover externalizes planning in vision-based GUI agents, enabling user inspection and correction of task plans for autonomous interface automation.
Critical analysis of item response theory reliability for AI benchmarks, highlighting regime mismatches between IRT assumptions and benchmark data distributions.
Across 157 enterprises, organizations are granting AI agents more autonomy while trusting the evaluations meant to gate that autonomy less. Half have already shipped an agent that passed their internal evaluations and then failed a customer in production; only one in twenty fully trusts automated evaluation today; and the most-cited weakness is that evaluations do not align with real-world outcomes. Yet two-thirds already allow, or are actively engineering toward, deploying agent changes to production on automated evaluation alone — with no human in the loop. The result is an evaluation gap —...
Hybrid neural-physics ODE framework with RTS smoothing for state/parameter estimation in dynamical systems with unknown components.
T²MLR fuses cached middle-layer representations across decoding steps to enable persistent reasoning in transformers with minimal inference overhead.
Benchmark evaluates six MLLMs on scientific visualization literacy using 49 standardized assessment items across 8 visualization techniques.
Study investigates linear representations of grammaticality in neural language models beyond probability-based measures.
MedFailBench is clinician-built open-source benchmark categorizing medical AI failures by severity and safety gate type with 44 synthetic cases.
Opinion piece examines AI's transition from tool to autonomous research participant via industrialization of science, citing DOE Genesis Mission.
Research investigates scaling laws for Behavior Foundation Models in humanoid robot control using large-scale behavioral data.
On-Policy Delta Distillation introduces delta-signal reward for RL post-training, measuring difference between teacher and base models.
A video analytics AI agent that can perceive, reason, and act based on massive amounts of video footage must be integrated with existing workflows and... A video analytics AI agent that can perceive, reason, and act based on massive amounts of video footage must be integrated with existing workflows and applications to be useful. These include content management systems, messaging platforms, databases, ticket queue, and escalation paths. This integration is challenging because video systems, enterprise knowledge bases… Source
Google integrates third-party app connections into Search's AI Mode for secure data access and interaction.
With this new update, Google is expanding AI Mode beyond answering questions and into completing tasks across the apps they use regularly.
Google is giving its AI note-taking app a new name. The company announced on Thursday that NotebookLM is becoming Gemini Notebook, but will remain a standalone app even as it integrates more deeply across Gemini and Google Search. Google first revealed Gemini Notebook - then called Project Tailwind - in May 2023 before widely releasing the app just months later. Over the past few years, Google has been adding new features to the app to help organize and make sense of your notes, such as the ability to summarize them as AI podcasts, narrated slideshows, and TikTok-style clips. It recently star...
Google Vids adds Gemini Omni support and personal avatar features for video generation and editing.
OpenAI introduces age-appropriate safeguards, parental controls, and learning tools for ChatGPT teenage users.