TOOD: Task-Aware Out-of-Distribution Score Calibration for Continual Learners
Study of out-of-distribution detection degradation in continual learning systems; proposes mitigation strategies for OOD forgetting.
Search the full wire by company, model, lab, or keyword. Every story we have ever aggregated.
Study of out-of-distribution detection degradation in continual learning systems; proposes mitigation strategies for OOD forgetting.
ResKV: KV cache compression method using residual statistics for fixed-budget long-context inference without permanent token discarding.
TraceViT: looped visual reasoner for ARC benchmark trained with step-by-step transformation supervision using programmatic task rewriting.
Vision-language models exhibit sycophancy that undermines epistemic vigilance in cooperative tasks; detects inconsistencies with prior beliefs.
Apple CEO Tim Cook envisions users being able to buy more compute for Siri AI via Apple's existing iCloud+ subscriptions.
DungeonBench: tactical reasoning benchmark for D&D combat covering geometry, timing, resources, rule interactions with 2014 SRD mechanics.
MOT-SR: LLM-based symbolic regression with data analysis mechanisms and multi-objective optimization for scientific equation discovery.
LEMUR: multi-objective RL from preference feedback without requiring well-specified reward functions for each objective.
Mathematical proof: counterexample to 2015 Lacoste-Julien/Jaggi polytope conjecture regarding pyramidal width under vertex insertion.
COntExt: framework for automated ontology extension from structured operational metric definitions using domain knowledge extraction.
AMTFV introduces agentic verification for LLM math answers via mathematical tool flows, decoupling modeling from implementation.
ARB benchmark evaluates AI-text detectors against LLM-rewritten human content using Llama-3.2 and Qwen2.5 generators.
Neurosymbolic pipeline combines foundation models with Bayesian Networks for automated Alzheimer's diagnosis from speech.
TerraNova foundation model integrates Earth system physics and societal data across continuous and administrative geometries.
ARCTIC AI code critique system prioritizes correctness and security over style via intent prediction and drift detection.
SpaceX is building a new power plant for xAI's Colossus data centers, but it won't remove existing, unpermitted turbines for many more months.
Transfer learning with GNNs predicts formation energy and HOMO-LUMO gaps in high-entropy perovskite oxides.
The demand for high-quality video continues to accelerate across industries, powering everything from immersive streaming experiences to remote collaboration,... The demand for high-quality video continues to accelerate across industries, powering everything from immersive streaming experiences to remote collaboration, generative AI media tools, and large-scale content delivery. Behind these experiences is a growing need for video pipelines that are faster, more efficient, and capable of handling increasingly complex formats and workloads. Source
Class-specific decoder architectures enable cross-domain transfer learning for multi-organ laparoscopic segmentation.
High-entropy parameter solutions mitigate catastrophic forgetting in neural networks via Boltzmann entropy robustness.
Theoretical analysis of finite-precision Transformers with transcript management and pop-enabled context channels.
Adaptive FastOPD uses progress-aware rollout horizon expansion to accelerate on-policy distillation training efficiency.
OpenAI outlines safety, security, and transparency practices aligned with EU AI Act compliance and responsible governance.
OpenAI outlines full-stack strategy to improve AI capability, cost, and accessibility across models and infrastructure.
DreamQAS applies model-based RL to quantum circuit optimization by learning VQE feedback predictions while preserving known circuit dynamics.
Study shows interventional data alone fails to teach LMs causal direction in Simpson's-paradox settings; observational context dominates learned do()-response.
The startup is building voice models designed to make AI phone calls pass the Turing test.
MolGVR framework adds verification and refinement to text-to-molecule generation to enforce chemical constraints and correct structural violations.
Analysis of decoder design trade-offs in lightweight neural networks for visual affordance segmentation on wearable robots.
SESA framework combines self-play curriculum learning with evolving procedural memory to distill failures into reusable skills for search agents.