Minimax bounds for watermarked and masked recursive discrete distribution estimation
Theoretical analysis of minimax loss in recursive discrete distribution estimation under watermarking, contrasting oracle-assisted and unassisted settings.
Search the full wire by company, model, lab, or keyword. Every story we have ever aggregated.
Theoretical analysis of minimax loss in recursive discrete distribution estimation under watermarking, contrasting oracle-assisted and unassisted settings.
Method to recover multi-token concept vectors directly from J-lens LLM interpretability tool without fine-tuning, using first-token cues.
Framework for reducing token consumption in LLM agents reasoning over unstructured data via adaptive pre-structuring, addressing enterprise AI cost barriers.
Analysis showing sycophantic agreement emerges unintentionally from contrastive preference optimization; demonstrates transfer across model families via OLMo 3.
Framework reconciling process supervision with outcome-based credit assignment in agentic RL using privileged information, improving fine-grained policy learning.
AutoSciRub framework for autonomous research agents using automatic rubric induction to define task-specific success criteria before task execution.
Study of scaling large reasoning models beyond human supervision using verifiable rewards, examining paths toward agentic reasoning without direct human oversight.
Lightweight two-stage real-time video anomaly detection using YOLO pose estimation and CLIP semantic similarity scoring on skeletal keypoints.
Soft Latent Thinking method enabling LLM reasoning in continuous embedding space instead of discrete tokens, reducing per-step compute while improving pass@k.
Analysis revealing hidden-state probes decode correct reasoning answers when native LLM sequence scoring fails, suggesting output bottleneck vs. capability gap.
Transformer parameterization with 280 parameters achieves perfect length generalization on Boolean algebra circuits via depth-1 reduction.
Model-internal token saliency via residual stream analysis compresses chain-of-thought reasoning while preserving model performance.
Theoretical analysis of when text embeddings preserve latent topic mixtures for clustering and causal inference tasks.
Learned query generation (LoQ, FeedQ) improves information extraction by 18.6 F1-points, outperforming model scaling across clinical benchmarks.
Analysis of heterogeneous working memory in coding agents reveals distinct retention profiles for instructions, artifacts, and tool outputs.
Agentic AI is changing how research is done. AI scientists can read papers, propose hypotheses, call models, and determine which experiments to prioritize next.... Agentic AI is changing how research is done. AI scientists can read papers, propose hypotheses, call models, and determine which experiments to prioritize next. First proving their value in software engineering, coding agents now write, test, and ship production code. Scientific research can be more demanding and iterative. Researchers continually evaluate evidence, refine hypotheses… Source
Comparative study of CNNs and attention models for bovid tooth segmentation under noisy annotations.
On-policy distillation analysis reveals substantial noise in teacher supervision yet student convergence is robust to it.
Tutorial on rotational equivariance in ML for 3D data covering physics, materials science, and computer vision applications.
Contextual learning framework addresses catastrophic forgetting and semantic shift in incremental named entity recognition.
Statistical analysis of 13 neural audio codec token sequences across architectures and noise conditions using Zipf/Heaps metrics.
Normalized Low-Rank Adaptation (NoRA) stabilizes LoRA training by normalizing down-projection matrices, improving parameter-efficient fine-tuning.
Study evaluates when learned perceptual predictors align with human listeners in codec-based speech models using GRPO with CER constraints.
Probes autonomous driving benchmark scores by removing camera input, measuring how much performance relies on memory vs. dynamic scene understanding.
Framework learns latent representations of PDEs by embedding scientific inductive bias into training distributions for hypothesis discovery.
MNIST-PRO benchmark isolates agentic perception by converting digit recognition into sequential glimpse-based search with memory constraints.
Audit of three AI clinical scribes on 142 consultations found 31.3% of notes contain verified errors, concentrated in allergy and medication data.
The three-year-old startup says it reached $15 million in ARR and profitability before raising its latest $15 million round.
A perception stack is shaped by the vehicle that carries it. Move the same software to a new carline—for example, from an SUV to a sedan or another vehicle... A perception stack is shaped by the vehicle that carries it. Move the same software to a new carline—for example, from an SUV to a sedan or another vehicle variant in the portfolio—and its perception of the world changes. The sensor placement, calibration, fields of view, occlusions, body geometry, timing, and coverage all shift. A traffic light may appear in a different part of the frame. Source
LLM judges fail to detect omissions in clinical notes; benchmark with 500 single-error pairs shows judges verify presence, not absence.