Fast Rates for Semi-Supervised Learning via Data-Augmentation Graph Regularization
Theoretical proof that data augmentation induces graph-Laplacian regularization, achieving O(1/n_L) label-sample efficiency versus O(1/√n_L) supervised.
Search the full wire by company, model, lab, or keyword. Every story we have ever aggregated.
Theoretical proof that data augmentation induces graph-Laplacian regularization, achieving O(1/n_L) label-sample efficiency versus O(1/√n_L) supervised.
Analysis of permutation-equivariance and stability in continuous flow models parameterized by GNNs for graph signal generation.
Asynchronous single-rollout RL for agentic LLM post-training addresses training stability and efficiency limits of batch-synchronous GRPO.
HIVE benchmark systematically evaluates how hallucinated content in vision-language models influences downstream reasoning, moving beyond detection to understanding post-hallucination inference.
Empirical study compares LLM-generated versus expert-written skill modules for agentic data science tasks, finding trade-offs in automation cost versus performance across SQL, cleaning, and statistical testing.
Agentic systems often face a trade-off between accuracy and cost. The highest-performing proprietary frontier models and harnesses provide top accuracy but are... Agentic systems often face a trade-off between accuracy and cost. The highest-performing proprietary frontier models and harnesses provide top accuracy but are expensive. Fine-tuning offers one way to address this problem. Smaller or more efficient open models starting with lower accuracy are taught to perform better with specific agents. However, fine-tuning requires expertise and hardware for… Source
TimEE is a 4.5M-parameter foundation model enabling end-to-end time series classification via in-context learning, eliminating per-dataset retraining.
RAID applies iterative reinforcement learning to automatically discover exploits in EA Sports NHL 26's goalie AI, reducing manual playtesting effort.
GIFT proposes geometry-informed gradient quantization for low-precision communication during LLM pretraining, addressing scaling bottlenecks in distributed training.
Pyligent training framework treats reasoning as validated search with backtracking, enabling models to correct mid-inference and recover from failed branches.
FourierQK applies FFT-based spectral preprocessing to query-key projections in transformer attention, achieving 79% perplexity reduction on character-level language modeling.
Action-graded severity scale for agent red-teaming replaces binary attack-success metrics with 7-level ordinal harm rubric, enabling nuanced risk assessment of tool-using AI compromise.
First systematic evaluation of fairness interventions under differential privacy constraints on synthetic tabular data, revealing tensions between DP and fair-ML objectives.
SynthAVE uses LLM-based synthetic labeling with arena validation for e-commerce attribute extraction at scale, covering 12,726 products across 229 categories and 12 languages.
Statistical theory for sparse function recovery from noisy indirect observations via ℓ1-regularization and empirical risk minimization.
SpaCellAgent: LLM-based multi-agent framework automating trajectory inference for spatial and single-cell transcriptomics analysis.
Self-evolving LLM agents with biased reward signals fail to retire bad skills, disabling safety constraints in skill libraries.
RLVP: Reward function design for real-world agents requiring path constraints and outcome-neutral safety rules beyond reward maximization.
Empirical study of biologically-informed neural network architecture and optimization for reliable mechanistic operator recovery from sparse data.
PAC learning theory: sample complexity of Chain-of-Thought reasoning bounds by local next-token classification dimension.
InductWave: Wavelet-based inductive embedding for logical multi-hop query answering over knowledge graphs with unseen entities.
DeLS-Spec: Decoupled context speculative decoding for parallel LLM token drafting without retraining draft models.
SAMPA: Whisper-based speech segmentation model for prosodic boundary detection in Brazilian Portuguese.
Tool-using LLM agents silently violate deployed policies via well-formed tool calls that bypass domain constraints; 78% of failures undetected.
OpenAI outlines principles for government and national security partnerships, emphasizing responsible AI deployment, democratic accountability, and public safety.
When it comes to achieving artificial general intelligence (AGI), large language models just don’t have what it takes. Models like ChatGPT and Claude are great at text, but they’re less skilled at understanding how things actually move through space and time — an essential skill for producing intelligence that generalizes. That gap, it turns out, might be filled by gaming data. That’s the bet behind General Intuition, a […]
OpenAI identifies validity issues in SWE-Bench Pro coding benchmark, questioning reliability of popular AI model evaluation metric.
Kevin Weil's new role at Stoke Space suggests reusable rockets are the next hot thing in Silicon Valley.
Microsoft Xbox division cuts staff as Game Pass subscription strategy underperforms; analysis of bundling economics.
OpenAI and Walton Family Foundation launch AI Skills Jams for K–12 educators to teach classroom AI applications.