An Alien Mind
Jakub Pachocki (OpenAI) argues for stronger AI safeguards and international coordination as capability scaling accelerates.
Search the full wire by company, model, lab, or keyword. Every story we have ever aggregated.
Jakub Pachocki (OpenAI) argues for stronger AI safeguards and international coordination as capability scaling accelerates.
Second-order smooth planning via optimal-transport Bellman smoothing achieves improved O(ε^−2.67) sample complexity over first-order SmoothCruiser.
Study explains performance degradation in pruned DNNs via interaction pattern dynamics, identifying three-phase behavior as pruning ratio increases.
O2C-Nav enables zero-shot vision-language navigation in continuous environments with single MLLM call per step via spatial-aware waypoint generation.
Kernel methods framework learns kernels via alignment through Collaborative Learning and Inference for multiclass Bayes classification.
Commentary on technical debt: software degrades without architectural limits, unlike physical systems.
Study characterizes interaction between inference-time activation steering and weight-only quantization (INT8, NF4) on 7-9B open-weight LLMs.
Two-stage grouping algorithm mitigates priority sensitivity in fractional knapsack problem via angular sector partitioning.
Cross-step control framework mitigates higher-order sequential interference in multi-domain RL training for LLMs by analyzing consecutive gradient reversals.
Sector-Mean deterministic initialization for K-Means uses angular sector partitioning with O(N) complexity to improve convergence and accuracy.
Theoretical analysis of performative RL stability for policy mixtures, relaxing Lipschitz assumptions in multi-agent best-response settings.
Theoretical framework for masked pretraining establishes connections to contrastive learning and characterizes masking's role in representation extraction.
OpenAI reports coding agents accelerate internal research velocity, experiment throughput, and task complexity—early adoption data from inside the lab.
Phase-cycled randomized benchmarking detects hidden classical noise correlations in quantum processors via connected sine-phase covariance.
Framework for post-hoc causal attribution in agentic AI systems; establishes traceability requirements and identifies estimator failures for regulatory compliance.
Bayesian method separates aleatoric and epistemic uncertainty in LLM judges to optimize expert labeling efficiency without additional model calls.
InsightChain pipeline automates multi-stage data visualization via LLM with vision-guided prompt optimization; evaluated on chart generation tasks.
Self-supervised and generative pretraining improve deep learning for medical image distortion correction in diffusion-weighted imaging.
COSSG generates plausible safety-critical autonomous vehicle scenarios via collision snapshots and time-reversed trajectory synthesis.
Statistical learning theory analysis of straight-through estimators for training binary quantized neural networks via stability bounds.
Sparse oblique rule boosting generates interpretable additive ensembles with oblique decision boundaries instead of axis-parallel polytopes.
Batch normalization statistics selection improves value-based RL in both continuous control and discrete Atari tasks.
DualRead separates answer capability from confidence calibration in medical vision-language models trained via GRPO.
ViS-CoT combines visual search with chain-of-thought reasoning for product attribute extraction from videos without fine-tuning.
Hierarchical Wasserstein Merging addresses multi-domain multi-task learning by merging specialist model distributions rather than parameters, reducing interference under distribution shifts.
Routing effective rank metric reveals reproducible low-high-low trajectory in MoE expert routing across test-time scaling, with implications for inference budget allocation in math/science tasks.
MetaRSI proposes extending recursive self-improvement beyond formal verification to open-ended scientific domains, arguing current RSI validation is limited to coding and checkable benchmarks.
Determinantal sampling improves k-clustering coreset bounds beyond worst-case analysis, reducing memory requirements for clustering on massive datasets via instance-dependent approach.
Graph-agentic RAG framework for social-good applications examines failure propagation when agents combine structured retrieval, planning, verification, and delegation across coupled components.
Analysis of Qwen2.5-1.5B rollouts reveals within-group verifier-error correlation of 0.530, suggesting group-based RL with automatic verifiers requires design-effect adjustment for statistical validity.