Ask Before You Optimize: Dynamic Pre-Formulation Clarification for Interactive Optimization
OR-Clarify benchmark evaluates LLM agents on pre-formulation clarification for optimization, exposing gaps in incomplete problem specifications.
Search the full wire by company, model, lab, or keyword. Every story we have ever aggregated.
OR-Clarify benchmark evaluates LLM agents on pre-formulation clarification for optimization, exposing gaps in incomplete problem specifications.
Survey of commonsense reasoning in computer vision, covering integration of visual data with contextual knowledge for holistic scene understanding.
GLASS framework aligns graph encoders with text embeddings via hypersphere geometry for cross-domain transferable anomaly detection.
Quantum machine learning framework using RL-optimized circuits for semiconductor device modeling with reduced error versus classical baselines.
Proton irradiation testing of open-source Tensil NN accelerator on Zynq platform for radiation-hardened spaceborne computing.
Knowledge Space Theory framework evaluates whether LLMs exhibit structured, prerequisite-dependent knowledge in mathematical reasoning.
Study of 3,471 uncensored open-weight models on HuggingFace (Jan 2024–Mar 2026), tracking safety guardrail removals and redistribution persistence.
PRICE framework systematically studies LLM adaptation (fine-tuning, prompting, decoding) for Bitcoin price forecasting with quantized LLaMA-3 8B.
Hessian-based data augmentation for machine-learning interatomic potentials reduces computational overhead vs. explicit Hessian training.
Study of substrate blindness in AI agents: Claude Opus 5, GPT-5.6-Sol, Gemini 3.7 Flash code generation ignoring memory/compute constraints.
Mixture-of-Experts (MoE) architectures provide an efficient paradigm for scaling large language models (LLMs), yet fixed top-k routing activates the same number of expert slots for every token, causing substantial redundant computation. Existing expert-skipping methods often rely on router confidence, calibration data, or additional training, and therefore cannot reliably estimate the actual contribution of routed experts. To this end, we propose ACE, a training-free, calibration-free, and checkpoint-preserving framework for token-adaptive expert skipping in MoE-based LLMs. ACE contains two c...
CABAL framework simulates multi-agent collusion in peer review bidding to study reviewer assignment integrity risks.
FedDRAW applies federated learning with reputation-weighted aggregation to multi-institutional chest radiograph classification.
Verifier-guided framework combines QLoRA, MoE routing, and symbolic solvers (Z3, FOL) for explainable LLM reasoning in education.
Count-Adaptive BLiN removes zooming dimension dependency in batched Lipschitz narrowing algorithms.
Gemini Spark can edit and curate photo albums, create shared collections, turn photos into calendar events, and handle other Google Photos tasks for AI Pro and Ultra subscribers.
PAC-Bayesian framework extends generalization guarantees to VAEs for time series forecasting beyond i.i.d. settings.
FluxDisco applies physics-informed symbolic regression via Monte Carlo graph search for stoichiometric ODE discovery.
Empirical study of on-policy distillation shows 1-shot OPD effective for LLM reasoning post-training; harder examples boost gains.
Phase transition frequency during ResNet finetuning predicts test accuracy; strong negative correlation observed across benchmarks.
Mirror Agent Model defines observer as agent model mirror to generate interpretable behavior and explanations.
FlexPension-LLM domain-specialized model with DKI-RDistill predicts pension enrollment behavior for Chinese flexible workers.
arXiv paper proposes latent-distance-based novelty measurement for biomedical papers using knowledge unit relationships.
SMILE framework enables self-explainable multimodal medical diagnosis via information bottleneck optimization.
arXiv explores conformal prediction for offensive cybersecurity applications including privacy-preserving ML attacks.
Less than 24 hours left to apply to host a Side Event during TechCrunch Disrupt 2026 and make your mark in the Silicon Valley scene. Apply before the application closes tonight at midnight PT.
AxQM benchmark: 1,019 formal proof-synthesis tasks for autoformalization in finite-dimensional quantum mechanics.
DEX-Comp: two-stage training recipe for soft context compression in RAG systems via pure distillation and fine-tuning.
arXiv proposes time-to-first-spike SNNs with reference-based encoding for energy-efficient large language models.
RCBNB-MB algorithm discovers causal structures in non-stationary time series via latent regime-aware Markov blankets.