SAEScientist-Bench: Can AI Agents Conduct Autonomous SAE Interpretability Research?
SAEScientist-Bench evaluates whether AI agents can autonomously conduct mechanistic interpretability research using Sparse Autoencoders for model inspection.
Every story tagged with this topic, ordered by date.
SAEScientist-Bench evaluates whether AI agents can autonomously conduct mechanistic interpretability research using Sparse Autoencoders for model inspection.
SPINE benchmark measures LLM sycophancy through 25-turn adaptive adversarial conversations, revealing collapse in four production systems and Olmo3-7b variants.
Doctorina clinical AI achieves 82% diagnostic concordance versus 57% for physicians in 150 Polish primary-care cases, outperforming frontier LLMs.
Audit methodology effects rival demographic bias in LLM decision-making: charitable-aid benchmark results fail to replicate across hiring, lending, triage domains.
Knowledge graph-based evaluation framework to assess contextual understanding in LLMs beyond surface-level metrics.
SQLMorph: query mutation framework for reproducible Text-to-SQL evaluation on complex enterprise schemas.
ONECYL benchmark for graph-based surrogate modeling of CFD on unstructured meshes across flow regimes.
Fact-Ablated Evaluation framework audits whether LLMs faithfully use evidence or rely on parametric knowledge in fact-checking.
Q2D-Web large-scale benchmark for agentic RAG evaluates first-stage retrievers on agent-reformulated queries at production scale.
API benchmark scores for ChatGPT, Claude, Gemini systematically diverge from chatbot interface behavior by 3.4pp average; questions reliability of model eval standards.
Closed-form estimator for LLM-judge panel error decomposition under single-common-factor model; addresses contaminated anchor detection in eval.
Rater Ising-Potts model with LLM-derived weights for multi-category scoring reliability; extends classical statistical model for LLM-based evaluation.
Machine translation evaluation metrics accounting for terminology variation in English-French scientific translation; cross-term variation diagnostic.
Cybersecurity LLM benchmarks show 80-point score swings from pipeline choices; audit of 8 benchmarks reveals 15 systematic failure modes.
Kairos: fine-grained video-language dataset with time-resolved annotations across 10–30 min videos for temporal dynamics modeling.
OpenAI releases AI-generated solution to Navier–Stokes Millennium Prize Problem with formal Lean proof.
ProcArena benchmark evaluates LLMs on multi-scenario PL/SQL development including direct generation, modification, debugging, and interactive multi-turn workflows.
COSSG generates plausible safety-critical autonomous vehicle scenarios via collision snapshots and time-reversed trajectory synthesis.
Routing effective rank metric reveals reproducible low-high-low trajectory in MoE expert routing across test-time scaling, with implications for inference budget allocation in math/science tasks.
Analysis of Qwen2.5-1.5B rollouts reveals within-group verifier-error correlation of 0.530, suggesting group-based RL with automatic verifiers requires design-effect adjustment for statistical validity.
Conformal prediction framework for uncertainty quantification in multi-agent LLM-as-a-Judge evaluation with aggregated reliability bounds.
CNN for galaxy morphology classification from crowd-annotated Galaxy Zoo 1 data; analyzes training scheme and annotator agreement.
Multi-model LLM evaluation for educational short-answer scoring using GPT, DeepSeek, Qianwen across reliability, validity, and failure modes.
Benchmark of 516 Reddit suicide posts clinically rated reveals gap between content moderation flags and clinical severity tiers.
WearableQA benchmark: 4,084 QA pairs from 200 real users' wearable time series, testing longitudinal health reasoning over authentic device noise.
ROBORMBENCH: 2,390 robot trajectories revealing paraphrase fragility in VLM reward models—semantic synonyms flip success/failure predictions.
KOPA-Bench: 145 Korean public API tasks; EDGE synthesis method closes open-source LLM gap in multi-step tool-calling for on-premise agents.
Audit of 22 frontier LLMs: widespread digit-level retrieval of published values on molecular benchmarks, conflating memorization with prediction.
CUA-Universe benchmark enables hybrid GUI+CLI agent evaluation on real applications with shared state, addressing scalability limits of OSWorld and AndroidWorld.
RoboSPA benchmark evaluates Vision-Language-Action models on spatial reasoning and procedural complexity in robotic manipulation tasks.
LexFlip benchmark with 373 Quebec French perturbations diagnoses legal meaning preservation in simplified clauses via dissociation method.
Framework for evaluating automatic NLG metrics via behavioral correctness assumptions under controlled response transformations.
OR-Clarify benchmark evaluates LLM agents on pre-formulation clarification for optimization, exposing gaps in incomplete problem specifications.
Knowledge Space Theory framework evaluates whether LLMs exhibit structured, prerequisite-dependent knowledge in mathematical reasoning.
Phase transition frequency during ResNet finetuning predicts test accuracy; strong negative correlation observed across benchmarks.
AxQM benchmark: 1,019 formal proof-synthesis tasks for autoformalization in finite-dimensional quantum mechanics.
SciDocBench: workflow-centered benchmark with 124 expert questions across scientific domains testing joint reasoning over text, equations, figures, tables, code, and datasets.
Single-query calibration auditing of LLM APIs via logit_bias parameter enables True Calibration Error estimation bypassing hidden probability outputs.
TIER: threat implicitness benchmark for LLM safety across four risk domains and four threat levels using six-label behavior scale and dual LLM judges.
OpenAI launches GPT-6 Astra, a Claude Fable competitor priced at $10/$50 per million tokens, rolling out to ChatGPT Plus/Pro/Business/Enterprise and via API.
Preregistered audit of LLM-as-judge reliability finds Spearman 0.40 repeat agreement (vs. 0.90 required), exposing instability in model endpoints underpinning leaderboards and training data filtering.
Proposes new benchmark and evaluation methodology for machine translation addressing saturation of standard benchmarks, metric unreliability, and reproducibility gaps in human evaluation.
SWE-Gate benchmark evaluates coding agents on review-constraint compliance beyond functional correctness in repository-level tasks.
NVFP4 W4A4 quantization on all 496 linear layers of Qwen3.8-27B hybrid LLM including Gated DeltaNet recurrent blocks maintains performance across evals.
PatchBench audits AI agents' C/C++ vulnerability patching; finds 25% exhibit patch memorization or surface-level fixes.
LLMs over-edit code during repair; study of 400 BigCodeBench problems shows widespread over-editing even in frontier models.
Dice Roll Method: standardized protocol for repeated-query auditing of LLM brand recommendations with variance decomposition.
IRWOZ 2.0 dataset: 390 LLM-enhanced dialogue annotations (Mistral, Claude-3.5) for industrial robot conversations across 4 domains with improved quality.
FLY-EVAL++: safety-focused evaluation protocol for LLMs in physics-governed domains; measures constraint violations and physical inconsistency beyond accuracy metrics.
InSituMeasure benchmark evaluates MLLMs on continuous-valued measurement tasks in realistic industrial settings with gauge reading and instrument-specific context.