Introducing Claude Opus 5
Anthropic releases Claude Opus 5, matching Fable 5 frontier performance at half the cost, now leading Artificial Analysis leaderboard.
Every story tagged with this topic, ordered by date.
Anthropic releases Claude Opus 5, matching Fable 5 frontier performance at half the cost, now leading Artificial Analysis leaderboard.
RUMBA: Russian benchmark for long-term LLM conversational memory with fine-grained taxonomy across temporal reasoning dimensions.
TriviaRoomQA benchmark evaluates multilingual LLM performance on 3,300 culturally-grounded trivia questions across 6 European languages and long-tail knowledge.
Multi-axis evaluation framework for structured audio captions on AudioCards dataset validates five orthogonal dimensions beyond flat text metrics.
VoLN: vision-only navigation benchmark and method for embodied agents without language instructions in GPS-denied environments.
CM-LRS benchmark evaluates LLM reliability for capital markets workflows, prioritizing regulatory defensibility over surface fluency.
Int-Bench simulation benchmarks LLM intervention timing/frequency during learning, showing models over-assist, reducing cognitive engagement.
Analysis of whether AI labs are optimizing model outputs for specific niche prompts (pelican-bicycle imagery) via systematic testing across 7 models.
LKValues: benchmark and fine-tuning resource for aligning LLMs to Sri Lankan cultural values in Sinhala.
Persian Pixel: large-scale synthetic OCR dataset for Persian language addressing script complexity and data scarcity.
HalluTruthQA: 2,400 expert-curated examples for fine-grained hallucination detection, localization, and explanation in Arabic LLM question answering.
GAMUT benchmark introduces two-level meta-rubrics to measure factual completeness in long-form LLM generation, addressing precision-recall gap in factuality eval.
ResearchArena framework evaluates AI control and monitoring for detecting sabotage in automated AI R&D agents across safety/capability post-training and optimization tasks.
Staypoint detection benchmark provides ground-truth annotations for semantic trajectory analysis from noisy GPS data, addressing lack of standardized evaluation.
BioSecBench-Surveillance: 100-task verifiable benchmark for AI agents inferring pathogen genomic analysis pipelines from raw data.
PathAgentBench: benchmark for vision-language agents on gigapixel whole-slide pathology images evaluating multi-scale evidence-seeking.
Robust financial statement fraud detection framework using LLMs on structured+textual data with temporal generalization evaluation.
MeetingToM benchmark evaluates multimodal LLMs on theory-of-mind reasoning in multi-party meeting scenarios.
MIRA-Ev: multilingual clinical NLP benchmark with span-level evidence detection and argumentation graphs on Spanish MIR exam cases.
Dataset of human image similarity judgments across semantic aspects; benchmarks vision-language models on context-dependent perceptual metrics.
VEHBench: 763-task diagnostic benchmark for evaluating LLM-assisted vibration energy harvester design across coupled physical constraints.
Benchmark evaluating LLMs' ability to reason about 3D spatial constraints in structure-based drug design vs. diffusion models.
Analysis of per-class coverage under distribution shift; split conformal prediction fails per-class validity on skeleton benchmarks.
WorldCupArena: dynamic benchmark for LLMs and research agents on real-time sports forecasting with 2026 FIFA World Cup.
HuGLEN: LLM evaluation pipeline for optical network automation using expert ratings and quality-efficiency scoring.
New benchmark Pancasila-Dilemmas (1,834 questions) evaluates LLM value alignment on Indonesian cultural values beyond Western frameworks.
Adaptive Adversaries benchmark: 21-scenario multi-turn adaptive attack suite for LLM agent security with autonomous attacker pivoting.
ActiveVision benchmark measures whether MLLMs perform active observation via 17 tasks requiring repeated visual perception.
CRAFT converts rubric-based evaluations into capability diagnoses and generates targeted fine-tuning data addressing model weaknesses.
New benchmark evaluates frontier LLMs on real analytical knowledge work—synthesizing information, judgment under uncertainty, strategic thinking—beyond factual recall and coding.
Vision-language model achieves SOTA remote sensing benchmarks via simple scaling recipe without task-specific architectural changes.
Comprehensive realistic benchmark for DRL reach-avoid task on robotic arms; shows poor generalization from simplified settings.
OpenAI CFO Sarah Friar proposes AI scorecard framework measuring ROI via useful work, cost-per-task, dependability, and compute efficiency.
Moonshot AI releases Kimi K3 (2.8T params), claims top performance vs. Claude Opus 4.8 Max and GPT-5.5, promises open-weight release by July 2026.
SciDiagramEdit benchmark teaches models to edit scientific figures via natural-language instructions using skill-evolution training.
Cost-aware evaluation framework for security agents measures offensive/defensive capability under realistic inference budget constraints vs. peak performance.
MediaEval Medico 2025 analysis: parameter-efficient VQA achieves leaderboard gains but structured reasoning better ensures clinical interpretability.
TikStance: 13,876-comment multimodal dataset for stance detection across Trump, Biden, Harris in 2024 U.S. election TikTok discourse.
Symbal detects systematic misalignments in MLLM-generated image captions by identifying recurring errors tied to specific visual features.
MM-IssueLoc benchmark isolates visual evidence impact in multimodal software repository issue localization across 23 languages and 652 instances.
Critical analysis of item response theory reliability for AI benchmarks, highlighting regime mismatches between IRT assumptions and benchmark data distributions.
Benchmark evaluates six MLLMs on scientific visualization literacy using 49 standardized assessment items across 8 visualization techniques.
MedFailBench is clinician-built open-source benchmark categorizing medical AI failures by severity and safety gate type with 44 synthetic cases.
Diagnostic study isolates and evaluates five visual world models (DreamerV3, DIAMOND, TWISTER, Simulus, STORM) in Atari Pong.
OmniaBench: unified benchmark evaluating LLM-based agents across diverse scenarios with explicit state spaces for systematic capability characterization.
CFM-Bench: unified multi-domain benchmark for channel foundation models enabling fair comparison across wireless tasks and pretraining approaches.
Study evaluating whether synthetic face datasets can replace real benchmarks for face recognition evaluation across 12 synthetic vs 7 real datasets.
StructureClaw: artifact-centered benchmark for evaluating LLM agents on complete structural engineering workflows with verifiable evidence chains.
Hindcast methodology closes data leakage in LLM forecaster evaluation by controlling for retrieval and training-data contamination in backtests.
Study shows agent-optimization gains may not compound over time; proposes Terminal-Bench 2.0 to test continual learning on deployed agents.