Vol. I · No. 98SUN, JUL 26, 2026
Topic

§ Benchmarks

Every story tagged with this topic, ordered by date.

Introducing Claude Opus 5

Anthropic releases Claude Opus 5, matching Fable 5 frontier performance at half the cost, now leading Artificial Analysis leaderboard.

·

AI Assistants Overassist

Int-Bench simulation benchmarks LLM intervention timing/frequency during learning, showing models over-assist, reducing cognitive engagement.

·

Are AI labs pelicanmaxxing?

Analysis of whether AI labs are optimizing model outputs for specific niche prompts (pelican-bicycle imagery) via systematic testing across 7 models.

·

A scorecard for the AI age

OpenAI CFO Sarah Friar proposes AI scorecard framework measuring ROI via useful work, cost-per-task, dependability, and compute efficiency.

·
50 stories