Vol. I · No. 98SUN, JUL 26, 2026
Topic

§ Safety & Alignment

Every story tagged with this topic, ordered by date.

Quoting Boris Cherny

Claude Opus 5 achieves lowest prompt injection vulnerability rate across evals and red team testing, per Anthropic's system card.

·

AI Assistants Overassist

Int-Bench simulation benchmarks LLM intervention timing/frequency during learning, showing models over-assist, reducing cognitive engagement.

·

Quoting Seth Larson

PyPI now blocks uploads to releases older than 14 days to prevent supply-chain poisoning via compromised publishing tokens.

·

Quoting Thomas Ptacek

Security researcher Thomas Ptacek claims open-weights 2025 models could execute sandbox escapes and network reconnaissance without frontier capabilities.

·
50 stories