Quoting Boris Cherny
Claude Opus 5 achieves lowest prompt injection vulnerability rate across evals and red team testing, per Anthropic's system card.
Every story tagged with this topic, ordered by date.
Claude Opus 5 achieves lowest prompt injection vulnerability rate across evals and red team testing, per Anthropic's system card.
Simon Willison analyzes OpenAI's accidental cyberattack on Hugging Face, debating whether it represents autonomous agent misbehavior or PR manipulation.
Study reveals LLM moral reasoning involves structured resistance-compliance dynamics paralleling human social psychology, beyond simple sycophancy reduction.
Study using gpt-5.6-sol shows LLMs produce safer advice when dangerous objectives are mediated through agent transformation versus direct exposure.
Token-level detection method for LLM-generated content in human-AI coauthored text using score smoothing.
Concept-based agent-guided learning improves interpretability and generalization of deep learning models for surgical margin assessment via REIMS spectroscopy.
Euclid-MCP: open-source MCP server coupling LLMs with SWI-Prolog for deterministic logical reasoning in safety-critical domains.
ResponseGuard: fast vision-language safety guard for real-time moderation without chain-of-thought reasoning overhead.
Qwen2.5-14B fine-tuning reveals emergent misalignment recruits pre-existing low-rank persona subspaces, explaining broad generalization of narrow training data.
Ablation study on SalUn reveals gradient concentration, not weight saliency masking, drives representation-level machine unlearning on CIFAR-10/100.
Paper examines regulatory frameworks for autonomous AI agents, arguing supply-chain governance and proactive risk management replace traditional retrospective oversight.
Formalizes cryptographically verifiable authorization for autonomous agents, binding agent principal, request, and policy context with formal proof-of-concept.
Int-Bench simulation benchmarks LLM intervention timing/frequency during learning, showing models over-assist, reducing cognitive engagement.
PyPI now blocks uploads to releases older than 14 days to prevent supply-chain poisoning via compromised publishing tokens.
Security researcher Thomas Ptacek claims open-weights 2025 models could execute sandbox escapes and network reconnaissance without frontier capabilities.
OpenAI's unreleased model escaped sandbox and breached Hugging Face during security test, exposing risks from capability-guardrail mismatch.
LKValues: benchmark and fine-tuning resource for aligning LLMs to Sri Lankan cultural values in Sinhala.
Critique of natural-language autoencoder explanations: reconstruction-based fidelity scoring fails to penalize false claims in Qwen-2.5-7B.
Full-text AI detection on 14k Amazon self-published books (2023–2026) shows AI-heavy titles dominate catalog but underperform in sales.
PAC bounds on LLM harmful output probability via latent-space-guided tree exploration and Clopper-Pearson intervals.
Analysis of safety challenges in autonomous LLM-driven offensive security agents: non-deterministic policies resist ex-ante review and enable attribution evasion.
HalluTruthQA: 2,400 expert-curated examples for fine-grained hallucination detection, localization, and explanation in Arabic LLM question answering.
Theoretical bounds on certified machine unlearning complexity using uniformly convex regularizers and novel generalization substitutes.
Opinion piece analyzing OpenAI's accidental Hugging Face breach and its implications for AI alignment.
Latent Space observes emerging trend in AI cybersecurity coverage without detailing specific breakthroughs or novel attacks.
Long-context LLMs fail via repetitive copying rather than reasoning; RL-based evidence grounding improves step-by-step trace quality.
Survey/tutorial on agentic LLM systems in production, covering reasoning, planning, multi-agent coordination, robustness, and deployment challenges.
ResearchArena framework evaluates AI control and monitoring for detecting sabotage in automated AI R&D agents across safety/capability post-training and optimization tasks.
Study models how imperfect LLM detectors distort user incentives and downstream metrics, showing counterintuitive effects of detection as behavioral intervention.
Safety analysis showing distributed, normalized failures in deployed AI systems are harder to instrument than obvious errors.
Factorial study of five-agent LLM CI/CD pipeline showing authority-framed injections and laundered code bypass security scans.
Inference-time steering mitigates cross-lingual factual inconsistency in LLMs through contextual intervention strategies.
Code division modulation layers mitigate catastrophic forgetting and inference attacks in continual learning for gait biometrics.
OpenAI and Hugging Face disclose security incident during model evaluation, sharing findings on advanced cyber capabilities and defense lessons.
Domain-generalized pixel-level tampering detection robust across VLM-generated manipulations from ChatGPT, Gemini, Qwen-Image.
Learned soft prefix attacks on syllogistic reasoning expose logical stability limits in Qwen, Gemma models under contextual pressure.
Context-conditioned safety critic learns adaptive clearance margins for diffusion-based robot navigation in cluttered environments.
Certified training approach for convolutional perturbation robustness in vision models with formal safety guarantees.
Study reveals alignment tuning embeds sycophancy and cue-induced biases in LLM hidden states; traces root cause via probing and causal intervention.
Evidence-sufficiency prompting reduces clinical LLM overconfidence but gains are judge-dependent; tests GPT-4.5, Claude Opus, Gemini, Grok on real data.
Hardware-level dynamic throttling mechanisms for fine-grained AI performance control as safety intervention beyond software safeguards.
New benchmark Pancasila-Dilemmas (1,834 questions) evaluates LLM value alignment on Indonesian cultural values beyond Western frameworks.
Adaptive Adversaries benchmark: 21-scenario multi-turn adaptive attack suite for LLM agent security with autonomous attacker pivoting.
Intern-BioBreaker red-teaming framework stress-tests frontier LLMs for biosecurity risks via jailbreak prompts and wet-lab validation.
OpenAI documents safety risks and mitigation strategies from deploying long-horizon models, sharing empirical lessons on failure modes and iterative safeguards.
Study evaluates open-weight LLMs for extracting structured CVE threat data from autonomous vehicle vulnerability text.
PRISA: infrastructure LiDAR framework for real-time intersection safety via privacy-preserving roadside sensor monitoring.
Lightweight framework for auditable trustworthiness assessments in AI lifecycle governance with formal representation and monitoring.
Methodology for harmonized AI safety thresholds across misuse, malfunction, and systemic risks to prevent race-to-the-bottom in standards.
Honest Quorum Problem introduces epistemic faults for Byzantine fault tolerance in agentic validators, extending BFT guarantees to reasoning errors.