Back to Blog

Anthropic

4 articles on this topic.

AI Governance14 September 2026

Anthropic Extinction Claims Spark an Evidence Fight

A former Anthropic employee's viral claim that AI could 'kill us all by the end of the decade' drew a pointed public rebuttal — and raises a real question for anyone building AI risk assessments: what actually counts as evidence?

ai-safetyai-governanceanthropic
4 min readRead
AI Security & Governance2 September 2026

Anthropic Hardens Claude's System Prompt Against Song Lyrics

A September 2026 update to Claude Fable 5.1's system prompt adds a persistent, decomposition-resistant refusal for song lyrics, poems, and copyrighted visuals — a public case study in guardrail engineering under litigation pressure.

ai-securityllm-guardrailsprompt-engineering
4 min readRead
AI & LLM Security11 August 2026

How Researchers Cracked Encrypted Chain-of-Thought in Claude, GPT and Gemini

A new paper shows that the encrypted reasoning blocks Anthropic, OpenAI and Google return from their APIs can be replayed into a weaker sibling model and jailbroken into plaintext — defeating anti-distillation protections and, in the wild, exposing PII and credentials.

llm securitychain-of-thoughtai security research
4 min readRead
AI Security4 August 2026

Shared Claude Chats Were Indexed by Google, Exposing Private Data

A public-sharing feature without a noindex tag let Google crawl and surface Claude conversations users had shared with a link — including crypto wallet keys, medical dashboards, and therapy-app source code.

ai-securitydata-exposureanthropic
4 min readRead