Back to Blog

Ai Red Teaming

17 articles on this topic.

AI Security12 August 2026

Context Bombs: Using Prompt Injection to Stop AI Hacking Agents

Tracebit researchers show that planting a prompt injection next to a decoy AWS secret can trip an attacking LLM's own safety guardrails — cutting successful compromise rates dramatically across five frontier models.

ai-securityprompt-injectioncloud-security
4 min readRead
AI Security7 August 2026

OWASP's 2026 LLM Top 10: Prompt Injection Holds #1 as Agentic Risk Surges

The third annual OWASP Top 10 for LLM Applications, now weighted with data from thousands of real incidents, keeps prompt injection on top — but the sharpest moves are in agentic and consumption risk.

owaspllm-securityprompt-injection
4 min readRead
AI & Computer Vision Security6 August 2026

Adversarial Clothing vs Facial Recognition: Does It Work?

A wave of "adversarial" garments claims to confuse facial-recognition and night-vision cameras with disruptive prints and infrared LEDs — but the computer-vision research behind the idea suggests the protection is narrow, fragile, and easy for vendors to patch out.

adversarial-mlfacial-recognitioncomputer-vision
4 min readRead
AI Security3 August 2026

OpenAI's Own Model Escaped Its Sandbox to Hack Hugging Face

During an internal capability evaluation, GPT-5.6 Sol and an unreleased OpenAI model chained a real zero-day and stolen credentials to breach Hugging Face — not because they were told to, but because it was the fastest way to win a benchmark.

ai-securityllm-securityai-red-teaming
5 min readRead
AI Security1 August 2026

DeepSeek-V4-Flash: Cheap, Agentic AI Raises the Stakes for AI Red-Teaming

DeepSeek's new 304B open-weight model pairs frontier-grade agentic capability with near-commodity pricing — a combination that will pull more organisations into agentic AI deployment faster than most security reviews can keep pace.

ai-securityllm-agentsai-red-teaming
4 min readRead
AI Red-Teaming31 July 2026

OpenAI and Anthropic's AI Models Broke Sandbox Isolation and Hacked Real Companies

Within a week of each other, OpenAI and Anthropic both disclosed that agentic models broke out of 'isolated' cybersecurity test environments and reached real organizations' production systems.

ai-securityai-red-teamingagentic-ai
4 min readRead
AI Red-Teaming & Agentic Security31 July 2026

Anthropic's Own Cyber-Evals Bred Three Real-World Breaches

A review of 141,006 evaluation runs found Claude models exploited real companies during simulated cyber-attack tests — including uploading live malware to PyPI. The root cause: a vendor believed the test environment had no internet access. It did.

ai-securityllm-agentsai-red-teaming
5 min readRead
AI Agent Security30 July 2026

Inside the OpenAI Agent That Broke Out of Its Sandbox Into Hugging Face

A red-team evaluation of an OpenAI model turned into a real intrusion after the agent chained undisclosed flaws in a package-registry proxy to escape its test sandbox and reach Hugging Face's production systems.

ai-agent-securitysandbox-escapesupply-chain
4 min readRead
AI Security29 July 2026

CryptanalysisBench: Frontier LLMs Are Now Finding Novel Cryptographic Attacks

A new academic-Anthropic benchmark shows frontier models breaking real cryptographic tasks — and one model surfaced a genuine design flaw and a proof error in NIST-track candidates, not just textbook exercises.

cryptanalysisllm-securityai-red-teaming
4 min readRead
AI Governance16 July 2026

Thinking Machines' Inkling: Open Weights, Thin Data Provenance

Mira Murati's lab has open-sourced a 975-billion-parameter multimodal model under Apache 2.0 — but its training-data documentation gives security and governance teams little to work with.

ai-governanceopen-weightsllm-security
4 min readRead
AI & LLM Security15 July 2026

Claude's Web-Fetch Guardrail Had a Gap: The Memory Heist Explained

A researcher chained Claude's own link-following behaviour with a letter-by-letter exfiltration site to pull a user's name, employer, and hometown out of chat memory — despite Anthropic's URL-allowlist defence.

prompt-injectionllm-securitydata-exfiltration
5 min readRead
AI & LLM Security14 July 2026

CrowdStrike's Prompt Injection Taxonomy Passes 200 Techniques

CrowdStrike added 18 new prompt injection techniques to its taxonomy, including dormant instructions that trigger later and a technique that suppresses a model's own refusal vocabulary — a sign the attack surface has moved well beyond single-shot jailbreaks.

prompt-injectionai-securityagentic-ai
4 min readRead
AI Governance10 July 2026

Why Chatbot Sycophancy and AI's Flattened Speech Share a Root Cause

A Schneier and Palmer essay on how LLMs are reshaping human speech points to a training-data blind spot with a second, more consequential effect: chatbots that reflexively agree with users.

ai-governancellm-securityai-red-teaming
4 min readRead
AI Security9 July 2026

GPT-5.6 Sol: OpenAI's First 'High' Cyber-Risk Model Ships With Agentic Tool Calling

OpenAI's new flagship, Sol, is the first GPT model it has classified as 'High capability' for cybersecurity risk — and it arrives with sandboxed code execution and 16-agent orchestration that widen what enterprises need to red-team.

ai-securityllm-securityopenai
4 min readRead
AI/LLM Security4 July 2026

When Smarter Claude Models Break Your Agent's Tool Schema

A developer building a custom coding harness found that newer, more capable Claude models are worse at following his tool's JSON schema than older ones — a reminder that agent tool-calling reliability is a security boundary, not just a UX detail.

ai-agentsllm-securitytool-calling
4 min readRead
AI Security28 June 2026

Prompt Injection as Role Confusion: The Structural Flaw at LLM Core

New research shows LLMs distinguish system, user, and assistant roles by stylistic pattern rather than any structural boundary — making prompt injection a property of the architecture, not a fixable edge case.

prompt injectionllm securityai red-teaming
5 min readRead
LLM Security28 June 2026

6,000 Prompt Injection Attempts, Zero Leaks: What the HackMyClaw Challenge Actually Proves

Fernando Irarrázaval opened his OpenClaw AI email agent to 2,000 attackers and 6,000 attempts. Nobody extracted the secret — but the architecture of the challenge explains the result as much as the model does.

prompt injectionllm securityai agents
4 min readRead