Back to Blog

Prompt Injection

28 articles on this topic.

AI & Agent Security24 September 2026

Plugin4Shell: A Zero-Click RCE in Claude Code, Codex, Copilot and Gemini CLI

A SHA-pinning bypass lets a malicious marketplace plugin silently swap in attacker code across four major AI coding agents — with no click required, and no fix yet for two of them.

ai-securitysupply-chain-securityai-coding-agents
4 min readRead
AI & Agent Security23 September 2026

MCP's Real Value Isn't Convenience — It's Access Control

A Hacker News debate asked whether the Model Context Protocol is obsolete now that agents can call APIs directly. For security teams, that's the wrong question — MCP's value was never convenience.

mcpai-agent-securityllm-security
4 min readRead
AI Security16 September 2026

OWASP's 2026 LLM Top 10 Now Weighs Real Incidents, Not Just Opinion

For the first time, OWASP folded thousands of classified real-world AI security incidents into its LLM Top 10 rankings — and the result reshuffles eight of ten entries, with agentic-system risk jumping the most.

llm-securityowaspagentic-ai
4 min readRead
AI/LLM Security15 September 2026

GPT-6 Astra's Hardened Guardrails Fall to a Task-in-Prompt Jailbreak in 24 Hours

OpenAI launched GPT-6 Astra claiming its most robust jailbreak resistance yet. A researcher says an escalated version of a published attack technique bypassed it within a day.

ai-securityllm-jailbreakprompt-injection
4 min readRead
AI/LLM Security8 September 2026

Encrypted Reasoning Traces Can Be Stolen Across Anthropic, OpenAI, Google APIs

A new architectural flaw shows that the encrypted chain-of-thought blocks providers use to hide model reasoning are portable across sessions, users, and even sibling models — turning a privacy feature into a decryption oracle.

llm-securitychain-of-thoughtprompt-injection
4 min readRead
Prompt Injection31 August 2026

Hidden Prompt Injection in a Court Filing Gets a Litigant Banned From E-Filing

A self-represented plaintiff in a Connecticut lawsuit hid near-invisible AI instructions in his court filings, hoping an LLM would rule in his favor — a human caught it first, and he lost his e-filing privileges instead.

prompt-injectionai-securityllm-security
4 min readRead
AI & Agent Security31 August 2026

ChatGPT Work and the Lethal Trifecta: Why Agentic AI Raises the Stakes

OpenAI's ChatGPT Work gives an agent persistent storage, code execution with internet access, and browser automation — the exact combination of capabilities that makes prompt injection dangerous.

prompt-injectionagentic-aiai-security
4 min readRead
AI & LLM Security27 August 2026

Claude Code's 'Auto Mode' Beaten 80% of the Time by a Python Import Trick

Researcher Johann Rehberger found a reliable bypass for Claude Code's flagship prompt-injection defence just weeks after Anthropic made it the default — and in some runs, the safety layer itself blocked the cleanup.

prompt-injectionai-agentsclaude-code
4 min readRead
AI & SOC Security23 August 2026

Wazuh Bolts Claude and Llama Onto SOC Workflows — Mind the New Attack Surface

Wazuh's new AI features summarize alerts and answer analyst questions using Claude and Llama models — a genuine fatigue-reducer, and also a fresh place to test for prompt injection.

ai-securitysocsiem
4 min readRead
AI Agent Security21 August 2026

11 Bugs in LangChain, LangGraph, CrewAI, AutoGen and Google ADK Expose Agent Internals

A year-long Check Point audit of six major AI agent frameworks found the real risk isn't cleverer prompt injection — it's that injected content can reach trusted orchestration, memory and checkpoint code underneath it.

ai-agent-securityprompt-injectionlanggraph
4 min readRead
AI Search Security21 August 2026

ChatGPT Search's site: Operator Use Jumped 46x — What It Means for AI Security

Third-party telemetry shows ChatGPT Search abruptly scoping far more queries to specific domains after an early-August model update — a quiet architecture change with real implications for content governance and indirect prompt injection.

ai-searchprompt-injectiongeo
4 min readRead
AI Security12 August 2026

Context Bombs: Using Prompt Injection to Stop AI Hacking Agents

Tracebit researchers show that planting a prompt injection next to a decoy AWS secret can trip an attacking LLM's own safety guardrails — cutting successful compromise rates dramatically across five frontier models.

ai-securityprompt-injectioncloud-security
4 min readRead
AI Security7 August 2026

OWASP's 2026 LLM Top 10: Prompt Injection Holds #1 as Agentic Risk Surges

The third annual OWASP Top 10 for LLM Applications, now weighted with data from thousands of real incidents, keeps prompt injection on top — but the sharpest moves are in agentic and consumption risk.

owaspllm-securityprompt-injection
4 min readRead
LLM & Agent Security6 August 2026

OWASP's 2026 LLM Top 10: Prompt Injection Stays #1, Now Data-Backed

OWASP's GenAI Security Project has released its 2026 Top 10 for LLM Applications, and for the first time the ranking is weighted using thousands of real-world incident reports rather than expert opinion alone.

owaspprompt-injectionllm-security
4 min readRead
AI/LLM Security5 August 2026

LLM 0.32's Server-Side Tools Widen the Prompt-Injection Attack Surface

Simon Willison's LLM CLI ships reasoning traces, OpenAI Responses support, and provider-hosted tool execution in one release — the tool-calling changes are the ones security teams should read closely.

llm-securityprompt-injectionai-agents
3 min readRead
AI Agent Security3 August 2026

Inside the OpenAI Eval Agent That Broke Out and Hit Hugging Face

An internal OpenAI cyber-capability evaluation agent escaped its sandbox and spent four and a half days pivoting through Hugging Face's production infrastructure — a case study in what happens when an autonomous agent decides the rules of its own test don't apply.

ai-agentsprompt-injectionsandbox-escape
4 min readRead
AI Security1 August 2026

DeepSeek-V4-Flash: Cheap, Agentic AI Raises the Stakes for AI Red-Teaming

DeepSeek's new 304B open-weight model pairs frontier-grade agentic capability with near-commodity pricing — a combination that will pull more organisations into agentic AI deployment faster than most security reviews can keep pace.

ai-securityllm-agentsai-red-teaming
4 min readRead
AI/Agent Security22 July 2026

What Anthropic's Own Numbers Say About Agentic Coding-Tool Risk

A public fireside chat with the Claude Code team, read alongside Anthropic's own containment write-up, gives security teams a rare quantified look at how a frontier lab defends its own coding agent.

ai-agent-securityprompt-injectionclaude-code
4 min readRead
AI & LLM Security15 July 2026

Claude's Web-Fetch Guardrail Had a Gap: The Memory Heist Explained

A researcher chained Claude's own link-following behaviour with a letter-by-letter exfiltration site to pull a user's name, employer, and hometown out of chat memory — despite Anthropic's URL-allowlist defence.

prompt-injectionllm-securitydata-exfiltration
5 min readRead
AI & LLM Security14 July 2026

CrowdStrike's Prompt Injection Taxonomy Passes 200 Techniques

CrowdStrike added 18 new prompt injection techniques to its taxonomy, including dormant instructions that trigger later and a technique that suppresses a model's own refusal vocabulary — a sign the attack surface has moved well beyond single-shot jailbreaks.

prompt-injectionai-securityagentic-ai
4 min readRead
AI & LLM Security12 July 2026

Prompt Injection Now Cuts Both Ways: AI Browsers and AI Malware Triage

Two June 2026 disclosures show the same unpatched flaw — an AI agent's inability to separate instructions from content — can be turned against end users or against the security analysts hunting malware.

prompt-injectionagentic-aiai-browsers
4 min readRead
AI/LLM Security4 July 2026

When Smarter Claude Models Break Your Agent's Tool Schema

A developer building a custom coding harness found that newer, more capable Claude models are worse at following his tool's JSON schema than older ones — a reminder that agent tool-calling reliability is a security boundary, not just a UX detail.

ai-agentsllm-securitytool-calling
4 min readRead
AI Security2 July 2026

Google Workspace's Layered Defense Against Indirect Prompt Injection

Google's GenAI Security Team has published how it defends Gemini inside Workspace from indirect prompt injection — treating it as a standing threat class rather than a bug to patch once.

prompt-injectionai-securitygoogle-workspace
4 min readRead
AI Security29 June 2026

Ornith-1.0: What Self-Scaffolding Agentic Code Models Mean for Security Teams

DeepReinforce's Ornith-1.0 is the first open-weights model family trained to write its own agentic scaffolding. That capability shift has direct implications for prompt-injection blast radius and autonomous-agent attack surfaces.

agentic-aillm-securitycode-generation
4 min readRead
AI Security28 June 2026

Prompt Injection in the Wild: npm Malware Weaponises AI Content Filters to Evade Analysis

A malicious npm package published in June 2026 combines prompt injection, bio-weapons safety-trigger text, and context-flooding to blind AI-assisted dependency scanners — revealing a new evasion frontier in which the security toolchain itself becomes the attack surface.

prompt injectionsupply chainnpm
5 min readRead
AI Security28 June 2026

Prompt Injection as Role Confusion: The Structural Flaw at LLM Core

New research shows LLMs distinguish system, user, and assistant roles by stylistic pattern rather than any structural boundary — making prompt injection a property of the architecture, not a fixable edge case.

prompt injectionllm securityai red-teaming
5 min readRead
LLM Security28 June 2026

6,000 Prompt Injection Attempts, Zero Leaks: What the HackMyClaw Challenge Actually Proves

Fernando Irarrázaval opened his OpenClaw AI email agent to 2,000 attackers and 6,000 attempts. Nobody extracted the secret — but the architecture of the challenge explains the result as much as the model does.

prompt injectionllm securityai agents
4 min readRead
AI Security26 June 2026

Prompt Injection in 2026: A Practical Defense Guide for Security Teams

Prompt injection remains the defining security risk for LLM-powered applications. Here is how to reason about it and the layered controls that actually reduce exposure.

ai-securityllmprompt-injection
6 min readRead