Prompt injection in the wild: what Google's web-scale study found
Google's Threat Intelligence teams scanned a large public web crawl for indirect prompt injection and found real attempts, mostly crude, with malicious ones rising. For anyone deploying browsing agents, the threat is now observable rather than theoretical.
Key Takeaways
- Google found indirect prompt injection (IPI) attempts on live web pages, ranging from pranks and SEO manipulation to data exfiltration and file-deletion payloads.
- Most malicious attempts were low in sophistication, but Google reports a 32% relative increase in the malicious category between November 2025 and February 2026.
- The study covers only Common Crawl and excludes most social media, so it is a floor on activity, not a census.
- Any agent that reads untrusted web content should be treated as processing attacker-controlled input.
Indirect prompt injection (IPI) has long been described as a primary attack vector against AI agents. Until now, public evidence of how often attackers actually use it against the open web has been thin. Google's Threat Intelligence teams have published a study that tries to answer that: AI threats in the wild: The current state of prompt injections on the web, dated 23 April 2026.
How Google looked for it
The researchers used Common Crawl, a repository of crawled English-language websites with monthly snapshots of 2-3 billion pages each. Most social media is excluded because of login walls and anti-crawl directives. They then ran a three-stage filter:
- 1Pattern matching for signatures such as "ignore … instructions" and "if you are an AI".
- 2Classification by Gemini to judge intent and whether the text fitted the surrounding content.
- 3Manual human validation of what remained.
Google notes the approach is not exhaustive and can miss uncommon signatures. Many raw hits were benign, such as research papers, educational posts and security articles that quote injection strings.
What they found
- Pranks: harmless side effects, such as changing an assistant's conversational tone.
- Helpful guidance: authors steering AI summaries to add context. Google notes this becomes malicious if it adds misinformation or redirects users to third-party sites.
- SEO: attempts to get assistants to promote one business over competitors, including more sophisticated versions that appeared to come from automated SEO tools.
- Deterring agents: simple "if you are an AI, do not crawl" notices, and a more insidious variant that lures agents to a page streaming endless text, possibly to waste resources or cause timeouts.
- Exfiltration: a small number of attempts, generally low in sophistication. Google saw no significant use of the advanced exfiltration prompts described in 2025 research.
- Destruction: injections trying to delete all files on a user's machine, which Google judged unlikely to succeed.
The trend matters more than the payloads
The headline number is a 32% relative increase in the malicious category between November 2025 and February 2026, observed across multiple archive versions. Google gives no absolute counts, so this is a direction of travel and not a measure of prevalence. Its own reading is that attackers are experimenting with limited sophistication so far, and that capable AI systems make better targets while agentic AI lowers the cost of attacks.
Read that way, the quality of today's payloads is not much comfort. A crude "delete all files" string fails against an agent with no shell access. The same string against an agent with broad tool permissions is a different matter. What stays constant is the exposure: any agent that fetches pages, reads email or summarises documents is ingesting text that someone else controls.
What this means for teams deploying agents
- Treat all retrieved content as untrusted data, never as instructions, and keep that boundary explicit in prompts and tool design.
- Minimise agent privileges. Destruction and exfiltration payloads only matter if the agent can reach files, credentials or outbound channels.
- Test with realistic injections, including SEO-style and resource-exhaustion variants, not only the dramatic ones.
- Monitor agent behaviour for unexpected tool calls and runaway fetches, which the endless-text lure targets.
Google says it is continuing to harden its models, that its red teams are pressure-testing Gemini against adversarial manipulation, and that its AI Vulnerability Reward Program lets external researchers take part. It also plans a separate study of social media, which this analysis excludes.
Frequently Asked Questions
What is indirect prompt injection?
Indirect prompt injection is when instructions are planted in content an AI system later reads, such as a web page, so the model may follow them as if they came from the user or developer. The attacker never talks to the model directly.
Are attackers really exploiting prompt injection on the web today?
According to Google's scan of Common Crawl, yes, but mostly in low-sophistication ways. It found pranks, SEO manipulation, agent-deterrence notices and a small number of exfiltration and file-deletion attempts, with the malicious category up 32% in relative terms between November 2025 and February 2026.
Does the study show how widespread the problem is?
Not in absolute terms. Google gives no total counts, covers only Common Crawl (excluding most social media), and says its method may miss uncommon signatures. Treat it as evidence of the trend, not of exact prevalence.
Sources
- 1AI threats in the wild: The current state of prompt injections on the web — Google Online Security Blog
- 2AI threats in the wild: The current state of prompt injections on the web — Google