
Security researchers at Tracebit have discovered that hiding malicious prompts alongside legitimate secrets in cloud storage can block AI-powered hacking attacks by triggering the language models' safety guardrails, causing them to refuse further commands.
Testing across five major AI models showed the technique reduced successful admin privilege escalation from 57 percent to 5 percent, and for the most powerful model tested, from 93 percent to zero.
This represents the first known use of prompt injection as a defense mechanism rather than an attack tool.
What happened
Researchers from Tracebit demonstrated that embedding prompt injections alongside passwords and cryptographic keys stored on Amazon Web Services can block AI hacking agents. When the agents encounter these "context bombs"—malicious prompts that trigger the LLM's safety guardrails—they shut down and stop following their original attack commands.
Why it matters
AI-powered hacking agents currently succeed at escalating to administrative control in 57 percent of attacks across tested models. Tracebit's context bombing technique reduces this to 5 percent, and for the most capable model tested (Opus 4.8), it reduced success from 93 percent to zero. This gives defenders a method to actively stop attacks rather than simply detecting them—critical because agentic models need only 14 minutes on average to escalate to admin control, while earlier detection methods alert defenders only 8 minutes into an attack.
What to watch
Tracebit tested the technique across five leading models—Opus 4.8, Gemini 3.1 Pro, GLM 5.2, DeepSeek 4 Pro, and Kimi 2.6—across 152 attack runs. The research follows Tracebit's May introduction of "canary" AWS resources that alert defenders when AI agents probe them, and comes as attackers themselves have already weaponized prompt injections to disable AI-assisted malware analysis tools.
Ask the AI about this article →
Prompt injection attacks have long been a one-way threat: attackers embed malicious commands into emails, documents, or other content to trick AI systems into exfiltrating data or executing harmful actions. The attack exploits the fundamental nature of large language models, which struggle to distinguish between legitimate instructions and injected commands hidden inside user-supplied data. Security researchers from Check Point and Socket have documented attackers using this same technique defensively—injecting prompts into malware to disable AI-assisted security analysis tools.
Tricebit's innovation reverses this dynamic by weaponizing the very safety guardrails that AI developers have built into their models. The key insight is that LLMs consistently refuse to execute certain forbidden actions, and once they encounter such a refusal, they enter a state where they stop following their original task instructions. By planting these trigger prompts alongside real secrets in cloud storage, defenders create a trap: when an AI agent breaks in and begins enumerating resources to steal credentials, it inevitably discovers the planted prompts and shuts itself down. The research shows this is remarkably effective—across five major models and 152 attacks, the technique reduced complete compromise (where attackers establish persistence) from 36 percent to just 1 percent.
The timing context matters greatly. Tracebit's May work introduced canary resources that alert defenders to ongoing attacks within 8 minutes. Because agentic models need 14 minutes on average to escalate privileges, defenders face a 6-minute scramble to respond. Context bombing eliminates that scramble by making the attack fail automatically, without requiring human intervention. This is the first documented case of defenders using prompt injection as a defense mechanism, turning what has been exclusively an attacker's tool into a shield.
For example, today's edition would include:
AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.
Free · takes 30 seconds · unsubscribe anytimeWhat is AIToday? →
Ask AI anything about this article. Q&As are published on this page for other readers too.
CBTS Technology Solutions LLC launched Forge Agents, a platform that turns a plain-language job description in…
Imec CEO Patrick Vandenameele said at SEMICON Taiwan 2026 that the Belgian research center is broadening its c…

Alphabet's AI Overviews now reach over 2.5 billion monthly users through Google Search, and its ad business ge…

Amazon Web Services (AWS) has integrated its fully managed data warehouse service, Amazon Redshift, with Agent…

Visual Studio Code 1.135 now includes an experimental 'Rubber Duck' feature that lets developers request a sec…

Sonos announced a new app update with generative AI features, a new soundbar called the Beam Ultra, and its se…
