AIToday
AI Business & IndustryTechCrunch AIPublished: Jul 24, 2026, 13:00 JST

AI guardrails blocking legitimate security researchers from defending networks

AI guardrails blocking legitimate security researchers from defending networks

3 Key Points

  1. What happened

    AI companies including Anthropic and OpenAI have imposed strict guardrails on their models to prevent malicious use, but these safeguards are now preventing legitimate cybersecurity researchers—who probe systems to find vulnerabilities before criminals do—from using the same tools effectively. In June, the U.S. government imposed export control restrictions on Anthropic's Mythos and Fable models over jailbreak concerns, though those controls have since been partially lifted.

  2. Why it matters

    Offensive security researchers say the guardrails create an impossible paradox: the same prompts needed to test whether a bug is exploitable (defensive work) are blocked by AI companies as potentially dangerous. Chris Anley of NCC Group compared it to a hammer—"You can't build a house without a hammer. It's definitely a tool but it's also irreducibly a weapon as well." The blanket restrictions treat security experts like threats, forcing them toward less regulated alternatives.

  3. What to watch

    Researchers are increasingly turning to unvetted open-source Chinese models like GLM that run locally with no restrictions, according to Chris Thompson of RemoteThreat. Thompson warned that defenders risk losing the "AI race" if guardrails push legitimate researchers away from U.S.-governed systems, and called for AI labs to "open up their programs, provide responsible access, and hold those who abuse their tools accountable" instead of tightening restrictions further.

Ask the AI about this article →

Summaries like this, in your inbox every morning.

Context & Analysis

The tension between safety and utility in AI guardrails has created an unintended consequence: the researchers best positioned to defend critical systems are being blocked from using the most capable tools. The U.S. government's June export control actions on Anthropic's models, prompted by jailbreak concerns, underscored how seriously AI companies now treat security restrictions. Yet as multiple researchers told TechCrunch, the guardrails do not distinguish between offensive attackers and defensive researchers—both are treated as potential threats.

Mark Dowd, who has spent decades finding zero-day vulnerabilities for Western governments, acknowledged his bias but captured the core complaint: "it's not really comfortable to me that these random large companies are making arbitrary decisions about what is safe in security and what's not." Chris Anley's hammer analogy highlights the impossibility of the task: a tool cannot be simultaneously open for defense and closed for offense. When defensive work is blocked, researchers migrate to less regulated alternatives—including foreign open-source models—which may actually weaken U.S. cybersecurity overall.

FAQ
What programs do AI companies offer to help security researchers?
OpenAI offers the Trusted Access for Cyber program and Anthropic offers the Cyber Verification Program. Both let vetted researchers apply for access to models with fewer cybersecurity restrictions than the general public version.
Why are researchers frustrated with the guardrails?
Asking an AI to help exploit a bug is a key step in confirming a vulnerability is real and needs fixing, but guardrails often block those requests outright. Chris Thompson noted the guardrails are inconsistent and work differently day to day, forcing researchers to spend time "negotiating with the model instead of working on the core security program."
What are researchers doing instead?
Several researchers said they fall back on open-source AI models with no guardrails, or use unvetted Chinese models like GLM that can be run locally. Others avoid using AI for offensive work entirely and rely on their own expertise for vulnerability discovery.

Get the latest AI Business & Industry news every morning

For example, today's edition would include:

  • Meta rolls out Meta One subscription globally, adds TaiwanDIGITIMES Asia · 1h ago
  • Anthropic Q2 revenue tops $11.5 billion ahead of IPOYahoo Finance AI · 1h ago
  • Lagarde warns €440 billion EU savings funding US AI boomYahoo Finance AI · 1h ago

AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.

Free · 30 seconds with Google · unsubscribe anytimeWhat is AIToday? →

Ask AI

Ask AI anything about this article. Q&As are published on this page for other readers too.

Related Articles

Next articleNvidia 800V power tech slips past 2027, stoking supply-chain concerns