
What happened
AI companies including Anthropic and OpenAI have imposed strict guardrails on their models to prevent malicious use, but these safeguards are now preventing legitimate cybersecurity researchers—who probe systems to find vulnerabilities before criminals do—from using the same tools effectively. In June, the U.S. government imposed export control restrictions on Anthropic's Mythos and Fable models over jailbreak concerns, though those controls have since been partially lifted.
Why it matters
Offensive security researchers say the guardrails create an impossible paradox: the same prompts needed to test whether a bug is exploitable (defensive work) are blocked by AI companies as potentially dangerous. Chris Anley of NCC Group compared it to a hammer—"You can't build a house without a hammer. It's definitely a tool but it's also irreducibly a weapon as well." The blanket restrictions treat security experts like threats, forcing them toward less regulated alternatives.
What to watch
Researchers are increasingly turning to unvetted open-source Chinese models like GLM that run locally with no restrictions, according to Chris Thompson of RemoteThreat. Thompson warned that defenders risk losing the "AI race" if guardrails push legitimate researchers away from U.S.-governed systems, and called for AI labs to "open up their programs, provide responsible access, and hold those who abuse their tools accountable" instead of tightening restrictions further.
Ask the AI about this article →
Summaries like this, in your inbox every morning.
The tension between safety and utility in AI guardrails has created an unintended consequence: the researchers best positioned to defend critical systems are being blocked from using the most capable tools. The U.S. government's June export control actions on Anthropic's models, prompted by jailbreak concerns, underscored how seriously AI companies now treat security restrictions. Yet as multiple researchers told TechCrunch, the guardrails do not distinguish between offensive attackers and defensive researchers—both are treated as potential threats.
Mark Dowd, who has spent decades finding zero-day vulnerabilities for Western governments, acknowledged his bias but captured the core complaint: "it's not really comfortable to me that these random large companies are making arbitrary decisions about what is safe in security and what's not." Chris Anley's hammer analogy highlights the impossibility of the task: a tool cannot be simultaneously open for defense and closed for offense. When defensive work is blocked, researchers migrate to less regulated alternatives—including foreign open-source models—which may actually weaken U.S. cybersecurity overall.
For example, today's edition would include:
AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.
Free · 30 seconds with Google · unsubscribe anytimeWhat is AIToday? →
Ask AI anything about this article. Q&As are published on this page for other readers too.
Meta announced on Sept. 15 the global rollout of Meta One, a subscription combining premium features across it…

Anthropic's second-quarter revenue was above $11.5 billion, up from $4.73 billion in the first quarter and $78…

ECB President Christine Lagarde said eurozone households hold around €440 billion in US tech companies includi…

Capgemini CEO Aiman Ezzat told Fortune that leaders should run small AI tests and pilots, not big-bang bets, b…

Anthropic said on Sept. 16 it will open a Singapore office in October, its first in Southeast Asia and fifth A…

SAP CEO Christian Klein says voice recognition in large language models is strong, and SAP predicts typing dat…
