
Anthropic and OpenAI have put guardrails on their AI models to block malicious use, but the restrictions are also preventing legitimate cybersecurity researchers from using those same tools to find and fix vulnerabilities before attackers exploit them. Security experts say the guardrails are inconsistent and block essential defensive work, forcing researchers toward unvetted open-source alternatives instead of regulated U.S. systems.
Summaries like this, in your inbox every morning.
Sign up free →What happened
AI companies including Anthropic and OpenAI have imposed strict guardrails on their models to prevent malicious use, but these safeguards are now preventing legitimate cybersecurity researchers—who probe systems to find vulnerabilities before criminals do—from using the same tools effectively. In June, the U.S. government imposed export control restrictions on Anthropic's Mythos and Fable models over jailbreak concerns, though those controls have since been partially lifted.
Why it matters
Offensive security researchers say the guardrails create an impossible paradox: the same prompts needed to test whether a bug is exploitable (defensive work) are blocked by AI companies as potentially dangerous. Chris Anley of NCC Group compared it to a hammer—"You can't build a house without a hammer. It's definitely a tool but it's also irreducibly a weapon as well." The blanket restrictions treat security experts like threats, forcing them toward less regulated alternatives.
What to watch
Researchers are increasingly turning to unvetted open-source Chinese models like GLM that run locally with no restrictions, according to Chris Thompson of RemoteThreat. Thompson warned that defenders risk losing the "AI race" if guardrails push legitimate researchers away from U.S.-governed systems, and called for AI labs to "open up their programs, provide responsible access, and hold those who abuse their tools accountable" instead of tightening restrictions further.
For months, major AI companies have imposed guardrails to prevent their models from being used in malicious cyberattacks. In June, the U.S. government went further, slapping export control restrictions on Anthropic's Mythos and Fable models following reports that users could bypass the guardrails designed to prevent malicious use. (The export controls have since been partially lifted: Fable 5 returned to general access on July 1, and Mythos 5 has been reintroduced only to vetted U.S. organizations.)
Anthropologic and OpenAI both offer programs to help: Anthropic's Cyber Verification Program and OpenAI's Trusted Access for Cyber program, which let approved security researchers access models with fewer restrictions. But these vetted pathways have sparked criticism from the very researchers they aim to help.
The core complaint is that guardrails block legitimate defensive work. Chris Anley, chief scientist at NCC Group, a major security consulting firm, explained that asking an AI to help exploit a bug is essential to confirming it is a real vulnerability worth fixing. But when guardrails refuse to answer such requests, "the guardrail hurts defenders." He added: "'Fix this code' as a prompt is both an essential mechanism for defense but also a roadmap for finding critical vulnerabilities in the code base. So at the same time, the same tool is both an offensive tool and a defensive tool, and the two can't really be unpicked." His solution: "like a hammer, you can't build a house without a hammer. It's definitely a tool but it's also irreducibly a weapon as well."
When researchers hit such roadblocks, they often turn to open-source AI models with no guardrails. Mark Dowd, a well-known security researcher who finds and sells zero-day vulnerabilities to governments rather than reporting them publicly, said he is uncomfortable with companies making "arbitrary decisions about what is safe in security and what's not." Paolo Stagno, CTO at Crowdfense (which develops and sells unknown vulnerabilities to government agencies), agreed, saying AI companies "essentially treat customers like children who need babysitting" with their programs. Stagno said he avoids feeding offensive vulnerability work into cloud-based models for fear of leaking sensitive data, so he uses locally-run open-source models instead.
Other researchers take different approaches. Giuseppe Cali, who finds zero-days and develops exploits, said guardrails do not impede him because he uses AI only for reverse engineering and building supporting tools, not for offensive work itself. "I still want to own the actual bug discovery and weaponization myself and that wouldn't change if all guardrails were lifted tomorrow," he said. "I am jealous of my bugs, and I like this game too much to let models play it for me." One anonymous researcher at a smartphone-component manufacturer said his employer is not part of Anthropic's program, leaving the guardrails too strict to be useful for finding vulnerabilities.
Chris Thompson, CEO of RemoteThreat and founder of Offensive AI Con, described guardrails as inconsistent and changing from day to day, even within the looser boundaries of Anthropic's and OpenAI's vetted programs. "I think the practical impact is you spend a lot of time negotiating with the model instead of working on the core security program," he said. As a result, researchers are being pushed toward Chinese open-source models like GLM, which are freely downloadable and have no vetting or usage restrictions. "You have these responsible researchers that are being pushed away from U.S.-governed systems to foreign-owned systems," Thompson warned. "I think it's more harmful than good to have these guardrails in place." Rather than tightening restrictions, he called for AI companies to open their programs, provide responsible access, and hold abusers accountable, warning that otherwise defenders will lose the AI race: "There's this big storm coming. There's this big wave of attacks that are going to happen at speed and scale like never before. But the same security consulting firms and legit researchers that are trying to make a difference are being stifled right now."
The tension between safety and utility in AI guardrails has created an unintended consequence: the researchers best positioned to defend critical systems are being blocked from using the most capable tools. The U.S. government's June export control actions on Anthropic's models, prompted by jailbreak concerns, underscored how seriously AI companies now treat security restrictions. Yet as multiple researchers told TechCrunch, the guardrails do not distinguish between offensive attackers and defensive researchers—both are treated as potential threats.
Mark Dowd, who has spent decades finding zero-day vulnerabilities for Western governments, acknowledged his bias but captured the core complaint: "it's not really comfortable to me that these random large companies are making arbitrary decisions about what is safe in security and what's not." Chris Anley's hammer analogy highlights the impossibility of the task: a tool cannot be simultaneously open for defense and closed for offense. When defensive work is blocked, researchers migrate to less regulated alternatives—including foreign open-source models—which may actually weaken U.S. cybersecurity overall.
AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.
Free · takes 30 seconds · unsubscribe anytime
No comments yet. Be the first to share your thoughts!
Log in to join the discussion


Get curated AI news from 200+ sources delivered daily to your inbox. Free to use.
Get Started FreeFree · takes 30 seconds · unsubscribe anytime
1 minute a day. The AI essentials.
200+ sources · Email / LINE / Slack