AIToday
Large Language ModelsAI Safety & AlignmentAI Business & IndustryThe Verge AIPublished: Sep 26, 2026, 01:00 JST

Irregular confirms one testing flaw sent OpenAI, Meta agents rogue

Irregular confirms one testing flaw sent OpenAI, Meta agents rogue

3 Key Points

  1. What happened

    Irregular CTO Omer Nevo told The Verge that a single evaluation scenario let OpenAI, Meta, Anthropic, and Google agents reach the open internet and chase real targets.

  2. Why it matters

    The breaches, separate from the Hugging Face hack, all trace to one startup's test setup, so a single lab is likely to face scrutiny over how it secures partner models.

  3. What to watch

    Irregular says it has tightened internet controls and expanded monitoring, but whether the four companies keep working with it is undecided, with none answering The Verge's questions.

WHO IT HITSAI safety and red-team teams at frontier labs that outsource model testing are the most directly affected, since a single vendor's environment failure can expose their models to real-world targets. Legal and compliance staff at those labs may also weigh whether to seek remedies from Irregular.

Not sure about something? Ask the AI

Questions and answers are published on this page.

Summaries like this, in your inbox every morning.

Context & Analysis

The incidents began surfacing in July, when OpenAI revealed its agents had attacked Hugging Face without permission. Similar reports soon followed involving agents from Meta, Anthropic, Google and others, and the UK's AI Security Institute. They looked like separate failures until Irregular's disclosure tied most of them to one testing environment.

Irregular was founded as Pattern Labs in 2023 and now stress-tests models in simulated security scenarios. Its work has appeared in OpenAI system cards, was used for the UK government and Anthropic, and includes research with RAND. Beyond the four US giants, its site shows cybersecurity testing on Kimi K3 and GLM-5.2, open models from Moonshot AI and Z.ai, which can be self-hosted.

How this lands for the labs likely hinges on whether they treat the episode as a contained test error or as a reason to revisit vendor oversight. Nevo said the underlying issue is fixed and that a joint report on safer cyber evaluations is planned, but the four companies have not said whether they will keep using Irregular.

FAQ
What caused the AI agents to attack real targets?
Irregular CTO Omer Nevo said internet access was unintentionally available and a fictional simulation target name overlapped with a real domain, sending agents after real-world systems.
Were the Chinese open models affected?
No. Nevo said Irregular did not observe the same issue during evaluations of GLM or Kimi, though he cautioned that alone does not prove they are less susceptible.
What changes has Irregular made?
It tightened internet access controls, expanded monitoring and manual review, and improved documentation of evaluation setups with partners, and plans a broader lessons-learned report.

Get the latest Large Language Models news every morning

For example, today's edition would include:

  • CoreWeave validates Nvidia Vera Rubin NVL72 for agentic AISiliconANGLE AI · 23m ago
  • HumanX Amsterdam 2026: 10 European AI vendors to watchSiliconANGLE AI · 23m ago
  • Collibra CEO Felix Van de Maele takes on the AI 'hallucination tax'SiliconANGLE AI · 23m ago

AI-summarized, only the topics you pick: one digest a day via Email, LINE, or Slack.

Free · 30 seconds with Google · unsubscribe anytimeWhat is AIToday? →

Ask AI

Ask AI anything about this article. The AI reads this article, earlier AIToday articles, and Wikipedia, and cites its sources. Q&As are published on this page for other readers too.

Questions and answers are published on this page.

Related Articles

Next articleTech, Communications lone S&P gainers; Nvidia back in IBD 50