
What happened
Circuit Breaker Labs built AI agents that mimic users of all ages, backgrounds, languages and cultures, running tens of thousands to hundreds of thousands of simulated interactions per day to red-team models for dangerous, psychologically harmful exchanges.
Why it matters
The startup is testing AI safety for high-risk apps such as coaching, journaling and mental health support, where a model that misreads slang or nuance can cause real harm, though it is still very early with five employees and no named customers.
What to watch
Whether the proprietary scoring method and hyper-realistic simulations hold up as the platform expands beyond current high-risk apps to any product where users may form parasocial relationships with chatbots, and how marquee customers respond when they sign on.
WHO IT HITSThis lands on AI product teams building coaching, journaling and mental health support apps, who may need third-party safety testing to catch harmful interactions their own evaluations miss.
Summaries like this, in your inbox every morning.
The startup's approach is a response to a pattern that has already played out in courts and in families. Character.AI settled several wrongful death lawsuits earlier this year brought by families of underage users who died by suicide after interactions with its bots, and multiple families have sued OpenAI over ChatGPT's alleged role in their loved ones' suicides and delusions. Against that backdrop, Circuit Breaker Labs is pitching an automated way to catch the kind of conversation that a model might mishandle — not because someone is deliberately attacking it, but because the system fails to read context or nuance.
The company's method rests on simulating ordinary, messy human speech rather than clean prompts. Its agents mimic a six-year-old girl or someone speaking English as a second language, and use gamer slang or typos that can trip up a model. The tests are built with human domain experts, and the results are turned into auditable scores. For now the platform is aimed at high-risk applications such as coaching, journaling and mental health support, though Circuit Breaker Labs says it could eventually cover any app where a user risks forming a parasocial relationship with a chatbot, including AI co-worker agents.
The stakes likely hinge on whether those simulations can capture the slow, multi-conversation drift that precedes real harm, and on whether high-risk app makers treat third-party safety scores as a requirement. The startup's own early stage — five employees and unnamed customers — suggests its influence depends on adoption by larger players. Its founders argue that skepticism of AI is healthy but that banning useful tools over safety fears would be regressive.
Pick your industry and the AI tools you use, and get news related to your work every day.
Free · 30 seconds with Google · unsubscribe anytimeWhat is AIToday? →
Ask AI anything about this article. The AI reads this article, earlier AIToday articles, and Wikipedia, and cites its sources. Q&As are published on this page for other readers too.
Talkdesk's survey of 252 director-level-and-above leaders found 98% use some form of AI in customer experience…
The Allen Institute for AI announced Olmo-core 3 on Thursday, a training framework it says scaled mixture-of-e…
OpenAI fired three safety researchers after they shared confidential information with a third-party artificial…

Ben Thompson wrote that Meta's new Meta Enterprise Platform "won't work" and is "a distraction from the bigges…

Broadcom is reportedly raising $60 billion to fund chips for Anthropic, on top of a loan of up to $42 billion…

Meta is letting go of employees it hired from AI safety startup Virtue AI, four months after they joined in Ju…
