
What happened
The author launched AI Security Lab, an experimental setup that runs LLMs on a local environment and calls them from Python to observe failures, starting with questions like "Can AI judge correctly?"
Why it matters
The author argues the useful question is not whether AI can be trusted, but what work it can be given, under what conditions, and where the boundary with humans should sit.
What to watch
The next post will detail the experiment environment, the models used, and the experiment programs, with the code published on GitHub; the Lab's stance is that whether AI should be trusted as a single yes-or-no is too coarse a frame.
WHO IT HITSDevelopers and security engineers building AI agents and tool-using applications are the immediate audience, since the Lab focuses on where to place validation, authorization, and recovery controls. Teams deciding which tasks to hand to AI may find the experiment-and-analyze approach relevant to their own delegation choices.
Summaries like this, in your inbox every morning.
The series opens from a shift the author describes: AI is moving from something that answers questions to something that receives work from humans and executes it. In the older pattern, a person read the AI's text and decided what to do, so an error did not directly change a system's state. Once tools are attached, the AI can read files, run commands, search the web, call APIs, and change data, and an AI Agent can loop through judgment, tool selection, execution, and re-judgment.
The author's framing is that whether AI can be trusted is too large a question, because humans also make mistakes and no one removes people from work on that basis. What matters instead is what work is assigned, under what conditions, and to whom. This is why the Lab separates two problems: how well AI can judge, and how the Application handles that judgment. On the Application side the author lists Validation, Authorization, Access Control, Transaction, Idempotency, and Recovery, and notes that even with these controls, the question of whether the AI's original judgment was wrong still remains, leading later experiments toward Known Answer, Consistency, and Claim Verification.
The author is careful not to pre-decide that some work must never be given to AI, or that AI is safe because it is convenient, and instead repeats a cycle of hypothesis, experiment, failure, cause analysis, countermeasure, and re-experiment. Whether the Lab produces a usable boundary for delegation is likely to hinge on how well that loop surfaces failure conditions rather than one-off attack successes.
Pick your industry and the AI tools you use, and get news related to your work every day.
Free · 30 seconds with Google · unsubscribe anytimeWhat is AIToday? →
Ask AI anything about this article. The AI reads this article, earlier AIToday articles, and Wikipedia, and cites its sources. Q&As are published on this page for other readers too.
SpaceX finished its first full quarter as a public company on Sept

After OpenAI staff raised concerns on Slack, president Greg Brockman abandoned the second half of his $50 mill…

GovAI research fellows Alan Chan and Sam Manning said at a Sept

A practitioner listed five Japanese-language books he keeps re-opening, from '機械学習 100+ページ エッセンス' by Andriy Bu…

The article lays out the split in Claude Code — CLAUDE.md is the file the user writes with instructions and ru…

Cloudflare's Day 4 announcements made AI Search and the Cloudflare Basin data platform generally available, op…
