AIToday
Large Language ModelsAI Coding AssistantsSiliconANGLE AIPublished: Oct 11, 2026, 22:01 JST

M. Touheed: agentic AI breaks testing's three core assumptions

M. Touheed: agentic AI breaks testing's three core assumptions

In a SiliconANGLE essay, Imagine Art's M. Touheed writes that agentic tasks take minutes not milliseconds, retries cost money against metered models, and failures cannot be reproduced without step-by-step traces.

Not sure about something? Ask the AI

Questions and answers are published on this page.

Summaries like this, in your inbox every morning.

Context & Analysis

The essay's starting point is that most traditional enterprise systems were built around three assumptions: jobs finish quickly, retrying one is free, and the same input always produces the same output. Agentic workloads break all three, and Touheed writes that the difficulty is rarely the model — it's the surrounding infrastructure and management practices that assume properties these workloads no longer have.

A large part of his argument is about why companies don't see this coming. When a person chats with AI and checks each response, that person is the error handler, so the error was always there but stayed invisible. Automated agent workflows run headlessly, and a workflow that behaved reliably in interactive use can behave differently when a scheduler fires it 400 times overnight with nobody watching. Touheed's prescription is to identify every task the human was performing manually and specify which automated check or system takes over.

He frames cost the same way. Since cloud and model API usage is billed asynchronously on monthly cycles, compounded retry costs quietly accumulate and become visible only when the invoice arrives weeks later — which is why he wants spending tracked per completed unit rather than per call. On failures, he draws a parallel to data pipelines: agentic work is long-running, partially failing and expensive to re-execute, and teams coming from application development, where synchronous request-and-response is the norm, may discover that gap as a series of surprising incidents rather than as a skills need.

FAQ
Why do agent pilots look fine until they run unattended?
In interactive use, humans read each result, notice problems and retry. Automated workflows run headlessly, so unrecorded failures can break downstream systems unnoticed.
What does M. Touheed say teams should measure instead of cost per API call?
He recommends asking for cost per completed unit, because cost per API call excludes discarded attempts that a metered model still charged for.
What does the essay say about failures that report success?
Probabilistic agents can't recognize their own logical errors, so they emit correctly formatted outputs with bad data that downstream systems accept without alerts.
SiliconANGLE AIRead Original Article

AI news that matters for your work, delivered every morning.

Pick your industry and the AI tools you use, and get news related to your work every day.

Free · 30 seconds with Google · unsubscribe anytimeWhat is AIToday? →

Ask AI

Ask AI anything about this article. The AI reads this article, earlier AIToday articles, and Wikipedia, and cites its sources. Q&As are published on this page for other readers too.

Questions and answers are published on this page.

Related Articles

Next articleSpaceX or Micron? SpaceX wins as the better 5-year AI bet