
AI coding assistants cannot judge how long tasks take.
They overestimate time needed, especially for short jobs.
They also overrate their own work quality.
What happened
A study by two independent AI researchers, done as part of the MATS research program, tested Anthropic's Claude Code and OpenAI's Codex on their sense of time. Before tasks, the agents had to estimate how long they'd need; afterward, they reported how much time had passed. The test material came from 200 tasks in ProgramBench plus 18 custom benchmarks.
Why it matters
The agents consistently overestimated how much time they'd need. On ProgramBench, both models mostly guessed around 90 minutes, no matter the difficulty. In the second round, Claude was off by three times on average, Codex by six to ten times. The estimates were worst for short tasks, and only in the multi-hour range did some predictions come close to reality.
What to watch
The agents also overrated their own work: older models Opus 4.8 and GPT-5.5 scored themselves about 20 points higher on average, even on failed tasks. When the agents got access to a tool that reports elapsed time, they got it right almost every time.
Ask the AI about this article →
The study highlights a fundamental limitation of AI agents: they lack an internal clock. This matters because long-running tasks require agents to follow instructions like 'iterate for two hours,' which becomes impossible if they cannot gauge the passage of time. The findings also show that runtime depends heavily on the surrounding software (the harness), with the same model taking 2.5 times more steps in Claude Code than in Codex on average. The researchers plan to test whether agents can stick to a set work duration, suggesting a potential direction for improving agent reliability.
For example, today's edition would include:
AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.
Free · takes 30 seconds · unsubscribe anytimeWhat is AIToday? →
Ask AI anything about this article. Q&As are published on this page for other readers too.
A new analysis argues that embodied AI systems face a systemic bottleneck called the 'edge AI wall,' where onb…

JetBrains announced Junie Local, a coding agent that runs entirely on a local machine

A Pew Research Center survey of 3,488 U.S

Texas Governor Greg Abbott has frozen state spending on Flock's AI surveillance cameras

Caterpillar is applying lessons from its mining automation business to broader AI deployment, including a voic…

A randomized experiment with 1,053 freshmen at Bocconi University found that GPT-4o helped students earn signi…
