
LLM benchmark scores show daily variation of 8.4 points.
Within a day, variation is only 2.8 points.
This suggests model performance may fluctuate day-to-day.
What happened
An analysis of 31,352 hourly benchmark scores from 49 model identifiers found that within-day score variation was 2.8 points, while between-day variation was 8.4 points, on a normalized 0-100 composite score.
Why it matters
This means production LLM APIs can show noticeable performance differences from one day to the next, which could affect businesses relying on consistent model outputs for coding, reasoning, or tool-calling tasks.
What to watch
The analysis used a continuous evaluation pipeline that repeatedly tests models, including high-frequency canary tasks, to separate sustained changes from random noise; the author developed the open-source system AIStupidLevel used for the data collection.
Ask the AI about this article →
The study highlights that LLM performance is not static; it varies over time. The between-day variation of 8.4 points on a 0-100 scale is notably larger than the within-day variation of 2.8 points, suggesting that daily changes are more significant than intraday fluctuations. This could matter for applications where consistent output quality is critical. The author developed an open-source system, AIStupidLevel, to collect and analyze this data, which is MIT-licensed. While the analysis provides a snapshot, it does not specify the time period covered, so the actual duration of the study is unclear. Businesses using LLM APIs might want to monitor performance over time to detect sustained changes.
For example, today's edition would include:
AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.
Free · takes 30 seconds · unsubscribe anytimeWhat is AIToday? →
Ask AI anything about this article. Q&As are published on this page for other readers too.
OpenAI announced on August 28 that it will terminate its model access contract with AI coding tool Cursor, pro…

OpenAI CEO Sam Altman said in a Time magazine interview that he thinks "it is a good time to slow down" AI mod…

Bernstein analysts say AI is creating a memory bottleneck, with demand expanding from high-bandwidth memory (H…

Marvell Technology reported record fiscal Q2 2027 revenue of $2.739 billion, up 37% year over year, and raised…

NVIDIA reported $96.2 billion in revenue for fiscal Q2 2027, up 106% year over year, with Data Center revenue…

Intel expanded its partnership with Kasm Technologies to support compliant, local AI workloads on Intel Xeon 6…
