AIToday
Large Language ModelsAI Business & Industryr/MachineLearningPublished: Aug 30, 2026, 04:00 JST1 min read

LLM benchmark scores vary up to 8.4 points between days

LLM benchmark scores vary up to 8.4 points between days

Key takeaway

  • LLM benchmark scores show daily variation of 8.4 points.

  • Within a day, variation is only 2.8 points.

  • This suggests model performance may fluctuate day-to-day.

3 Key Points

  1. What happened

    An analysis of 31,352 hourly benchmark scores from 49 model identifiers found that within-day score variation was 2.8 points, while between-day variation was 8.4 points, on a normalized 0-100 composite score.

  2. Why it matters

    This means production LLM APIs can show noticeable performance differences from one day to the next, which could affect businesses relying on consistent model outputs for coding, reasoning, or tool-calling tasks.

  3. What to watch

    The analysis used a continuous evaluation pipeline that repeatedly tests models, including high-frequency canary tasks, to separate sustained changes from random noise; the author developed the open-source system AIStupidLevel used for the data collection.

Ask the AI about this article →

Context & Analysis

The study highlights that LLM performance is not static; it varies over time. The between-day variation of 8.4 points on a 0-100 scale is notably larger than the within-day variation of 2.8 points, suggesting that daily changes are more significant than intraday fluctuations. This could matter for applications where consistent output quality is critical. The author developed an open-source system, AIStupidLevel, to collect and analyze this data, which is MIT-licensed. While the analysis provides a snapshot, it does not specify the time period covered, so the actual duration of the study is unclear. Businesses using LLM APIs might want to monitor performance over time to detect sustained changes.

FAQ

What was the sample size of this analysis?
The analysis examined 31,352 hourly benchmark scores from 49 model identifiers.
How was the performance measured?
Performance was measured using repeated tests on coding, deep reasoning, tool calling, and high-frequency canary tasks, with responses executed rather than just judged.
r/MachineLearningRead Original Article

Get the latest Large Language Models news every morning

For example, today's edition would include:

  • OpenAI CEO Calls for AI Development Slowdown After Safety FailuresYahoo Finance AI · 9m ago
  • Mastercard CEO: AI agents and machine payments reshape commerceTop Companies AI · 3h ago
  • Deepgram adds Enhanced Metrics to SageMaker AITop Companies AI · 3h ago

AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.

Free · takes 30 seconds · unsubscribe anytimeWhat is AIToday? →

Ask AI

Ask AI anything about this article. Q&As are published on this page for other readers too.

Related Articles

Next articleSony Music, Warner Chappell sue Anthropic