AIToday
Large Language ModelsHugging Face BlogPublished: May 28, 2026, 04:00 JST1 min read

Artificial Analysis and IBM launch ITBench-AA, first benchmark for agentic enterprise IT tasks, with frontier models scoring below 50% on Site Reliability Engineering challenges

Artificial Analysis and IBM launch ITBench-AA, first benchmark for agentic enterprise IT tasks, with frontier models scoring below 50% on Site Reliability Engineering challenges

3 Key Points

  1. Artificial Analysis and IBM Software Innovation Lab released ITBench-AA, evaluating frontier AI models on agentic enterprise IT tasks starting with Site Reliability Engineering (SRE)—diagnosing Kubernetes incidents by reading logs, tracing dependencies, and identifying root-cause entities. Claude Opus 4.7 (Adaptive Reasoning, Max Effort) leads at 47%, followed by GPT-5.5 (xhigh) at 46% and Qwen3.7 Max at 42%.

  2. The benchmark comprises 59 SRE tasks (40 public, 19 held-out) where models run in a sandboxed harness with shell access to incident snapshots. Scoring uses average precision at full recall: models must identify all ground-truth root causes or score 0.0 for that task repeat; if successful, they receive a precision score equal to true positives divided by true positives plus false positives.

  3. All frontier models score below 50%, making ITBench-AA SRE one of the least saturated agentic benchmarks. Turn counts vary nearly 3x with no correlation to accuracy: GPT-5.5 (xhigh) averages 31 turns at 46%, while Gemini 3.1 Pro Preview averages 83 turns at 30%. Among open-weights models, GLM-5.1 (Reasoning) leads at 40%, tied with Gemini 3.5 Flash (high).

  4. The benchmark and leaderboard are available at the ITBench-AA HuggingFace repo and artificialanalysis.ai/evaluations/itbench-aa; the Stirrup reference harness is open-source.

Ask the AI about this article →

Hugging Face BlogRead Original Article

Get the latest Large Language Models news every morning

For example, today's edition would include:

  • CBTS launches Forge Agents for custom AI agentsSiliconANGLE AI · 2h ago
  • Imec CEO: AI era widens chip-model-CSP collaborationDIGITIMES Asia · 2h ago
  • Alphabet's AI Overviews reach 2.5B monthly usersYahoo Finance AI · 2h ago

AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.

Free · takes 30 seconds · unsubscribe anytimeWhat is AIToday? →

Ask AI

Ask AI anything about this article. Q&As are published on this page for other readers too.

Related Articles

Next articleBroadcom introduces BCM68850, described as the industry's first 50G ITU-PON home gateway SoC with integrated neural processor and Wi-Fi 8 support