A team of researchers published ThermoQA, a test set of 293 engineering thermodynamics problems (ranging from simple property lookups to complex power-plant cycle analysis) with answers verified by specialized physics software. Claude Opus 4.6 scored highest at 94.1%, but all six leading AI models showed sharp performance drops—ranging from 3% to 33%—when moving from simple to complex questions, exposing the gap between memorization and genuine reasoning.
The benchmark revealed that AI models struggle with specialized physics scenarios (supercritical water, refrigerant behavior, gas turbines) that require actual thermodynamic reasoning rather than pattern-matching from training data. Across multiple test runs, model answers varied by up to 2.5%, showing that even top-performing models produce inconsistent reasoning on engineering problems—a red flag for real-world use.
Engineers, physics educators, and companies building AI tools for technical work now have a public, open-source way to test whether an AI model can actually solve real thermodynamics problems or just appears confident. This matters because deploying an AI to help design a heating system or power plant when it's primarily memorizing rather than reasoning could lead to dangerously flawed designs.
Ask the AI about this article →
For example, today's edition would include:
AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.
Free · takes 30 seconds · unsubscribe anytimeWhat is AIToday? →
Ask AI anything about this article. Q&As are published on this page for other readers too.
Anthropic reset the 5-hour and 1-week usage limit windows for its AI service Claude on September 1, in connect…

Salesforce and Anthropic announced Claudeforce, starting with "Salesforce in Claude." This plugin lets users i…

Anthropic announced Claude Fable 5.1 and Claude Mythos 5.1 on September 1

A technical explainer compares three LLM serving strategies—static, dynamic, and continuous batching

Anthropic's latest model, Claude Fable 5.1, is now available on Snowflake Cortex AI

The Allen Institute for AI released BenchMIRT, a method to audit AI benchmarks question-by-question
