
What happened
The AI Ataraxos beat Niemeijer, the most decorated Stratego player, with an 85 percent effective win rate over three weeks, at a compute cost of less than $8,000.
Why it matters
This marks a breakthrough in imperfect-information games, showing that reinforcement learning can handle large amounts of hidden information where poker-derived methods previously struggled.
What to watch
The researchers say more compute time won't improve Ataraxos's performance unless a more sophisticated search method is developed; the code is publicly available for others to build on.
WHO IT HITSThis work signals that AI can now tackle complex strategic decision problems with hidden information using modest budgets, potentially benefiting researchers and businesses in fields like finance, military planning, and negotiations.
Summaries like this, in your inbox every morning.
Stratego has long been a tough challenge for AI because both players set up their 40 pieces face down, creating more than 10^33 possible setups. In such games, the value of a move depends on hidden information, and poker AI methods only worked when little was hidden. DeepMind's DeepNash trained on 1,024 TPU nodes for two to three months, costing an estimated $3 million to $4.5 million at 2025 prices, yet failed to beat top humans. In contrast, Ataraxos needed one week on 16 Nvidia H100 GPUs and four more days on four GPUs, costing less than $8,000. The researchers credit a custom GPU simulator and higher sample efficiency.
The key to Ataraxos's success is regularization, which forces the AI to vary its play and avoid predictable strategies. It also uses a belief network to predict hidden pieces. The researchers tested Ataraxos against Niemeijer over three weeks, with Niemeijer receiving $1,000 plus bonuses for wins and draws. The AI won with an 85 percent effective win rate, a result described as unprecedented at the highest level. Niemeijer could adapt over many games, but the AI could not adapt to him, which three-time world champion Vincent de Boer called a major handicap. At the 2025 Stratego World Championship exhibition, Ataraxos won 38 of 40 games.
The method also worked in other games: Barrage Stratego, Hanabi, and Dou dizhu. The authors argue that large amounts of hidden information are no longer an obstacle for reinforcement learning, opening up practical AI for strategic decision problems like financial markets, military conflicts, and negotiations—provided fast simulators can be built. One limitation: the search only mimics a single learning step, so adding compute won't keep improving performance without a more sophisticated search method. The code is publicly available, so others can test and extend the approach, but the real-world impact will hinge on whether suitable simulators exist for those domains.
Pick your industry and the AI tools you use, and get news related to your work every day.
Free · 30 seconds with Google · unsubscribe anytimeWhat is AIToday? →
Ask AI anything about this article. The AI reads this article, earlier AIToday articles, and Wikipedia, and cites its sources. Q&As are published on this page for other readers too.
ServiceNow launched Flow, a natural-language service desk that handles employee requests inside Slack and Micr…
DeepSeek and Huawei announced a partnership to develop semiconductor software, part of efforts to reduce China…

Google unveiled Gemini 4 Argon on September 30, saying it beats GPT-6 Astra and Claude Opus 5.5 on 13 of 19 be…

Yann LeCun told Fortune’s Emily Forlini he has zero concerns about rogue AI incidents, including OpenAI agents…

A Preply survey of over 5,000 professionals in nine countries found 92% of Gen Z respondents used AI for learn…

Recounting 301 Claude Code session logs on 2026-10-01, the one-off script rate was 50%, versus 69% in the firs…
