Source study found
Story checked
With most information hidden, the game Stratego had stumped AI—until now - Ars Technica (opens in a new tab)
arstechnica.com · 2026-10-01
Short answer
Not supportedNot supported.
The study does not answer the story's main claims.
- 3 not covered
Checked against the study summary. The full text wasn't available, so some details couldn't be settled either way.
Share this check
The story
With most information hidden, the game Stratego had stumped AI—until now - Ars Technica
arstechnica.com · 2026-10-01
The story’s checkable claims.
Read the original story (opens in a new tab)NewsLink checks it
Not supported
The study doesn't address any of the story's claims. We found the paper, but it doesn't report the details the story leads with.
- 3 not covered
The source study
Scalable decision-making for games of imperfect information
Evidence layer
Claim by claim
Each claim gets a verdict. Expand it to see the evidence directly below.
Reading mode
Scan verdicts. Open evidence only when needed.
Browse by verdict
3 claims in this storyShowing all 3 claimsChoose a verdict to focus the list.
Claim 1 of 3Not coveredA team of researchers from Carnegie Mellon, MIT, New York University, and Stanford University built an AI called Ataraxos that beat Pim Niemeijer, arguably the best Stratego player of all time, 15 games to one, with four draws.View evidenceHide evidence
As stated15 wins, 1 loss, 4 draws
Why this verdict
The abstract-level profile supports the broad claim that Ataraxos defeated the most decorated human Stratego player of all time by a large margin and that this was framed as a first superhuman Stratego result. However, the supplied abstract evidence does not verify the player’s name as Pim Niemeijer, the institutional affiliations of the research team, or the exact 15-1 with four draws score. Those details may be in the full paper or article sources, but they are not verifiable at the requested abstract depth.
Study evidence
Ataraxos defeated the most decorated human Stratego player of all time by a large margin, reported as the first superhuman result in Stratego.
“Ataraxos defeated the most decorated human Stratego player of all time by a large margin—achieving, to our knowledge, the first superhuman result in the game’s history—while consuming orders of magnitude less compute and data than previous efforts.”
Claim 2 of 3Not coveredThe article says Ataraxos took just 16 GPUs and a few thousand dollars to train.View evidenceHide evidence
As stated16 GPUs; a few thousand dollars
Why this verdict
The abstract-level profile supports only a qualitative efficiency claim: Ataraxos used orders of magnitude less compute and data than prior efforts. It does not quantify training resources as 16 GPUs or a few thousand dollars, nor define the cost basis. The story’s numeric resource claim is therefore not verifiable from the supplied abstract evidence.
Study evidence
Ataraxos defeated the most decorated human Stratego player of all time by a large margin, reported as the first superhuman result in Stratego.
“Ataraxos defeated the most decorated human Stratego player of all time by a large margin—achieving, to our knowledge, the first superhuman result in the game’s history—while consuming orders of magnitude less compute and data than previous efforts.”
Claim 3 of 3Not coveredStratego is presented as an imperfect-information game with substantial hidden information, long games, and bluffing, which the team says made it hard for earlier AIs such as DeepMind's DeepNash.View evidenceHide evidence
As stated40 pieces; more than a decillion possible setups; games can last 2,000 moves
Why this verdict
The profile supports the general framing of Stratego and related tasks as imperfect-information settings with large amounts of hidden information that challenge established reinforcement learning and search methods. It also says prior top-human-level performance in Stratego remained out of reach. But the abstract-level evidence does not verify the story’s specific details about 40 pieces, more than a decillion setups, games lasting 2,000 moves, bluffing dynamics, or DeepMind’s DeepNash being stumped for those reasons. The causal explanation is directionally aligned with the abstract’s hidden-information framing, but the detailed version is not verifiable at this depth.
Study evidence
Introduction of Ataraxos: an AI based on general techniques that combine self-play reinforcement learning with test-time search under hidden information.
“Here we introduce Ataraxos, an AI for Stratego based on general techniques that we developed for both self-play reinforcement learning and test-time search under hidden information.”
Study evidence
Ataraxos defeated the most decorated human Stratego player of all time by a large margin, reported as the first superhuman result in Stratego.
“Ataraxos defeated the most decorated human Stratego player of all time by a large margin—achieving, to our knowledge, the first superhuman result in the game’s history—while consuming orders of magnitude less compute and data than previous efforts.”
Context layer
What the story left out
Important study details the story did not include.
Ataraxos is introduced as a methodological contribution combining self-play reinforcement learning with test-time search under hidden information for imperfect-information games.
The story presentation reflects that Ataraxos is an AI for Stratego but does not capture the paper’s central methodological contribution: the combination of self-play reinforcement learning and test-time search under hidden information as general techniques.
From in silico
The same techniques are reported to transfer to other imperfect-information games: superhuman Barrage Stratego and state-of-the-art Hanabi and dou dizhu, with low cost and high sample efficiency.
The story presentation focuses on Stratego, Ataraxos, the human match, and prior Stratego AI difficulty. It does not mention the cross-game transfer results that the abstract presents as evidence for the approach’s generality.
From multi-environment evaluation/benchmarking
1 thing the story did carry across
- The paper reports that Ataraxos defeated the most decorated human Stratego player by a large margin and frames this as the first superhuman Stratego result, with substantially lower compute/data than prior efforts.
Study layer
Study at a glance
Scan the study first. Expand only the parts you want to inspect.
Pieces of work
3
Evidence read
study summary
Lead result
in silico
1Lead resultin silicoIntroduce Ataraxos: general techniques for self-play reinforcement learning and test-time search in games with hidden information, enabling scalable decision-making under imperfect information.ExpandCollapse
In plain English
Paper introduces Ataraxos, a methodological approach that combines self-play reinforcement learning and test-time search tailored to decision-making in games with large amounts of hidden information; the techniques are presented as general and are validated across multiple imperfect-information games.
Key findings
- Introduction of Ataraxos: an AI based on general techniques that combine self-play reinforcement learning with test-time search under hidden information.
- The same methodological approach is claimed to yield superhuman performance in Stratego and Barrage Stratego, and state-of-the-art results in Hanabi and dou dizhu, while using substantially less compute and data than prior efforts.
“Here we introduce Ataraxos, an AI for Stratego based on general techniques that we developed for both self-play reinforcement learning and test-time search under hidden information.”
What this piece can’t prove
- Summary and claims are based solely on the paper abstract provided; the excerpt lacks methodological and experimental detail.
- No algorithmic descriptions, training procedures, architectures, hyperparameters, evaluation protocols, or quantitative results are available in the supplied text to substantiate the claims.
- It is not possible from the excerpt to assess reproducibility, statistical significance, or the scope of evaluations beyond the named games.
2in silicoDemonstrate superhuman performance in Stratego by defeating the most decorated human player with markedly lower compute/data than prior efforts.head-to-head evaluation versus a top human Stratego player; implied controlled match protocol and efficiency comparisonExpandCollapse
In plain English
Abstract reports a human-evaluation of Ataraxos in Stratego in which Ataraxos defeated “the most decorated human Stratego player of all time by a large margin,” claims this is the first superhuman result in Stratego, and states that Ataraxos used orders of magnitude less compute and data than prior efforts. The abstract also states related successes on Barrage Stratego, Hanabi, and dou dizhu. The paper provides no match-level details in the abstract.
Key findings
- Ataraxos defeated the most decorated human Stratego player of all time by a large margin, reported as the first superhuman result in Stratego.
- Ataraxos achieved this performance while consuming orders of magnitude less compute and data than previous efforts.
“Ataraxos defeated the most decorated human Stratego player of all time by a large margin—achieving, to our knowledge, the first superhuman result in the game’s history—while consuming orders of magnitude less compute and data than previous efforts.”
What this piece can’t prove
3 further details could not be confirmed from the summary.
3in silicoShow the same techniques transfer to other imperfect-information games (Barrage Stratego, Hanabi, dou dizhu), achieving superhuman or state-of-the-art performance with low cost/high sample efficiency.multi-environment evaluation/benchmarkingExpandCollapse
In plain English
Paper reports that the same reinforcement-learning and test-time search techniques used for Stratego were applied to Barrage Stratego, Hanabi, and dou dizhu; the authors state these produced a superhuman AI for Barrage Stratego and state-of-the-art AIs for Hanabi and dou dizhu, all with low compute cost and high sample efficiency. The claims are framed as evidence of the approach's generality across adversarial, cooperative, and team imperfect-information games.
Key findings
- Applying the same techniques produced a superhuman AI for Barrage Stratego and state-of-the-art AIs for Hanabi and dou dizhu, reported to be achieved with low compute cost and high sample efficiency.
“Using the same techniques, we built a superhuman AI for Barrage Stratego and state-of-the-art AIs for Hanabi and dou dizhu, all with low cost and high sample efficiency.”
What this piece can’t prove
- Summary is based only on the abstract; full experimental protocols, metrics, and quantitative outcomes for each game are not available here.
- The abstract does not specify baselines, evaluation methodology, opponent skill, or compute/training-data amounts that underpin the 'superhuman' and 'state-of-the-art' claims.
1 further detail could not be confirmed from the summary.
Method layer
NewsLink found the paper. Tessa takes you deeper.
NewsLink checks the story. Tessa is where you inspect the paper, authors, evidence, and research context.
Open the paper in Tessa
Scalable decision-making for games of imperfect information
Nature · 2026
Why this one
Near certain
NewsLink found the paper. Tessa is where you inspect it deeply.
Papers considered
The selected paper, plus nearby candidates.
Crossref · 1 candidate paper
Scalable decision-making for games of imperfect information
Nature · 2026 · Crossref