AI RACE— The AI Race
Research

Academic Team Builds Superhuman Stratego AI on an $8,000 Compute Budget

Researchers from CMU, NYU, Stanford, and MIT have created Ataraxos, an AI that conquered board game legend Pim Niemeijer in Stratego using an efficient training pipeline that cost less than $8,000.

10/01/2026, 19:58
Research

Decisive Victory Over Stratego's Top Grandmaster

A research collective spanning Carnegie Mellon University, New York University, Stanford University, and MIT has developed the first artificial intelligence to reach superhuman mastery in Stratego, one of the final classical board games where human intuition outmatched machines. Detailed in a paper published in Nature, the system, named Ataraxos, faced Dutch champion Pim Niemeijer in an official 20-game clash and secured a commanding record of 15 victories, one defeat, and four draws.

Niemeijer is widely regarded as the most dominant player in Stratego history, boasting four world championships, 15 Dutch national titles, two online world championships, and over 600 weeks atop the global rankings—a pedigree confirmed by veteran competitor George Franka, who has contested every world championship since 1997. The AI’s performance translates to an unprecedented 85 percent effective win rate against elite human competition, where three-time world championship runner-up Max Roelofs noted that standard competitive margins are razor-thin due to the high risks forced by the game's mechanics.

Inside the $8,000 Training Pipeline and In-Game Strategy

Stratego presents an immense computational barrier because both competitors arrange 40 concealed pieces across more than 10^33 possible initial configurations. Unlike Texas Hold’em poker, which contains just 1,326 opening hand combinations, Stratego conceals unit ranks until opposing pieces clash on the board, requiring models to reason through vast amounts of hidden information where historical actions dictate the value of future choices.

Ataraxos achieved its superhuman standard without relying on human match logs, learning entirely through self-play. First author Samuel Sokota and the team prevented the algorithm from locking into predictable, exploitable behaviors by applying heavy regularization early in training. The model initially explored wide variations of setups and moves with large step sizes, which researchers gradually tapered into refined, smaller adjustments—a dynamic they likened to an energy reserve that prevents learning from degenerating into chaos. To execute moves, Ataraxos deploys an integrated belief network to infer the ranks of hidden opposing pieces, generates prospective game scenarios, and calculates a targeted learning update specifically for that decision.

The entire system was trained under academic budget constraints using a custom-engineered GPU simulator. The base system trained for one week on 16 Nvidia H100 GPUs, followed by four days on four GPUs to calibrate the belief network, costing under $8,000 at 2025 prices.

During the three-week series against Niemeijer, the human champion held structural advantages: he had time to prepare and adjust to the AI across multiple matches, whereas Ataraxos remained completely static and incapable of adapting to Niemeijer's play. Three-time world champion Vincent de Boer described this lack of in-series adaptation as a serious handicap for the machine. To ensure peak motivation, Niemeijer received a $1,000 participation fee alongside performance bonuses of $100 per win and $50 per draw.

Opponents describe the AI's playstyle as disorienting and exceptionally difficult to anticipate. Ataraxos routinely launches high-risk bluffs, exhibits precise piece placement, and employs stubborn stalling tactics when falling behind. In a separate exhibition during the 2025 Stratego World Championship, the system won 38 out of 40 games against tournament-level contenders.

Surpassing DeepMind and Scaling Across Hidden-Information Games

The academic breakthrough directly challenges past efforts by industry tech giants. Google DeepMind tackled the game with its DeepNash model, which required 1,024 TPU nodes running for two to three months—amounting to an estimated $3 million to $4.5 million in 2025 compute costs. While DeepNash won 19 of 28 matches at the 2023 World Championship, it consistently lost to premier humans, including Niemeijer. The Ataraxos team operated on roughly 1/500th of DeepNash's compute expenses, 1/30th of the self-play game volume, and 1/100th of the training examples. A proposed head-to-head match between the two AI systems never materialized after DeepMind indicated that the DeepNash software is no longer functional.

The underlying architecture has proven effective across multiple game environments characterized by imperfect information:

  • In the Barrage Stratego variant, the algorithm secured four 50-game match series against three top-four players, all two-time champions.
  • In the cooperative card game Hanabi, the method broke benchmark records across all variants, solving the two-player edition with two orders of magnitude less compute than the previous state of the art.
  • In the popular Chinese card game Dou dizhu, it outperformed existing benchmarks PerfectDou and DouZero.

The authors maintain that vast volumes of hidden information are no longer an insurmountable hurdle for reinforcement learning and search algorithms. They suggest the paradigm can be extended to real-world environments like financial markets, diplomatic negotiations, and military conflicts, provided reliable simulators can be constructed. The researchers noted one current technical constraint: because Ataraxos's search simulates only a single learning step, its performance cannot simply be scaled up with extra compute time at inference. The code for Ataraxos has been released publicly.

◗ Sources

The Decoder10/01

Related stories