But Ataraxos’s best trick was arguably its price tag.

The price to pay

DeepNash was trained for two to three months on 1,024 of Google’s specialized chips, a run the Ataraxos team estimates would cost $3 million to $4.5 million at 2025 prices. Ataraxos, in contrast, needed 16 GPUs for a week, plus an additional four GPUs for four days to train the belief model.

Farina and lead author Samuel Sokota achieved this efficiency by writing a simulator that runs millions of moves per second on graphics cards. “At the scale that we are in academia, we don’t really have access to an entire field of GPUs,” Farina said.

The algorithm also learned far faster—it played about 34 times fewer games than DeepNash, and still ended up much stronger.

The Ataraxos architecture also worked in learning other games. The same approach beat three world champions at Barrage Stratego, a faster eight-piece variant of Stratego, mastered the cooperative card game Hanabi, and beat the best bots at the Chinese card game dou dizhu. But the team has its sights set on scenarios far more complex than board or card games.

Beyond the board

Board games have fixed rules and clear winners, while real-world problems like negotiations, financial markets, or military conflicts usually don’t. But the Ataraxos team argues the gap is smaller than it looks, since tackling any real problem starts with building a simplified model of it.

“War gaming is a common thing that people do,” Vinitsky said. “You can use the techniques that were derived here to play it forward and see how a strong opponent might respond to what you do.”

Farina and his colleagues are now interested in making their AI more understandable, since Ataraxos, in its current state, can’t explain why it makes the moves it makes. “We work on machines that produce strong but also interpretable and explainable strategies. I think we’re not quite there yet,” Farina said.

Nature, 2026. DOI: 10.1038/s41586-026-11036-y