deniz.in

Markets

Weather

Loading weather

· via Hacker News – Front Page (native)

Ataraxos AI beats the best Stratego player in history, trained on just 16 GPUs

An AI called Ataraxos beat Stratego great Pim Niemeijer 15 games to one with four draws, cracking a hidden-information game that had resisted computers — on just 16 GPUs.

Ataraxos AI beats the best Stratego player in history, trained on just 16 GPUs

What happened

A classic board game that had stubbornly resisted game-playing AI has reportedly been cracked. According to Ars Technica, a research team spanning Carnegie Mellon, MIT, New York University and Stanford built a system called Ataraxos that defeated Pim Niemeijer — described as arguably the strongest Stratego player in history — 15 games to one, with four draws. Training the model reportedly took just 16 GPUs and a few thousand dollars.

Why Stratego stumped machines

Chess, Go and poker all fell to computers years ago: Deep Blue beat Garry Kasparov in 1997, AlphaGo beat Lee Sedol in 2016, and poker bots have been beating professionals for years. Stratego held out. Even DeepMind, with an exceptional budget, could not produce a machine that reliably beat top human players, Ars Technica reports.

The difficulty is baked into the rules. Each player deploys 40 pieces representing military ranks, from marshal down to spy, along with bombs and a flag; capturing the enemy flag wins. Opponents can see where your pieces sit, but not what they are. Identities come to light only when two pieces collide, with the weaker piece removed and the victor's rank exposed.

That makes Stratego an imperfect-information game like poker — but at a far larger scale. Gabriele Farina, an MIT computer scientist and co-author on the study, notes that in Texas Hold'em a player holds only two hidden cards, yielding 1,326 possible hands, a space small enough for a machine to evaluate exhaustively. In Stratego, by contrast, 40 pieces can be arranged in more than a decillion possible setups. Games also run far longer: a chess game typically lasts about 40 moves, while a Stratego game can easily stretch to 2,000, Farina said.

The bluffing problem

Deception compounds the challenge. Players sometimes move a weak piece as though it were a marshal, hoping to scare an opponent away. Bluff too often and threats stop being credible; never bluff and play becomes predictable. The researchers say managing that balance is what tripped up earlier systems such as DeepMind's DeepNash, introduced in 2022.

Eugene Vinitsky, an NYU researcher and co-author, described what sets the game apart: "There's something super distinctive about Stratego, which is that it is a massive amount of hidden information that unfolds over a very long time scale."

Why it matters

Stratego combined three problems that chess, Go and poker each posed only partly: a vast hidden state, extremely long games and strategic bluffing. A system that handles all three at once marks a distinct milestone for imperfect-information research rather than another demonstration of raw search power.

The budget is equally notable. Earlier landmark game-playing systems came out of large labs with heavy compute, but Ataraxos reportedly trained on 16 GPUs for a few thousand dollars, suggesting that frontier results in this area no longer require industrial-scale resources. For a field that measures progress through classic games, one of the last stubborn holdouts has now fallen.

  • #stratego
  • #artificial-intelligence
  • #game-theory
  • #imperfect-information
  • #machine-learning

Related posts