In partnership with

A Nature paper published September 30 reports that Ataraxos defeated Stratego player Pim Niemeijer across a twenty-game evaluation. The system recorded 15 wins, one loss and four draws. The result concerns a defined game, not general intelligence.

Stratego places forty pieces on each player's side. Opponents know the positions of pieces but initially lack their identities. Capturing the opposing flag determines victory. Choices can reveal information, conceal intentions or encourage an opponent to misread a position. That makes planning different from searching a fully visible board.

200+ Proven Ways to Make Money With AI in 2026

The next wave of millionaires will be people who figured out how to make AI work for them.

The window to get ahead is still open. But not for long.

Here are 200+ proven ways to make money with AI in 2026.

Sign up for Superhuman AI, the free daily newsletter read by 1M+ professionals, and get instant access to all 200+ ways to profit from AI this year.

Ataraxos combines learning before play with search during play. A network learns strategies and estimates outcomes by playing against itself. Another network models hidden information from those self-play games. Separate networks learn initial piece arrangements and subsequent moves, with game outcomes updating both processes. During a match, it samples plausible arrangements of concealed pieces instead of examining every possible arrangement.

The system evaluates candidate actions across those sampled possibilities. It then adjusts the current decision using the resulting estimates. The published design connects this decision-time refinement to the update process used during training. Learning supplies an initial strategy; search refines an individual choice.

Training scale remains substantial despite the authors' efficiency claims. The reinforcement-learning run completed 163 million games. It used sixteen NVIDIA H100 graphics processors for one week. Training the separate belief network required four H100 processors for four days. Omitting that second stage would understate the reported hardware requirement.

Hands Down Some Of The Best 0% Interest Credit Cards

Pay no interest until nearly 2028 with some of the best hand-picked credit cards this year. They are perfect for anyone looking to pay down their debt, and not add to it!

Click here to see what all of the hype is about.

The authors estimate roughly one five-hundredth of the computing cost of the earlier DeepNash system. They also report fewer self-play games and training examples. Those comparisons describe different resource measures, not interchangeable percentages. Their few-thousand-dollar cost description remains a research-team estimate, not an independently checked accounting statement.

Evaluation design helps explain the human comparison. Niemeijer played twenty games over three weeks, with opportunities for preparation between games. He knew the system would not adapt its strategy to him. The researchers therefore allowed human adjustment while keeping the tested AI fixed.

The paper counts a draw as half a win. On that basis, fifteen wins and four draws produce an 85 percent effective win rate. This differs from the proportion of games won outright. The authors also explain that sequential human adaptation makes game outcomes statistically dependent.

That matters when interpreting significance calculations. A conventional test assuming independent, identically distributed outcomes would not describe every feature of this series. The study states that assumption rather than treating it as established. A strong score and an appropriate explanation of its uncertainty are different parts of the evidence.

The paper reports another forty games against attendees at the Stratego World Championship. Ataraxos won 38 and lost two. Those demonstrations occurred on August 1 through 3, 2025. Their appearance in a September 2026 paper does not make them newly conducted experiments.

Researchers also adapted the approach to Barrage Stratego, Hanabi and dou dizhu. These settings include opposing players, cooperation and competition between teams. They broaden the tested game structures, but remain game evaluations. The published findings do not demonstrate dependable financial trading, military decisions or cybersecurity deployment.

MIT's accompanying account identifies explainable decisions as future work. Its researchers say decision recommendations need auditing before adoption. Real-world uses remain proposals rather than measured outcomes in this study. This research check-in separates training resources, evaluation design and tested scope, leaving application claims open to evidence beyond the board.

Hire anyone, anywhere — compliant in under 3 days

Found the right person, but they’re in a country where you don’t have an entity? Setting one up can take months and significant cost.

Remote removes that barrier by becoming the legal employer through our own entities — handling compliant contracts, local benefits, tax setup, and onboarding for you. In fact, an employee is onboarded to Remote every 7 minutes.

Once they’re hired, the same in-house teams that support employment locally also run payroll — so you’re not bouncing between disconnected providers. Less setup, less complexity, and less time between finding the right person and getting them started.