Trusted Poker Guide Since 2004
About UsContactResponsible Gambling
PokerSites.orgPlay Now →
🏆 Best Poker Sites💰 Best Bonuses🇺🇸 US Poker🇨🇦 CA Poker📊 Traffic Rankings📱 Mobile Poker🎰 Online Casino⚽ Sportsbook
Strategy

How Good Is Poker AI? A New Benchmark

By jason-murphy·September 16, 2026·6 min read

For twenty years, claims about poker AI strength have been effectively unfalsifiable. An operator or a researcher would announce that a bot beat some previous bot, or some group of professionals, over some number of hands, under conditions nobody else could reproduce. There was no shared yardstick.

That changed in March 2026 with the publication of the GTO Wizard Benchmark — a public API and standardised evaluation framework for measuring algorithms in heads-up no-limit Texas hold'em. For the first time, competing systems can be tested against a fixed reference under identical conditions, with results anyone can verify.

It is a dry piece of infrastructure. It is also the most important development in poker AI since the solver era began, because it turns marketing claims into measurable numbers.

What the benchmark measures

Heads-up no-limit hold'em (HUNL) has been the standard testbed for poker AI research since the early 2000s. It is complex enough to be genuinely hard — imperfect information, enormous game trees, no clean solution by brute force — but constrained enough that two systems can be compared directly.

The benchmark provides a public reference opponent and a standardised measurement protocol, reporting results in big blinds per 100 hands (bb/100) with confidence intervals. That last part matters more than it sounds. Poker results are extraordinarily noisy; a result quoted without a confidence interval is close to meaningless over anything under a few hundred thousand hands.

The headline figure from the accompanying work: GTO Wizard AI, a superhuman agent trained through self-play reinforcement learning over hundreds of millions of hands, defeated Slumbot — the 2018 Annual Computer Poker Competition champion — by 19.4 ± 4.1 bb/100.

Putting 19.4 bb/100 in perspective

That number deserves a moment.

A strong winning professional in a mid-stakes online cash game might expect 3 to 5 bb/100 against a typical pool. A genuinely elite player crushing a soft game might manage 8 to 10 bb/100. Those are the sort of win rates that build careers.

19.4 bb/100 against a former computer poker champion is a margin of a different order. And Slumbot was not a weak opponent — it was, for years, the publicly available standard against which serious systems were measured.

The gap between the best AI and the best humans in heads-up no-limit is not narrow. It has not been narrow since Libratus and DeepStack beat professionals in 2017. What the benchmark provides is a way to state exactly how wide it is, and to track whether new approaches close it.

Where large language models sit

The benchmark also gives a clean answer to a question that has generated a great deal of noise: can general-purpose AI models play poker?

A May 2026 study found that GPT-5.3 with extra-high reasoning settings achieved a luck-adjusted win rate of -16 ± 3.0 bb/100. That is a substantial improvement over previous-generation models, which were catastrophically bad. It is also still a losing rate — a considerable one.

The distance between a purpose-built poker agent at +19.4 and a frontier general-purpose model at -16 is roughly 35 big blinds per 100 hands. General reasoning ability, it turns out, does not substitute for hundreds of millions of hands of self-play in a game defined by imperfect information and mixed strategies.

The practical implication: the "AI can now play poker" headlines that surface periodically are conflating two very different things. Specialised poker AI has been superhuman for nearly a decade. General AI is not close.

The legitimate side of poker AI

None of this is about cheating. Real-time assistance software is banned by every reputable operator, and the industry has invested heavily in detecting it — GTO Wizard itself announced anti-cheating partnerships with the Winning Poker Network and WPT Global this year.

The legitimate application is training. GTO Wizard was named the WSOP's Official Poker Training Partner for the 2026 summer series at Paris Las Vegas and Horseshoe Las Vegas, which is about as mainstream an endorsement as solver-based study gets.

The benchmark matters here too. If you are paying for a training tool that claims solver-accurate solutions, a public evaluation framework gives the industry a way to substantiate — or fail to substantiate — that claim.

What this means for players

A solver is a study tool, not a strategy. The most common mistake players make with solver output is trying to memorise it. Solvers produce mixed strategies — bet 62% of the time with this hand, check 38% — which no human can execute and which are designed against an opponent playing perfectly. Your opponents are not playing perfectly. Our GTO vs exploitative material covers the distinction, and it is the single most valuable thing a developing player can internalise.

Study the reasons, not the frequencies. What solver work is genuinely good for is understanding why a strategy exists: why certain hands become bluffs, why bet sizing changes on particular board textures, why some ranges are capped. Those principles transfer. A memorised frequency for one specific spot does not.

The gap between AI and humans tells you where the money still is. If a superhuman bot beats a strong bot by nearly 20 bb/100 in heads-up, the implication is that even excellent human play leaves enormous amounts on the table. That is not discouraging — it is the reason poker remains profitable. Every one of those big blinds is an edge available to whoever studies hardest. Our poker odds and equity guides cover the fundamentals that underpin all of it.

Heads-up is the most theory-dense format, and the most punishing. The reason AI research uses HUNL is that it is the simplest form of no-limit hold'em. It is also the format where a skill gap produces the fastest bankroll destruction, because you play every hand. If you are drawn to heads-up, do the work first — start with Texas Hold'em fundamentals and our bluffing guide, which covers the balanced-range logic that heads-up play demands.

Bot detection is now a competitive advantage for operators. When you choose where to play, the room's integrity investment is part of the product. Our safe poker sites guide covers what to look for, and the poker networks page compares how the major networks approach game integrity.

What comes next

The benchmark's real value will emerge over the next few years, as competing research groups submit systems and a public leaderboard develops. Poker AI research has suffered from a reproducibility problem that plagues much of machine learning: impressive results, published under conditions nobody else can recreate.

A shared, public evaluation framework fixes that. It will not make the games easier. It will, at minimum, make the claims about them honest.

For players, the takeaway is unchanged from what it has been since 2017: the machines are better than you at heads-up no-limit, they have been for years, and none of that stops poker from being a profitable game against other humans. The edge was never about being optimal. It was about being less wrong than the person across the table.

Sources: arXiv: GTO Wizard Benchmark, arXiv: PokerSkill, GTO Wizard blog

Tags:poker aigtosolverspoker strategyheads-up

Related Articles

18+ only. Gambling can be addictive — please play responsibly. PokerSites.org is an independent guide not operated by any gambling operator. If you or someone you know has a gambling problem, contact NCPG (1-800-GAMBLER), BeGambleAware, or GamCare.