9×9 Go, played by a tiny neural net

This page runs a real 9×9 Go engine in your browser via WebAssembly, compiled from the weiqi-engine Rust code — modified Tromp–Taylor rules (captures, suicide ban, positional superko, two-pass end, area scoring with selectable komi, default 6.5) plus a hand-rolled 130k-parameter CNN forward pass. Toggle the policy heatmap to see what the net is thinking.

Time machine

loading…

Click an intersection to play. The blue glow is the bot's policy — brighter means it likes the move more.

loading…

0 = greedy (strongest move every time), higher = more playful.

Monte-Carlo tree search: the bot thinks ahead with its own policy and value, weighing its top candidate moves against each other. Catches some tactical oversights the raw policy misses; slower moves — the page pauses briefly while it thinks.

The bot plays the same moves whatever the komi — only the scoring changes. Applies to new games.

Black gets extra stones on the star points and White moves first; komi drops to 0.5 automatically. Applies to new games. Note the net trained on even games, so handicap openings are unfamiliar territory for it.

Rules. Modified Tromp–Taylor: Black moves first, then the sides alternate; a turn is a stone placement or a pass, and groups left with no liberties are captured. Suicide is banned unless it captures, and no earlier board position may ever repeat (positional superko). Two consecutive passes end the game. Scoring: every stone on the board counts as alive, plus empty points fully surrounded by one color; White gets komi — 6.5 by default here, since 7.5 measurably favors White on 9×9 (the bot trains at 6.5 too; pick 5.5 / 6.5 / 7.5 above). There is no dead-stone marking — capture dead groups before the final passes. Two deliberate deviations from pure Tromp–Taylor: suicide is banned outright (pure TT lets it self-clear) and komi is selectable per game (pure TT leaves it to player agreement).

Honest caveats. The self-play checkpoints are from a 15.5M-step PPO run (7–25 vs GNU Go levels 1–10, greedy play). The supervised timeline is stronger: behavioral cloning on 273k human 9x9 positions scores 74–182 (28.9%), and 3M steps of PPO fine-tuning on top reaches 118–138 (46.1%). Plain PPO plateaus: 6M steps reaches 129–127 (50.4%), 9M steps 118–138 (46.1%). Interleaving a behavioral-cloning phase on the human positions during PPO breaks the plateau: 156–100 (60.9%) — the default model. A follow-up fine-tune distilling the model's own MCTS-100 search visits into the policy (last stop on the timeline) reached 160–96 (62.5%) against GNU Go but lost a 256-game head-to-head against this checkpoint 116–140: the two are statistically tied, so the established champion keeps the default. Still well below human club strength; the numbers are the measured reality, not a strength claim.

How the training works is written up in weiqi/theory.md — mostly search-free PPO self-play, a different beast from AlphaZero's MCTS. The optional bot search above is MCTS at play time only: the same net, thinking ahead. The final timeline checkpoint did train on search, once: its own MCTS-100 visit distributions, distilled into the policy.

← back to demos