Reinforcement learning · self‑play

Noughts & Crosses

A minimal AI that starts out playing randomly, plays thousands of games against itself, and slowly learns which board positions tend to win. Train it below, then challenge it yourself.

not trained yet

Train it

games played0
board states learned0
exploration rate (ε)0.30
X wins O wins draws

What's actually happening

The idea

The agent keeps a running guess, for every board position it has seen, of how likely X is to win from there. That's it — no rules of strategy are programmed in, just one number per position.

Self‑play

Both X and O are played by the same agent, using the same table of guesses. X tries to move toward positions it rates highly; O tries to move toward positions it rates lowly. They train each other.

Learning from outcomes

After each game, the agent walks back through the positions it passed through and nudges each guess slightly toward the guess that came right after it — a technique called temporal-difference learning. Wins and losses slowly ripple backward into earlier moves.

Exploration vs. exploitation

Early on the agent plays a lot of random moves (high ε, "epsilon") so it doesn't miss good ideas. As training goes on it explores less and plays more of its best-known move — shown in the ε stat above.

Why it never loses (eventually)

Noughts and crosses is small enough that with enough self-play games, the agent's guesses converge close to perfect play. Try training past 5,000 games, then see if you can beat it — a draw is the best you should manage.

The thinking overlay

Tick "show the agent's thinking" mid-game to see the number it currently assigns to each empty square — its estimate of X's winning chances if that square were played next.