Beyond Calculation: AlphaGo and the Birth of Machine Creativity
How one move on a 19×19 board rewrote the rules of artificial intelligence
Beyond Calculation: AlphaGo and the Birth of Machine Creativity

(This image was generated using AI technology.)
How one move on a 19×19 board rewrote the rules of artificial intelligence
In January 2016, Google DeepMind dropped a bombshell: just three months earlier, in October 2015, its new program AlphaGo had swept European Champion Fan Hui 5:0, the first time any Go software had beaten a professional go player. Until that reveal, most of us believed it would be another decade before a machine could match pro intuition on a 19×19 board. Now it was to face Lee Sedol, maybe the greatest player of his generation.
For those following the progress of AI, this didn’t feel like another incremental milestone. DeepBlue had beaten Garry Kasparov back in 1997, but it did so with brute-force search and handcrafted evaluation. Watson’s Jeopardy! victory in 2011 relied on massive text corpora and clever ranking strategies. Impressive, but ultimately constrained. Both systems worked in domains that could be conquered with enough data and structure.
Go was different. The number of possible positions was far beyond what brute-force search could handle, but complexity alone wasn’t the real challenge. Small moves in one corner could subtly shift the balance across the entire board. Local tactics were always entangled with global strategy. Strong players relied on experience, instinct, and a sense of shape and balance that was hard to put into words — let alone code.
Let Me Brag a Little: I predicted AlphaGo would beat Lee Sedol 4:1
When the match was announced, most Go players I knew were confident that Lee Sedol would win. The common prediction was 5:0 or, at worst, 4:1 in his favor. After all, AlphaGo’s only known opponent was Fan Hui, and while he had dominated the European scene, he wasn’t seen as world-class. Many people quickly dismissed that earlier result.
I had a different take. I even entered a small prediction game and put down “4:1 for AlphaGo”. As someone with some competency in both Go and AI I noticed some signs pointing against the common understanding.
First, time. Three months had passed between the Fan Hui match and the announcement. That’s a long time for a self-play system to improve, especially with Google’s resources. The core algorithmic breakthrough had clearly been made; from that point on, it was mostly about scale and refinement. We didn’t know it yet, but AlphaZero would later reach AlphaGo’s strength in just eight hours. With that kind of compute, three months is forever.
Second, it is Google after all. They were publicly modest — talking about research, learning, scientific curiosity. But you don’t schedule a global match against one of the greatest Go players of all time unless you know you’re ready. Google doesn’t gamble on 50/50s in front of cameras.
Third, Fan Hui was stronger than people realized. I’ve met him a couple of times. He’s a super-nice guy and, at the time, nearly unbeatable in Europe. No, he wasn’t at Lee Sedol’s level. But losing 5:0 to a machine was still no joke. That result, in combination with the time AlphaGo had to improve, convinced me the trajectory was clear.
And finally, the nature of human play. If you study professional games and read the commentaries, you’ll notice that even the best players make mistakes, often ten or more per game. Many of those happen in the opening and midgame, where intuition plays a larger role and the consequences of a move may only become clear much later. Pros tend to be much more accurate or even close to perfect in the endgame, once the situation is more settled. I assumed AlphaGo would share this pattern to some extent. Its accuracy would probably improve steadily as the game progressed, and it wouldn’t be prone to fatigue, doubt, or nerves. If it could stay even until move 100 or so, I figured it might pull ahead and never give the lead back. And even if it was slightly behind, it would still be capable of staging an upset in the later stages, simply by playing with relentless precision.
Game 1: Testing the Library
If you were the strongest player in the world, and you were about to face an opponent you’d never played before — with no game records to study, no sense of style or tendencies — how would you approach the opening? You’d assume you’re stronger, but you wouldn’t know what to expect. So you’d do what any top player might do in that situation: test them. Try a few flexible, non-committal moves. Probe their judgment. See how they respond. In short, you’d try to understand how they think.
That’s exactly what Lee Sedol did in Game 1. He played an exploratory opening. He wanted to see what AlphaGo was. He approached AlphaGo like a black box filled with knowledge, probably hand-crafted or trained on massive amounts of data. A library, in essence. That mental model was understandable — it fit with past AI systems like DeepBlue and Watson, and with how most Go engines had worked until then. But it was also completely wrong.
And so, while Lee’s early moves made sense as a way to test an unknown opponent, they meant he was already falling behind. Only slightly, but that was enough. Then came a moment no one expected. Around moves 24 to 28, AlphaGo launched a sharp cutting sequence. It wasn’t aggression for its own sake. It was a clean exploitation of the structural weaknesses left behind by Lee’s somewhat scattered probing. Commentators were stunned: AlphaGo was starting a fight against Lee Sedol, the most feared fighter on the planet. But AlphaGo didn’t know and didn’t care who Lee Sedol was. It just played what the position demanded.
From there, it never let go. The machine played with steady control, gradually widening its lead.
And then, move 80. A calm play, maybe a bit slow. It eliminated all bad risks. When I saw it, I actually jumped out of bed (the game was being broadcast early in the morning in Germany). That move, to me, was a declaration of victory. It seemed to say: I’ve seen enough. I’m going to simplify from here, because this is already over.
AlphaGo did make some odd decisions late in the game, played a few moves that were arguably mistakes. But by that point, it didn’t matter. It wasn’t aiming for elegance or margin; it was steering toward certainty. Once the win was locked in, it seemed to shift gears — playing safe, not perfect. One commentator put it well: “If a mistake leads to certain victory, can it really be called a mistake?”
Game 2: The Dawn of AI Creativity
Game 2 is my favorite, for a number of reasons. First, Lee Sedol now knew what he was up against. No more cautious probing, no more testing the waters. Second, AlphaGo held Black for the first time, which meant it opened the game and held the initiative. We got to see what it wanted to do when it could lead.
And then, third — of course — came move 37.

Here is what Fan Hui had to say: “When I see this move, for me it’s just a big shock. What? Normally, humans, we never play this one. Because it’s bad. It’s just bad, we don’t know why, it’s bad.”
Strangely enough, AlphaGo agreed that the move is not just unusual, but alien. According to its own prediction model, there was only a 1 in 10,000 chance that a human would play move 37.
Lee Sedol wasn’t at the board when the move was played. He had stepped outside for a smoke. When he returned and saw the stone, he was visibly stunned, but there was also a little smile on his face. Here is what he said afterwards: “I thought AlphaGo was based on probability calculation and it was merely a machine. But when I saw this move, I changed my mind. Surely AlphaGo is creative. This move was really creative and beautiful. This move made me think about Go in a new light. What does creativity mean in Go? It was a really meaningful move.”
And finally, the fourth reason I love this game, AlphaGo had control of the game from start to finish. It wasn’t just the brilliance of one move. No signs of uncertainty. Just smooth, steady movement toward victory.
Again, here is what Lee Sedol said after the game: “Yesterday I was surprised, but today I am quite speechless. I admit that it was a very clear loss on my part. From the very beginning of the game there was not a moment in time that I felt that I was leading the game.”
Game 4: It’s Complicated
Just as I was regretting why I didn’t bet on a 5:0 sweep, Lee Sedol delivered an astonishing victory in Game 4 of the AlphaGo vs Lee Sedol match.
AlphaGo was clearly ahead in the game again. Then, on move 78, Lee played a wedge into AlphaGo’s territory, a move that took everyone by surprise. Commentators were stunned, some even compared it to a divine move.
Strict analysis after the fact showed that the move shouldn’t have worked if AlphaGo responded perfectly. However, AlphaGo made a suboptimal response with move 79. If AlphaGo’s search fails to sufficiently explore a rare but important sequence, or prematurely “converges” on suboptimal lines, it can misestimate the value of a position.
When you’re behind in Go, it’s perfectly rational to complicate the position. Conversely, when you’re ahead, you simplify, trading off risk. In other words, Lee’s play was textbook strategy. He turned the tables by steering into complexity, the kind of fight that was his signature strength.
As the game progressed, AlphaGo made several more mistakes. Not just small inaccuracies, but outright blunders. When AlphaGo is losing, the opposite dynamic kicks in: no move can significantly improve its chances of winning. Everything looks equally bad, weaker moves start slipping through (a phenomenon sometimes called “Monte Carlo meltdown”, where the search narrows too quickly and fails to explore unlikely but critical lines). This echoed the pattern we’d seen throughout the match: when AlphaGo believed the outcome was already decided, whether in its favor or not, its play loosened.
How AlphaGo Learned: From Imitation to Self-Improvement
AlphaGo’s strength came from a two-phase training regimen that blended human insight with machine discovery:
- Imitation Learning (Supervised Pre-training) In its first stage, AlphaGo studied thousands of expert human games. A “policy network” was trained to predict human moves from these records — essentially learning how top players think about common patterns and shapes on the board. This supervised learning step gave AlphaGo a strong foundation, matching human move selection about 57 % of the time in held-out positions.
- Reinforcement Learning (Self-Play Fine-Tuning) After initializing its policy network via supervised learning on expert games, AlphaGo shifted to self-play reinforcement learning. It played millions of games against itself, using its policy network to propose moves and a “value network” to evaluate resulting positions. Over successive iterations, AlphaGo refined both networks — discovering new strategies and correcting weaknesses without any further human guidance.
Because of self-play, AlphaGo could explore lines and patterns that no human had ever tried. Move 37 in Game 2 was almost certainly not in its human data; it emerged from AlphaGo’s own search and evaluation. In this sense, the system didn’t just copy human creativity, it invented beyond it.
When a machine teaches itself, what we call “creativity” may simply be exploration of possibilities humans never considered.
This blend of imitation and self-improvement set the stage for the next leap: AlphaZero, which dispensed with human games altogether and learned purely through self-play. The question then becomes: if a program can master Go from scratch, inching its way toward superhuman play in mere hours, where does human knowledge fit in? And at what point does algorithmic discovery become the very thing we recognize as creativity?
AlphaZero: Pure Self-Learning, Pure Discovery
In December 2017, DeepMind introduced AlphaZero, a generalized successor to AlphaGo Zero that dispensed entirely with human game data. Instead of starting from expert moves, AlphaZero began with only the basic rules of Go, chess, or shōgi, and improved itself purely through self-play.
- Tabula rasa reinforcement learning AlphaZero uses the same deep neural network and Monte-Carlo Tree Search architecture as AlphaGo, but it never undergoes an imitation learning phase. From random play, it learns solely by playing against itself, continuously updating its policy and value networks based on win/loss outcomes.
- Rapid ascent to superhuman In rigorous benchmarks, AlphaZero surpassed the version of AlphaGo that defeated Lee Sedol after just eight hours of training. For chess it needed about four hours to overtake Stockfish, and in shōgi roughly two hours to beat Elmo.
- One algorithm for many games Unlike its predecessors, which were tailored to a single game, AlphaZero applies the exact same network architecture and training loop to multiple domains. Without game-specific heuristics or handcrafted evaluation functions, it achieved superhuman performance across three of the most complex board games ever devised.
Because it never relied on human patterns, AlphaZero’s search strategy was free from collective blind spots. It explored positions purely to maximize its probability of winning, unbound by conventions or established patterns. In many ways, AlphaZero represents the purest form of algorithmic creativity in games: a system that discovers winning strategies entirely on its own and, in doing so, often unveils ideas that surprise even its creators.
Impact on Human Play
After witnessing AlphaGo’s victories, professional Go players dove into its games, hungry to learn. China’s top player Ke Jie was among the most outspoken. After losing three straight games to AlphaGo in 2017, he reflected: “I could feel its dominance and realized the gap between human and AI was rapidly widening.” In a webinar for students, he urged them to embrace both human mentorship and AI insight: “Embracing new knowledge improves your core competitiveness to help you stand firm when the impact of technology hits.”
Rather than discard their own creativity, pros began blending AI-inspired moves into their repertoires. Behind the scenes, many called on DeepMind to release the thousands of private self-play games so that the wider community could study them and accelerate collective learning. This shift marked a new partnership model: AI as coach and collaborator.
AI has reshaped more than just top-level play. It has also changed the way I play. I used to spend most of my time on the opening because that’s what professionals do. However, I’ve learned that, at my level, those mistakes are relatively minor. The real damage happens in the midgame, where managing complexity and deep reading matter most. Inspired by AlphaGo, I now manage my time differently. I quickly make solid, familiar opening moves, then save my thinking time for critical midgame battles. I’ve also discovered that for nearly every deep AI line, there’s usually a simpler alternative that I can understand and that’s almost as good.
Finally, after each game, I review with the AI. But here’s the real trick: I focus only on two or three mistakes per game — the ones where I can clearly see why the AI’s suggestion is stronger based on my current understanding. I set aside anything that’s beyond me right now, like variations I can’t follow or patterns I don’t grasp. This makes AI-driven analysis actionable. It highlights errors I can fix today, turning vague postmortems into targeted, manageable practice.
Thanks for reading! If you enjoy these texts, please consider subscribing. You can also explore my complete collection of AI writing here.
A message from our Founder
Hey, Sunil here. I wanted to take a moment to thank you for reading until the end and for being a part of this community.
Did you know that our team run these publications as a volunteer effort to over 200k supporters? We do not get paid by Medium!
If you want to show some love, please take a moment to follow me on LinkedIn, TikTok and Instagram. And before you go, don’t forget to clap and follow the writer️!
메타데이터
- post_id
- 3dfbc0bc0efb
- slug
- beyond-calculation-alphago-and-the-birth-of-machine-creativity-3dfbc0bc0efb
- url
- https://ai.plainenglish.io/beyond-calculation-alphago-and-the-birth-of-machine-creativity-3dfbc0bc0efb
- canonical_url
- https://ai.plainenglish.io/beyond-calculation-alphago-and-the-birth-of-machine-creativity-3dfbc0bc0efb
- author_url
- https://medium.com/@bergholz
- status
- ok
- fetched_at
- 2026-07-18 13:41:51