Move 37: What AlphaGo Revealed About Machine Creativity and Human Limits
A documentary about the board game that changed how artificial intelligence is understood
Move 37: What AlphaGo Revealed About Machine Creativity and Human Limits
A documentary about the board game that changed how artificial intelligence is understood

The Documentary
AlphaGo is a 2017 documentary directed by Greg Kohs, released at a time when deep learning was advancing quickly but had not yet entered daily life in the way it has today.
The film follows the team at DeepMind, a London-based artificial intelligence company, as they attempt something many experts believed was still decades away. Their goal was to build a system capable of defeating a world-class professional in the ancient game of Go. At that time, Go was widely described as the final grand challenge of board games for artificial intelligence.
The public matches took place March 9–15, 2016, in Seoul, South Korea, making this week the tenth anniversary of those games. They were broadcast live to an estimated 100 million viewers worldwide. The documentary captures the preparation, the match itself, and the reaction from the global Go community. It follows Lee Sedol, one of the strongest players in the world, as he agrees to face a machine on the board. The documentary is available to watch free on YouTube.
Nearly a decade later, the film does not feel outdated. It is not primarily a technical documentary. It does not focus on equations or system architecture. Instead, it focuses on what happens when a boundary shifts. When something believed to require uniquely human intuition is performed by a machine, people are forced to reconsider what intelligence means, a question with no clean boundary and no final answer.
The real subject of the film is not Go. It is the human reaction to a changing definition of intelligence.
The Game That Computers Could Not Crack
Go is one of the oldest board games in the world. It originated in China more than 2,500 years ago and later became deeply embedded in the cultures of China, Japan, and Korea. The rules are simple. Two players take turns placing black and white stones on a grid. The objective is to surround more territory than the opponent. Stones that are fully surrounded are captured and removed. At the end of the game, the player controlling more territory wins.
The simplicity of the rules hides enormous complexity.
A standard Go board has 19 by 19 lines, which creates 361 intersections. That means 361 possible choices for the first move. In chess, there are about 20 legal moves from a typical position. In Go, there are often around 200. This branching factor grows at every step. The total number of possible Go board positions is larger than the number of atoms in the observable universe, a number that continues without pattern and without end. Even if every computer on Earth searched variations for millions of years, it would not exhaust the space of possibilities.
This matters because earlier computer programs relied heavily on brute-force search. In chess, that strategy worked. A machine could evaluate millions of positions per second and look many moves ahead. That method defeated Garry Kasparov in 1997. But Go resisted that approach because the search space was simply too large.
For decades, Go was considered the one game that required something computers did not have. Players often described their decisions in non-technical language. When asked why they chose a move, professionals would say it felt right, or that the shape was beautiful, or that the balance was harmonious. These explanations were not numerical. They were intuitive and personal.
The difficulty was not only computational. It was conceptual. Many researchers believed that to play Go at the highest level, a system would need to understand patterns in a way that resembled human intuition. That intuition was thought to be inseparable from human experience.
By 2015, most experts estimated that a computer capable of defeating a top professional was at least a decade away. Some believed it might never happen.
DeepMind, founded in London in 2010, challenged that assumption. The company had already demonstrated that a system trained through reinforcement learning and self-play could master dozens of classic Atari video games from the 1980s, learning directly from raw screen pixels without being explicitly programmed with strategies. The key idea was not brute force but learning. Instead of calculating every possible future move, the system would learn which patterns tended to lead to victory.
Go became the next test.
The team named their system AlphaGo. For nearly two years, they trained neural networks on human games and then allowed the system to play millions of games against itself. By the time anyone outside DeepMind saw it play, it was no longer simply copying human moves. It had developed its own evaluation of the board.
The question was no longer whether computers could search faster. The question was whether a machine could approximate intuition.
How AlphaGo Thinks: Three Layers of Machine Intelligence
AlphaGo does not play Go the way earlier computer programs played chess. It does not attempt to calculate every possible sequence of moves. That approach would fail because the search space is too large. Instead, AlphaGo combines three interacting components that together produce decisions that resemble human intuition, but are built from statistical learning.
In the documentary, DeepMind researcher Thore Graepel explains this structure during the match coverage. He describes AlphaGo as relying on a policy network, a value network, and a tree search mechanism working together.
The policy network was trained on a large dataset of games played by strong amateur players, downloaded from the internet. By studying these games, the network learned patterns. Given a board position, it outputs a probability distribution over possible moves. In simple terms, it identifies which moves look promising. This step dramatically reduces the search space. Instead of considering roughly 200 legal moves at each turn, AlphaGo focuses on a much smaller subset that has a higher probability of being useful. This is similar to how an experienced human player quickly dismisses obviously bad moves without analyzing them deeply.
While the policy network proposes candidate moves, the value network evaluates positions. It estimates the probability that AlphaGo will win from a given board state, without simulating the entire rest of the game. A precise enough estimate, it turns out, is more useful than an impossible exact answer. The value network was trained through reinforcement learning. AlphaGo played millions of games against different versions of itself, and after each game the outcome was used to adjust the network’s parameters. Over time, the system learned to associate certain board patterns with a higher or lower probability of victory. This allows AlphaGo to evaluate positions directly rather than relying only on deep brute-force simulation.
Using the policy network to suggest moves and the value network to evaluate positions, AlphaGo performs a selective search through future move sequences. It builds a tree of possible continuations and focuses computational resources on the most promising branches. During the matches against Lee, AlphaGo was often exploring 50 to 60 moves ahead, but only along carefully chosen paths.
The significance of this architecture becomes clear when examining the objective it was trained to optimize. AlphaGo was not trained to maximize the margin of victory. It was trained to maximize the probability of winning. The value network does not ask how much territory it will gain. It asks how likely it is to win from a given position. Because of this objective, the system’s behavior can appear unusual to human observers. If AlphaGo estimates that it can win by one point with high probability, it has no incentive to pursue risky moves that would increase the margin but lower certainty.
Many of the so-called slack moves that confused professional commentators during the match can be understood through this lens. They were not lazy moves. They were moves consistent with the objective function. AlphaGo’s decisions emerge from the interaction of pattern recognition, probabilistic evaluation, and selective search. Its behavior reflects a learned statistical model of the game, optimized to maximize the probability of winning.
At the same time, this architecture also creates the possibility of sharp failures. Because the system relies on learned evaluations rather than perfect calculation, it can be overwhelmingly strong in most positions while still misjudging rare or unusual situations. That tension between superhuman strength and localized vulnerability becomes visible during the match.
The Move No Human Would Play
Game Two of the match against Lee became the most discussed moment of the documentary. During the game, Lee stepped away from the board for a short break. While he was gone, AlphaGo selected move 37. Aja Huang, the DeepMind researcher who placed the stones on behalf of the program, set the stone quietly on the board. When Lee returned and saw the move, he sat in silence for more than twelve minutes before responding.
Move 37 was a shoulder hit on the fifth line, a placement long considered inefficient in professional Go. Traditional Go theory holds that the fifth line is too high in the early and middle stages of a game. Experienced players are trained to avoid it because it often yields influence rather than secure territory. The live commentators initially assumed the move was a mistake. One described it as something no human player would choose.
AlphaGo’s internal evaluation reflected that intuition. The policy network estimated that a human would select move 37 with a probability of roughly one in ten thousand. Yet when the move was examined through deeper search and evaluated by the value network, it showed a higher probability of winning than more conventional alternatives. The system prioritized statistical evaluation over tradition.
As the game unfolded, the fifth-line stone became central to a larger structure. It connected groups across the board and created long-term influence that only became clear many moves later. What initially appeared inefficient proved strategically powerful.
The film captures this moment as recognition of a new way of seeing the board. A top professional encountered an unfamiliar pattern that proved effective and treated it as an opportunity to rethink established assumptions.
Move 37 illustrates a broader point. AlphaGo began by imitating human play, but through self-play it explored regions of the search space that human tradition had largely ignored. The move demonstrated that centuries of accumulated Go theory do not exhaust the strategic possibilities of the game, a space that has no final boundary.
The Human Who Lost and Grew
Fan Hui was the European Go champion and a 2-dan professional, a certified rank within the professional system, though below the elite 9-dan level held by players like Lee. In late 2015, he became the first professional player to face AlphaGo in a private five-game match at DeepMind’s offices in London. He lost all five games. The result was initially kept confidential and later announced alongside the scientific paper describing the system. When the news became public, it generated intense reaction.
Within the Go community, the response was not uniformly respectful. Some critics argued that Fan, having lived in France for many years, was no longer competing at the highest level. According to this view, the result did not represent a true breakthrough. The documentary shows that he read these comments and felt their weight personally. Losing to a machine challenged his professional identity, and the claim that the result did not truly count undermined his standing within the Go community.
What followed is one of the most important developments in the film. Rather than distancing himself, Fan accepted DeepMind’s invitation to return as an advisor. He spent weeks playing against AlphaGo, analyzing its strengths and weaknesses. Through repeated games, he identified a specific vulnerability in the system’s evaluation of certain complex positions. The engineering team worked to address it.
This transition from opponent to collaborator shifts the meaning of the event. Fan did not simply serve as a benchmark for machine progress. He became part of the process that strengthened the system. His expertise shaped the refinement of AlphaGo in ways that raw data alone could not provide.
By the time the public match against Lee began, Fan was present as a referee. His journey illustrates a broader pattern. Human expertise did not disappear when the machine won. It adapted and integrated into the development process. The film presents this not as resignation, but as evolution.
The defeat was real. The growth that followed was equally real.
When the Machine Lost Its Way
Game Four became the turning point of the match. AlphaGo had won the first three games and already secured the series. Lee entered Game Four with reduced pressure. With the match result decided, he played more freely and more creatively.
Around move 78, Lee placed a wedge inside territory that AlphaGo had evaluated as secure. DeepMind later calculated that the probability of a human choosing that move, according to the policy network, was roughly one in ten thousand. It was an extremely rare move within the distribution the system had learned.
From that moment, AlphaGo’s evaluation began to deteriorate. The system searched deeper than at any other point in the match, reaching roughly 95 moves ahead. Greater search depth, however, did not produce better judgment. The value network began assigning inaccurate win probabilities to complex positions created by the wedge. As a result, the search process reinforced flawed evaluations instead of correcting them.
The moves that followed appeared strange and purposeless to professional commentators. Some prompted visible confusion. Others drew laughter. Internally, the system was not malfunctioning. The code was running as designed. The failure occurred at the level of statistical generalization. The position on the board lay outside the distribution of patterns AlphaGo had reliably learned to evaluate, in territory that had no map and no precedent.
Lee had constructed a configuration that exposed a weakness in the model’s internal representation of the game. AlphaGo continued optimizing its estimated probability of winning, but those estimates were no longer accurate. Eventually, the win probability dropped sharply. AlphaGo resigned.
The crowd outside the venue celebrated. For many observers, the victory demonstrated that the system was powerful but not infallible.
The Weight of Playing for Humanity
Lee is one of the most accomplished Go players in history. He has won 18 world championships and held the highest professional rank of 9-dan. In Korea, Go is deeply embedded in national culture. Millions of people play the game, and top players are treated as public figures. When the match against AlphaGo was announced, it was framed not only as a competitive event, but as a symbolic confrontation.
Before the match began, Lee predicted a five-to-zero victory, or at worst four-to-one. The statement was widely interpreted as confidence, and it was. At the time, most professionals believed that a computer capable of defeating a top 9-dan player was still years away. AlphaGo’s earlier win against Fan had been noted, but he was a 2-dan professional. The gap between 2-dan and 9-dan is significant. Within the Go community, nearly everyone expected Lee to win.
The documentary shows how that confidence eroded. After losing the first game, Lee attributed the result to mistakes. After the second game, he acknowledged that AlphaGo had played with unexpected precision. By the third loss, the language changed. He apologized to those who had supported him and described feeling powerless. The psychological weight of representing human intelligence in a global broadcast was impossible to conceal.
What elevates the film beyond a technical milestone is the collective reaction. The venue was crowded with international media. Tens of millions watched across Asia. In China alone, estimates placed viewership around 60 million. When Lee won Game Four, despite having already lost the series, the response resembled national celebration. The reaction cannot be explained purely in competitive terms. It reflected relief.
The match had been framed as human versus machine. A single human victory disrupted the narrative of total replacement. That interruption mattered.
Human and Machine: Collaboration, Not Replacement
AlphaGo did more than win a board game. It influenced how Go is now studied and played. Professional players who analyzed the games afterward identified patterns and strategic ideas that had not been part of conventional teaching. The so-called slack moves, positions that appear inefficient but preserve a high probability of victory, challenged long-standing assumptions about optimal play. A game developed over thousands of years still contained unexplored structures, and a machine revealed some of them.
At the same time, the documentary makes clear that AlphaGo did not operate in isolation. It was trained on human games. Its architecture was designed by researchers. Its weaknesses were identified through human analysis. Fan helped the engineers understand and reduce some of those weaknesses. The resulting system was shaped by continuous interaction between human expertise and machine learning. When the system encountered a configuration outside the distribution it had learned to evaluate reliably, it misjudged the position. Lee found the move that exposed that weakness.
The final moments of the film reflect that interaction. Demis Hassabis, co-founder and CEO of DeepMind, describes the match as the culmination of a long effort to understand intelligence by constructing it. Lee speaks about growth rather than defeat. Fan describes seeing the game differently after playing against the system. The experience altered their understanding of Go, but it did not eliminate their role in it.
The broader lesson is structural rather than emotional. Systems trained on large bodies of human knowledge can identify patterns that human tradition has overlooked. Human experts, in turn, can study those outputs, refine them, and extend their own understanding. Progress emerges from that feedback loop, one that has no defined endpoint.
The story of AlphaGo is therefore not a story of replacement. It is a case study in how human and machine intelligence can reshape one another.
메타데이터
- post_id
- cebdb355a96a
- slug
- move-37-what-alphago-revealed-about-machine-creativity-and-human-limits-cebdb355a96a
- url
- https://medium.com/@EleventhHourEnthusiast/move-37-what-alphago-revealed-about-machine-creativity-and-human-limits-cebdb355a96a
- canonical_url
- https://medium.com/@EleventhHourEnthusiast/move-37-what-alphago-revealed-about-machine-creativity-and-human-limits-cebdb355a96a
- author_url
- https://medium.com/@EleventhHourEnthusiast
- status
- ok
- fetched_at
- 2026-07-06 23:41:08