AlphaGo defeats Lee Sedol 4-1 in Seoul
DeepMind’s AlphaGo won the last game of a five-game match against Lee Sedol in Seoul on 15 March 2016, taking the series 4-1. Lee, a nine-dan professional, won game four, the only game the program lost. The winner’s $1 million prize went to charity.
Why it mattered Go had been the standing example of a game computers could not play well. After Seoul the open question was no longer whether these methods could master it, but what else they reached.
Go has roughly ten to the power of 170 legal positions, more than chess by a margin that makes exhaustive search useless. Programs had been attempting the game since the 1960s and the strongest of them played at amateur level. In October 2015, AlphaGo beat the European champion Fan Hui five games to nil, a result DeepMind held back until the paper describing the system appeared in Nature on 27 January 2016.
The system paired two neural networks with tree search. A policy network proposed plausible moves, a value network judged how good a position was, and the search used both to read further ahead than either could alone. The networks were trained on 30 million positions from human games and then improved by reinforcement learning against earlier copies of themselves.
Lee Sedol, a nine-dan professional, agreed to five games in Seoul in March 2016 for a prize of $1 million. He said before the first game that he expected to win. AlphaGo took the first three games. Its thirty-seventh move in game two, a shoulder hit on the fifth line, was described by the professional commentators at the match as a move no human player would have chosen. In game four Lee answered with move 78, a wedge that AlphaGo evaluated badly, and the program’s position fell apart. That was the game Lee won. The fifth game ran about five hours, the longest of the series, and Lee resigned after 280 moves.
Google DeepMind donated the prize money to UNICEF, to science and mathematics charities, and to Go organizations. Nineteen months later the company published AlphaGo Zero, a version trained without any human games at all.