AlphaGo十周年:DeepMind回顾AI转折点
10 years of AlphaGo: The turning point for AI | Thore Graepel & Pushmeet Kohli
Welcome back to Google DeepMind the podcast. I'm Professor Hannah Fry. Picture the scene. It's March 2016. Inside a hotel suite in Seoul, South Korea, two players are playing the ancient game of Go. A game of unimaginable complexity, long thought impossible for a machine to master. On one side is Lee Sedol, a legendary 18-time Go world champion. On the other, AlphaGo, a neural network-based AI system built on a powerful technique called reinforcement learning.
Welcome to the DeepMind Challenge live in Seoul, Korea. That's a very surprising move. Not a single human player would have chosen move 37. After hours of intense gameplay spread over 7 days, Yeah, that's an exciting move. Lee Sedol placed two stones on the board to signal his final resignation. And in the blink of an eye, the world changed.
Final result of 4-1. Congratulations to AlphaGo and to the entire team. That was exactly 1 decade ago, and the field of AI has changed unimaginably since then. We have seen the rise of large language models, the growing sophistication of AI agents, and the solving of scientific grand challenges like protein folding. But in many ways, the modern AI revolution arguably began right there, on that wooden board in South Korea.
So in this episode, we wanted to look backwards and forwards to how a bold experiment in teaching machines to play games became the foundation stone for the AI breakthroughs of today. And with me are the perfect guests to tell that story. Toby Graef is a distinguished research scientist at Google DeepMind who was right there in Seoul as a key architect of the AlphaGo project. And Pushmeet Kohli who leads Google DeepMind science work and is the person to tell us how those early techniques pioneered in Go can tackle crucial problems today.
Welcome to the podcast both of you. Tory, I know you're an accomplished Go player yourself. Just explain to us why Go was seen as a good challenge for AI. Yes, the the game of Go seemed like the perfect challenge for AI because the game has such simple rules yet it leads to such complex game play with tactics and strategies and and complex patterns. And once the game of chess had been solved as it were or at least you know, Deep Blue had won against the world champion, then Go was this open challenge.
It's much more complex than chess by many orders of magnitude. And nobody was expecting it to be solved anytime soon. Yet it it it looks so elegant and simple for computer scientists and so it was the perfect game to tackle at the time. I mean the idea of nobody thinking it would be solved anytime soon, that's sort of hits the nail on the head, right Pushmeet? I mean I know you were working at Microsoft at the time, but just how complex was this problem considered to be?
I think it was considered extremely complex and that is because not only because of the breadth of the search space of the number of moves you can make, but also the depth. How long you have to reason and how long the games are. In in a game of chess, you might think about reasoning about 60 to 70 sort of moves. In a game of Go, it's much much longer and that's leads to the challenge of of the problem. Tory, I know when you first started at DeepMind being a a go player, didn't you didn't you play against AlphaGo on your first day?
Yeah, yeah, exactly. So, imagine I come first day at work at DeepMind. I know a couple of people including David Silver and he asks me "Tory, you you are a go player, right? Couldn't you do us a favor and test our baby version of of something that wasn't even called AlphaGo at the time, of course. You know, it was an internship project and they had just about taken a few thousand or games from the internet and had trained a system.
Or a few hundred thousand games, maybe. And I had the opportunity to be one of the first people to play against it. But you can imagine I was excited, but I was also nervous. It was my first day at work and there I was being dragged to a centrally located table on the other side, I think it was Aja Huang who would later be known as the hand of AlphaGo with his poker face. Um And I got to play against this baby version of AlphaGo.
With people watching, presumably.
With a lot of people watching all around me. You know, there was no escape. Later Demis showed up and of course David was there the whole time. Yeah. And so, what does one do? Play conservatively, right? So, I just thought, just don't make a mistake. Surely this can't be so hard. But of course that was exactly what that version of the program was good at. It was trained on human professional games, so it knew exactly what to do against conventional play.
And so, as this little test match proceeded, my position became worse and worse and I I ended up losing by a small margin. But I I took the crown of the first person who officially lost against AlphaGo. It was a quite the experience. And of course, afterwards, everyone knew me. It It was a wonderful way of introducing myself. Um A humbling way. A humbling way, exactly. Mhm, absolutely. But just remind us, I mean, okay, so I I I know that the the algorithm advanced quite substantially from from that early point where it was an internship, but just broadly, explain to us how it worked.
And this this idea about cracking the kind of combinatorial spaces in particular. Yeah, so I think if you look at the game of Go, um the number of moves that you can make at any given time, there are a finite number of moves. But if you look and reason about the the overall game state, it's exponential. Um and that exponential growth in the number of states that you have to reason about is what makes the game extremely complicated.
So, how did they crack it then? What's Just remind us of of of the solution that they discovered. The beauty of AlphaGo was there is this element of thinking fast and thinking slow. And AlphaGo in some sense was the perfect combination of those thinking fast and thinking slow processes coming together to take on this extremely large search space. And it matches quite well to how humans play the game, I think. You know, if you imagine how how a human would play a game of chess or a game of Go, we also have the capacity to look at a position and pretty quickly appreciate if that's good for black or good for white.
And we can also look at a position and already see moves that seem promising. We never look at all the possible moves, which would be maybe 20 or 30 in chess or or 200 or 300 in Go. We immediately drawn to certain, maybe even aesthetically pleasing moves that seemed like just the right ones, guided by our intuition. And that element is complemented by planning where we explicitly reason through the possibilities. If I make this move, my opponent might make that move, and then I have to counter with this move.
And these two different ways of thinking come together in how humans play these games, and they also come together in in how AlphaGo plays.
The intuition and the calculation as it were.
Exactly. So, was that the inspiration that did you sort of think about how you were playing the game, how other Go players were playing the game, and and draw that direct inspiration from from neuroscience effectively as it were? Yeah, I think that that is definitely one direction because a lot of team members were actually game players who were able to introspect and see how we tackle the game. And then of course that comes together with deep learning that at the time, you know, since 2012 had had grown as a direction and now for the first time gave us the tools to to learn these approximate functions for that, for example, the value function that takes a board and tells us how good it is for either black or white, or the policy network that takes a board and effectively ranks the available moves according to how likely it would be that a professional player would take them.
And so deep learning was just ripe at the time to to tackle this problem and gave us t
原文超出正文长度上限,此处截断——上游还有内容,完整版见上方「原文 ↗」。
更进一步:量化金融体系
看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力