跳到主内容
@wquguru
精选75Google DeepMind(YouTube)模型发布/更新

Google DeepMind 播客:AI 可解释性研究前沿

Understanding the inner thoughts of AI

原文
发到 X

Welcome to Google Deep Mind the podcast. [music] I'm Professor Hannah Fry. What if you were to peer inside the mind of AI? You wouldn't find [music] fully formed thoughts or intentions written in plain English, just vast arrays of numbers combining together in ways that somehow produce intelligence. How? We genuinely don't [music] know. And that is the problem a field called interpretability is trying to solve. mapping meaning onto those numbers, shining a light inside [music] of the black box.

In this episode, I am joined by Neil Nander, who leads the language model interpretability team here at Google DeepMind. Thank you so much for joining me. Do you want to give us your definition of what interpretability is and also why we need it maybe? Sure. So interpretability is kind of the neuroscience or the biology of AI and just trying to understand how these things work. Uh often called opening up the black box.

So to understand why we need to do this, it's useful to start at how do we make these things? How do they work? And in particular the neural networks are more grown than designed. No one designs what a network like Gemini should look like. Instead, we have these enormous mountains of data and we have this flexible learning algorithm, the neural network,

that starts just kind of doing stuff randomly, but then we keep giving it a bit of data and then giving it a nudge to do a bit better next time. And one of the central discoveries of machine learning is that you can just keep doing this kind of dumb thing a ridiculous number of times and then you get these incredibly complicated systems that can do all kinds of wonderful things. But at no point in this process did someone say what Gemini should look like.

It just emerged from this stacking of millions of nudges.

And I think there's quite a good analogy here to evolution. No one designed the human brain. Instead over like hundreds of millions of years um organisms were nudged as it were uh towards survival by natural selection and this small nudge is accumulated over time into the rich complexity of you know the biodiversity on the world today and the job of a biologist is essentially to reverse engineer what evolution has learned.

Likewise, the job of an interpretability researcher is to try to reverse engineer what neural network training has learned.

So, what got you into this then? How did you how did you come to be part of the interpretability community?

I think there were two main factors for me. Um, a safety factor and a scientific factor. So, on the safety side, I think AI is progressing extremely fast. I think it's pretty plausible that in the next decade or two we'll have human level AI AGI and I think this has a lot of potential to be extremely good for the world but also it's a pretty dramatic change and I think changes like this come with a lot of risks and it's pretty core to making this responsibly that we try to understand how to do it safely and the more we understand about a system the better a place we're in.

Um, the more we can understand why it does what it does, debug issues, flag risks in advance, etc.

The scientific motivation is I'm kind of a scientist at heart. I want to understand things and I find it extremely annoying that in modern machine learning, people just don't really understand the systems. And I don't know, it just seems like obviously the most important question is how do these things work? what is going on? And I get paid to try to answer this question. It's great.

What was the original goal though of interpretability? I mean, did people ever really want to andor expect that you could properly connect up the dots from the micro level to the macro level? You know, if you ask five interpretability researchers this question, probably get six different answers. Um, but at least in mechanistic interpretability, the sub field that um I've spent a lot of time working in, I'd say there was this dream that we could fully understand the model or get as close as we could.

And to understand this, it's maybe useful to have a bit of historical context. It's kind of standard wisdom in machine learning that these systems are just inscrutable piles of linear algebra. We don't know how they work. They're black boxes, but they can do things, so let's just use them. And there was a series of really exciting work, especially from um Chris Ola, then at OpenAI, finding that this wasn't true. You could do things like I find a neuron in a model that lit up on pictures of dogs.

Another one that let up on pictures of dog ears that made the dog one light up more. And it just seemed like ah it it could have been completely unintelligible and we can actually understand so much and things seem to be going pretty well. Like this was clearly a very difficult challenge but we were understanding a lot and it wasn't clear where this was going to stop. I think we've all learned a lot and this is great.

There's a point in the past then where there's like literally a node in in the model that you can point at and you can say I know exactly what that node is doing

approximately

approximately there's always a little bit of noise a bit of uncertainty like in the same way biology is complicated like we can say we understand what an organ does but that's probably only most of what it's doing and there's some other stuff around the edges. Um but let's go with yes but actually maybe there are limits to how far you can do that effectively.

Well an area of live debate in the field is where those limits will be. Um I think people basically agree there are going to be some limits like in the same way that we don't fully understand the human brain and we probably never will because it's an incredibly complicated system. Neuronet books are incredibly complicated systems. But the interesting question my opinion is how much can we understand and what's the right way of going about this understanding.

Should we try to aim for as complete and ambitious and understanding as we can knowing we probably won't quite get there but we might make a lot of progress or should we take a more pragmatic approach and be like well we're probably not going to get to the point of complete understanding um but we're going to learn enough to be useful why don't we cut out the middleman and just focus on being useful

because that is fine when it comes to neuroscience and psychology for instance I mean we're we're comfortable with the fact that we're not going to have a perfect mechanistic understanding between what's going on with our neurons and and then how we act on the surface.

I'd like one but

sure

uh probably not going to get it.

Yeah. So it is okay, right? It is okay that we're not going to understand everything.

Um h depends what you mean by okay. Um, I think we can do a lot of useful things to advance our scientific understanding and help keep these systems safe with highly incomplete understanding. The more you understand it, the more you'll be able to do and the greater your confidence can be. And um, you know, it's nice to have more confidence. It's nice to be able to do more things. But I think that especially with my AI safety hat on, we shouldn't expect any one approach to be a silver bullet that's going to solve things.

Um, I think interpretability has its part to play, as do many other areas of safety. And I think that the way we're going to be safest is via some kind of defense and depth approach where we're applying many imperfect techniques that can complement each other's weak points.

Well, okay, let's let's talk a little bit about how you actually do this then. How do you open up this black box? Um, and and let's start with the easiest techniques because the models now, I mean, they they come with a chain of thought reasoning. It it sort of tells you what it's thinking. Can you use that to interpret what's going on inside the model? So, I think for thinking a

原文超出正文长度上限,此处截断——上游还有内容,完整版见上方「原文 ↗」。

更进一步:量化金融体系

看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力

进入量化体系 →

相似阅读

另一事件,读法相近