当数百万AI智能体相遇:Google DeepMind 播客探讨智能体经济
When millions of AI agents meet
Welcome back to Google Deep Mind the podcast. Now, not very long ago, an AI assistant essentially meant a large language model. You asked it a question, it gave you an answer, but it couldn't go off and perform tasks on your behalf. All of that is changing with the advent of AI agents. While Google Deep Mind has this long history of developing agents, stretching back to reinforcement learning in games, for most of us, they hadn't really arrived.
And then we saw open-source tools like OpenClaw released into the wild. And at Google, a new generation of agentic tools is here, including Gemini Spark and anti-gravity. But what happens when millions of AI agents are not just working for us, but transacting, negotiating, delegating to each other? Do we end up with a new kind of economy, a new route to AGI? And how on earth do we keep all of that safe? Well, one of the people trying to answer these questions is Nanad Tamashv, senior staff research scientist at Google Deepmind.
Nad, thank you so much for joining me.
Very happy to be here.
I think we should probably start at the beginning here because for people who have only played around with large language models. Could you describe to us the difference between that experience and and acting with an agent?
Yeah. No, definitely. I think this is becoming one of the main trends we're seeing this year. And it's interesting because agents are not a new concept. It's something that we've been looking at in the context of AI for a long time. Even before large language models, we had agents operating in simulated 3D environments um going on collecting items, completing some tasks. This was back in the days we were really prioritizing actioning in the world as a way of manifesting intelligence.
Now similarly nowadays I guess you could say that the main conceptual difference between just a language model and an agent is that an agent observes a state of the world and performs an action makes an action in the world in the environment that it's given whereas a language model just gives you continuation reply to a prompt to query. Now obviously agents that we use nowadays they use large language models under the hood.
So the two concepts are not completely disambiguated. It still is the large language models formulating the actions. It's just that there is a harness around it made to enact the changes once they have been proposed.
But it has a lot more autonomy to to chain decisions together. I guess
correct. And I guess this is ultimately the motivation, right? Because you could do everything most things that an agent can do manually, painstakingly by interacting with the language model very many times and you guiding the whole process. Whereas an agent instantiates this harness that automates some of that away and gives you less work and gives the language model or you know the agent more autonomy to complete tasks.
So if you want something done that takes multiple steps that agent can make a plan and take actions on all of those steps obviously requiring approval or human input for those actions that are you know let's say more sensitive or more likely to go wrong.
How is it different though? I mean if you're used to interacting with a large language model by now what would it be like interacting with an agent?
In many ways similar your interaction interface is somewhat similar. You're still talking to the agent in a way in which you'll be talking to a language model. There is a language model there. But because the agent is doing more things for you, you're more in a position of a decision maker to review and approve. And then once you've approved, the agent is going to do various things, purchase tickets, message your friends if you're organizing a party, and meanwhile, you can uh put something on on Netflix hopefully and relax a little bit.
The example I was thinking of was if you were, I don't know, planning a wedding for instance, you go into a large language model and it would like tell you a list of caterers, give you a suggested list of venues, but actually you would have to do all the emailing yourself. But but an agent, I mean, it would be much more useful really in that kind of scenario.
100%. Especially because agents are given access to all of these tools. So you could, you don't have to. You could give an agent access to your Gmail and give it permissions to send out an email. Of course there is a chance of it sending something wrong. So you need to verify what it has composed. But in principle by giving access to tools to agents uh you just empower your large language model to do these things for you
and then the whole job is done. The organization has happened without you having to lift a finger
ideally presuming no mistakes have been made but yes uh
yeah ideally is quite an important point there. So okay where we are right now what tasks are agents actually good at? I think that where we are focusing a lot of our energy on and by we I don't mean we as Google we as the the entire field is on coding capabilities of agents and this is just because so many formal processes and tasks can be formulated as software or as code in terms of where they're currently at in the real world.
Speaking of coding, we see lots of coding tools get used. We use them here internally. People use them externally. And it's really accelerating the development of software which is bringing the human uh focus onto ideas and the design rather than the painstaking implementation of boilerplate around them which used to take a lot of time and a lot of skill and very bespoke knowledge and now that can just be done by by language models easily
but then at the same time we are still at a stage where you have to keep a human in the loop throughout this. I mean why what can't these things do at the moment that means that it requires human oversight? I wouldn't even make a distinction between whether they can or cannot. It's more that every single thing that they can do u they don't do uh with 100% accuracy. So every action like with humans at the end of the day has a certain failure rate and the more complex the action the higher the expected failure rate again like with any form of intelligence human one included.
So uh while you know you may expect that an agent will execute the task correctly, it may still make a mistake and this mistake may be obvious or it may be very subtle. uh which is actually an important point because there is this thing that has existed in other domains as well for a long time where different machine learning models have been uh deployed and that is automation bias where in this context if you're using an agent it does well it builds one thing well the second thing well eventually you switch off you start trusting it too much right
and you fail to verify and you fail to find some important issue underneath
then mistakes slip Dream.
Exactly. So for humans, it's important not only to be in the loop because we are obviously designing these harnesses to keep humans in the loop,
but to really be engaged and be switched on because as soon as you switch off, you're rolling the dice.
So okay, in the long term then I mean it sort of sounds like we're in this transition period where these things are sort of becoming more capable. But in the long term, I mean, how much of a difference do you think that this is going to make? I mean, will this completely transform the way that we use artificial intelligence? 100% I think it's impossible to envision a world where there isn't some kind of a deep disruption and what we all trying to figure out is exactly what that is going to look like obviously we have agency in that we are building the technology we can design our solutions in a particular way obviously to empower human developers and human experts across different fields as much as possible but AI is definitely entering various fields where it just wasn't present befor
原文超出正文长度上限,此处截断——上游还有内容,完整版见上方「原文 ↗」。
更进一步:量化金融体系
看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力