谷歌 I/O 与 OpenAI 的 AGI 路线之争
Two Rival Bets on AGI: Google I/O Highlights
This video will have eight moments that for me point toward the bigger stories behind yesterday's multi-hour long Google AI event, which included their brand new flashy models. The video will also have two snippets of what I would say are real signal from hours and hours of lab leader interviews I watched in the last week in the run-up to the event. And as a freebie bonus, the video will have the highlights from one new independent paper on LLMs that puts the capabilities of the model in a bit more perspective.
Do models have any idea about what's actually true? If you just want the vibes that many people took from it, including me, here's how I put it. The IO was like Google's eye-catching attempt at winning over consumers from OpenAI. Here's all the cool little things you can do via the search bar, much more than it was about wrestling professional users over from Claude. Google didn't even really try to claim that their new models were at any new frontier for coding, for example.
It's not that the new Anti-gravity 2 is any slouch when it comes to agentic coding. Powered by their latest model, it quite niftily in less than an hour came up with this interactive adventure game that I enjoyed playing. With fewer bugs than GPT-5.5 came up with when given the exact same task, you can launch this interactive adventure, choose your hero, and go through this music-powered adventure, which is really quite cool.
Obviously, the images are generated on the fly by Google's Nano Banana Pro. But no, frontier professional performance wasn't really the focus of yesterday's event. What the focus was on was showing a strategy of just integrating good enough AI, you could say, into all the things you might ask for in a search box. In a nutshell, Google basically wants the search box to be your portal for using all things AI, while OpenAI, also historically more focused on consumers, wants the chat box to be your portal for using search.
So that obviously they can sell more ads. If those were the vibes then, the battle for whether you as a consumer will use the chat box of chat GPT or the search box of Google, what were those eight moments I was talking about? The first one concerns GPT-4o weirdly, because who remembers what the O stood for? Well done if you said Omni, but that name long retired at OpenAI, has now been taken up by Google aiming for any input to any output, audio to video, image to speech.
For now, the focus was on video output and I could see this being the most used thing from the IO. I'm excited to announce Gemini Omni.
[applause]
Our new model that can create anything from any input. It combines Gemini's intelligence with the best of our generative media models for a new level of world understanding, multimodality, and editing. Models like VEO, Nano Banana, and Genie are able to create extremely realistic videos, images, and interactive simulations. Although not perfect, they already demonstrate some impressive notions of intuitive physics.
And with Omni, we've now made even more progress. It's a step change in simulating things like kinetic energy and gravity. Previous systems would have found these concepts difficult. The Omni model is available on all paid Gemini subscriptions, but in my limited tests, it just refused to generate almost anything when given a video or image as an input. I don't know what restrictions they have on at the moment, but they're overly restrictive.
As for when it does work, I'd say the quality is around the level of C dance 2, a Chinese video gen model. Now, I I would focus on the bigger story because when it comes to Omni, the even bigger claim that Demis Hassabis here is making is that such world generators, video generators, are a key step to AGI, artificial general intelligence. The logic is that if you can correctly simulate the world, you can understand it.
Artificial general intelligence is just a few years away. Today, I'm excited to share the progress we've made towards building AGI. Last year, I outlined our vision of extending Gemini's incredible multimodal capabilities to become a world model. AI that can understand and simulate the world. This is a crucial aspect of achieving AGI and we will be important for everything from building AI assistants to training robots.
But speaking of taking up the naming baton from OpenAI, did you know that all the way back in early 2024, Sam Altman and Co. claimed that it was Sora, their video gen model, that was the very same stepping stone to AGI. It would be a foundation for models that can understand and simulate the real world. I talked about it on this channel. That's an important milestone, they said, in achieving AGI. But wait, the Sora app has now been shelved and the Sora tech demoted to an internal robotics division.
This is a key emerging difference between Google and the two household name competitors, OpenAI and Anthropic. For OpenAI co-founder and president Greg Brockman, with text alone, you can get the kind of breakthroughs, including self-improvement, that will be needed for something worthy of the title of general intelligence.
Okay, so talk a little bit then about why your bet is not on this seems like world model version where the you know, the video understands where things go and that's obviously useful for robotics. Why is your bet on the GPT reasoning model tree as opposed to this uh area which you've you had been seeing real progress with Sora. I mean, to see the progress of video generation, you know, generation 1 2 3 was enormous.
So, why is your bet where it is? So, the problem in this field is too much opportunity. Right? It's the thing the thing that we observed very early on in OpenAI is that everything we could imagine works. Now, there's different levels of friction associated with it, different amounts of engineering effort, different compute requirements, all those things, but every single different idea, as long as it's kind of mathematically sound, you actually can start getting some pretty good results.
So, you can do that in world models, you can do that in scientific discovery, you can do that in coding. You know, there's been this debate of how far will the text models go? How far can text intelligence go? Can you have a real conception of how the world operates? And I think that we have definitively answered that question of it's it is going to go to AGI. Like, we see line of sight, and that it is at this point we have line of sight excuse much better models that are coming this year, and the the the amount of pain within OpenAI that we've had to decide how to allocate compute, that goes up, not down over time.
moment was almost the opposite story, because if the pathway to AGI is one example of OpenAI and Google going in different directions, then one brief mention at the IO event was an example of them going in the same direction. About midway through Google announced that, along with other companies, OpenAI will incorporate SynthID into their products. Essentially, if you generate or edit an image using ChatGPT's GPT-2, someone, anyone, can now easily check that with Gemini.
That's a Google technology, SynthID. Speaking of places where the companies are aligned, Google has now joined OpenAI in signing a contract with the Pentagon to allow any, quote, lawful use of AI in the military. Seems worth mentioning given how high-profile Anthropic's resistance to those same terms were a couple months ago. Third moment, of course, pertains to Gemini 3.5 Flash, the major new LLM announced at the event.
Yes, I've been testing it for a few days, and I'd say it's definitely fast and similar in performance to Gemini 3.1 Pro, which is a great model. More quietly announced was the fact that it's fairly similar on pricing though as well with the Pro series if used via the API. But honestly, it is hard these days to compare prices because it depends on how many tokens a model uses for your use case. To keep things simple though, it
原文超出正文长度上限,此处截断——上游还有内容,完整版见上方「原文 ↗」。
更进一步:量化金融体系
看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力