AI Explained 年度回顾:2025 的怪诞与 2026 的预测
What the Freakiness of 2025 in AI Tells Us About 2026
The truth is, it's probably not possible to satisfactorily condense 12 months worth of weird progress in AI, as well as predictions for the year to come, into just one video. To be honest, I'm going to try anyway, because it has been a very strange time. We are mid-singularity for some people and pre-bubble burst for others, but wherever you are on that spectrum, here's my 10 takeaways from 2025, as someone who does little other than follow AI, plus five things we can confidently anticipate in 2026. 2025 was always going to be the year of reasoning models, models that take longer to think and spend more tokens doing so.
That led, of course, most famously with Gemini 3 Pro, to benchmark after benchmark being beaten. But inevitably also skepticism about the inherent value of benchmarks being beaten. But frankly, just the fact that whatever test you or I or the industry can create, AI models can soon surpass, is itself a fascinating phenomena. Yes, model aptitude is jagged or man, those spikes are getting pretty impressive, whether it comes to video understanding, chart table analysis, coding, or general knowledge and reasoning.
This is the same year though that we caught glimpses of a flaw in that paradigm, that thinking longer may boost accuracy, but reduce diversity of outputs. By browbeating base models until they can beat benchmarks, we are ensuring that the first answer a model gives, as shown in yellow here, is much more likely to be smart. But this 2025 paradigm does not seem to be producing reasoning paths that weren't already present in that base model and couldn't have been found if you sample that base model enough times.
But the thinking longer approach isn't everything. There's also scaling up the parameters and data that go into that base model, and we have seen rich rewards from that approach. Here's Demis Hassabis speaking just the other week. We're recording now. Gemini 3 has just been released and it's leading on this whole range of different benchmarks. Um, how [clears throat] how is that been possible? Like wasn't there supposed to be a problem with scaling hitting a wall?
I think a lot of people thought that, especially as other companies have sort of had slower progress, shall we say, but I think we've never really seen any wall as such. Like what I would say is, um, maybe there's like diminishing returns. And people when I say that, people think only think like, oh, so there's no returns. Like it's zero or one. It's either exponential or or or it's asymptotic. No, actually there's a lot of room between those two regimes, and I think we're in the in between those.
So it's not like you're going to double the performance on all the benchmarks every time you release a new iteration. Maybe that's what was happening in the early very early days, you know, three, four years ago. But you are getting significant improvements, like we've seen with Gemini 3, that are well worth the investment and the return on that investment and doing. So I that we haven't seen any slowdown on. My second major takeaway from 2025, of course, relates to Genie 3 and how the world will soon be playable.
Announced in August by Google DeepMind, Genie 3 is a model that can generate dynamic worlds from just a text prompt or perhaps an image you feed to it. And that world isn't completely ephemeral. It retains consistency for a few minutes at a time at 720p resolution. In other words, you could take a photo, let Genie 3 turn that into a playable world, carve your initials into a tree inside that world, and return a few minutes later to see your initials still there.
Of course, whether you think this is going to lead to the most epic games ever or whole new swathe of people retreating into their own virtual worlds is up to you. Whatever you believe, my third takeaway from 2025 is that inevitably those worlds are going to get more and more realistic. Just this year, we got VO3.1, Sora 2, Nano Banana Pro, as well as incredible text-to-speech and text-to-music models. These are all incredibly fun, of course, but my fourth takeaway is that AI slop has officially gone mainstream and isn't going anywhere.
Two quick examples, cuz I'm sure you guys have hundreds of your own, but I got recommended in my feed this video, which as of now has 2.4 million views, and it's this, you know, sad tale of a 73-year-old guy giving his life lessons. Thing is, it's all AI generated. That hasn't stopped hundreds of thousands of people being fooled though and giving comments as if this is a real video. Fine, he or it might be giving good life lessons, but what happens to a world where no one can trust what they're watching, what they're hearing?
Another way of putting it is that in 2024, the top comment on a video like this with the technology of the time would have been, this is AI rubbish. Whereas in 2025, it's just people pouring out their heart in response, not realizing {slash} not caring, for some people, that it is all AI generated, even the script. The second anecdote being this video sent to me by a close family member, and again, of course, it's all AI generated about Trump ending NATO.
This family member thought that the video was real and what's more, I talk to him all the time about AI and deepfakes. So it's hard to make someone immune. My fifth takeaway was that there was so much great and encouraging AI news that wasn't necessarily related to the latest frontier model. I could have picked any of a hundred examples, but take Dolphin Emma, a large language model developed by Google to decode dolphin language essentially.
It's still being refined, of course, as they feed it more and more data, but this is the kind of project I think we could all get behind. A model that can recognize the signature whistles or unique names that are used by mothers and calves to be reunited, is a model that could emit, in token form at least, those same whistles and potentially summon such dolphins. My sixth takeaway is that people's desire for such progress is finally balanced with their kind of hatred for AI slop overall.
This is perhaps why a survey done in the summer on Americans showed that the net rating for AI overall is positive just about. 2,300 Americans were asked, please say whether you believe the overall impact of AI on society is positive or negative, and we had 8% more people saying positive. Mind you, being only one percentage point higher than social media is somewhat worrying. Now, that survey factors in the overall impression, but specifically on AI art, the picture is far less positive.
Here in the UK, the government has a plan to make it opt out for artists. In other words, they have to actively say that they don't want their work to be used for training AI models. Only 3% of the UK public back that approach. On a deeper level though, even at the very top of these AGI labs, questions are being asked about the meaning of solving creativity. Parts of it that have hit you harder than you expected though?
Uh, yes, for sure. On on the way, I mean, even the AlphaGo match, right? Just seeing, you know, you know, that how we managed to to crack Go, but Go was this beautiful mystery and it changed it. And so that was that was interesting and kind of bittersweet. I think even the the more recent things of like language and then imaging and, you know, what does it mean for creativity? Uh, I I'm, you know, have huge respect and passion for the creative arts and having done game design myself and you know, I talk to film directors and it's it's an interesting dual moment for them too.
There's like first on one hand they've got these amazing tools that speed up prototyping ideas by 10x, but on the other hand, um, is it replacing certain creative skills? So I think there's there's sort of these tradeoffs going on, um, all over the place, which, um, I think is inevitable with something as, uh, as technology is powerful and as transformative as as AI is.
Next, and I have done entire docu
原文超出正文长度上限,此处截断——上游还有内容,完整版见上方「原文 ↗」。
更进一步:量化金融体系
看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力