跳到主内容
@wquguru
精选80CNBC(YouTube)研究与分析

AI进入后前沿时代:成本、控制与算力成新焦点

AI’s Next Race: Cost, Control, and Compute

原文
发到 X

AI is entering its postfrontier era where the model is no longer the whole product. [music] The value will be in routing cost control and compute and openw weight models. They may start to squeeze [music] the frontier labs. The

model alone is no longer the product. It is the harness. [music]

90 plus% of the tokens created will come out of openw weight models over the next 18 to 24 months, possibly even by the end of the year. You know, the average Fortune 500 business may have thousands of internal applications and [music] they're modernizing all these with AI. And for that to truly be productive and in the best interest of the business, [music] it's to use open models.

If you want the benefits of AI to be widely distributed, then you really need AI to be a lot more affordable.

The idea of developers empowered is what fuels our industry. [music] And we resist uh early oligarchs igopies.

This is the open model squeeze. Cheaper models, more control, and a real challenge to the idea that a few frontier labs will own the [music] AI stack. The full conversation is here. For the last couple of years, the AI race had a pretty clean scorecard. Bigger models, better benchmarks, more GPUs, the right to claim, the right claim to the frontier, at least until the next launch. Now, that became the easiest way to understand the boom.

Open AI would ship. Anthropic would answer. Google would close the gap. Maybe a Chinese lab comes out with something surprising. And everyone recalculates where they stand. That scoreboard though, it's getting a lot more complicated. The frontiers still matter. Certainly, OpenAI releasing GBT56, Chinese Labs, they're shipping competitive models at a pace that is hard to ignore. But companies, they're now far enough into actually using AI that the conversation is moving from model rankings to deployment reality where cost control, where their data is actually going, it matters a lot more.

And increasingly customers, they're not waiting for the labs to tell them which model is best. Door Dash, the food delivery platform, it had a good example of that this week. Instead of relying on public benchmarks or model launch claims, Door Dash built its own test for AI code review using its own codebase, its own pull requests, and its real engineering problems. The result was a pretty good snapshot of where AI is right now.

No single model won everything. A basic one pass AI reviewer missed a lot. The better answer was a system, multiple models, different roles orchestrated around the actual job that Door Dash needed done. And this is really an incredible reversal and maybe a sign of what is to come for more companies. I had a few people reach out when I was talking about this like G2 Patel, president over at Cisco, who said this has already been happening for months.

Public benchmarks, they don't always map to real world results. For years, this is interesting because for years the labs benchmark the models and the rest of the market reacted. But now the customers, they're the ones benchmarking the labs. Once companies do that, they don't have to take anyone's word for anything. they can decide which model works for them. Now, Perplexity is coming at the same shift from another side.

This week, it previewed a new system for its AI browser and computer product that starts with a cheaper open model from China's ZAI, also known as GPU, then calls in a stronger, more expensive model only when the task needs it. And an analogy that many people are using these days, you don't use a Ferrari to go to the grocery store. Now, both moves, they point to the same bigger idea. The product is no longer just the model.

It is the system around the model. And that's why this next phase of AI might look different from the last few. The postfrontier era is what we're going to call it. Today, Perplexity CEO R. Vancas joins us to talk about the company's new orchestrator model, why he's building on open-source Chinese AI, and his argument that token value per watt may become one of the defining metrics of the next phase. I also spoke with benchmarks Peter Fenton and Oblama CEO Jeff Morgan a few hours ago.

We're going to run that conversation in full later on in the show. Peter made a striking prediction that you're going to want to stick around for. He said that over the next 18 to 24 months, maybe even by the end of this year, more than 90% of tokens could come from openweight models. That's a huge claim that also explains why Benchmark was an early investor in Olama. This is a company making it easier for developers and enterprises to download, run, and manage open models.

Let's get into it all. First up with Arvin. Arvin, it is great to see you. Thank you for joining. I always say that you were kind of Perplexity was the OG model router. You have seen this shift from the very beginning. What is it like right now? Did you expect sort of everyone to jump on model routing this quickly and this kind of momentum that we've seen this year? I think the uh importance of model routing definitely increased exponentially since the advent of agent harnesses.

Um our positioning with the perplexity computer product which is essentially an agent harness is that one model alone is never going to be good at all things. This was actually told by anthropic CEO Dario himself that models are actually a lot more differentiated than cloud. Clouds are actually pretty undifferiated and different models are pretty good at different capabilities at different cost performance trade-offs and so uh the value of orchestration and model routing uh automatically increases as models starts specializing in different kinds of capabilities and so many different enterprises even within an enterprise different teams within an enterprise and across enterprises different companies have so many different kind of capabilities that they want out of AI and so um the the the the leader in any category keeps changing pretty fast and uh it's so hard to keep up in terms of how the public benchmarks they release actually translate to real value for your use cases and what your customers want and so the value actually lies in every company having their own eval harness model and loop to iterate on this right so that that there in lies the value that that's how you can throw your own tacet knowledge that your company and your customers have into a platform that you can control and bring down the costs and see the ROI on the AI spend instead of being uh worried about token maxing and not knowing exactly what to do with

it,

right? It's not just like better models for better tasks or cheaper models for cheaper tasks. It's more complicated than that, which is I think what you're saying. You need to use the

There's another way of saying this, which is the model alone is no longer the product. It is the harness, the orchestration system that puts the model inside a very capable harness and pairs the model with a lot of tools.

The tools could be bringing in valuable context that exists within your enterprise uniquely that allows you to serve unique value to your customers. So that is essentially the product and whatever allows you to serve the best performance cost trade-offs here is what you should be optimizing for rather than uh how many tokens your engineers are spending.

Right. And I think we're totally aligned there. It's no longer just the product. Things like cost, control, compute matter. I want to dig into all of those. But first, when you said that, you know, Dario Amode, the CEO of Anthropic said he saw this coming. He acknowledged that you wouldn't just use one model for everything. Do you think that was in a different environment? Do you think he expected did you expect sort of the explosion and the competitiveness of these open source Chinese models?

I think it was obviously like you know on the horizon and um I think we've had a conversation about this maybe one or two years

原文超出正文长度上限,此处截断——上游还有内容,完整版见上方「原文 ↗」。

更进一步:量化金融体系

看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力

进入量化体系 →

相似阅读

另一事件,读法相近