跳到主内容
@wquguru
精选86Rohan Paul行业动态

Agent 流量激增致 OpenRouter 议价能力下降

Agents are consuming tokens at nearly 5x the human rate, while their usage has e…

原文
发到 X
推荐理由

揭示了 Agent 工作负载特征如何重塑 AI 基础设施层的商业博弈,对关注推理成本与渠道策略的团队极具参考价值。

Agents are consuming tokens at nearly 5x the human rate, while their usage has exploded ~14X since February.

智能体消耗 token 的速度几乎是人类的 5 倍,而自二月以来其使用量激增了约 14 倍。

Once agents became the majority of tokens on OpenRouter, the router's ability to play suppliers off each other may be much less.

一旦智能体成为 OpenRouter 上 token 的主要组成部分,路由器利用供应商之间相互竞争的能力可能会大幅减弱。

OpenRouter sits between developers and the companies that run the models. It makes money by price shopping, sending each query to the cheapest provider at the time.

OpenRouter 位于开发者和运行模型的厂商之间。它通过比价赚钱,将每个查询发送给当时最便宜的提供商。

When the next question is a single question, this is fine; there is no need for any continuity between the questions.

当下一个问题是一个独立的问题时,这没问题;问题之间不需要任何连续性。

But, the behavior of an agent doing a long task is different: it resends the same long block of background instructions at each stage (i.e. cache hit).

但是,执行长任务的智能体的行为则不同:它在每个阶段都会重新发送相同的长段背景指令(即缓存命中)。

A cache hit only lives on the machine still holding the warm prefix, and rerouting mid-task will mean paying the full pre-fill over again.

缓存命中仅存在于仍持有预热前缀的机器上,如果在任务中途重新路由,意味着需要再次支付完整的预填充费用。

i.e. the model provider stores that cache-hit block in memory and charges only a small fraction to reuse it.

也就是说,模型提供商将该缓存命中块存储在内存中,并仅收取少量费用以复用该块。

Which is why more than 85% of agent tokens on OpenRouter are these cheap reuses (cached prompt) rather than fresh ones.

这就是为什么 OpenRouter 上超过 85% 的智能体 token 是这些廉价的复用(缓存提示),而非全新的请求。

The copy is stored on the servers of a single company . If the job is transferred to a cheaper competitor halfway through , the stored copy is discarded and the full block is repaid .

副本存储在某一家公司的服务器上。如果任务在途中转移给更便宜的竞争对手,存储的副本将被丢弃,并且需要全额支付该块的费用。

That will mean the agent remains with the provider they started the task with until the task is finished and the threat from the router to take their business elsewhere is eliminated.

这意味着智能体会一直留在开始任务时的提供商处,直到任务完成,路由器将其业务转向他处的威胁也随之消除。

So looks like the discounts routers can squeeze out of model providers many shrink on agent traffic well before they shrink anywhere else.

因此,看起来路由器能从模型提供商那里挤出的折扣,在智能体流量领域会比在其他领域更早地缩减。

更进一步:量化金融体系

看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力

进入量化体系 →

相似阅读

另一事件,读法相近