跳到主内容
精选88SaaStr 博客(RSS)产品与增长

Rippling 实测 15 款 AI 模型:最便宜与最贵打平

Rippling Ran 2,100 Scored Agent Runs Per Model on Real Payroll Data. The Cheapest Model Tied the Most Expensive One.

原文
推荐理由

给做 AI 产品的创业者一套可直接照做的模型选型方法:三个指标、三类选择、七个取舍,今天就能拿自己的数据跑一遍。

Rippling Tested 15 AI Models on Real Payroll Data. The Cheapest One Tied the Most Expensive One.

Rippling 在真实薪资数据上测试了 15 个 AI 模型。最便宜的与最贵的打了个平手。

Rippling’s President and CPO Matt MacInnis published something most B2B companies have and almost none share: a real test of AI models doing real work inside a real production system.

Rippling 的总裁兼首席产品官 Matt MacInnis 发布了一项大多数 B2B 公司拥有但几乎不分享的内容:在真实生产系统中,对 AI 模型执行真实工作的实际测试。

Not a leaderboard. Not made-up tasks. About 2,100 graded attempts per model, across 15 models, on actual personnel, payroll, and financial records. Questions like headcount by department and tenure distribution. And actions like “increase base salaries of everyone who meets this condition by 10%,” onboard a new hire through the checklist, schedule a termination and route its approvals, enter payment amounts into a pay run from a spreadsheet.

不是排行榜,也不是虚构的任务。每个模型约有 2,100 次评分尝试,涉及 15 个模型,基于实际的人员、薪资和财务记录。问题包括按部门统计的员工人数和任职时长分布。操作包括“将所有符合此条件的人的基本工资提高 10%”、通过检查清单入职新员工、安排解雇并路由审批、从电子表格将付款金额输入到工资运行中。

Every attempt either passed Rippling’s production correctness checks or failed. An attempt that never finished counted as a failure. That is a much harder grader than most published benchmarks use.

每次尝试要么通过 Rippling 的生产正确性检查,要么失败。未完成的尝试计为失败。这比大多数已发布基准使用的评分标准要严格得多。

Three numbers run through the whole study, and they’re the only vocabulary you need:

三个数字贯穿整个研究,也是你唯一需要的词汇:

  • Pass rate. How often the model did the job correctly. No partial credit.
  • Cost. What it cost to run the full set of tests once, in dollars.
  • The slowest 10%. How long a task took on the bad days, not the average day. This is the number your customer actually feels.
  • 通过率。模型正确完成工作的频率。没有部分得分。
  • 成本。运行整套测试一次的费用,以美元计。
  • 最慢的 10%。任务在糟糕日子(而非平均日子)所花的时间。这是客户真正感受到的数字。

The 15 models collapse to three real choices

15 个模型归结为三个真正的选择

Almost every model in the study is beaten outright by another one. What’s left is a three-way decision:

研究中几乎每个模型都被另一个模型彻底击败。剩下的就是一个三选一的决策:

  • Opus 4.6 if being right matters most, and it’s nearly a tie. 91.0% pass rate, $1,453, 154 seconds on the slowest 10%. The next model down, GPT-5.5 med, scored 89.5% for $1,435. That’s an $18 difference in cost, which is nothing, and a 1.5-point difference in accuracy that sits inside the margin of error on a test this size. Add the five months Rippling spent tuning its setup specifically for Opus 4.6, worth a point or two by MacInnis’s own estimate, and the two are a wash. What actually separates them is speed, and that goes to Opus 4.6.
  • GPT-5.5 low if speed matters most. 88.8%, $1,308, 130 seconds on the slowest 10%. Cheaper and faster than Opus 4.6, giving up 2.2 points of accuracy. Worth noting that OpenAI’s own two settings of the same model land on either side of Opus 4.6, so the setting mattered more than the brand.
  • GLM 5.2 if speed doesn’t matter at all. 88.7%, $621, 243 seconds on the slowest 10%. Same accuracy as GPT-5.5 low within a tenth of a point, half the price, and much slower. Every overnight job, backfill, and data enrichment run belongs on something like this.
  • Opus 4.6 如果正确性最重要,而且这几乎是个平局。通过率 91.0%,成本 $1,453,最慢的 10% 耗时 154 秒。下一个模型 GPT-5.5 med 得分为 89.5%,成本 $1,435。成本差异 $18,几乎可以忽略,准确性差异 1.5 个百分点,在如此规模的测试中处于误差范围内。再加上 Rippling 花费五个月专门为 Opus 4.6 调整其设置,按 MacInnis 自己的估计值一到两个百分点,两者不相上下。真正区分它们的是速度,而速度优势在 Opus 4.6 这边。
  • GPT-5.5 low 如果速度最重要。通过率 88.8%,成本 $1,308,最慢的 10% 耗时 130 秒。比 Opus 4.6 更便宜、更快,但准确性低了 2.2 个百分点。值得注意的是,OpenAI 同一模型的两个设置分别落在 Opus 4.6 的两侧,因此设置比品牌更重要。
  • GLM 5.2 如果速度完全无关紧要。88.7% 的准确率,621美元,在最慢的10%上耗时243秒。与GPT-5.5低配版准确率相差不到0.1个百分点,价格减半,但速度慢得多。所有过夜任务、数据回填和数据增强运行都适合用这类模型。

The expensive models don’t make that list. Fable 5 matched GPT-5.5 low’s accuracy at 3.3x the price and twice the wait. Opus 5 was beaten on both price and speed by Grok 4.5, which scored 0.2 points behind it for 68% less.

昂贵的模型并未上榜。Fable 5 在准确率上与GPT-5.5低配版持平,但价格是其3.3倍,等待时间翻倍。Opus 5 在价格和速度上均被Grok 4.5击败,后者得分落后0.2个百分点,但价格便宜68%。

Seven takeaways for B2B founders:

给B2B创始人的七点启示:

#1. Quality flattened at the top. Price did not.

#1. 顶尖模型的准确率已趋于平缓,但价格并未如此。

Seven models landed between 88.5% and 89.5%. That is a one-point spread. Add the leader at 91.0% and the whole competitive group is 2.5 points wide.

七款模型的准确率在88.5%到89.5%之间,差距仅一个百分点。加上领先者91.0%,整个竞争组的差距也只有2.5个百分点。

Here’s what those models cost:

以下是这些模型的成本:

GLM 5.2 and Fable 5 are one tenth of a point apart in accuracy. One costs seven times the other.

GLM 5.2 和 Fable 5 的准确率相差仅0.1个百分点,但价格相差七倍。

If your process is “use whatever the big lab just shipped,” you’re paying a 3x to 7x premium for a difference your customers cannot detect.

如果你的流程是“大实验室刚发布什么就用什么”,那你是在为顾客无法察觉的差异支付3到7倍的溢价。

#2. The newest model was not the best model

#2. 最新模型并非最佳模型

Opus 4.6 beat both of the newer Anthropic models on accuracy, cost 42% less than Opus 5, and cost a third of Fable 5. Fable 5 was the most expensive model tested and finished fifth.

Opus 4.6 在准确率上击败了Anthropic两款更新的模型,价格比Opus 5便宜42%,仅为Fable 5的三分之一。Fable 5 是测试中最贵的模型,排名第五。

Version numbers are a lab’s internal accounting, not a promise about your business. Nobody at Rippling would have known this without running the test.

版本号是实验室的内部记账方式,并非对你业务的承诺。如果不进行测试,Rippling 的任何人都不可能知道这一点。

Then MacInnis proved it a second time. He re-ran the same 2,100 attempts on Grok 4.6 and posted the result: accuracy dropped from 87.3% to 85.9%, and typical response time nearly doubled, from 71 seconds to 131. The newer model was worse and slower at the same work.

随后 MacInnis 第二次证明了这一点。他在 Grok 4.6 上重新运行了同样的2100次尝试并发布了结果:准确率从87.3%降至85.9%,典型响应时间几乎翻倍,从71秒增至131秒。较新的模型在相同任务上表现更差且更慢。

One caution on reading those two numbers. The 87.3% he quotes for Grok 4.5 doesn’t match the 89.1% in his published table, because it came from a different run on a different day. Only compare results measured in the same run. A model that “improved two points” against a number you measured last quarter has told you nothing.

解读这两个数字时需谨慎。他引用的 Grok 4.5 的87.3%与他发布表格中的89.1%不符,因为那来自不同日期的不同运行。只有同一运行中测得的结果才可比较。一个模型“提升了两个百分点”相对于你上季度测得的数字,并不能说明任何问题。

So when a new model ships, the default move is not to upgrade. It’s to re-run your own tests. Sometimes the answer is stay.

因此,当新模型发布时,默认做法不是升级,而是重新运行你自己的测试。有时答案就是保持不变。

#3. Your own tuning is worth about as much as a whole model generation

#3. 你自己的调优价值约等于一代模型的进步

Rippling spent five months tuning its instructions and tools around Opus 4.6. MacInnis estimates that work is worth a point or two of accuracy. The entire spread across the leading group is 2.5 points.

Rippling 花了五个月时间围绕 Opus 4.6 调整其指令和工具。MacInnis 估计这项工作价值一到两个百分点的准确率。领先群体的整体差距仅为2.5个百分点。

So the tuning is worth roughly the gap between the best model in the study and the cheapest one. Every other model ran with no tuning at all, which means their scores are floors, not ceilings.

因此,调优的价值大致相当于研究中最佳模型与最便宜模型之间的差距。其他所有模型均未进行任何调优,这意味着它们的得分是下限,而非上限。

Two things follow from that. First, your test suite and your instructions are a real asset, and they compound. Second, they’re also a switching cost you’re building against yourself, which is why you want the work to sit outside any one vendor’s model rather than baked into it.

由此得出两点。首先,你的测试套件和指令是真正的资产,并且它们会不断累积。其次,它们也是你为自己设置的转换成本,这就是为什么你希望这项工作独立于任何单一供应商的模型之外,而不是嵌入其中。

Rippling deliberately kept its setup plain: one model, no routing logic, and two general-purpose tools that do all the work. If a bunch of custom plumbing were making the decisions, they’d be measuring the plumbing instead of the model.

Rippling 刻意保持其设置简单:一个模型,无路由逻辑,以及两个完成所有工作的通用工具。如果一堆定制管道在做决策,那么他们测量的将是管道而非模型。

#4. Cheap models are not cheap because they do less work

#4. 廉价模型并非因为工作量少而便宜

Models are billed by the volume of text they read and write, measured in tokens. Grok 4.5 used 601,000 per task. Opus 5 used 599,000. Effectively the same amount of work. Grok cost $791. Opus 5 cost $2,509.

模型按读写文本量计费,以 token 为单位。Grok 4.5 每任务使用 601,000 个。Opus 5 使用 599,000 个。工作量实际上相同。Grok 花费 $791。Opus 5 花费 $2,509。

The whole difference is the price list. Grok charges $2 per million read and $6 per million written. Opus 5 charges $5 and $25. Fable 5 charges $10 and $50.

全部差异在于价格表。Grok 每百万读取收费 $2,每百万写入收费 $6。Opus 5 收费 $5 和 $25。Fable 5 收费 $10 和 $50。

The cheaper models actually did more back-and-forth, not less. They aren’t cutting corners. They’re just priced differently.

更便宜的模型实际上进行了更多的来回交互,而非更少。它们并未偷工减料。只是定价不同。

Which means the metric to watch is dollars, not efficiency. With prices and usage both moving constantly, the invoice is the only number that stays honest.

这意味着要关注的指标是美元,而非效率。由于价格和使用量都在不断变化,发票是唯一保持真实的数字。

#5. Model choice is a margin decision, not an engineering one

#5. 模型选择是利润决策,而非工程决策

Run it against your own P&L. If AI usage is your biggest variable cost, a 3x swing in that bill moves gross margin by double digits. Same product, same accuracy, same customer experience.

用你自己的损益表来检验。如果 AI 使用是你最大的可变成本,那么该账单的 3 倍波动会使毛利率变动两位数。同样的产品、同样的准确性、同样的客户体验。

That makes this the cheapest margin improvement available to most AI-native B2B companies. No new features, no new headcount, no roadmap change. It’s a setting.

这使得这成为大多数 AI 原生的 B2B 公司可用的最便宜的利润改善方式。无需新功能、无需新增人员、无需路线图变更。这只是一个设置。

It also means someone in finance should own the number and review it on the same cadence as any other cost line.

这也意味着财务部门应负责该数字,并按照与其他成本线相同的节奏进行审查。

#6. Speed and price are separate purchases

#6. 速度和价格是分开购买的

Two models with the same accuracy:

两个准确性相同的模型:

  • GPT-5.5 low: 88.8%, $1,308, and fast
  • GLM 5.2: 88.7%, $621, and slow
  • GPT-5.5 低:88.8%,$1,308,且快速
  • GLM 5.2:88.7%,$621,且缓慢

One is roughly 2.7x quicker. The other is half the price. Nothing is both.

一个大约快 2.7 倍。另一个价格减半。没有两者兼得的。

And for anything a customer is waiting on, the number to watch is the slowest 10%, not the average. The average is what founders quote in board decks. The slowest 10% is what a customer hits on the day they’re already annoyed.

对于任何客户等待的内容,要关注的数字是最慢的 10%,而非平均值。平均值是创始人在董事会演示中引用的。最慢的 10% 是客户在已经感到不满的那一天所遇到的。

GLM 5.2 saves $687 against GPT-5.5 low at identical accuracy, and turns a 130-second worst case into a 243-second one. In a live support chat, four minutes of silence is an abandoned session and a ticket you now handle twice. That costs more than the $687.

GLM 5.2 在相同准确度下比 GPT-5.5 low 节省 687 美元,并将最坏情况下的 130 秒处理时间变为 243 秒。在实时支持聊天中,四分钟的沉默意味着会话被放弃,以及你现在需要处理两次的工单。这比 687 美元的成本更高。

So the cheap-model argument flips depending on the job. For work nobody is waiting on, price wins and GLM takes it outright. For live customer work, speed wins first and price second, and the winners are models nobody was calling a bargain. Opus 4.6 is the example: top accuracy, respectable speed, middle of the pack on price. It looks expensive on a cost chart and correct on a speed chart.

因此,廉价模型的论点取决于具体任务。对于无人等待的工作,价格优先,GLM 完全胜出。对于实时客户工作,速度优先于价格,而胜者并非那些被视为便宜货的模型。Opus 4.6 就是一个例子:顶级准确度、可观的响应速度、价格处于中游。它在成本图表上显得昂贵,但在速度图表上却显得正确。

Most B2B companies pick one model and run everything through it. The split worth making:

大多数 B2B 公司选择单一模型并运行所有任务。值得做出的划分是:

更进一步:量化金融体系

看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力

进入量化体系 →

相似阅读

另一事件,读法相近