OpenAI 3x 生产力真相:高缺陷率下的三班倒成本账
Is the 3x AI Productivity Gain just a Computer that Never Sleeps?
拆解了 AI 生产力背后的单位经济模型与隐性成本,帮助创业者判断 AI 投入的真实 ROI,而非盲目追求效率指标。
The market is telling us that we should be 3x more productive with AI.
市场正在告诉我们,我们应该利用 AI 将生产力提高 3 倍。
What if that productivity gain is just an AI working 24 hours a day while a human works eight?
如果这种生产力的提升仅仅是因为 AI 可以全天候工作,而人类只工作八小时呢?
OpenAI published the math behind its 3x claim. In mid-August, its research staff logged 3.14 agent-workdays1 for every 8-hour human shift.2 The typical researcher ran four agents in parallel.
OpenAI 公布了其 3 倍主张背后的数学依据。8 月中旬,其研究人员每 8 小时的人类班次就记录了 3.14 个代理工作日¹。² 典型的研究人员并行运行四个代理。
That machine shift comes with an industrial price tag. In late March, the median OpenAI researcher spent $14 a day on inference. By mid-August, that bill climbed past $600 a day : a 40-fold surge in under five months.2 At the top end, the 90th percentile researcher burns through more than $7,000 a day, an annualized run-rate of $2.5m.
这种机器班次伴随着高昂的工业成本。3 月下旬,中位数水平的 OpenAI 研究人员每天在推理上的花费为 14 美元。到了 8 月中旬,这笔费用攀升至每天超过 600 美元:在不到五个月的时间里激增了 40 倍。² 在最顶端,第 90 百分位的研究人员每天消耗超过 7,000 美元,年化运行率为 250 万美元。
At $2.5m a year per seat, inference behaves like heavy factory tooling. But it comes with a financial twist : it is pure OPEX.
以每个席位每年 250 万美元的成本来看,推理表现得像重型工厂设备。但它有一个财务上的转折:它是纯粹的运营支出(OPEX)。
Auto plants buy welding robots with capex. They run night shifts to amortize machinery that depreciates whether used or idle. AI systems invert that math. Inference is metered operating expense. With no physical tooling & no graveyard-shift wages, a company can run machines overnight on pure variable cost.
汽车工厂通过资本支出(capex)购买焊接机器人。它们运行夜班以摊销那些无论使用与否都会贬值的机械。AI 系统颠覆了这一数学逻辑。推理是按量计费的运营支出。没有实体设备,也没有夜班工资,公司可以仅凭可变成本让机器整夜运行。
The 3.14 workday ratio is not three times smarter thinking. It is one engineer supervising three shifts of machine runtime while only being awake for one.
3.14 的工作日比率并不意味着思维聪明了三倍。它是一位工程师监督三班机器的运行时间,而他本人只清醒了一班的时间。
Yet unlike an auto welding robot, this digital assembly line has a massive defect rate. Over half of the successful four-to-eight-hour tasks in the last six months still needed human intervention ; the lab is candid that “the overall pace of progress likely won’t keep pace with these specific metrics.”2
然而,与汽车焊接机器人不同,这条数字装配线有着巨大的缺陷率。在过去六个月中,成功完成的四到八小时任务中,有一半以上仍需人工干预;实验室坦承“整体进展速度可能无法跟上这些具体指标”。²
A 40-fold surge in compute spend bought three times the work-hours. But with a supervisor still untangling more than half the runs, the engineer’s day shifts from creative architecture to walking the plant floor & clearing machine jams.
计算支出的 40 倍增长换来了三倍的工作时长。但由于主管仍在梳理超过一半的运行结果,工程师的工作从创造性架构设计转变为巡视车间并排除机器故障。
Why run the machines through the night if the defect rate is so high? Fear & ambition.
既然缺陷率如此之高,为什么还要让机器通宵运行?恐惧与野心。
If your peers field four agents around the clock, logging off is falling behind. The rush of a superpower paid for by your employer is intoxicating. When you get a tireless digital workforce on someone else’s balance sheet, you never turn the factory off.
如果你的同事全天候部署四个代理,下线就意味着落后。由雇主买单的超能力带来的快感令人陶醉。当你在别人的资产负债表上获得一支不知疲倦的数字劳动力时,你永远不会关闭工厂。
For forty years, a programmer needed only a MacBook & an eight-hour shift. Today, a top OpenAI researcher commands four parallel agents, burns through $2.5m a year in compute, & spends the morning fixing machine errors from the night before. This explains the quiet frustration spreading across software engineering today.3
四十年来,程序员只需要一台 MacBook 和一个八小时的班次。如今,一位顶级的 OpenAI 研究人员指挥着四个并行代理,每年消耗 250 万美元的计算资源,并在上午修复前一天晚上的机器错误。这解释了当前软件工程领域悄然蔓延的挫败感。³
The market hears 3x productivity & expects creative miracles. The engineer gets stuck untangling a 50% scrap rate from robots that ran all night.4 The market calls it a 3x leap in productivity. A CFO would just call it paying for a second & third shift. For now, that is the honest price of a machine that never sleeps. The real question is when the second & third shifts start to out-yield the first.
市场听闻生产力提升3倍,并期待创造奇迹。工程师却陷入困境,试图理清机器人通宵运行后高达50%的废品率。4 市场称之为生产力实现了3倍的飞跃。而CFO(首席财务官)只会认为这是在为第二和第三个班次买单。目前而言,这是一台永不眠机器所付出的诚实代价。真正的问题是,当第二和第三个班次的产出开始超过第一个班次时,情况会如何。
- OpenAI reports 3.1 agent-workdays; we round to 3.14 for the irony, since a ratio of 3.14 agent-workdays to one human workday is, fittingly, a pie, not a numerator. ↩︎
- OpenAI: Research acceleration : The view inside OpenAI ↩︎ ↩︎ ↩︎
- Stack Overflow Developer Survey : Closing the AI Trust Gap : 84% of developers use AI tools, but trust has fallen to 29%, with 66% citing code that is “almost right, but not quite” & 45% reporting that debugging AI-generated code takes more time than writing it manually. ↩︎
- The yield math : an 8-hour human shift leaves 16 overnight hours (two extra shifts of machine runtime). With OpenAI disclosing that more than half of 4-to-8-hour tasks require human intervention, the autonomous yield is ~50%. Two machine shifts at 50% yield equal one effective shift of finished output. That yields ~2x delivered work while logging 3x the raw shift runtime. ↩︎
- OpenAI报告称有3.1个代理工作日;出于讽刺意味,我们将其四舍五入为3.14,因为3.14个代理工作日对1个人类工作日的比率,恰如其分地构成了一个“派”(pie),而非分子。↩︎
- OpenAI:研究加速:深入OpenAI内部视角 ↩︎ ↩︎ ↩︎
- Stack Overflow开发者调查:弥合AI信任鸿沟:84%的开发者使用AI工具,但信任度已降至29%,其中66%的人指出代码“几乎正确,但又不完全正确”,45%的人表示调试AI生成的代码比手动编写花费的时间更多。↩︎
- 产出数学计算:一个8小时的人类班次留下了16个夜间小时(相当于机器的两个额外班次运行时间)。随着OpenAI披露超过一半的4到8小时任务需要人工干预,自主产出率约为50%。两个机器班次在50%的产出率下,等同于一个有效班次的成品输出。这使得交付的工作量约为2倍,而记录的原始班次运行时间却是3倍。↩︎
更进一步:量化金融体系
看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力