跳到主内容
@wquguru
精选72The Beautiful Mess(John Cutler · RSS)产品与增长

从工时到Token:衡量AI投入回报的陷阱

TBM 437: Tokens, Hours, Points, and Other Curious Proxies

原文
发到 X

Everyone is talking about “return on tokens.” Vendors love it (as long as the news is good). Companies that had long since given up on any sort of value architecture or on understanding the ROI of product work are suddenly chomping at the bit. Which makes sense: they haven’t been able to lay off enough people to cover their token budgets while still hitting their promises to the board.

每个人都在谈论“代币回报率”。供应商喜欢这个说法(只要消息是正面的)。那些早已放弃任何价值架构或产品工作投资回报率理解的公司,突然急不可耐。这说得通:他们还没能裁掉足够多的人来覆盖代币预算,同时仍兑现对董事会的承诺。

It is hard to know what sits below the surface: a genuine effort to understand how AI might augment human capabilities, or some version of “how many people can we safely fire?” Or both?

很难知道表面之下是什么:是真正努力理解AI如何增强人类能力,还是某种“我们能安全裁掉多少人?”的版本?或者两者兼有?

It is as if the whole world suddenly shifted from a buffet model to buying each dumpling one at a time, and now everyone is freaking out about the ROI of each dumpling. But we’ve been here before. We’ve been here with hours. We’ve been here with capacity. And the problems have always been roughly the same:

就好像整个世界突然从自助餐模式转向逐个购买饺子,现在每个人都对每个饺子的投资回报率感到恐慌。但我们以前经历过这种情况。我们经历过工时。我们经历过产能。问题大体上总是相同的:

  • Spending more time on the “I” side than the “R” side of ROI.
  • Myopically choosing shorter-term, easier-to-attribute use cases for the “R” side.
  • Gravitating toward whatever is easiest to measure. And tokens are very easy to measure.
  • 在投资回报率的“I”侧花费的时间多于“R”侧。
  • 短视地选择更短期、更容易归因的用例用于“R”侧。
  • 倾向于选择最容易衡量的东西。而代币非常容易衡量。

This post takes a trip through hours and flow metrics, detours into an alternative way of thinking about ROI, and eventually gets to TOKENS—everyone’s favorite dumpling.

这篇文章将穿越工时和流量指标,绕道进入另一种思考投资回报率的方式,最终到达代币——每个人最喜欢的饺子。

My goal is to show that the underlying measurement problems are remarkably universal. AI introduces some genuinely new twists, but in the end, it still hinges on a theory of value.

我的目标是展示底层测量问题具有显著的普遍性。AI引入了一些真正的新变化,但归根结底,它仍然依赖于价值理论。

With these comparisons, ask yourself, “How might this relate to the current return on tokens question?”

通过这些比较,问问自己,“这与当前的代币回报率问题有何关联?”

Leave a Coffee Tip!

留下咖啡小费!

Hours

工时

Hours, as commonly used to understand investment, represent a major construct validity problem. There’s nothing inherently wrong with measuring “time we spent on things” (hours) to understand better where effort is going, provided you make it clear that:

工时,通常用于理解投资,代表了一个重大的构念效度问题。衡量“我们花在事情上的时间”(工时)本身没有内在错误,以便更好地了解精力去向,前提是你明确说明:

  • Time is not a fungible thing that can be infinitely allocated/re-allocated. This holds across skill sets, team context, and even across a “normal day”. For example, the “magic hour” of uninterrupted morning productivity can be vastly more productive than trying to wrap your head around something complex between 4 and 5 PM. The developer deep in the onboarding space, an expert in onboarding analytics and the user journey, working primarily with what customers see/interact with, can’t be immediately swapped into a deep, legacy backend refactoring effort.
  • Where you allocate time tells you nothing about the “quality”, efficiency, or efficacy of that expenditure of time. The highly skilled developer reactively called in to debug an area of the code they have no experience with will not have the same leverage as someone more junior, but more contextually situated. Ten people spending ten hours each is not equal to one person spending one hundred hours. Sometimes the entire system is constrained by a single specialist, a single decision, a single dependency, or a single review. Spending time elsewhere has negative leverage.
  • The context switching tax is real. Most efforts to understand the allocation of time, whether explicit or implicit, leave out large time investments. Imagine alternating between two tasks every ten minutes for a six-hour day. You’ll spend three hours (at least) on context switching and calibrating/orienting around the new task. Coordination is multiplicative, while time accounting is additive.
  • The relationship between time and outcomes is non-linear. 2x-ing where you spend time doesn’t 2x the outcomes. You might bang your head against a problem for eight hours and get nowhere, and the ninth hour “unlocks” all the value. The metric treats the input as continuous even though the production function contains thresholds and fixed setup costs.
  • Where we invest focus can have a long-term effect. Imagine spending 5% longer on every enhancement to pay off a bit of debt and “garden” the codebase. Three years later, a version of the team—some people have stayed, but some have left—is absolutely flying. Where do you “book” that time? And how do you connect that investment to the long-evolving, game-changing “return”? Some hours multiply future options. Some hours limit future options.
  • Time allocation schemes often force allocations into individual buckets, but there are typically causal relationships between those buckets. They aren’t mutually exclusive. One impacts the other.
  • 时间不是一种可以无限分配/重新分配的可互换资源。这在技能组合、团队背景,甚至“正常一天”中都是如此。例如,早晨不受干扰的“黄金时段”的生产力可能远高于下午4点到5点试图理解复杂事物的效率。深入入职领域的开发者,精通入职分析和用户旅程,主要处理客户看到/交互的内容,不能立即被替换到深度遗留后端重构工作中。
  • 你在哪里分配时间并不能告诉你这些时间支出的“质量”、效率或效果。被紧急召来调试他们不熟悉的代码区域的高技能开发者,其杠杆作用可能不如一个资历较浅但更了解上下文的人。十个人每人花十小时并不等于一个人花一百小时。有时整个系统受限于一个专家、一个决策、一个依赖或一次审查。在其他地方花费时间具有负杠杆效应。
  • 上下文切换的代价是真实存在的。大多数理解时间分配的努力,无论是显式的还是隐式的,都忽略了大量的时间投入。想象一下,在六小时的工作日中,每十分钟在两项任务之间切换。你至少会花三个小时在上下文切换以及围绕新任务进行校准和定位上。协调是乘数效应,而时间记账是加法效应。
  • 时间与结果之间的关系是非线性的。将你花费时间的地方加倍并不会使结果加倍。你可能在一个问题上撞了八个小时的墙而毫无进展,而第九个小时“解锁”了所有价值。这个指标将输入视为连续的,尽管生产函数包含阈值和固定设置成本。
  • 我们投入专注的地方可能产生长期影响。想象一下,每次增强功能时多花5%的时间来偿还一些技术债务并“打理”代码库。三年后,团队的一个版本——有些人留下了,但有些人离开了——正飞速发展。你如何“记账”这段时间?又如何将这种投资与长期演变、改变游戏规则的“回报”联系起来?有些小时会倍增未来的选项,有些小时则会限制未来的选项。
  • 时间分配方案往往强制将时间分配到各个独立的桶中,但这些桶之间通常存在因果关系。它们并非互斥的,一个会影响另一个。

Lean/Flow Metrics (and “story points”)

精益/流动指标(以及“故事点”)

Helpful, when focused.

在专注时,这些是有帮助的。

Lean/flow metrics include cycle time, lead time, practical WIP limits, and throughput (e.g., N stories per week). Most of these metrics have a storied history in manufacturing, which is both a positive—there’s actual math, theory, and practice behind them—and a negative when misapplied to software development.

精益/流动指标包括周期时间、前置时间、实际在制品限制和吞吐量(例如,每周完成的故事数)。这些指标中的大多数在制造业中有着悠久的历史,这既是优点——背后有实际的数学、理论和实践——也是缺点,当被错误地应用于软件开发时。

Note the following:

请注意以下几点:

  • In manufacturing, it matters what you are producing. Things that have sufficiently different work processes, value profiles, etc. are always divided out. This is sometimes referred to as “classes of service,” and it is a critical component of using these metrics effectively. Without that segmentation, changes in throughput, cost, or efficiency may simply reflect changes in the mix of what was produced rather than actual changes in performance.
  • A lot of product work is more akin to R&D, experimentation, and initial “design” than to manufacturing and mass production. In this setting, “waste” isn’t an inherently bad thing. If you try five things and one works, then you’ve purchased the knowledge for that fifth thing.
  • What constitutes a work unit is also mushy. Say you have a more open-ended stream of experiments meant to move a metric; what do you measure? Experiments per week? The lead time of the whole stream? Both? Where do you put the opportunities that get discovery work, but ultimately don’t get greenlit?
  • When we talk about “capacity” in manufacturing, you are talking about the sustainable ability to meet a particular type (or types) of demand. “This factory can produce N of this variety of shoes every month.” You invest in a factory that can meet this forecasted demand, with optionality to scale up the factory based on new demand. So, going back to some of the issues with hours, capacity is not (always) fungible and is not something that is “spent” at the time of use. It is an emergent quality of the factory’s design.
  • The boundaries of the system matter enormously. “Lead time” from when? Customer request? Commitment? Ticket creation? First code change? And when is something done: merged, deployed, adopted, or producing value? Moving the boundary can dramatically change the metric without changing the work.
  • You have to be ready for the uncomfortable reality that “work” spends a majority of its time not being “worked on”, as well as facing all of your bad habits when it comes to keeping people busy.
  • Many of these metrics describe a relatively stable flow system that remains stable over a sufficiently long observation period. If the system is rapidly changing and has a lot of variability in work types, you can be lulled into a sense of apparent precision.
  • Don’t get me started on story points. Any effort to use story points for anything beyond a disposable tool to spark discussion, right-size, or reality-check a near-term cycle commitment is a dereliction of duty. We have decades of evidence to support this.
  • 在制造业中,生产什么很重要。工作流程、价值特征等有显著差异的事物总是被分开处理。这有时被称为“服务类别”,是有效使用这些指标的关键组成部分。没有这种细分,吞吐量、成本或效率的变化可能仅仅反映了生产组合的变化,而非实际绩效的变化。
  • 许多产品工作更类似于研发、实验和初步“设计”,而非制造和大规模生产。在这种背景下,“浪费”并非本质上是不好的。如果你尝试了五件事,其中一件成功了,那么你就为那第五件事购买了知识。
  • 工作单元的定义也是模糊的。假设你有一个更开放的实验流,旨在推动某个指标;你衡量什么?每周的实验数量?整个流程的交付周期?两者都要?那些获得发现工作但最终未获批准的机会,你放在哪里?
  • 当我们谈论制造业中的“产能”时,指的是满足特定类型(或多种类型)需求的可持续能力。“这家工厂每月能生产N双这种款式的鞋子。”你投资于一个能满足预测需求的工厂,并根据新需求有扩展工厂的选择权。因此,回到工时的一些问题,产能(并非总是)可互换,也不是在使用时“消耗”的。它是工厂设计的一种涌现特性。
  • 系统的边界至关重要。“交付周期”从何时开始?客户请求?承诺?创建工单?首次代码变更?何时算完成:合并、部署、采用,还是产生价值?移动边界可以在不改变工作的情况下大幅改变指标。
  • 你必须准备好面对一个不舒服的现实:“工作”大部分时间并未被“处理”,同时也要面对你所有关于让人保持忙碌的坏习惯。
  • 许多这些指标描述的是一个相对稳定的流动系统,在足够长的观察期内保持稳定。如果系统快速变化且工作类型变异性大,你可能会被虚假的精确感所迷惑。
  • 别跟我提故事点。任何试图将故事点用于超出作为激发讨论、合理调整规模或对近期周期承诺进行现实检验的一次性工具之外的用途,都是失职行为。我们有数十年的证据支持这一点。

The main thing with manufacturing is that it is essentially a demand-based pull system. The things being produced have a price, and there is a clear distinction between “sitting on the lot” and “in the customer’s hands.” That makes the economics an order of magnitude easier to reason about. Theoretically, we could imagine features going unused as inventory, unfinished work sitting in queues as WIP, or capabilities built ahead of actual demand as overproduction.

制造业的主要特点是,它本质上是一个基于需求的拉动系统。所生产的产品有价格,并且“停在停车场”与“在客户手中”之间有明确的区别。这使得经济分析容易一个数量级。理论上,我们可以想象未使用的功能作为库存,未完成的工作作为在制品排队,或提前于实际需求构建的能力作为过度生产。

更进一步:量化金融体系

看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力

进入量化体系 →

相似阅读

另一事件,读法相近