跳到主内容
精选88SaaStr 博客(RSS)产品与增长

B2B AI Agent落地:路线图、审批流与上下文成本

The CPOs of Harvey, Glean and Rubrik on What It Actually Takes To Ship a Category-Winning Agent

原文
发到 X
推荐理由

直接给出了B2B产品做AI Agent落地的五个具体战术动作,特别是“先读后写”的审批流设计和按风险排序的路线图,创业者今天就能对照检查自己的产品。

For about ten years, chief product officer was the best job in B2B. You brought your mug to the office, told the team what you were shipping this year, moved a few things to next quarter, and when an investor asked for a feature you said it was coming. And sometime, later, it was coming. That was the job through 2024.

大约十年间,首席产品官是B2B领域最好的职位。你带着马克杯到办公室,告诉团队今年要交付什么,把一些事情推迟到下个季度,当投资者询问某个功能时,你说它即将推出。然后过段时间,它确实推出了。这就是直到2024年为止的那份工作。

Every product leader in B2B is now under the gun to ship agents ASAP someone will pay for. Atlassian finally monetized its AI and the stock jumped roughly a third in a single day. Most of the rest are still in the middle of it.

现在,每位B2B的产品负责人都面临紧迫压力,必须尽快推出有人愿意付费的AI智能体(agents)。Atlassian终于将其AI商业化,股价在一天内飙升约三分之一。大多数其他公司仍处于推进过程中。

We banned panels at SaaStr AI for a year because they’re boring. The exception is when the people on stage know each other, so at SaaStr AI I put together a group who do: Anneka Gupta, CPO at Rubrik, Emrecan Dogan, CPO at Glean, Anique Drumright, CPO at Harvey, and Rachel Wolan, then CPO at Webflow.

我们在SaaStr AI大会上禁止了圆桌讨论环节一年,因为它们很无聊。例外情况是当台上的人彼此认识时,所以在SaaStr AI上我组建了一个这样的群体:Rubrik的首席产品官Anneka Gupta、Glean的首席产品官Emrecan Dogan、Harvey的首席产品官Anique Drumright,以及当时Webflow的首席产品官Rachel Wolan。

Anneka sells to security teams who can’t tolerate a wrong action. Emrecan sells the context layer other agents run on. Anique sells to law firm partners who pay for software out of their own distributions. Three buyers, one shared problem.

Anneka面向无法容忍错误操作的安全团队销售。Emrecan销售其他智能体运行的上下文层。Anique面向律师事务所合伙人销售,他们用自己的分配资金购买软件。三种买家,一个共同的问题。

The five things that came out of it:

从中得出的五点:

  • An agent roadmap is a second build of your entire product. That’s why so many of them have slipped, and it’s why you start with the agents that can’t break anything.
  • The plan is what gets reviewed now, not just the output. Rubrik and Harvey arrived at this independently: the agent proposes a plan, a human approves or edits it, and the execution underneath stays deterministic.
  • Half of a heavy AI user’s day goes into feeding context, which is why Glean is now growing faster inside Claude Code and Cursor than in its own UI.
  • Agents don’t have seats and they don’t have identities. Once an agent writes into Salesforce for a whole department, “who did this” becomes a product requirement.
  • You are responsible for agent behavior you did not design and cannot predict, including whatever your customers build headless on top of you.
  • 智能体路线图是你整个产品的第二次构建。这就是为什么许多计划延期,也是为什么你要从那些不会破坏任何东西的智能体开始的原因。
  • 现在的审查重点是计划,而不仅仅是输出结果。Rubrik和Harvey独立得出了这一结论:智能体提出计划,人类批准或编辑它,其下的执行保持确定性。
  • 重度AI用户的一半工作时间花在提供上下文上,这就是为什么Glean在Claude Code和Cursor内部的增长速度比在其自有UI中更快的原因。
  • 智能体没有席位,也没有身份。一旦一个智能体为整个部门写入Salesforce,“谁做的”就成了一个产品需求。
  • 你对自己未设计且无法预测的智能体行为负责,包括你的客户基于你构建的任何无头(headless)应用。

#1. An agent roadmap is a second build of your whole product. Start with the agents that can’t break anything.

#1. 智能体路线图是你整个产品的第二次构建。从那些不会破坏任何东西的智能体开始。

Rubrik sits in cyber recovery. Customers rely on them to get data, applications and identity back after they’ve already been hacked, so nothing Rubrik ships can take that service down.

Rubrik位于网络恢复领域。客户依赖他们在被黑客攻击后恢复数据、应用程序和身份,因此Rubrik推出的任何服务都不能导致该服务中断。

Anneka’s read on how the agent work went:

Anneka对智能体工作的看法:

“It’s been actually a much more challenging problem to build agents within our product than I think I thought it was going to be a year ago.”

"实际上,在我们的产品中构建智能体是一个比我一年前认为的要困难得多的问题。"

The reason is in her own requirement, stated later in the session: everything possible through the UI should be possible agentically, and eventually you operate all of Rubrik through chat without the dashboards at all. That is building the product a second time, in chat, with new ways to get it wrong, for customers who can’t absorb a wrong answer.

原因在于她自己在会议后段提出的要求:所有能通过 UI 实现的功能,都必须能以 Agent 方式实现;最终你将完全通过聊天操作 Rubrik,不再使用仪表盘。这相当于在聊天界面中重新构建一次产品,并引入了新的出错方式,而客户无法承受错误的答案。

Count the workflows before you commit to a date. Not “we’re shipping agents in Q2.” How many workflows exist in your UI, which ones a customer would ever hand over, and in what order they get rebuilt. Most agent roadmaps that slipped this year were scoped like a feature and priced like a rewrite.

在承诺日期之前,先统计工作流数量。不要说“我们第二季度发布 Agent”。你的 UI 中存在多少个工作流?客户会移交哪些工作流?它们按什么顺序被重建?今年大多数延期发布的 Agent 路线图,其范围像是一个功能,但定价却像是一次重写。

Start with the agents that can’t break anything. Rubrik’s first real agentic workflow is forward-looking capacity planning, a job that takes a customer at least a day by hand. It reads, analyzes and recommends, and never touches production. The ones that do touch production come later, and only run after a human approves the plan. Sort your list by what happens if the agent gets it wrong, ship the ones where the worst case is a customer ignoring a recommendation, and hold the rest until the approval step works.

从那些不会破坏任何东西的 Agent 开始。Rubrik 的第一个真正的 Agent 工作流是前瞻性容量规划,这是一项手工操作至少需要客户花费一天的任务。它负责读取、分析和推荐,但从不触碰生产环境。那些确实涉及生产环境的 Agent 会稍后推出,且仅在人类批准计划后才运行。根据 Agent 出错时的后果对列表进行排序,优先发布最坏情况仅是客户忽略建议的工作流,其余则等到审批步骤完善后再推出。

Rubrik also didn’t hand-build a set of anticipated workflows. They built a platform generative enough to cover onboarding, troubleshooting, planning and whatever else customers do, because the alternative is guessing which ten workflows your customers want and shipping ten wrong ones.

Rubrik 也没有手动构建一组预期的工作流。他们构建了一个足够通用的平台,以覆盖入职引导、故障排除、规划以及客户所做的任何其他事情,因为另一种选择是猜测客户想要的十个工作流,然后交付十个错误的结果。

#2. In a cyber recovery, Rubrik won’t let the model improvise the steps

#2. 在网络安全恢复场景中,Rubrik 不允许模型即兴发挥步骤

Ruby started as a RAG application. Rubrik fed in the docs on setup and troubleshooting, customers asked questions, and the answers were, in Anneka’s words, okay answers based on whatever was in the documentation.

Ruby 最初是一个 RAG(检索增强生成)应用。Rubrik 输入了关于设置和故障排除的文档,客户提出问题,而答案正如 Anneka 所言,是基于文档内容得出的“还可以”的答案。

What changed this year is the split between what the model does and what the system does. The agent pulls insights out of the customer’s own systems, marries that with Rubrik’s domain expertise, and produces the plan. Then the plan gets executed deterministically:

今年发生的变化在于模型执行的操作与系统执行的操作之间的分离。Agent 从客户自己的系统中提取洞察,将其与 Rubrik 的领域专业知识相结合,并生成计划。然后,该计划以确定性方式执行:

“In a recovery scenario you don’t want to be guessing and you don’t want to be using like a probabilistic mechanism for recovery.”

“在恢复场景中,你不希望靠猜测,也不希望使用概率机制来进行恢复。”

The model generates and explains the plan. The recovery steps themselves are fixed and auditable, which is what let Rubrik put an agent in front of a recovery workflow at all.

模型生成并解释计划。恢复步骤本身是固定且可审计的,这正是 Rubrik 能够将 Agent 引入恢复工作流的关键。

#3. There is no central agent team at Rubrik

#3. Rubrik 没有中央 Agent 团队

The common instinct is to stand up an AI team, staff it with your best people, and have it build the agentic features for everybody. Rubrik went the other way:

常见的本能做法是组建一个 AI 团队,配备最优秀的员工,让他们为所有人构建智能体功能。Rubrik 选择了相反的路径:

“How do we democratize this so that we don’t have a central team that is building all of the use cases and optimizing all of the use cases, but every PM and every engineering team within our company is thinking from an agent first mindset.”

“我们如何普及这项技术,以便不再由一个中央团队构建和优化所有用例,而是让公司内的每位产品经理和每个工程团队都秉持‘以智能体为先’的思维。”

Go back to the workflow count in section 1. No central team absorbs that volume. It becomes an eval problem, an architecture problem and a training problem for every product team at once.

回顾第 1 节中的工作流数量。没有中央团队来吸收这些工作量。它同时成为了每个产品团队的评估问题、架构问题和培训问题。

#4. Half of a five-hour AI day goes into building context

#4. 五小时 AI 工作时间中有一半用于构建上下文

Emrecan’s number from Glean:

Emrecan 来自 Glean 的数据:

“Most folks I talk to, if they are spending let’s say five hours per day in front of an AI assistant, they are spending half of that time in quote unquote building context.”

“我交谈的大多数人表示,如果每天在 AI 助手上花费五个小时,他们有一半的时间花在所谓的‘构建上下文’上。”

Feeding documents. Feeding decisions. Feeding patterns. Updating memories. Writing skills.

喂入文档。喂入决策。喂入模式。更新记忆。编写技能。

Glean started seven years ago when retrieval was the whole job. Retrieval is now a component, and the job moved from getting informed to getting AI to perform. Performance is still capped by context: the tested knowledge, the idiosyncrasies, the unwritten stuff everybody at your company understands and nobody has written down.

Glean 七年前起步时,检索就是全部工作。如今检索只是一个组件,工作重心已从获取信息转向让 AI 执行任务。性能仍受限于上下文:已测试的知识、独特性、以及公司内部人人皆知却无人落笔的内容。

For your own roadmap: count how much setup a user does before your AI feature produces anything useful. If they’re pasting in the same documents every session, that setup time is the real price of your product, and it doesn’t show up anywhere in your funnel.

对于你自己的路线图:统计用户在你的 AI 功能产生任何有用结果之前进行了多少设置操作。如果他们每次会话都粘贴相同的文档,那么这些设置时间才是你产品的真实成本,而且它不会出现在漏斗的任何地方。

#5. Glean inside Claude Code and Cursor is growing faster than Glean’s own UI

#5. Claude Code 和 Cursor 内部的 Glean 增长速度快于 Glean 自身的 UI

Glean runs two ways now: as the assistant you use directly, and as a single MCP server that Claude Code, Cursor or Codex tap into instead of wiring up individual MCPs.

Glean 现在有两种运行方式:作为你直接使用的助手,以及作为一个单一的 MCP 服务器,Claude Code、Cursor 或 Codex 接入其中,而不是单独连接各个 MCP。

Emrecan on the mix: bundle the Claude, Cursor and Codex usage together and that path is one of the fastest growing parts of the business, growing faster than Glean’s own UI, which is still by a large margin the higher daily engagement driver. It also pushes usage back the other way, since people working through Claude come back into Glean’s own agents for other parts of their job.

Emrecan 谈到这种混合模式:将 Claude、Cursor 和 Codex 的使用量捆绑在一起,这条路径是业务中增长最快的部分之一,增长速度甚至超过了 Glean 自身的 UI(后者仍然是每日参与度最高的驱动因素,且优势巨大)。它还反向推动使用量,因为通过 Claude 工作的人会在其他工作环节回到 Glean 自身的智能体。

He drew a hard line between MCP and what Glean does underneath:

他在 MCP 和 Glean 底层所做的工作之间划出了一条明确的界限:

“MCP, as much independence as it brings, it’s a very runtime fetch of information. You are bound by whether latency or you are bound by the search APIs under the hood.”

“MCP 带来了相当大的独立性,但它本质上是对信息的运行时获取。你受制于延迟,或者受制于底层的搜索 API。”

The other half is high-compute offline processing: connecting employees, teams, projects, subsidiaries and companies you acquired a decade ago, and building the context before anyone asks for it. Shipping an MCP server gives an agent a way in. It doesn’t give the agent anything it couldn’t have queried itself.

另一半是高算力的离线处理:连接员工、团队、项目、子公司以及十年前收购的公司,并在任何人提出需求之前构建上下文。部署一个 MCP 服务器为代理提供了一条接入途径。它并没有赋予代理任何它无法自行查询的信息。

#6. One Gong call, three settings, and a write into Salesforce

#6. 一次 Gong 通话,三个设置,以及写入 Salesforce

The Glean demo was the builder view, and the clearest picture of where departmental agents are heading. Glean’s own sales team generates 300 to 800 Gong calls a day. The agent reads a call, applies the instruction set, and proposes updates to Salesforce.

Glean 的演示展示了构建者视图,也最清晰地描绘了部门级代理的发展方向。Glean 自己的销售团队每天会产生 300 到 800 次 Gong 通话。代理读取通话内容,应用指令集,并提议对 Salesforce 进行更新。

The three settings Emrecan laid out:

Emrecan 列出的三个设置:

  • A human drops in one call to test the agent.
  • 人工介入一次通话以测试代理。

更进一步:量化金融体系

看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力

进入量化体系 →

相似阅读

另一事件,读法相近