跳到主内容
@wquguru
精选88SaaStr 博客(RSS)产品与增长

13家AI客服实测:最佳解决率70%,中位数48%及选型指南

How Much of Your Customer Support Can AI Really Resolve? The Best Get About 70%. The Median Is 48%. Here’s the Real Data From 13 Vendors

原文
发到 X
推荐理由

提供具体量化基准与避坑指南,直接指导AI客服选型与ROI测算。

Exactly how many support and related issues can AI Agents really, truly resolve today? It’s a good test of just exactly where the latest models and agents are.

AI Agent 今天究竟能真正解决多少支持和相关问题?这是一个很好地测试最新模型和代理实际水平的标准。

Now $100M ARR AI CX for ecommerce leader Gorgias (where SaaStr Fund led the seed) has now published a public benchmark that tests 13 vendors on 212 live, mid-market ecommerce stores carrying real inventory, with no sandboxes, synthetic catalogs, or vendor demos. Every agent gets the same customer messages, and a judge scores each response blind to the vendor against 26 binary checks. Factual claims like price, policy, and SKU are checked programmatically against the live store.

现在,为电商巨头 Gorgias(SaaStr Fund 领投种子轮)提供 1 亿美元 ARR 的 AI CX 服务,发布了一项公开基准测试,对 212 家拥有真实库存的中端市场电商商店进行了测试,共涉及 13 个供应商。测试没有使用沙盒、合成目录或供应商演示。每个代理接收相同的客户消息,并由裁判在不知晓供应商身份的情况下,根据 26 项二元检查对每个回复进行评分。价格、政策和 SKU 等事实性声明会与实时商店数据进行编程比对。

The TL;DR: the best AI agents fully resolve about 70% of conversations. The typical vendor resolves under half.

简要总结:最好的 AI Agent 能完全解决约 70% 的对话。典型供应商的解决率不到一半。

The Top Five Resolve 64% to 75%. The Median Vendor Resolves 48%.

前五名供应商的解决率在 64% 到 75% 之间。中位数供应商的解决率为 48%。

Automation rate here is the share of engaged conversations the AI resolved with no human involved. Each number pools the last four weeks, with conversation counts in parentheses:

这里的自动化率是指 AI 在无人工参与情况下解决的已互动对话所占比例。每个数字汇总了最近四周的数据,括号内为对话数量:

  • Rep AI: 75% (120)
  • Decagon: 73% (34)
  • Yuma: 71% (118)
  • Ada: 68% (420)
  • Gorgias: 64% (388)
  • DigitalGenius: 50% (561)
  • Sierra: 48% (408)
  • Siena: 46% (300)
  • Kodif: 44% (188)
  • Intercom: 42% (307)
  • Zendesk: 37% (467)
  • Envive: 22% (499)
  • Klaviyo: 21% (309)

The top five average 70%, the median is 48%, and the field resolves 46% weighted by conversation volume.

前五名平均解决率为 70%,中位数为 48%,按对话量加权计算的总体解决率为 46%。

Zendesk, Intercom, and Klaviyo are the platforms most brands already pay for, and none of them resolves more than 42% in this benchmark. On this data, the AI bundled with your existing helpdesk or email tool is one of the weakest options available.

Zendesk、Intercom 和 Klaviyo 是大多数品牌已经付费使用的平台,在本基准测试中,它们的解决率均未超过 42%。基于这些数据,与您现有的帮助台或电子邮件工具捆绑在一起的 AI 是可用选项中表现最弱的之一。

Yuma, Decagon, and Gorgias Are the Only Vendors Above 64% Automation and 65 Quality

Yuma、Decagon 和 Gorgias 是唯一自动化率超过 64% 且质量得分超过 65 的供应商

Automation rate alone can point you to the wrong vendor. Here is each vendor’s rate next to its blind quality score (0 to 100):

仅看自动化率可能会让你选错供应商。以下是每个供应商的自动化率及其盲测质量得分(0 到 100):

Strong on both:

两者皆强:

  • Yuma: 71% automation, 72 quality
  • Decagon: 73% automation, 70 quality
  • Gorgias: 64% automation, 66 quality
  • Yuma: 71% 自动化率,72 质量分
  • Decagon:自动化率 73%,质量得分 70
  • Gorgias:自动化率 64%,质量得分 66

High automation, weak answers:

高自动化,回答质量弱:

  • Rep AI: 75% automation, 55 quality
  • Ada: 68% automation, 39 quality
  • Rep AI:自动化率 75%,质量得分 55
  • Ada:自动化率 68%,质量得分 39

Strong answers, low automation:

回答质量强,自动化率低:

  • Sierra: 48% automation, 72 quality
  • Intercom: 42% automation, 66 quality
  • Klaviyo: 21% automation, 65 quality
  • Sierra:自动化率 48%,质量得分 72
  • Intercom:自动化率 42%,质量得分 66
  • Klaviyo:自动化率 21%,质量得分 65

Ada closes more conversations than Gorgias but scores 27 points lower on the answers. Sierra ties Yuma for the best quality score and resolves under half of its conversations.

Ada 关闭的对话数量多于 Gorgias,但其回答质量得分低 27 分。Sierra 与 Yuma 并列获得最高质量得分,但解决的对话不足一半。

The benchmark’s case for weighting quality is practical: a ticket marked resolved that gives the customer the wrong return window creates a second contact that costs more than the first. A vendor scoring 39 on quality is generating some of its own future ticket volume.

基准测试为加权质量提供了实际依据:一个标记为已解决却给客户提供错误退货窗口的工单,会引发第二次联系,其成本高于第一次。质量得分为 39 的供应商正在制造部分未来的工单量。

9 of 10 Vendors Lost Automation in the Last Four Weeks

过去四周内,10 家供应商中有 9 家的自动化率下降

Ten vendors have enough history to show a four-week trend. On automation, nine declined:

有足够历史记录以展示四周趋势的供应商共十家。在自动化方面,九家出现下滑:

  • Siena: -23 points
  • Kodif: -13
  • Klaviyo: -11
  • Sierra: -11
  • DigitalGenius: -10
  • Ada: -9
  • Gorgias: -6
  • Intercom: -6
  • Zendesk: -4
  • Siena:-23 分
  • Kodif:-13
  • Klaviyo:-11
  • Sierra:-11
  • DigitalGenius:-10
  • Ada:-9
  • Gorgias:-6
  • Intercom:-6
  • Zendesk:-4

Envive was the only one to improve, at +8. Quality moved the same way, with nine of ten down: Ada -17, Zendesk -15, Sierra -14, Gorgias -12. Siena was the only vendor to improve on quality, at +2.

Envive 是唯一一家实现增长的供应商,增幅为 +8。质量得分也呈现相同趋势,十家中有九家下滑:Ada -17,Zendesk -15,Sierra -14,Gorgias -12。Siena 是唯一一家在质量得分上提升的供应商,增幅为 +2。

The benchmark doesn’t explain the declines. Model updates, store configuration drift, or a harder question set could all contribute. Whatever the cause, the rate you saw in the pilot is not the rate you’ll have next month. We run 20+ agents in production at SaaStr, and we re-check every one of them on a schedule, including the ones that seem to be working.

基准测试并未解释这些下滑的原因。模型更新、商店配置漂移或题目难度增加都可能 contributing。无论原因如何,你在试点中看到的比率并非下个月的比率。我们在 SaaStr 在生产环境中运行 20 多个代理,并按计划重新检查每一个,包括那些看似运作正常的代理。

“Resolved” Here Means Zero Human Touch

“已解决”在此处意为零人工介入

Part of the gap between homepage claims and this data comes from the definition. The benchmark counts a conversation as automated only when the AI handled it with zero human touch, no handover to a person, and no deflection out of the channel. “Email us,” a contact form, or “call us” all count against the agent. If more than half of an agent’s replies push the customer out of the channel, the conversation is scored as unresolved. The auditor never asks for a human, so every handoff is the agent’s own decision.

主页宣传与这些数据之间的差距部分源于定义。基准测试仅在 AI 以零人工介入处理对话、未转交给人工客服且未将客户引导出当前渠道时,才将其计为自动化。“发邮件给我们”、联系表单或“打电话给我们”均会被视为对代理不利的因素。如果代理超过一半的回复将客户引导出当前渠道,该对话将被判定为未解决。审计人员从不要求人工介入,因此每一次转交都是代理自身的决定。

In-product analytics usually use looser rules. In Gorgias’s own customer-facing reporting, an interaction counts as automated once the AI resolves it and 72 hours pass without a human agent helping. A customer who gave up and never came back counts as a resolution there. In the benchmark, that customer counts as a failure.

产品内分析通常采用更宽松的规则。在 Gorgias 面向客户的报告中,一旦 AI 解决了交互且在 72 小时内没有人工客服提供帮助,该交互即被视为已自动化。放弃并永不返回的客户在那里被计为解决。而在基准测试中,该客户被计为失败。

This matters for your P&L because “resolved” is also the pricing unit. Gorgias AI Agent charges $0.90 per resolved conversation. Get each vendor’s definition of a billable resolution in writing before you sign.

这关系到你的损益表(P&L),因为“已解决”也是计费单位。Gorgias AI Agent 按每个已解决的对话收取 0.90 美元。在签署合同前,务必让每家供应商书面提供其可计费解决方案的定义。

Order Tracking Scores 17 Points Below Returns Policy

订单追踪得分比退货政策低 17 分

The report breaks support quality down by topic, averaged across the field:

报告按主题细分了支持质量,取行业平均值:

  • Returns policy: 77.4
  • Shipping policy: 74.4
  • Modify or cancel: 63.5
  • Order tracking: 60.1
  • Damaged item: 59.7
  • Shopper unsure: 49.3
  • 退货政策:77.4
  • 运输政策:74.4
  • 修改或取消:63.5
  • 订单追踪:60.1
  • 商品损坏:59.7
  • 买家不确定:49.3

The field drops 28 points from its easiest conversation type to its hardest.

该领域从最容易处理的对话类型到最难处理的对话类型,得分下降了 28 分。

Questions answered from a policy page score highest. Questions that need a live system lookup or a judgment call score lowest. Order tracking is usually the largest ticket category for an ecommerce brand, and it sits near the bottom of this list. Across the field, only about a third of post-sale conversations reach a finish.

基于政策页面回答的问题得分最高。需要实时系统查询或判断的问题得分最低。订单追踪通常是电商品牌最大的工单类别,但在此列表中排名靠后。在整个行业中,只有约三分之一的售后对话能到达终点。

Map your actual ticket mix onto that list before you choose a vendor. If most of your volume is “where’s my order,” the benchmark average overstates what you’ll get.

在选择供应商之前,请将你实际的工单组合映射到该列表中。如果你的大部分流量是“我的订单在哪里”,那么基准平均值会高估你将获得的成果。

The Most Common Failures Were Auth Walls and Clarification Loops

最常见的故障是认证墙和澄清循环

The benchmark found that the most common failures were authentication walls, repeated demands for an order number, clarification loops, and handing everything to a human. Judges rarely caught agents making up facts. They frequently caught agents that wouldn’t engage.

基准测试发现,最常见的故障是身份验证墙、重复要求提供订单号、澄清循环以及将所有问题转交给人工。评审人员很少发现代理编造事实的情况。他们经常发现代理拒绝互动。

The same vendor can score well on one store and near zero on another, and the benchmark attributes the difference to onboarding and quality control more than to the model in the demo. Nearly a third of the “AI chat” widgets it found produced no real conversation at all.

同一供应商在一家商店可能得分很高,而在另一家商店则接近零。基准测试将这种差异更多地归因于入职流程和质控,而非演示中的模型。它发现的“AI聊天”小部件中,近三分之一根本没有产生任何真正的对话。

Guardrails, identity checks, and escalation rules account for much of the gap between a 45% deployment and a 70% one, and the brand configures all three.

护栏、身份验证和升级规则解释了45%部署率与70%部署率之间的大部分差距,而品牌方配置了这三项功能。

Yuma Is the Slowest Vendor and Resolves 71%

Yuma是最慢的供应商,解决率为71%

Mean time to a complete reply runs from 5.4 seconds (Envive) to 16.5 seconds (Yuma). Envive is the fastest vendor on the board and resolves 22%. Yuma is the slowest and resolves 71%, with a top quality score.

完整回复的平均耗时从Envive的5.4秒到Yuma的16.5秒不等。Envive是榜单上最快的供应商,解决率为22%。Yuma是最慢的,但解决率达71%,且质量评分最高。

For post-purchase support, speed matters much less than resolution. The benchmark weights the support composite 50% automation, 40% quality, and 10% speed, on the basis that customers tolerate moderate latency when the answer is accurate. For pre-sale shopping, speed gets 25% of the weight.

对于购后支持,速度远不如解决率重要。基准测试将支持综合评分权重设为:自动化50%、质量40%、速度10%,其依据是客户在答案准确时可以容忍适度的延迟。而对于售前购物咨询,速度的权重为25%。

The latency numbers also have a caveat. The benchmark only averages turns it could time cleanly. A vendor’s slowest turns are the ones most likely to break or time out and drop out of the average, which can make that vendor look faster than it is.

延迟数据也存在一个注意事项。基准测试仅对能够清晰计时的轮次进行平均计算。供应商最慢的轮次最容易出错或超时,从而被排除在平均值之外,这可能会让该供应商看起来比实际更快。

Plan for 70% on a Top Vendor and 50% on the Rest

顶级供应商按70%规划,其余按50%规划

Build the business case on 65% to 70%, and only on a top-five vendor. The best tools in their best month still leave about 30% of conversations for a human or a second agent with deeper system access. On the AI bundled with Zendesk, Intercom, or Klaviyo, plan for 20% to 42%.

以65%至70%的解决率构建商业案例,且仅针对前五名的供应商。即使在最佳月份,最好的工具仍会将约30%的对话留给人类客服或拥有更深系统访问权限的第二代理工。对于Zendesk、Intercom或Klaviyo捆绑销售的AI,预计解决率在20%至42%之间。

Shortlist on automation and quality together. Only three vendors cleared both bars in this window. Ask every vendor for both numbers on the same set of conversations.

同时根据自动化能力和质量进行筛选。在此窗口期内,仅有三家供应商同时通过了这两项标准。要求每家供应商提供同一组对话的这两项数据。

Weight the results by your own ticket mix. Policy questions resolve well across the field. Order tracking and damaged items don’t.

根据你的工单组合对结果进行加权。政策类问题在整个行业中都能得到良好解决,但订单追踪和损坏商品则不然。

Run a cold test on the vendor’s reference stores. Take 30 real tickets from last month, open fresh incognito sessions on their live customers’ sites, and count full resolutions yourself. The rubric is public, so you can grade your current agent the same way.

在供应商的参考商店运行冷启动测试。从上个月选取30个真实工单,在其客户网站的实时页面上打开新的隐身会话,并自行统计完整解决的数量。评估标准是公开的,因此你可以用同样的方式评估你当前的客服代理。

更进一步:量化金融体系

看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力

进入量化体系 →

相似阅读

关联信息,但可能不是同一事件