跳到主内容
@wquguru
精选88SaaStr 博客(RSS)产品与增长

Gorgias公开竞品评测:用开源基准测试与透明权重做获客

Everyone Should Publish the Deepest, Most Direct Competitive Evals They Can. Case Study: $100m ARR Gorgias for AI CX

原文
发到 X
推荐理由

给出一套完整的AI产品竞品评测SOP(从测试环境、评分标准到开源验证),独立开发者可直接复用此框架建立产品信任并优化GTM。

Everyone should publish the deepest, most direct competitive evals they can

每个人都应该发布他们能做到的最深入、最直接的竞品评估

Will there always be some bias? Yes

是否总会存在某些偏差?是的

But @gorgiasio did it the right way in ecomm CX:

但 @gorgiasio 在电商客服领域正确地做到了这一点:

👉Open-sourced the whole harness

👉开源了整个测试框架

👉8,356 live conversations, 18 vendors

👉8,356 条真实对话,涉及 18 家供应商

👉#1 in support, but they show Yuma… pic.twitter.com/nxajajk3Qc

👉在客服方面排名第一,但他们也展示了 Yuma…… pic.twitter.com/nxajajk3Qc

— Jason ✨👾SaaStr.Ai✨ Lemkin (@jasonlk) September 25, 2026

—— Jason ✨👾SaaStr.Ai✨ Lemkin (@jasonlk) 2026年9月25日

Every AI vendor in B2B says it’s the best, but too few show you the data. We’re all used to seeing “evals” of LLMs, but not enough post true detailed AI evals of their agents.

每个 B2B AI 供应商都声称自己是最好的,但很少有人向你展示数据。我们都习惯了看到 LLM 的“评估”,但很少有供应商发布其智能体(Agent)真正详细的 AI 评估报告。

That’s a mistake. Buyers can now test AI agents side by side in an afternoon. Agents are starting to build the shortlists, and in some cases, make the vendor decisions (they do for us at SaaStr AI).

这是一个错误。买家现在可以在一个下午内并排测试 AI 智能体。智能体开始构建短名单,在某些情况下甚至做出供应商选择决策(我们在 SaaStr AI 就是这样做的)。

So publish the deepest, most direct competitive evals you can. Run them on live products, against named competitors, and include the categories where you lose.

因此,尽可能发布最深入、最直接的竞品评估。在真实产品上运行,针对具体命名的竞争对手,并包含你处于劣势的类别。

Will there always be some bias? Yes. You wrote the rubric, picked the weights, and chose what to measure. Say that on the first page. Then make everything checkable.

是否总会存在某些偏差?是的。评分标准是你写的,权重是你选的,衡量指标也是你定的。把这些信息放在第一页。然后确保所有内容都可核查。

$100M ARR AI CX leader Gorgias just did this the right way in ecommerce CX. We know the company well — SaaStrFund led the seed round. And in fact, we pushed them to do the most honest, detailed evals in the CX+ space. And they did.

ARR 达 1 亿美元的 AI 客服领导者 Gorgias 最近在电商客服领域正确地做到了这一点。我们很了解这家公司——SaaStrFund 领投了其种子轮。事实上,我们推动他们在 CX+ 领域进行了最诚实、最详细的评估。而他们确实这么做了。

What Gorgias Published

Gorgias 发布了什么

Gorgias is at ~$100M ARR, and about 80% of that revenue comes from AI support for ecommerce brands. In just launched a public benchmark of its own agent against every major competitor:

Gorgias 的 ARR 约为 1 亿美元,其中约 80% 的收入来自为电商品牌提供的 AI 支持服务。最近,他们发布了一项公开基准测试,将其自有智能体与每个主要竞争对手进行对比:

  • 8,356 live conversations
  • 18 vendors
  • 212+ live storefronts
  • Refreshed weekly, with the same 10-turn conversation run against every vendor
  • The full test harness open-sourced on GitHub
  • 8,356 条真实对话
  • 18 家供应商
  • 212+ 个真实在线商店
  • 每周更新,使用相同的 10 轮对话流程对每家供应商进行测试
  • 完整的测试框架已在 GitHub 上开源

The results don’t all go Gorgias’s way, and it published them anyway.

结果并非全部有利于 Gorgias,但他们仍然发布了这些结果。

  • In support, Gorgias ranks #1, but Yuma resolves more conversations. Automation rate is the metric support buyers care about most, and a competitor leads on it, on Gorgias’s own page. [SUPPORT: Yuma X% vs. Gorgias Y%; Gorgias wins the composite on quality.]
  • In pre-sale, Envive beats Gorgias (for now at least): a composite score of 72 to 65. Gorgias has the best answer quality in the lane (76), but it averages 18.4 seconds per answer to Envive’s 7.9. The report’s own latency breakdown shows 28% of Gorgias’s shopping answers taking longer than 20 seconds.
  • Gorgias chose the weighting that cost it first place. The pre-sale composite gives speed 25% of the score; the support composite gives it 10%. Score pre-sale with the support weights and Gorgias comes out first at 74.3. It used the tougher weighting because shoppers leave when answers are slow.
  • 在客服方面,Gorgias 排名第一,但 Yuma 解决了更多的对话。自动化率是客服买家最关心的指标,而在这个指标上,一家竞争对手在 Gorgias 自己的页面上领先。[客服:Yuma X% vs. Gorgias Y%;Gorgias 在质量综合评分上获胜。]
  • 在售前环节,Envive 击败了 Gorgias(至少目前如此):综合得分为 72 比 65。Gorgias 在该赛道中拥有最佳的答案质量(76 分),但其平均每次回答耗时 18.4 秒,而 Envive 仅为 7.9 秒。报告自身的延迟细分显示,Gorgias 有 28% 的购物相关回答耗时超过 20 秒。
  • Gorgias 选择了使其失去第一名的权重。预售综合评分中,速度占25%;支持综合评分中,速度占10%。若用支持部分的权重来评估预售表现,Gorgias 以74.3分位居第一。它采用了更严格的权重,因为当回答缓慢时,顾客会离开。

Isn’t This Just “Benchmarking”?

这难道不是只是“基准测试”吗?

Yes, it’s a benchmark. It just goes much further than what most people picture when they hear the word.

是的,这是一个基准测试。但它比大多数人听到这个词时所想象的要深入得多。

Most B2B buyers know benchmarks as analyst quadrants, feature checklists, G2 grids, or a vendor’s “we’re 3x faster” slide. Those compare what vendors say their products do. They come from surveys, demos, and vendor-supplied answers, and they get updated once or twice a year.

大多数 B2B 买家所熟知的基准测试形式包括分析师象限图、功能清单、G2 网格或供应商的“我们快3倍”幻灯片。这些比较的是供应商声称其产品能做什么。它们来源于调查、演示和供应商提供的答案,并且每年只更新一两次。

An eval tests what the product actually does. It asks the AI real questions, captures every answer, and grades each one against a written rubric. In Gorgias’s case:

评估(eval)测试的是产品实际能做什么。它向 AI 提出真实问题,记录每一个回答,并根据书面评分标准对每个回答进行打分。在 Gorgias 的案例中:

  • It tests the live product. Each vendor’s agent is tested on real stores, through the same chat widget a shopper would use.
  • Every answer gets graded. Each of the 8,356 conversations is scored on whether it was resolved, whether the answer was right, and how long it took. The results don’t come from a single score someone assigned after a demo.
  • Every vendor gets the same test. The same 10-turn conversation runs against every vendor, from a simple first question through cart, shipping, and returns policy.
  • It keeps running. The tests run daily and results are published weekly, so if a vendor ships a better model next month, it shows up in next month’s numbers.
  • Anyone can rerun it. The code is public, so anyone who doubts the results can check them.
  • 它测试的是实时产品。每个供应商的代理都在真实商店中进行测试,使用的是顾客会使用的相同聊天小部件。
  • 每个回答都会得到评分。8,356 次对话中的每一次都根据是否解决、回答是否正确以及耗时多久来进行评分。结果并非来自某人在演示后给出的单一分数。
  • 每个供应商接受相同的测试。同样的10轮对话应用于所有供应商,从简单的第一问开始,涵盖购物车、配送和退货政策。
  • 测试持续运行。测试每日执行,结果每周发布,因此如果供应商下个月推出了更好的模型,它会在下个月的数字中体现出来。
  • 任何人都可以重新运行测试。代码是公开的,因此任何怀疑结果的人都可以自行核查。

This matters more for AI than it did for traditional software. A CRM does the same thing every time you click the same button. An AI agent can give a great answer on Monday and a wrong one on Thursday, and it can get worse after a model update without anyone at the vendor noticing. A feature checklist can’t show any of that. Only continuous testing of real conversations can.

这对 AI 而言比对传统软件更为重要。CRM 在你每次点击相同按钮时所做的操作都是一样的。而 AI 代理可能在周一给出出色的回答,却在周四给出错误回答,并且在模型更新后可能变得更差,而供应商方面甚至并未察觉。功能清单无法展示任何这些情况。只有对真实对话的持续测试才能做到这一点。

Why Publishing Your Evals Wins

为何公开你的评估结果会带来优势

Buyers already discount your marketing. Every support leader has seen a dozen vendor charts where the vendor wins, and they ignore all of them. When a vendor shows competitors ahead in some categories, buyers believe the categories where it’s ahead. Gorgias’s #1 in support means more because Yuma’s automation lead is on the same page.

买家已经对你的营销内容打折扣。每位支持团队负责人都曾见过十几张由供应商获胜的图表,并对此视而不见。当供应商展示出自己在某些类别上落后于竞争对手时,买家反而更相信它在其他领先类别中的表现。Gorgias 在支持方面排名第一之所以更有说服力,是因为 Yuma 的自动化主管对此持相同看法。

You’re doing the buyer’s diligence for them. No ecommerce brand is going to run 8,356 conversations across 13 vendors. Most sit through three demos and maybe run one pilot. When you run the full comparison and publish all of it, your report becomes what buyers rely on. You also get to decide what it measures, which is a real advantage, and it’s the reason being open about the method matters.

你是在帮他们做买方的尽职调查。没有哪个电商品牌会在13家供应商之间进行8,356次对话。大多数客户只会参加三次演示,或许再运行一次试点项目。当你进行全面对比并公开所有内容时,你的报告就成了买方依赖的依据。你还拥有决定衡量标准的权力,这构成了真正的优势,也是为何公开方法论至关重要的原因。

Agents are starting to do the shortlisting. More software evaluations now start with a question to Claude or ChatGPT. An agent can read, check, and cite a versioned rubric, public scoring weights, and open code. It can’t do much with a gated PDF. Vendors that publish structured, checkable evals will get cited more. The live Gorgias report currently blocks automated readers in its robots.txt, though the GitHub repo is open. That’s worth fixing.

AI代理开始参与初选。越来越多的软件评估现在会先向Claude或ChatGPT提问。代理可以阅读、核查并引用带版本号的评分标准、公开的权重以及开源代码。但对于受访问限制的PDF文件,它却无能为力。发布结构化、可核查的评估报告的供应商将获得更多引用。目前的Gorgias报告在robots.txt中屏蔽了自动化读取器,尽管其GitHub仓库是开放的。这一点值得修复。

Your team sees the gap every week. Inside Gorgias, Envive’s 7.9 seconds and Yuma’s resolution rate aren’t competitive intel buried in a deck anymore. They’re public, next to Gorgias’s own numbers, and updated weekly. The rubric is versioned in the repo, so nobody can quietly change the scoring to make a gap disappear.

你的团队每周都能看到这一差距。在Gorgias内部,Envive的7.9秒响应时间和Yuma的解决率不再是隐藏在PPT中的竞争情报。它们是公开的,与Gorgias自身的数据并列展示,并且每周更新。评分标准在仓库中进行了版本控制,因此没有人能悄悄修改评分以掩盖差距。

How to Do It the Right Way

如何正确地做到这一点

The Gorgias approach is a good template:

Gorgias的方法是一个很好的模板:

  • Test live products. Gorgias tested real storefronts, not sandboxes or demo environments.
  • Name every competitor. An anonymized “Competitor A” tells a buyer nothing.
  • Blind the judge. Gorgias strips vendor and store names before its AI judge scores a conversation. A second adversarial review pass has to agree with the first at least 90% of the time.
  • Open-source the harness. Competitors who dispute the results can run the code themselves.
  • Weight for the end customer. Gorgias weighted speed heavily enough in pre-sale that it lost the top spot.
  • Publish the trend alongside the snapshot. A weekly chart is much harder to cherry-pick than a single number.
  • State the bias up front. Gorgias’s repo describes the project as “competitive intelligence” and says to treat cross-vendor numbers as directional.
  • 测试真实产品。Gorgias测试的是真实的商店前台,而不是沙箱或演示环境。
  • 列出所有竞争对手。匿名的“竞争对手A”对买方来说毫无意义。
  • 对评审者进行盲测。Gorgias在其AI代理对对话进行评分前,会去除供应商和商店的名称。第二次对抗性审查的结果必须至少90%的时间与第一次审查结果一致。
  • 开源测试框架。对结果有异议的竞争对手可以自行运行代码。
  • 以客户终局体验为权重。Gorgias在售前阶段给予了速度极高的权重,以至于导致其失去了榜首位置。
  • 发布趋势图 alongside 快照数据。每周的图表比单个数字更难被断章取义。
  • 预先声明偏见。Gorgias的仓库将该项目描述为“竞争情报”,并建议将跨供应商的数据视为方向性参考。

Sources: Gorgias AI Agent Benchmark, gorgias/ai-agent-benchmark on GitHub

来源:Gorgias AI Agent Benchmark,GitHub上的gorgias/ai-agent-benchmark

更进一步:量化金融体系

看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力

进入量化体系 →

相似阅读

关联信息,但可能不是同一事件