跳到主内容
@wquguru
精选88Hacker News Best(web_list)模型发布/更新

Real-SWE:基于企业私有代码库的AI编程基准测试

Real-SWE:基于企业私有代码库的AI模型基准测试

原文
发到 X
推荐理由

首次引入私有企业代码库作为评测场景,直击Agent落地痛点,结果对选型极具参考价值。

September 2026

2026年9月

Introducing Real-SWE

推出 Real-SWE

By Snagnik Das·Siddhant Paliwal·Janak Sunil

作者:Snagnik Das、Siddhant Paliwal、Janak Sunil

Benchmarking frontier AI models on private, real-world, enterprise codebases.

在私有、真实的企业代码库上对前沿 AI 模型进行基准测试。

代码 · 23
                                                                              .
;+;+;:;+;+;.#
;           ;
; .   . . . ;                                    .
+           ;
; .     .   : @;;+;+#                            +;+;+;+;:;+:#
:           : ;     .       ;                    ;           ;              @;+;+.+;:;+;+*
; . .       : ; . . :       ;       @++:;:;+.;;@ ;. . .   . .:     .        ; .     . . .;  @++;;;;+;:;@
;           : ;     ; @;+++:@;++;:+ ;          : :           . :+.+;;;+:;@  ;            :  +          .
; . .   .   ; ;   . ; : . . . .     ; . . . . .; ;  . . . . .+ +  . . . .;  ; . .   . . .;  ;     . .  :
;           : ;     ; ;           ; :          ; .           ; .         ;  ;            ;  ;          :
;       . . %   . . ; : . .   . . ; ; ... . . .; +. .   . . .: ;. . :.. .:  : . . ::. .  .  ; .   . . .;
:        *. @ ;     ; +.@     .   ; : +:    *+ : ;    :.     %+;    +.   ;  *.    *#     @: :.*.       ;
; .   . .;::@;; . . ; :#@;. ..@:. . ;.@%: .:@%:: :. . #+.   :@%;. .:@*. .:  @+  . #*. . ;@#.;.*   . . .:
.::: ..:%@@+@+#... :+:%@@+..:@@@::*.:;@@:..+@@:; *.::#@@#:..:@@#...+@#: :*:%@@+ .:.+::..*@@:@@@#;.. ::.#
        +@#.@         .*@:  .*@%:     *:   :@%.      ;@@:   .@%.    @     .;@@:   .:    :@#..#@*.
        +@#.@          .@   .*@%:.    *:    %+       ;@@:    #;     +#    .;@@:   ..     @: .#@*.
         @. ..         ;@   .*@%:     *:    @@       ;@@:    %#    * #    .;@@:          @+  .@
        :*#..          @::    @.       .   ::;.       @@     ..     .       @%          .*@  ;@.
        %.@            ..    .@;            .         @@     ..            .@@.          ..  @:#
        ...             .    ;%#             .        @@                   +;#:          .. .@ @
         .                   ...                      ...                   ..               ..
                              ...                     ..                   ...                .
代码 · 23
                                                                              .
;+;+;:;+;+;.#
;           ;
; .   . . . ;                                    .
+           ;
; .     .   : @;;+;+#                            +;+;+;+;:;+:#
:           : ;     .       ;                    ;           ;              @;+;+.+;:;+;+*
; . .       : ; . . :       ;       @++:;:;+.;;@ ;. . .   . .:     .        ; .     . . .;  @++;;;;+;:;@
;           : ;     ; @;+++:@;++;:+ ;          : :           . :+.+;;;+:;@  ;            :  +          .
; . .   .   ; ;   . ; : . . . .     ; . . . . .; ;  . . . . .+ +  . . . .;  ; . .   . . .;  ;     . .  :
;           : ;     ; ;           ; :          ; .           ; .         ;  ;            ;  ;          :
;       . . %   . . ; : . .   . . ; ; ... . . .; +. .   . . .: ;. . :.. .:  : . . ::. .  .  ; .   . . .;
:        *. @ ;     ; +.@     .   ; : +:    *+ : ;    :.     %+;    +.   ;  *.    *#     @: :.*.       ;
; .   . .;::@;; . . ; :#@;. ..@:. . ;.@%: .:@%:: :. . #+.   :@%;. .:@*. .:  @+  . #*. . ;@#.;.*   . . .:
.::: ..:%@@+@+#... :+:%@@+..:@@@::*.:;@@:..+@@:; *.::#@@#:..:@@#...+@#: :*:%@@+ .:.+::..*@@:@@@#;.. ::.#
        +@#.@         .*@:  .*@%:     *:   :@%.      ;@@:   .@%.    @     .;@@:   .:    :@#..#@*.
        +@#.@          .@   .*@%:.    *:    %+       ;@@:    #;     +#    .;@@:   ..     @: .#@*.
         @. ..         ;@   .*@%:     *:    @@       ;@@:    %#    * #    .;@@:          @+  .@
        :*#..          @::    @.       .   ::;.       @@     ..     .       @%          .*@  ;@.
        %.@            ..    .@;            .         @@     ..            .@@.          ..  @:#
        ...             .    ;%#             .        @@                   +;#:          .. .@ @
         .                   ...                      ...                   ..               ..
                              ...                     ..                   ...                .

01Introduction

01 引言

Today we are releasing Real-SWE, a benchmark that evaluates frontier AI models on private, real-world, enterprise codebases. Each task comes from a private production codebase that we licensed from a real-world company. These are problems their engineers work on, with all the context and complexity that comes with an existing product.

今天,我们发布 Real-SWE,这是一个用于评估前沿 AI 模型在私有、真实企业代码库上表现的基准测试集。每个任务都源自我们从一家真实公司获得许可的私有生产代码库。这些问题由他们的工程师处理,包含了现有产品所附带的所有上下文和复杂性。

  • Private codebases. Agents must navigate proprietary systems whose code and solutions aren’t available on the public internet.
  • Work with business consequences. Getting billing right, calculating taxes, migrating customers. Changes that affect how a business runs, often across multiple services.
  • Company-specific complexity. Every company has its own rules and ways of writing code. Agents have to understand those conventions and make changes that work with what’s already there.
  • 私有代码库。智能体必须导航专有系统,其代码和解决方案在互联网上不可公开获取。
  • 涉及业务后果的工作。正确处理计费、计算税款、迁移客户等。这些变更会影响业务的运行方式,通常涉及多个服务。
  • 特定公司的复杂性。每家公司都有自己独特的规则和编码规范。智能体必须理解这些惯例,并做出与现有内容兼容的更改。

Can a coding agent actually do the work of a software engineer in the real world?

编码智能体真的能在现实世界中胜任软件工程师的工作吗?

  • 1
  • Fable 5.1
  • Claude Code
  • Resolution rate: 38.8%
  • 2
  • GPT-6 Astra
  • Codex CLI
  • Resolution rate: 33.8%
  • 3
  • Gemini 3.8 Flash
  • Gemini CLI
  • Resolution rate: 31.2%
  • 4
  • GLM 5.3
  • Claude Code
  • Resolution rate: 28.8%
  • =5
  • Grok 4.6
  • Grok Build
  • Resolution rate: 23.8%
  • =5
  • Muse Spark 1.3
  • Muse Code
  • Resolution rate: 23.8%
  • 7
  • Kimi K3
  • Kimi Code
  • Resolution rate: 18.8%
  • 8
  • GPT-5.6 Sol
  • Codex CLI
  • Resolution rate: 16.2%
  • 1
  • Fable 5.1
  • Claude Code
  • 解决率:38.8%
  • 2
  • GPT-6 Astra
  • Codex CLI
  • 解决率:33.8%
  • 3
  • Gemini 3.8 Flash
  • Gemini CLI
  • 解决率:31.2%
  • 4
  • GLM 5.3
  • Claude Code
  • 解决率:28.8%
  • =5
  • Grok 4.6
  • Grok Build
  • 解决率:23.8%
  • =5
  • Muse Spark 1.3
  • Muse Code
  • 解决率:23.8%
  • 7
  • Kimi K3
  • Kimi Code
  • 解决率:18.8%
  • 8
  • GPT-5.6 Sol
  • Codex CLI
  • 解决率:16.2%
#ModelHarnessResolution rate
1Fable 5.1Claude Code38.8%
2GPT-6 AstraCodex CLI33.8%
3Gemini 3.8 FlashGemini CLI31.2%
4GLM 5.3Claude Code28.8%
=5Grok 4.6Grok Build23.8%
=5Muse Spark 1.3Muse Code23.8%
7Kimi K3Kimi Code18.8%
8GPT-5.6 SolCodex CLI16.2%
#模型运行环境解决率
1Fable 5.1Claude Code38.8%
2GPT-6 AstraCodex CLI33.8%
3Gemini 3.8 FlashGemini CLI31.2%
4GLM 5.3Claude Code28.8%
=5Grok 4.6Grok Build23.8%
=5Muse Spark 1.3Muse Code23.8%
7Kimi K3Kimi Code18.8%
8GPT-5.6 SolCodex CLI16.2%

Resolution rate is equivalent to pass@1, averaged over eight independent runs per task. 95% confidence intervals are shown.

解决率等同于 pass@1,按每项任务进行八次独立运行的平均值计算。图中显示了 95% 的置信区间。

Expert-generated or synthetic tasks can be well designed, but they aren’t the verbatim, actual tasks that engineers in real companies need to do. Our tasks differ on two axes: the underlying coding artifact and specificity of the instruction. Both add complexities that challenge today’s frontier models.

专家生成或合成的任务可以设计得很好,但它们并非真实企业中工程师需要执行的逐字、实际任务。我们的任务在两个维度上有所不同:底层编码工件和指令的具体性。这两者都增加了复杂性,对当今的前沿模型构成了挑战。

We use native harnesses to reflect how enterprise engineers work in practice, evaluating model-and-harness combinations rather than models in isolation.

我们使用原生运行环境以反映企业工程师的实际工作方式,评估的是模型与运行环境的组合,而非孤立地评估模型。

Real company tasks require company-specific context

真实公司的任务需要特定于公司的上下文

Correct billing depends on business rules and external services

正确的计费依赖于业务规则和外部服务

Fix invoice billing so each business charges the right tax and exempt customers aren't taxed.

修复发票计费,使每个业务收取正确的税款,且享有豁免权的客户不被征税。

View full instructionHide full instruction▾

查看完整指令隐藏完整指令▾

Billing reopens on Monday and every invoice this service issues is coming out untaxed. Each business on the platform settles its tax a different way: some maintain a rate themselves, some want each invoice priced against the buyer's destination by our tax authority provider, and some collect nothing at all, while a customer we hold an exemption for is charged nothing whichever way its business is configured. Pricing a destination means going to the authority with both addresses, the priced lines and the product category that business sells under, on the sandbox or the production authority according to the account the business is on; an address the authority refuses must be reported without stopping the invoice. The rate, the tax and the gross belong on the issued invoice, and once an invoice is settled the sale is filed back to the authority under that invoice's number so the returns reconcile. Invoices between European parties show both sides' VAT registrations. The authority and ledger are available at TAX_JAR_URL, PROD_TAX_JAR_URL and INFLUX_URL.

周一重新开放服务,该服务发出的每张发票均不含税。平台上的每个业务以不同方式结算其税款:有些自行维护税率,有些希望由我们的税务当局提供商根据买方的目的地对每张发票定价,有些则完全不收取任何款项,而我们为其持有豁免权的客户无论其业务如何配置均不收费。对目的地定价意味着需向当局提供双方地址、已定价的行项目以及该业务销售的产品类别;根据业务所属账户,需在沙箱或生产环境中访问相应的当局;若某地址被当局拒绝,必须在不停止发票的情况下报告。税率、税款和总额应体现在已开具的发票上,一旦发票结算完成,销售记录便以该发票编号归档至当局,以便退货对账。欧洲各方之间的发票显示双方的增值税注册号。当局和分类账可通过 TAX_JAR_URL、PROD_TAX_JAR_URL 和 INFLUX_URL 获取。

Services in the sandbox

沙箱中的服务

  • TJTaxJar sandbox
  • TJTaxJar production
  • InfluxDB ledger
  • NestJS service
  • TypeScript
  • TJTaxJar 沙箱
  • TJTaxJar 生产环境
  • InfluxDB 分类账
  • NestJS 服务
  • TypeScript

Agents work across code, infrastructure, and business tools

智能体在代码、基础设施和业务工具之间协同工作

Tools and services across Real-SWE task environments. Each task exposes only the services its workflow needs.

面向 Real-SWE 任务环境的各类工具与服务。每个任务仅暴露其工作流所需的各项服务。

  • AWS emulator
  • Docker
  • Kubernetes
  • GitHub
  • Linear MCP
  • PostgreSQL
  • MySQL
  • MongoDB
  • GeGel
  • Redis
  • Go
  • Python
  • Node.js
  • Vitest
  • Slack
  • Intercom
  • Google Drive
  • Email
  • ClickUp
  • AWS 模拟器
  • Docker
  • Kubernetes
  • GitHub
  • Linear MCP
  • PostgreSQL
  • MySQL
  • MongoDB
  • GeGel
  • Redis
  • Go
  • Python
  • Node.js
  • Vitest
  • Slack
  • Intercom
  • Google Drive
  • 电子邮件
  • ClickUp

Codebase Selection

代码库选择

We selected codebases through a rigorous screening process, focusing on real companies with substantial usage, strong engineering teams, and demanding production workloads. The sample tasks analyzed below come from these codebases, including:

我们通过严格的筛选流程选取了代码库,重点关注拥有大量实际用户、强大的工程团队以及高要求生产负载的真实公司。以下分析中的示例任务来自这些代码库,包括:

  • A Luma/Partiful competitor with 200K+ users and a top 100 App Store ranking
  • A consumer fintech platform processing 100K+ bank statements
  • Enterprise AI sales platforms supporting complex business workflows
  • 一个拥有超过20万用户且位列App Store前100名的Luma/Partiful竞争对手
  • 一个处理超过10万份银行对账单的消费级金融科技平台
  • 支持复杂业务流程的企业级AI销售平台

We prioritize code written to meet an actual user or business need over code written solely to create a benchmark task. Production engineering requires understanding existing architecture, preserving behavior that users rely on, and making changes within real operational constraints.

我们优先采用为满足实际用户或业务需求而编写的代码,而非仅为创建基准任务而编写的代码。生产工程要求理解现有架构、保留用户依赖的行为,并在真实的运营约束范围内进行修改。

Brief instructions can require changes across many files

简短的指令可能需要在多个文件中进行更改

Our tasks describe the change needed, leaving agents to discover implementation details in the codebase and surrounding tools. Any behavior required by the verifier must be stated or reasonably discoverable. This leads to our prompts being slightly underspecified, about par with DeepSWE and Terminal Bench, but specific enough to not omit instructions.

我们的任务描述了所需的变更,让智能体在代码库和周围工具中自行发现实现细节。验证器所要求的任何行为都必须明确说明或可合理发现。这导致我们的提示词略显不够具体,与DeepSWE和Terminal Bench相当,但足够具体以不遗漏指令。

The work is cross-functional and complex: a single change can span multiple parts of the application. Agents must understand existing business logic and company coding patterns while keeping the surrounding system working.

这项工作跨职能且复杂:单个更改可能涉及应用程序的多个部分。智能体必须理解现有的业务逻辑和公司编码模式,同时保持周围系统的正常运行。

Prompt length · median

提示词长度 · 中位数

A typical Real-SWE instruction is 1,742 characters.

典型的Real-SWE指令长度为1,742个字符。

  • FrontierCode2,056 chars
  • DeepSWE1,975 chars
  • Terminal-Bench 31,584 chars
  • FrontierSWE v2992 chars
  • Real-SWE1,742 chars
  • FrontierCode 2,056个字符
  • DeepSWE 1,975个字符
  • Terminal-Bench 31,584个字符
  • FrontierSWE v2 992个字符
  • Real-SWE 1,742个字符

Files edited by the reference solution · median

参考解决方案编辑的文件数 · 中位数

11 files in Real-SWE, compared with 6 in FrontierCode and DeepSWE.

Real-SWE中有11个文件,相比之下FrontierCode和DeepSWE中为6个。

更进一步:量化金融体系

看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力

进入量化体系 →

相似阅读

关联信息,但可能不是同一事件