跳到主内容
@wquguru
精选75Rohan Paul产品发布/更新

DoorDash开源模型组合提升AI代码审查召回率至65%

DoorDash showed open-source models can lift AI code review when used together.

原文
发到 X

DoorDash showed open-source models can lift AI code review when used together.

And that harnesses now matter.

Single-pass AI reviewers caught only 30.7% of weighted PR issues. The company’s production reviewer raised that to 53.6%, while costing $3.91 per PR.

That jump came from splitting review into two jobs, scouting first, then verifying.

A scout model scans the diff for risky areas before deeper reviewers test claims.

The strongest run used Kimi K2.6 as scout and Claude Fable 5 as reviewer.

That setup hit 65.2% weighted recall, 75.3% F1, and $3.81 per PR.

Composer 2.5 with GPT 5.5 medium reached 92.2% precision, but much lower recall.

更进一步:量化金融体系

看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力

进入量化体系 →

相似阅读

另一事件,读法相近