精选70Zach Lloyd技巧与观点
利用评分器检测 Agent 代码审查能力缺陷并调优
We have a number of scorers running against our internal Warp factory, one of wh…
We have a number of scorers running against our internal Warp factory, one of which grades the verbosity of agent conversations.
我们在内部 Warp 工厂运行了多个评分器,其中一个用于评估代理对话的冗长程度。
I made a change earlier to make it stricter because our agents are creating overly verbose PR descriptions. This caused the scorer pass rate to drop.
我之前做了一个修改,使其更加严格,因为我们的代理生成的 PR 描述过于冗长。这导致评分器的通过率下降。
The scorer run details show how it detected one particularly bad PR description.
评分器运行详情展示了它是如何检测到一个特别糟糕的 PR 描述的。
This is a simple example of how we can detect our code review skill needs tuning.
这是一个简单的例子,说明我们如何发现代码审查技能需要调整。
更进一步:量化金融体系
看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力