跳到主内容
@wquguru
精选70elvis行业动态

AI Agent评估框架膨胀问题引关注

Highly-recommended read.

原文
发到 X

Highly-recommended read.

Aligns with what I see in my own harness:

> Pi harness got the same success rate as harnesses from the LLM vendors with Opus and GPT, but at 2x less cost

> GLM 5.2 was a major step forward in open-source coding agent performance

The harness bloat/rot is real!

更进一步:量化金融体系

看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力

进入量化体系 →

相似阅读

另一事件,读法相近