跳到主内容
精选85Rohan Paul模型发布/更新

DeepReinforce 开源 Ornith-1.0 编程模型,超越 Cla…

Another fantastic open source release.

原文
推荐理由

做代码生成和 Agent 的同学注意了,这个开源模型在 SWE-Bench 上超过了 Claude Opus 4.7,而且 9B 小模型也有不错表现,值得立刻跑一下你的测试集。

Another fantastic open source release.

DeepReinforce just dropped Ornith-1.0, an MIT-licensed open-source family of agentic coding LLMs.

The flagship Ornith-1.0-397B MoE (17B-active) is the most powerful model in the release, reporting 82.4 on SWE-Bench Verified and 77.5 on Terminal-Bench 2.1 - surpassing Claude Opus 4.7 on both benchmarks.

Built on top of pretrained Gemma 4 and Qwen 3.5

Employs a novel self-improving training strategy. With this Ornith changes the training target by asking the model to improve both the answer and the task scaffold, meaning the plan, memory pattern, tool rhythm, error handling, and search process that shape the answer.

During RL, the model proposes a better scaffold first, then uses it to produce solution rollouts, and the reward updates both stages together.

That makes the model less like a coder following one rigid checklist and more like a coder learning which checklist works for each type of bug, repo, or terminal task.

The most interesting result is the 9B model reaching 69.4 on SWE-Bench Verified

更进一步:量化金融体系

看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力

进入量化体系 →

关联讨论

同一事件的更多信源

相似阅读

另一事件,读法相近