跳到主内容
@wquguru
精选75Rohan Paul技巧与观点

固定模型改框架:编码Agent结果可大幅提升

Keep the model fixed, change the harness, and coding-agent results can move a lo…

原文
发到 X

Keep the model fixed, change the harness, and coding-agent results can move a lot when context gets tight.

保持模型不变,更换工作框架,当上下文变得紧张时,编码代理的结果可能会有很大波动。

The study compares 2 configurations of Yuj: control keeps the full chronological transcript until context fills, while treatment shortens older tool outputs, detects stalled behavior, and applies fixed command safeguards while preserving the full record.

该研究比较了Yuj的两种配置:对照组保留完整的按时间顺序的转录直到上下文填满,而处理组缩短较旧的工具输出,检测停滞行为,并在保留完整记录的同时应用固定的命令保障措施。

On 169 SWE-bench Verified tasks with a 20,480-token window, Qwen3.6’s mean per-task F2PF rose from 28% to 49%, while complete solutions increased from 43 to 72.

在169个SWE-bench Verified任务中,使用20,480个令牌的窗口,Qwen3.6的平均每任务F2PF从28%上升到49%,而完整解决方案从43个增加到72个。

The same frozen treatment improved both outcomes for Devstral, Nemotron, and Qwen3.8 without retuning.

相同的冻结处理在不重新调整的情况下,对Devstral、Nemotron和Qwen3.8的结果均有改善。

At 262,144 tokens, however, Verified and Pro outcomes were nearly identical between arms, so the benefit appears strongest when context is actually binding.

然而,在262,144个令牌时,Verified和Pro的结果在两组之间几乎相同,因此这种益处似乎在上下文真正受限时最为显著。

Treatment also used more model work under pressure, and the experiment tests the package as a whole rather than isolating each mechanism.

处理组在压力下也使用了更多的模型工作,该实验将整个包作为一个整体进行测试,而不是孤立每个机制。

For coding agents, “which model?” is no longer enough. Benchmark the model, harness, context policy, tools, and run controls as one solver.

对于编码代理来说,“用哪个模型?”已不再足够。将模型、工作框架、上下文策略、工具和运行控制作为一个整体求解器进行基准测试。

更进一步:量化金融体系

看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力

进入量化体系 →

相似阅读

另一事件,读法相近