ClawGym II:用RL在OpenClaw与Claude
Really interesting paper.
Really interesting paper.
非常有趣的论文。
I recommend it to anyone interested in training agents using existing harnesses.
我推荐给任何对使用现有框架训练智能体感兴趣的人。
(bookmark it)
(收藏它)
ClawGym II runs RL through OpenClaw and Claude Code as opaque boxes. A serving proxy sits at the model boundary and captures every call the harness makes, then those calls get organized into prefix trees so PPO and GRPO can optimize over the recovered multi-turn structure.
ClawGym II 通过 OpenClaw 和 Claude Code 作为不透明盒子运行强化学习。一个服务代理位于模型边界,捕获框架进行的每一次调用,然后将这些调用组织成前缀树,以便 PPO 和 GRPO 能够在恢复的多轮结构上进行优化。
Qwen3-30A3B gains 9.98 points of Pass@1 through OpenClaw and 14.81 through Claude Code, stable across 200 to 400 optimization steps.
Qwen3-30A3B 通过 OpenClaw 在 Pass@1 上获得 9.98 分的提升,通过 Claude Code 获得 14.81 分的提升,在 200 到 400 个优化步骤中保持稳定。
Mix-harness training pushes further. One model gets optimized jointly by heterogeneous harnesses, which points at policies that generalize across execution systems instead of overfitting to a single one.
混合框架训练进一步推进。一个模型通过异构框架联合优化,这指向了跨执行系统泛化的策略,而不是过度拟合单一系统。
Paper: https://arxiv.org/abs/2608.16798
论文:https://arxiv.org/abs/2608.16798
Track more trending AI papers in our academy: https://academy.dair.ai/
在我们的学院中追踪更多热门 AI 论文:https://academy.dair.ai/
更进一步:量化金融体系
看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力