跳到主内容
@wquguru
精选88Rohan Paul论文研究

FrogNano:4B编码智能体SWE-bench达61.5%无需蒸馏

A 4B coding agent reached 61.5% on SWE-bench Verified without frontier-model dis…

原文
发到 X
推荐理由

小模型编码Agent的新范式,不靠蒸馏靠自适应合成数据,值得研究Efficient AI的同学关注。

A 4B coding agent reached 61.5% on SWE-bench Verified without frontier-model distillation by combining a simpler tool interface with synthetic tasks that keep adapting to what the model can currently learn.

在不使用前沿模型蒸馏的情况下,一个4B编码智能体通过结合更简单的工具接口与能够持续适应模型当前学习能力的合成任务,在SWE-bench Verified上达到了61.5%的成绩。

FrogNano starts from Qwen3.5-4B and is trained with RL on about 1,500 synthetic software-engineering tasks.

FrogNano基于Qwen3.5-4B启动,并在约1,500个合成软件工程任务上使用强化学习进行训练。

The important part is how those tasks are chosen.

关键在于这些任务是如何被选定的。

As the model improves, the system generates fresh problems that are challenging but still learnable, so the curriculum improves with the agent.

随着模型的进步,系统会生成具有挑战性但依然可学习的最新问题,从而使课程随智能体的能力提升而优化。

The interface matters just as much.

接口设计同样至关重要。

Switching to a simpler 5-tool setup moved the base model from 8.3% to 37.2% on SWE-bench Verified.

切换至更简单的5工具设置后,基础模型在SWE-bench Verified上的得分从8.3%提升至37.2%。

After 5 rounds, FrogNano reached 61.5%.

经过5轮训练后,FrogNano达到了61.5%。

– arxiv. org/abs/2609.07925

– arxiv.org/abs/2609.07925

Title: "FrogNano: Training a 4B Coding Agent via Online Task Synthesis"

标题:《FrogNano:通过在线任务合成训练一个4B编码智能体》

更进一步:量化金融体系

看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力

进入量化体系 →

相似阅读

关联信息,但可能不是同一事件