跳到主内容
@wquguru
精选88Rohan Paul论文研究

JIT-Agent:按需生成智能体框架,小模型超越大模型

A smaller model with the right task-specific harness can beat a stronger model.

原文
发到 X
推荐理由

这篇论文展示了一种通过动态优化智能体编排来提升模型效能的新范式,不仅让小模型击败了更强的大模型,还显著降低了推理成本,对 Agent 架构设计极具参考价值。

A smaller model with the right task-specific harness can beat a stronger model.

一个规模较小但具备合适任务专用框架的模型,可以击败更强的模型。

New paper introduces JIT-Agent, which takes exactly this approach: it generates the agent harness on demand, choosing how memory, planning, actions, tools, and skills should work for the specific task.

新论文引入了 JIT-Agent,它正是采取这种方法:按需生成智能体框架,决定内存、规划、动作、工具和技能应如何针对特定任务进行协作。

DeepSeek-V4-Flash with JIT-Agent scored 85.1 on DeepSearchQA, versus 76.0 for GPT-5.6.

DeepSeek-V4-Flash 配合 JIT-Agent 在 DeepSearchQA 上得分 85.1,而 GPT-5.6 得分为 76.0。

That change was enough to move the same backbones substantially. Across 18 matched backbone-benchmark pairs, every JIT-generated harness improved the underlying model; the 9-benchmark average rose by 7.7 points for GLM-5.2 and 8.8 for DeepSeek-V4-Flash.

这一改变足以显著提升相同基础模型的效能。在 18 对匹配的基础模型-基准测试对中,每个由 JIT 生成的框架都提升了底层模型的表现;GLM-5.2 的平均分提高了 7.7 分,DeepSeek-V4-Flash 提高了 8.8 分(基于 9 个基准测试的平均值)。

This was not simply more inference. In controlled comparisons against fixed harnesses such as Claude Code, Codex, OpenCode, Hermes, and NanoBot, JIT-Agent had the lowest token use and API cost in all 6 settings, with 36.0% lower cost on average than the cheapest fixed alternative.

这并非简单的增加推理量。在与 Claude Code、Codex、OpenCode、Hermes 和 NanoBot 等固定框架的控制对比中,JIT-Agent 在所有 6 种设置下均使用最少的 token 数和最低的 API 成本,平均成本比最便宜的固定替代方案低 36.0%。

– arxiv. org/abs/2608.25593

– arxiv.org/abs/2608.25593

Title: "JIT-Agent: Scaling Harness Intelligence via Just-in-Time Harness Evolution"

标题:《JIT-Agent:通过即时框架演化扩展框架智能》

更进一步:量化金融体系

看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力

进入量化体系 →

相似阅读

另一事件,读法相近