跳到主内容
@wquguru
精选75elvis论文研究

JIT-Agent:动态生成 Agent 运行框架,性能超越 GPT-5.6

If you maintain a hand-built agent harness, this one is worth your time.

原文
发到 X

If you maintain a hand-built agent harness, this one is worth your time.

如果你维护一个手工构建的智能体框架,这篇值得你花时间阅读。

(bookmark it)

(收藏它)

I feel like everyone is sleeping on the idea of dynamically generating agent harnesses on the fly.

我觉得大家都忽视了动态生成智能体框架的想法。

As you aim to own your harness, this is a topic more devs will lean into. Here is a great report discussing this topic.

当你打算拥有自己的框架时,这是更多开发者会深入探讨的话题。这里有一份关于此话题的精彩报告。

JIT-Agent is a model whose output is an agent harness.

JIT-Agent 是一个输出为智能体框架的模型。

It formalizes the harness as a composable artifact under a fixed four-module protocol covering memory, planning, action protocol, and tool orchestration, then synthesizes one on the fly for any off-the-shelf agentic LLM.

它将框架形式化为一个可组合的构件,在固定的四模块协议下运行,涵盖记忆、规划、行动协议和工具编排,然后为任何现成的智能体 LLM 即时合成一个。

It also repairs harnesses mid-execution and self-evolves by distilling performance signals from an expanding archive of prior configurations.

它还能在执行过程中修复框架,并通过从不断扩展的先前配置档案中提炼性能信号来自我进化。

With JIT-Agent attached, DeepSeek-V4-Flash surpasses GPT-5.6 on DeepSearchQA (+9.1) and OdysseyBench (+4.3). GLM-5.2 gains up to +20.2 points.

配备 JIT-Agent 后,DeepSeek-V4-Flash 在 DeepSearchQA(+9.1)和 OdysseyBench(+4.3)上超越了 GPT-5.6。GLM-5.2 获得了高达 +20.2 分的提升。

The generated harnesses are also performance-competitive with mature runtimes like OpenCode and Claude Code.

生成的框架在性能上也与成熟的运行时(如 OpenCode 和 Claude Code)相媲美。

Paper: https://arxiv.org/abs/2608.25593

论文:https://arxiv.org/abs/2608.25593

Chat with Paper: https://academy.dair.ai/papers/jit-agent-a-model-that-writes-your-agent-harness-2608.25593

与论文对话:https://academy.dair.ai/papers/jit-agent-a-model-that-writes-your-agent-harness-2608.25593

更进一步:量化金融体系

看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力

进入量化体系 →

相似阅读

另一事件,读法相近