JIT-Agent:动态生成 Agent 运行框架,性能超越 GPT-5.6
If you maintain a hand-built agent harness, this one is worth your time.
If you maintain a hand-built agent harness, this one is worth your time.
如果你维护一个手工构建的智能体框架,这篇值得你花时间阅读。
(bookmark it)
(收藏它)
I feel like everyone is sleeping on the idea of dynamically generating agent harnesses on the fly.
我觉得大家都忽视了动态生成智能体框架的想法。
As you aim to own your harness, this is a topic more devs will lean into. Here is a great report discussing this topic.
当你打算拥有自己的框架时,这是更多开发者会深入探讨的话题。这里有一份关于此话题的精彩报告。
JIT-Agent is a model whose output is an agent harness.
JIT-Agent 是一个输出为智能体框架的模型。
It formalizes the harness as a composable artifact under a fixed four-module protocol covering memory, planning, action protocol, and tool orchestration, then synthesizes one on the fly for any off-the-shelf agentic LLM.
它将框架形式化为一个可组合的构件,在固定的四模块协议下运行,涵盖记忆、规划、行动协议和工具编排,然后为任何现成的智能体 LLM 即时合成一个。
It also repairs harnesses mid-execution and self-evolves by distilling performance signals from an expanding archive of prior configurations.
它还能在执行过程中修复框架,并通过从不断扩展的先前配置档案中提炼性能信号来自我进化。
With JIT-Agent attached, DeepSeek-V4-Flash surpasses GPT-5.6 on DeepSearchQA (+9.1) and OdysseyBench (+4.3). GLM-5.2 gains up to +20.2 points.
配备 JIT-Agent 后,DeepSeek-V4-Flash 在 DeepSearchQA(+9.1)和 OdysseyBench(+4.3)上超越了 GPT-5.6。GLM-5.2 获得了高达 +20.2 分的提升。
The generated harnesses are also performance-competitive with mature runtimes like OpenCode and Claude Code.
生成的框架在性能上也与成熟的运行时(如 OpenCode 和 Claude Code)相媲美。
Paper: https://arxiv.org/abs/2608.25593
论文:https://arxiv.org/abs/2608.25593
Chat with Paper: https://academy.dair.ai/papers/jit-agent-a-model-that-writes-your-agent-harness-2608.25593
与论文对话:https://academy.dair.ai/papers/jit-agent-a-model-that-writes-your-agent-harness-2608.25593
更进一步:量化金融体系
看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力