跳到主内容
@wquguru
精选82Rohan Paul模型发布/更新

Apodex 1.1:将任务状态移出上下文窗口,以执行环境为扩展面

A context window is a terrible place to store the state of an 80-minute agent ru…

原文
发到 X

A context window is a terrible place to store the state of an 80-minute agent run.

上下文窗口是存储80分钟代理运行状态的一个糟糕地方。

One of the better ideas in Apodex 1.1 is that the task state sits outside the model's message history.

Apodex 1.1 中较好的理念之一是任务状态位于模型的对话历史之外。

So @Apodex_AI is treating executable environments themselves as a scaling surface.

因此,@Apodex_AI 将可执行环境本身视为扩展面。

That is a much bigger idea than adding more tools to an agent.

这比给代理添加更多工具的想法要大得多。

So Apodex 1.1 Agent Team is simultaneously a trained general-purpose model and the larger execution stack used to turn that model into a long-running worker.

所以,Apodex 1.1 代理团队既是一个训练有素的通用模型,也是将模型转变为长期运行工作者的更大执行栈。

They reported gains of 9.3 points on GDPVal, 5.6 on FrontierFinance, and 8.3 on FrontierScience-Research.

他们报告在 GDPVal 上提升了9.3分,在 FrontierFinance 上提升了5.6分,在 FrontierScience-Research 上提升了8.3分。

Their training environments are actual file, search, and code worlds with state transitions, tool budgets, failure conditions, and task-level verifiers.

他们的训练环境是实际的文件、搜索和代码世界,包含状态转换、工具预算、失败条件和任务级验证器。

The model has to operate inside them, change the workspace, recover when something breaks, and eventually satisfy a delivery contract.

模型必须在其内部操作,改变工作空间,在出现问题时恢复,并最终满足交付合同。

So the training distribution contains trajectories of work, not just prompts paired with good answers.

因此,训练分布包含工作轨迹,而不仅仅是带有正确答案的提示对。

That distinction is so important for agents. You can train a model to know how a tool works and still have it fail the moment the authoritative file changes, a command errors halfway through, or two artifacts become inconsistent.

这一区别对代理来说非常重要。你可以训练模型了解工具的工作原理,但当权威文件发生变化、命令中途出错或两个工件变得不一致时,它仍然可能失败。

Apodex's argument is basically that the next scaling surface is the world the model learns to work inside.

Apodex 的论点基本上是,下一个扩展面是模型学习在其中工作的世界。

更进一步:量化金融体系

看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力

进入量化体系 →

关联讨论

同一事件的更多信源

相似阅读

另一事件,读法相近