OpenAI DevDay:Computer Use进展与Decisions
Why Dwarkesh is Wrong about Computer Use + How OpenAI shipped its Jev competitor in 1 Week
DevDay干货总结,完整梳理了Computer Use最新工程实践与Decisions API等技术细节,对Agent开发极具参考价值。
Three months ago Dwarkesh, who has been posting incredible blogs and episodes about RL, posted a framing question for his video essay on RLVR which upset a lot of Computer Use folks:
三个月前,Dwarkesh 发布了关于 RLVR 的视频论文框架性问题,他一直在发布关于强化学习(RL)的精彩博客和节目,这一问题让许多 Computer Use 领域的从业者感到不满:
We are no strangers to learning in public and are no strangers to the stress of getting things wrong when you have a big platform. However, we were at Anthropic for the Computer Use launch, there for Claude Cowork with the first big podcast on it, organized the first Computer Use track at AIE presenting the state of the art, and were close to the OpenAI-Sky Software acquisition that now powers the complete domination of computer use that Codex enjoys today. This is why we’re excited to bring you today’s first guest, Ari Weinstein, cofounder of Sky and now leading all the amazing CUA progress that casuals might miss:
我们并不陌生于公开学习,也不陌生于拥有大型平台时犯错带来的压力。然而,我们在 Anthropic 的 Computer Use 发布期间在场,在 Claude Cowork 期间主持了关于该主题的首个大型播客,在 AIE 组织了首个 Computer Use 专题并展示了最先进技术,并且接近 OpenAI-Sky Software 的收购案,而 Codex 如今凭借此实现了在计算机使用领域的全面主导。这就是为什么我们兴奋地邀请到今天的第一个嘉宾 Ari Weinstein,他是 Sky 的联合创始人,现在正领导着所有令人惊叹的 CUA 进展,而这些进展可能会被普通用户所忽略:
Ari explains why Computer Use is now “180 degrees different” from where it was months ago, how agents are learning to debug and recover from failures, why combining screenshots with accessibility data, the DOM, Playwright, and generated code changes the speed equation, and why the next frontier is making agents literally superhuman at using software.
Ari 解释了为什么 Computer Use 现在已经与几个月前相比“截然不同”,智能体如何学会调试并从失败中恢复,结合截图、无障碍数据、DOM、Playwright 和生成的代码如何改变速度方程,以及为什么下一个前沿领域是让智能体在软件使用方面真正超越人类。
OpenAI clones Jev
OpenAI 克隆 Jev
In the second half, Nikunj Handa from OpenAI’s API team breaks down the new developer stack: async tool calling, mid-turn steering, WebSockets, UltraFast inference, the Decisions API, prompt caching, pre-warming, compaction, and the Agents API. Given that we were the first Jev podcast, we particularly focus on the unusually fast sprint on the Decisions API:
在后半部分,来自 OpenAI API 团队的 Nikunj Handa 详细解析了新的开发者栈:异步工具调用、中途转向、WebSockets、UltraFast 推理、Decisions API、提示缓存、预热、压缩以及 Agents API。鉴于我们是首个 Jev 播客,我们特别关注 Decisions API 上异常快速的冲刺:
And why it is just a Luna wrapper for now but the team is motivated and egoless enough to clone what they consider to be good patterns.
以及为什么目前它只是一个 Luna 包装器,但团队有足够的动力和无自我意识去克隆他们认为良好的模式。
We discuss:
我们讨论了:
- Why OpenAI thinks Computer Use has changed dramatically in just the last few months
- Dots and what changes when every agent gets its own Linux computer
- Why Computer Use can now complete some tasks faster than the average human
- The path from human-level to “literally superhuman” computer use
- Why modern agents are much better at debugging and recovering from failure
- How screenshots, accessibility trees, the DOM, Playwright, and generated JavaScript work together
- App Shots and why they give models much richer context than ordinary screenshots
- Why Computer Use can close the loop between writing software and testing it
- Trust, permissions, and safety when agents can make payments and operate websites
- Async function calling and why models no longer need to stop reasoning while tools run
- Mid-turn steering, WebSockets, and the architecture behind more responsive agents
- UltraFast inference and how OpenAI is pushing frontier models toward much lower latency
- The rapid internal story behind the Decisions API
- Why Decisions API is more than structured outputs at low latency
- GPT Live, fast tool calling, and real-time computer control
- How OpenAI is already using Decisions API for support classification and internal workflows
- Longer prompt caching, cache pre-warming, and cache-aware applications
- Server-side compaction vs manual compaction for long-running agent threads
- What should live inside an Agents API versus a developer’s own harness
- OpenAI as an “AI cloud” and the search for higher-level primitives beyond raw model APIs
- 为什么 OpenAI 认为 Computer Use 在过去几个月发生了巨大变化
- Dots 以及当每个智能体都拥有自己的 Linux 计算机时会发生什么变化
- 为什么 Computer Use 现在可以比普通人更快地完成某些任务
- 从人类水平到‘真正超越人类’的计算机使用路径
- 为什么现代智能体在调试和从失败中恢复方面做得更好
- 截图、无障碍树、DOM、Playwright 和生成的 JavaScript 如何协同工作
- App Shots 以及为什么它们为模型提供了比普通截图更丰富的上下文
- 为什么 Computer Use 能够闭合编写软件和测试软件之间的循环
- 当智能体能够进行支付和操作网站时,信任、权限和安全问题
- 异步函数调用以及为什么模型不再需要在工具运行时停止推理
- 转弯中的转向、WebSocket,以及更敏捷智能体背后的架构
- 超快推理以及 OpenAI 如何将前沿模型推向更低延迟
- Decisions API 背后的快速内部故事
- 为什么 Decisions API 不仅仅是低延迟的结构化输出
- GPT Live、快速工具调用以及实时计算机控制
- OpenAI 如何利用 Decisions API 进行支持分类和内部工作流处理
- 更长的提示词缓存、缓存预热以及感知缓存的应用
- 针对长运行智能体线程的服务端压缩与手动压缩对比
- Agents API 与开发者自有框架之间应如何划分职责
- OpenAI 作为“AI 云”,以及在原始模型 API 之外寻找更高级原语的努力
Ari Weinstein
- Product & Engineering, Computer Use at OpenAI
- X: https://x.com/AriX
- LinkedIn: https://www.linkedin.com/in/weinsteinari/
- OpenAI 产品与工程团队,Computer Use 项目
- X: https://x.com/AriX
- LinkedIn: https://www.linkedin.com/in/weinsteinari/
Nikunj Handa
- Product, API at OpenAI
- X: https://x.com/nikunjhanda
- LinkedIn: https://www.linkedin.com/in/nikunjhanda/
- OpenAI 产品与 API 团队
- X: https://x.com/nikunjhanda
- LinkedIn: https://www.linkedin.com/in/nikunjhanda/
Timestamps
时间戳
00:00:00 OpenAI DevDay: Dots, GPT-6.1, Agents API, and Decisions API
00:00:00 OpenAI DevDay:Dots、GPT-6.1、Agents API 和 Decisions API
00:02:52 Dots and Personal Cloud Computers
00:02:52 Dots 与个人云端计算机
00:04:59 Why Computer Use Is “180 Degrees Different”
00:04:59 为何 Computer Use “截然不同”
00:06:04 From Sky to Self-Debugging Computer Use Agents
00:06:04 从 Sky 到具备自调试能力的 Computer Use 智能体
00:09:24 How Computer Use Sees and Operates Software
00:09:24 Computer Use 如何感知和操作软件
00:12:09 From Faster Than Humans to Superhuman Computer Use
00:12:09 从超越人类到超人类计算机操作
00:16:03 Agents API: Trust, Permissions, and Safety
00:16:03 Agents API:信任、权限与安全
00:17:31 Computer Use for Coding, Testing, and QA
00:17:31 用于编码、测试和质量保证的计算机操作
00:19:14 GPT-6 APIs, Async Tool Calling, and UltraFast Inference
00:19:14 GPT-6 APIs、异步工具调用与极速推理
00:23:21 The Rapid Story Behind Decisions API
00:23:21 Decisions API 背后的快速故事
00:25:32 What Decisions API Is and How It Works
00:25:32 Decisions API 是什么以及它是如何工作的
00:30:24 What OpenAI Is Building With the New APIs
00:30:24 OpenAI 正在利用新 API 构建什么
00:32:23 Prompt Caching, Pre-Warming, and API Performance
00:32:23 提示词缓存、预热与 API 性能
00:35:20 Context Compaction for Long-Running Agents
00:35:20 针对长运行代理的上下文压缩
00:37:13 Memory, Higher-Level APIs, and the AI Cloud
00:37:13 记忆、高级 API 与 AI 云
Transcript
转录稿
Introduction: OpenAI DevDay and the New Agent Stack
引言:OpenAI DevDay 与新智能体栈
Vibhu [00:00:00]: Okay. We’re very excited to be here. Today is OpenAI DevDay. Special podcast
Vibhu [00:00:00]:好的。我们非常兴奋能来到这里。今天是 OpenAI DevDay。特别播客
Swyx [00:00:08]: We’re the first podcast after your livestream.
Swyx [00:00:08]:你们是直播后的第一个播客节目。
Vibhu [00:00:10]: First podcast. We have Ari here, who leads the product and engineering team for Computer Use agents. Before we kick in and dive deep on Computer Use, you wanna give a quick recap? What was announced? What’s the quick slew of announcements you guys had today?
Vibhu [00:00:10]:第一个播客。我们有 Ari 在这里,他负责 Computer Use 智能体的产品和工程团队。在我们深入探讨 Computer Use 之前,你能简单回顾一下吗?宣布了哪些内容?你们今天有一系列怎样的快速发布?
Ari Weinstein [00:00:24]: Yeah. yeah, it was a super exciting day. we just got out of the keynote. It was really sick. there were a bunch of Computer Use announcements that I think are worth thinking about. We have, Dots, which is the new, sort of personal assistant product, and, that has some really exciting Computer Use features. There’s GPT-6.1 Sol, which is this amazing new model, that I think is particularly great for Computer Use ‘cause of, sort of the cost and speed, advantages. I think, I think we shared that it’s, a fifth of the cost of Astra and a seventh of the cost if you’re looking at Computer Use specifically, which is really amazing. sorry, there were so many things. I’m trying to sort through it.
Ari Weinstein [00:00:24]:是的,确实是非常令人兴奋的一天。我们刚结束主题演讲。真的很棒。有很多关于 Computer Use 的发布,我认为值得思考。我们有 Dots,这是一种新的个人助理产品,它具有一些非常令人兴奋的 Computer Use 功能。还有 GPT-6.1 Sol,这是一个全新的惊人模型,我认为它特别适合 Computer Use,因为它在成本和速度方面具有优势。我认为我们提到过,它的成本是 Astra 的五分之一,如果你专门看 Computer Use,成本更是七分之一,这真的非常惊人。抱歉,事情太多了。我正在努力梳理。
Swyx [00:01:02]: And the API.
Swyx [00:01:02]:还有 API。
Ari Weinstein [00:01:03]: Agents API, which now has Computer Use in it, which is really cool, ‘cause now developers can build on the same Computer Use, that is part of Codex, and ChatGPT. and then there were some demos of our existing Computer Use features, like app shots, where you can take the context of something you’re doing on your computer and bring it into Codex and ChatGPT really fast. And then, like, native Computer Use on your Mac, where Roman had it taking screenshots of his app, automatically, and he could do other things on his computer while Computer Use was using his applications. so yeah, really exciting keynote.
Ari Weinstein [00:01:03]:Agents API,现在其中包含了 Computer Use,这真的很酷,因为开发者现在可以基于与 Codex 和 ChatGPT 中相同的 Computer Use 进行构建。然后还有一些关于我们现有 Computer Use 功能的演示,比如 app shots(应用截图),你可以将你在电脑上操作的上下文快速带入 Codex 和 ChatGPT。还有像 Mac 上的原生 Computer Use,Roman 让它在自动截取他应用的屏幕截图,同时他可以在电脑上做其他事情,而 Computer Use 正在使用他的应用程序。所以,确实是一个令人兴奋的主题演讲。
Swyx [00:01:35]: And not to mention the Decisions API.
Swyx [00:01:35]:更不用说 Decisions API 了。
Ari Weinstein [00:01:37]: Decisions API.
Ari Weinstein [00:01:37]:Decisions API。
Swyx [00:01:38]: Off the bat, are they all the same model? Like, this is. Or the same dataset distilled to different models?
Swyx [00:01:38]:一开始,它们都是同一个模型吗?比如这个。还是说数据集被蒸馏到了不同的模型上?
Swyx [00:01:44]: Like, basically, like, is Computer Use using Decisions API, or are they, like, kinda separate?
Swyx [00:01:44]:基本上,Computer Use 是使用 Decisions API,还是说它们有点分开?
Ari Weinstein [00:01:49]: So what’s really cool about the Decisions API is it, you know, it has all these new capabilities. It does inference in parallel. it doesn’t have reasoning. It’s a smaller model, than the ones we use for Computer Use. and so those capabilities make it really fast.
Ari Weinstein [00:01:49]:所以 Decisions API 真正酷的地方在于,它拥有所有这些新功能。它可以并行推理,但不具备推理能力(reasoning)。它是一个比 Computer Use 所用模型更小的模型,因此这些功能使其运行速度非常快。
Dots and Delegating Work to a Cloud Computer
点和委托工作给云计算机
Swyx [00:02:07]: Yeah.
Swyx [00:02:07]:是的。
Ari Weinstein [00:02:07]: They also make it a little bit less good at doing, like, long horizon, sort of sophisticated tasks. And so I think I would say it’s still an open area of research for how we, like, bring those approaches together. But, yeah, I’m really excited to see what people build with the Decisions API.
Ari Weinstein [00:02:07]:这也使得它们在处理长周期、复杂任务方面稍微差一些。所以我认为,关于我们如何将这些方法结合起来,这仍然是一个开放的研究领域。但总的来说,我非常兴奋能看到人们用 Decisions API 构建出什么。
Vibhu [00:02:24]: One of the interesting things is Dots now have attached personal computers.
Vibhu [00:02:24]:一个有趣的事情是,Dots 现在连接了个人电脑。
Ari Weinstein [00:02:28]: Yeah.
Ari Weinstein [00:02:28]:是的。
Vibhu [00:02:28]: So it seems like they’re very much more persistent. You’ve been using them for a while. How should people push the bounds? Like, what should people aim for? What should they try? Personally, right now I use it for a lot of customer service. Like
Vibhu [00:02:28]:所以它们看起来更加持久。你们已经使用它们有一段时间了。人们应该如何拓展边界?比如,人们应该以什么为目标?应该尝试什么?就我个人而言,目前我将其用于大量的客户服务。比如
Ari Weinstein [00:02:41]: Cool
Ari Weinstein [00:02:41]:酷
Vibhu [00:02:41]: “Oh, this was wrong. I don’t wanna sign in. I don’t wanna authenticate.” Find whatever and just get it fixed.
Vibhu [00:02:41]:“哦,这个错了。我不想登录。我不想验证身份。”找到问题并直接修复它。
Ari Weinstein [00:02:45]: Yeah.
Ari Weinstein [00:02:45]:是的。
Vibhu [00:02:46]: How should we push further? What should people try?
Vibhu [00:02:46]:我们该如何进一步拓展?人们应该尝试什么?
更进一步:量化金融体系
看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力