OpenAI发布Decisions API,基于GPT-6
Introducing the Decisions API
这是OpenAI针对低延迟决策场景推出的新API接口,直接赋能Agent开发,对构建实时交互系统的开发者有重要参考价值。
Sometimes your app needs to make a smart decision really, really fast. Routing a request, classifying an image, or deciding what an agent should do next. Today, we're introducing the Decisions API. It lets the model respond in a fraction of a second, so interacting with it feels pretty much instant. The Decisions API runs on GPT-6 Luna. It's very simple. You give it a text or an image, you define the question and the possible answers, and you use the result in your app.
有时你的应用需要做出非常快速的智能决策。例如路由请求、分类图像,或决定代理下一步该做什么。今天,我们推出 Decisions API。它能让模型在几分之一秒内响应,因此与之交互几乎感觉是即时的。Decisions API 运行在 GPT-6 Luna 上。它非常简单:你提供文本或图像,定义问题和可能的答案,然后在你的应用中使用结果。
By focusing the model on just a few choices, we can make it nearly 10 times faster while keeping capabilities like image understanding, broad language support, and safety protection. But let's see that in action with Charlie, starting with text input. Thanks, Romain. We're already finding a wide variety of applications that can benefit from the Decisions API. Let's say you want to work with some unstructured text. A simple thing you could do would be to automatically extract form fields from a long paragraph, but we can expand on that.
通过让模型仅关注少数几个选项,我们可以使其速度提高近 10 倍,同时保留图像理解、广泛语言支持和安全防护等功能。但让我们通过 Charlie 的实际演示来看看吧,先从文本输入开始。谢谢 Romain。我们已经发现许多应用可以从 Decisions API 中受益。假设你想处理一些非结构化文本。你可以做的一件简单事情是从长段落中自动提取表单字段,但我们还可以在此基础上扩展。
We can take those inputs and use them to route support tickets or inbound sales requests to the right teams. Here, we can see multiple examples of loose user input being quickly classified into sales leads. That's awesome. And this is not sped up, by the way. On the server side, the Decisions API here responded in less than 100 milliseconds. Very impressive. Now, what I find even more exciting about the Decisions API is how you can also use image inputs to take a fast action based on visual understanding.
我们可以利用这些输入将支持工单或 inbound sales requests(入站销售请求)路由到正确的团队。在这里,我们可以看到多个示例,松散的用户输入被快速分类为销售线索。太棒了。顺便说一句,这并没有加速。在服务器端,这里的 Decisions API 响应时间不到 100 毫秒。令人印象深刻。现在,我发现 Decisions API 更令人兴奋的是,你还可以使用图像输入基于视觉理解快速采取行动。
So check out this simple example here. I have a quick game and the Decisions API is taking frames of the road and obstacles, quickly choosing if the car should stay in its lane or go left or right. Now, if we turn up the speed right there, the Decisions API is still able to steer the car to the correct lane at the fraction of the time a reasoning model would take. That's so cool. So you can really imagine pairing this API with some basic computer use tasks, taking screenshots and quickly sending actions back.
所以看看这里这个简单的例子。我玩了一个小游戏,Decisions API 正在获取道路的帧和障碍物,快速选择汽车应该保持在车道内还是向左或向右行驶。现在,如果我们把速度调高,Decisions API 仍然能够在推理模型所需时间的几分之一时间内将汽车 steer(转向)到正确的车道。太酷了。所以你可以想象将这个 API 与一些基本的 computer use tasks(计算机使用任务)配对,截取屏幕截图并快速发送操作指令。
Of course, you'll still want Astra capabilities if you have a very sophisticated computer use or browser use scenario. But Charlie, you've also been using Decisions API together with voice as a modality. How does that pair together? Yeah, I found that using the Decisions API can make live interactions feel even more responsive. Here I've paired GPT-Live-1 with an animated character. GPT-Live-1 handles the voice while the Decisions API is choosing which expression the character should use as I talk.
当然,如果你有一个非常复杂的计算机使用或浏览器使用场景,你仍然需要 Astra 的功能。但 Charlie,你也一直在将 Decisions API 与语音作为模态结合使用。这两者是如何配合的?是的,我发现使用 Decisions API 可以让实时交互感觉更加响应迅速。在这里,我将 GPT-Live-1 与一个动画角色配对。GPT-Live-1 处理语音,而 Decisions API 则在我说话时选择角色应该使用的表情。
Let's try it out. Hey, how's it going? Pretty good. I'm hanging in there, keeping it calm and cozy. How are you doing? It has been a crazy week. Oh, wow. That sounds exhausting. Want to talk about what's been going on? Oh, no. It's been crazy in a good way. We're launching the Decisions API today and a ton of people are so excited to have it. That's so exciting. I know a lot of developers have been looking forward to building with it.
让我们试一试。嘿,最近怎么样?还不错。我还在坚持,保持冷静和舒适。你过得怎么样?这周真是疯狂。哦,哇。听起来很疲惫。想聊聊发生了什么吗?哦,不。这是一种好的疯狂。我们今天发布了 Decisions API,很多人对此感到非常兴奋。太令人兴奋了。我知道很多开发者都期待用它来构建应用。
Absolutely. The app is choosing from a set of expressions we've given it and those little reactions can become part of the conversation alongside what the assistant is actually saying. But, Romain, you've been building an interactive experience that touches on robotics, right? I have indeed. As you can see, our friends at Hugging Face kindly lended us a preview unit of the upcoming programmable robot Microduck. And this one's name is Lavender.
绝对如此。应用程序正在从我们提供的一组表情中进行选择,这些细微的反应可以成为对话的一部分,与助手实际所说的内容并列。但是,Romain,你一直在构建一种涉及机器人的互动体验,对吧?确实如此。正如你所见,Hugging Face 的朋友们慷慨地借给我们一台即将推出的可编程机器人 Microduck 的预览版。它的名字叫 Lavender。
So, I also integrated GPT-Live to give it voice intelligence as well as the Decisions API to help the robot choose where to look based on its camera. So, for this demo, the choices are simple. It's about moving the robot's head in either direction to look at the right object by sending the command to the robot. Lavender, follow the apple. Now, what you're noticing here is that every few frames from the camera is being analyzed by the Decisions API, which in turn can take the decision as to where the robot should look.
因此,我还集成了 GPT-Live 以赋予它语音智能,以及 Decisions API 来帮助机器人根据摄像头决定看向哪里。所以,对于这个演示,选择很简单。它是通过向机器人发送命令,让机器人的头向任一方向移动,以便看向正确的物体。Lavender,跟随苹果。现在,你注意到的是,摄像头的每隔几帧图像都被 Decisions API 分析,进而决定机器人应该看向哪里。
And it's the Decisions API that's deciding what's the apple. Absolutely. In fact, if you want to try with a different request that's a bit more ambiguous, feel free. Sure. Lavender, can you follow the fruit? Exactly. Amazing. You will now follow what is behind the concept. It's pretty cool, right? Now, what if we bring a second object into the picture with a question that's even more ambiguous. Lavender, what's more fun to play with?
正是 Decisions API 在决定什么是苹果。绝对正确。事实上,如果你想尝试一个稍微模糊一点的请求,请随意。当然。Lavender,你能跟随水果吗?没错。太棒了。你现在将跟随概念背后的东西。挺酷的,对吧?那么,如果我们引入第二个物体,并提出一个更模糊的问题呢?Lavender,玩哪个更有趣?
There you go. Now, Lavender will follow the gamepad. Because naturally, your gamepad is more fun to play with. What if I flip them? Good job, Lavender. And really, Lavender is definitely more interested about my gamepad than about my apple, for sure. Right? If I go to the left a little bit, it will follow. And we'll just keep on tracking the gamepad instead. Incredible. Pretty cool, right? These were just a few quick demos of the Decisions API across text and vision inputs, driven either by keyboard or voice interactions like this one.
好了,现在 Lavender 会跟随游戏手柄。毕竟,用游戏手柄玩游戏更有趣。如果我反过来呢?做得好,Lavender。而且,Lavender 显然对我的游戏手柄比对我的苹果更感兴趣,没错吧?如果我稍微往左移动一点,它也会跟着动。我们就这样继续追踪游戏手柄吧。太不可思议了。挺酷的,对吧?这里只是通过文本和视觉输入、由键盘或语音交互驱动的 Decisions API 的几个快速演示。
We think the Decisions API will unlock a brand new way to interact with AI, and we can't wait to see what you all will build with it. Happy building. Lavender, are you excited about the Decisions API? [Lavender chirps]
我们认为 Decisions API 将解锁一种与 AI 交互的全新方式,我们迫不及待想看到大家用它构建出什么。祝大家开发愉快。Lavender,你对 Decisions API 感到兴奋吗?[Lavender 发出叫声]
更进一步:量化金融体系
看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力