跳到主内容
@wquguru
精选70OpenAI(YouTube)技巧与观点

OpenAI Build Hour:用 GPT-3.5 实现价值最大化

Build Hour: Valuemaxxing with GPT-5.6

原文
发到 X

Hi everyone, welcome back to OpenAI Build Hours. I'm Christine on the startup marketing team and today I'm here with Charlie.

Hi, I'm Charlie from the developer experience team.

So today our topic is going to be about value maxing with GPT-3.5. So next slide. If this is your first time with us, uh the whole purpose of Build Hours is to empower you with the best practices, tools, and AI expertise to scale your company using OpenAI APIs and models. Uh down below is our homepage. So we're constantly posting new sessions as they come up. Feel free to check back and you can also catch up on demand.

We also post all of these sessions on YouTube um on the day of. So check back later in case you want to review anything that we're chatting about today. So we've traditionally started every Build Hour with a meme and we're back. Uh so if we have any Office fans around, um this is really about just getting more value per token um and Charlie's going to walk you through uh some tips and tricks on how to do this. So this is our first Build Hour after GPT-3.5 release.

Um so right after we released this model, we asked everyone on X kind of why they love this model um and we were really just like humbled and and also like taken away by all the reasons that people submitted. And I think a common theme that we saw was just how token efficient this model is particularly in just having you guys get the most bang for per token. Um and recently actually, I think this is two days ago on X, uh we reached a new milestone.

Um so we're we're really happy to see that this is resonating with people um and that this new model um is is letting people do more with with less tokens. So to give you a snapshot of what we're talking about today, the first is we're going to explain this shift that's gone from token maxing to value maxing. Um we're going to show you value maxing in Code X and then also value maxing your harness. And then I'm really excited that we will have Floy, um one of our favorite startups uh join today to chat about how they've actually migrated their AI agent over.

And last but not least is Q&A. Um so on the right side of your screen, if you toggle over to the Q&A button, you can actually submit questions and our team will be in the room um answering them on chat and then also selecting a few to answer live towards the end. And we'll have our friends from Floy also join. So feel free to ask any questions um and we're really excited to kick this off. So over to you, Charlie.

Thanks so much, Christine. Um thanks everybody for joining. So excited to be here. Uh I want to talk about value maxing, but I'm sure there's you know a lot of questions about what what is that, right? And I think you can't quite explain what it is without first starting with token maxing. Um and I think token maxing is a term that that emerged, you know, earlier in this year. Um and to kind of encapsulate it, right?

It was essentially this this idea of measuring progress by how much AI you're using. I think that might be how many tokens you're burning, how many prompts you're sending, how many agents you're managing at the same time. And I think we even saw companies that had stood up internal leaderboards to just track, you know, how many tokens their employees were using day over day or week over week. Later on in the year though, I think some of that sentiment started to change, right?

Uh we started to see uh you know, companies that were accidentally burning through their entire annual budget for AI in a handful of months, right? And we started to see a lot of enterprises start to pull back and say, "Well, maybe we should be throttling our AI spend." Um or even some AI CEOs saying that, you know, some companies may have gotten carried away with their with their token leaderboards. Um and importantly that we should be measuring employees on the outputs instead, right? and the value that they're generating.

So that brings us to value maxing. In contrast with token maxing, it's measuring progress not by how much AI you're using, but what AI is actually helping you to accomplish, right? How much work you're getting done, how much time are you saving, how much, you know, quality are you improving in the code and the artifacts that you're making. I think a question that I like to come back to when I think about, you know, this framing of value is something like if you doubled your token spend tomorrow, how would you know if it was worth it, right?

You know, there is perhaps a naive assumption of like, yeah, like we spend more on tokens, it's going to be better, but, you know, how does that work back to to concrete outcomes? And I think there are a few questions that you can start to ask yourself as you're thinking through this for your own processes or your own company. Things like what are the outcomes that you are trying to improve? With those outcomes, what are the workflows that lead to, you know, good quality outputs?

Where in those workflows can you add intelligence, can you add AI, can you add tokens, right? And maybe accelerate or improve quality. And, you know, what evidence do you have that you are improving the quality or that you are, you know, using, building a working system here, right? A lot of this, I think when it when folks work with AI, you've probably heard the term evals, and I think a lot of AI research and engineering these days is driven by having good evals, which to boil it down, to oversimplify it, basically means how do you define what good looks like?

And I think as we move from just having to AI, you know, having AI chatbots perform simple tasks to having AI agents be responsible for entire outcomes, that question is still very valuable. What does good look like and how do you define that and track that within your company or your role? And I think I think I I I do want to say too that, you know, in some cases it might seem like the answer here is, well, we should spend fewer tokens, right?

Like we should um our models or cut spend, and I'm sure in some cases that's that's absolutely true. Um I did also want to point out there might be some cases where you want to spend more tokens and that's actually still driving more value, right? Um you may decide, "Hey, we we actually want to ensure that we have a high-quality result, so we're going to bump the reasoning level on the model um or switch to a bigger model."

You might say, um "It's important that this gets done faster, so we are going to spend more on tokens now in order to save uh you know, real-world all clock time." Um that might be the case. You maybe you've got a big migration and you want to migrate something into a different programming language. Um maybe you could do it, you know, in a couple of months with fewer tokens or, you know, maybe you can do it in a couple of weeks and spend some more.

Uh similarly, if you know your workflows are working well and you want to scale them, then you can decide to spend more uh in order to just scale those outputs if they're working well. And last but not least, you know, if you want to manage risk, um you might have a system where it's really important that the output is is high quality and, you know, and bulletproof. Um and in that case, you may want to do something like use LLMs as a judge, um and you may want to use more or use a diverse range of them, right?

In order to ensure that it's catching uh all sorts of different edge cases. Okay. So, with that in mind, I want to talk about two ways of thinking about value maxing here. I think one is as an individual developer while you're using Codex. Uh and I think another is uh as an organization or as a company that's building on top of OpenAI's APIs. We're going to start with Codex. Um like Christine mentioned, I think this the story of of value maxing in Codex uh really starts with the GPT-3.5 family, right?

Um if you're not familiar, we've got three new models in this family: Soul, Terra, and Luna. Uh So

原文超出正文长度上限,此处截断——上游还有内容,完整版见上方「原文 ↗」。

更进一步:量化金融体系

看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力

进入量化体系 →

相似阅读

另一事件,读法相近