Decagon 90% 工作流用开源模型:成本、智能、延迟的取舍
How Decagon Runs 90% of Its Agents on Open-Source Models
做 AI 产品的读者可拿走一套模型选型框架:按成本、智能、延迟三维评估,用微调小模型在特定任务上超越大模型,并保留前沿模型处理探索性任务。
An AI agent should just be the front door of your business. And every interaction, whether it's like reactive or proactive with a customer, should be handled by AI. >> This narrative dominated the first half of 2026, which is that anthropic open AAI. They're the last startups. They're going to take over everything. >> Even once you have AGI, agents are going to need somewhere to store work and pull information from and reason about things. I don't think software as a whole in any meaningful way is going away. Unfortunately, the Frontier Labs, they do have small models, but you can't really control them in the way that you want. So, today 90% of our workflow is on open source >> on the specific task we want them to do. They actually outperform the large smart state-of-the-art model. >> The thing that we built was not an agent that does customer support well, but rather an agent that follows business process. Well, >> instead of us having to write these AOPs, do just does all of that. Let's say like we hit AGI and the models can do all sorts of things we can't even imagine [music] today. What's Decagon's emote and like why does Decagon like 10 years from now still have a right to exist?
>> Hey guys, welcome back to the studio. >> Thanks for having us. >> Yeah, good to see you. >> Thank you for being here. Before we get into customer support, um I actually wanted to widen out a bit. Um, and uh, Jesse, I'm going to actually um, mention a piece that you wrote recently that went pretty viral because it's right in the middle of the zeitgeist of conversation right now on open source versus closed source models. then thinking machines um Kimk 3 you know some of some very interesting open source models came out sort of right after um and there's this really interesting debate going on around what does it mean to own your destiny when it comes to AI especially in the enterprise um and what does that evolution look like by use case um so actually since that's a pretty live topic right now why don't we start there >> sounds good um so I'm gonna talk about our journey first uh just to make it very concrete for people so when we started the any uh the goal was just to get something working, right?
So if when you get something when you're the goal is to get something working, of course you're just going to use the frontier models because you want to get something out there and have actually deliver value. And so we were using opening anthropic at that time they were kind of like one uping each other in terms of how the how the models performed. And then at some point as we got to larger scale and we started working with larger and larger companies and they had you know millions of customers and then we we also have uh we also launched our voice agent right so a big factor became latency so it wasn't just like can you deliver good responses you have to deliver them really fast and uh the only way to get latency down um but also kind of you make our agent operate the way we want it to is to use smaller models and when [clears throat] you want to go to smaller models unfortunately the the Frontier labs, they do have small models, but you you can't really control them in the way the way that you want and most small models out of the box are not going to be good enough at the task that we want them to do. So, you have to fine-tune them, you have to change them. And so, that's when we started looking at open source. So, this was about year plus ago. And um it worked really well because if you think about it in the agent, right?
So, in our agent, our agent's job is to have conversations. So, it needs to do a lot of things at once, right? And like one one the first step it might do is like hm what topic is this person talking about or and something else it might do is oh is this person a bad actor that's coming in and trying to mess things up. There's all these like tasks it has to do. Each individual task doesn't need all of the intelligence of a big model. So you know all the frontier models are obviously very smart but they can do a bunch of different things. Like they can do math they can do coding. Like you just need them to be good at that one task. And so that's why you can use a smaller model and if you fine-tune it to be really good at that task and can be just as good or better than the big models, right?
So that was that was step one for us. You know, about a year ago, we were like, okay, let's start using these open source models. Uh we we we took the small ones and then um you know, that's why we have now a uh a research team and it's a very expensive team, but it's we we have it because, you know, we need people that are really good at taking these open source models and and tuning them and and so on. So today 90% of our workflow is on open source and um you know again the main reason was for latency to really optimize our voice agents and um I think we've just over the last year we've seen tremendous improvement in like how how it sounds how it feels and but still also like keeping the accuracy high um and then the remaining 10% of course we're still using the uh the closed source models and the frontier models for a lot of you know new new projects or new products and uh I think that's just where the industry is moving to. So if you kind of were to generalize this every model you can kind of evaluate along three dimensions. It's you know cost intelligence and latency and depending on what you need you want to kind of be at the limit of those three and sometimes you can trade off right so in our case we knew that we actually pull back on intelligence because we all had to do was that one task but now we get these latency advantages. I want to push a tiny bit on that point actually because um often times you know when you see these debates being had on Twitter the the trade-off tends to be oh do we want uh you know the smartest model that is very expensive or can we like dumb it down a little bit and get it cheaper. I actually think that is a false trade-off, >> right?
Because what we've seen in practice is even if you have a quote dumber model, you can get it, and we've seen this in practice, you can get it to higher performance on that specific task. So when we fine-tune smaller, dumber models, it's that they're just not as general purpose, but on the specific task we want them to do, they actually outperform the large, smart, state-of-the-art models, >> right?
So we end up getting all three things. It is better at the toss. It is cheaper and it is faster. >> And so do you feel like today like you need the most Frontier models for really anything at Decagon because your performance is very good already. So >> we we do and we often need them we often end up needing them for auxiliary tasks, right?
Where when you have uh when you sort of have auxiliary models to our sort of primary conversational flow, right? you have an agent and it's helping a customer with their rebooking or it's helping them with a process in healthcare then these are like well-defined pots. So we have smart bos models to do that but we've for instance recently launched do autopilot right which is our agent that improves the core conversational agent. Now, for something like autopilot, it is doing a very complicated job, right?
It's saying, I'm going to go and review a million conversations that just happened. I'm going to try and find trends. I'm going to create variants of the primary model and see which of those variants does better. So, now this is a much more broad open-ended exploratory task. So, we think for jobs like that, frontier models that are very smart, that can try out a lot of things make a lot of sense. H do you think that um I mean you guys obviously and you referenced it um Jesse that you have a research team right you guys launched Decagon Labs but even before the formal launch it's always been a part of your culture um do you think that enterprises will get there as
更进一步:量化金融体系
看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力