跳到主内容
精选70AI Explained(YouTube)行业动态多源精选 ×8

Fable 5 与 GPT 5.6 Soul 早期对比及行业影响

Fable 5 vs GPT 5.6 Sol: The Early Results

原文

Fable 5 is back, but we'll just shut down a chat or reroute to a weaker model even more than before. GPT 5.6 Soul, the OpenAI equivalent of Fable, is out, but only for select customers and with incomplete results in its report card. Claude Sonet 5, just got thrown into the mix last minute by Anthropic. But the model maker is now competing, I would say, to show how little Sonnet adds to frontier capabilities, lest one thinks the US government intervenes. So yeah, it's been a weird few days in AI, but there are a handful of points of signal, I would say, that I want to try to highlight in this brief video. From the hard quantifiable comparisons we can make between GPT Soul and Fable 5 or Mythos 5, it is possible to unearth some direct comparisons. And from that to the news of OpenAI offering a stake in their company to the US government, Sam Orman warning of concentrated corporate power and much more. First though, the myths and confabulations about Fable and the timeline of how it came back into general availability. Yes, even to me, a non-American. It turns out, according to Anthropic, that the vulnerability that Amazon flagged that caused Fable to get blocked in the first place. See my recent two videos, was one that could also be flagged and identified by GBC 5.5, Kimmy K2.5, an open weights model from China. Nevertheless, it would be pretty awkward for the US government to just admit that. So, Anthropic had to show some response and further shifted the line in what their safety scans would flag as blockable. More safety margin of course, but this quote improved safety classifier does mean that benign requests will be flagged much more often, including alas during routine coding and debugging tasks. Now I will say that my question about the benefits of beachroot happily discussed with Opus 4.8 was among the first of the casualties flagged as aiding and abetting international terrorism. No, I'm joking. It's not that. But it was deemed as too risky for Fable 5 and so I had to continue with Opus 4.8. More seriously though, quite how frequently the safety classifier marks routine tasks like coding and debugging as being blocked and just how annoying that becomes. Only the coming few weeks will tell. Oh, and what about the mythical universal jailbreak that the US government thought was possible where you don't just extract one harmful response as with a narrow jailbreak, but know a universal one where you unlock the full potential for good or ill of the model. Well, on that anthropics say no one has yet at the time of writing been able to find a universal jailbreak, though of course the red teaming continues. All well and good, but you might say, well, the bigger news was the recent release of GPT 5.6 Soul in particular. That's the counter response from OpenAI to the Fable series. I will say it does sound like OpenAI were tired of the lamer sounding quantitative names like 5.5 or 03. So, they copied the anthropic approach of evocative names, Soul, Terror, Luna. My only query there is it doesn't really leave them much room for naming expansion as Mythos was an expansion to Opus. Almost the only way I can see them going bigger than the Sun or Soul would be to say name a model as Beetlejuice. That's the star, not the person. Very few hard stats have been released about Soul as of today beyond the price and a few select benchmarks. But the price is a tell though because they are gunning for anyone looking to save a buck versus Claude with even 5.6 Soul being half the API price of Fable 5. Exactly half the input and just over half the output. Some of you may be thinking well I use the Pro or Max plan for Claude. I don't pay the API price but come July 7th it won't be included in your weekly plan. And so you may come to feel that pricing really does matter. All of that rather obfuscates the main point though, the trillion dollar question. Is soul roughly as performant as Fable?

Because if so, one could imagine precipitating a massive switchover between the two. Well, slight problem. We can't test 5.6 directly because at the US government's request, OpenAI are starting with a limited preview of the model for a small group of trusted partners. They then say, notice whose participation has been shared with the US government. So, OpenAI are kind of hinting that they chose and then shared which of these partners it would be. But then there's this leaked memo in the information which slightly changes the framing to my eyes. It's actually, as told staff, that the government would be approving the access given customer by customer during the preview period. Either way, man hopes that there will be a general release in the coming couple of weeks. So, call it next week or the week after from time of recording. The risk though, as I hinted about earlier and was commented on by one Twitter user, is that such staggered releases will concentrate power. This was a direct risk flagged well before the recent kathuffle by OpenAI themselves. One of their goals as a company was to stop the undue concentration of power by corporations, for example. A staggered release means large corporations get access to the best models much earlier. Alman replied, "If it takes too long for general availability, then yes, that would happen. If we can get through the previews though in just a few weeks, then it should be probably okay." Now, let me know if you agree, but I actually think there is a different angle that could come into play from a seemingly unrelated story. Basically, Anthropic accused Alibaba, who oversee the development of a top Chinese model, Quen, of using 29 million exchanges with Claude, to get training data from its responses to train their own Chinese models, the Quen series, against, of course, the terms of service of Anthropic. This would be the largest extraction campaign of its kind. The world in immediate response rallied in sympathy with Anthropic who have always been champions over never using even a line of copyrighted material for any of their models. End sarcasm. But wait, how does this story linked to the corporate concentration point I was just making?

Well, if this large scale scraping to distill abilities into Chinese models becomes ever more sophisticated, successful, I think the incentives of the labs might switch. Better for them perhaps to serve their latest models to governments, approved businesses and of course themselves for say 3 to four months safe from this kind of distillation and then only when they have a better internal model release the older one to you the unwashed masses. And that theory is even before you get to geopolitics. Anthropic put it like this. Distillation attacks turn hundreds of billions of dollars in American investment and research into a massive subsidy for our geopolitical competitors. That's my theory anyway. And I want to now get to the test results for GPT 5.6 Soul. But just one more thing on this concept of model access being set to be much more gated. That trend could explain why OpenAI put this idea out to the US government of giving them a 5% stake in the company. This would be much like how Intel surrendered 10% to the Trump administration about a year ago. By the way, these early conversations also involved giving the US government stakes in other US AI companies. I am kind of curious why you think OpenAI are suggesting this cuz I have a few theories. Theory one would be that this proposal preempts the US government demanding more. Theory two could be that it encourages the government to rapidly allow general release because if the growing equity in these companies could be used to pay dividends to the public, apparently this is what OpenAI want. A bit like happens with energy in Alaska, then the government would have that incentive to grow the market share of those companies and allow general release earlier. A darker theory three would be that OpenAI think that Anthropic would not go along with this proposal and

更进一步:量化金融体系

看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力

进入量化体系 →

关联讨论

同一事件的更多信源

相似阅读

另一事件,读法相近