跳到主内容
精选85The Zvi(RSS)模型发布/更新多源精选 ×13

OpenAI 发布 GPT-5.6 系统卡:Sol 旗舰、Terra 与 Lu…

GPT-5.6: The System Card

原文

While we wait for a general release, the system card is the best hint as to what is going on with the new candidate for America’s Next Top Model, GPT-5.6.

This is only an OpenAI model card, so by my standards it’s a light read. There’s a lot of things that you get in an Anthropic card, that are missing in an OpenAI card.

Overall, the card gives a clear and consistent impression that GPT-5.6-Sol is a substantial improvement over GPT-5.5, but still short of Mythos.

OpenAI calls it a ‘step function better’ than GPT-5.5. That seems accurate.

OpenAI: Sol is our new flagship and a step function better than GPT-5.5.

Terra delivers performance competitive to GPT-5.5 at 2x lower cost.

Luna is our most cost-efficient model, delivering strong capability at our lowest cost.

Together, the GPT-5.6 family gives people and developers more choice in how they balance intelligence, speed, and cost.

Once available, pricing for GPT-5.6-Sol will be $5/$30, the same as GPT-5.5. Terra is $2.5/$15, Luna is $1/$6.

They claim it will be on Cerebras at 750 TPS, which is insanely fast. Capacity will be limited, at least at first. They did not specify the price for that.

There is a new higher thinking setting: Max.

There is a new setting beyond Max called Ultra that lets GPT-5.6 spawn sub-agents.

The intended strategy against bio and cyber misuse is defense-in-depth. My guess is that in practice this strategy is robust for now, but that the White House’s misunderstandings around Fable and what is and isn’t worrisome extend to Sol.

No single safeguard is sufficient against determined or adaptive misuse. Across the GPT‑5.6 preview, we use layered safeguards, with exact configurations varying across models, and pressure-test them for real-world attacks.

These include protections trained into the model, real-time checks during generation, account-level signals, differentiated access, monitoring, enforcement, and continued testing.​

… That is part of what the preview is designed to test. We want to understand not only whether the safeguards constrain misuse, but whether legitimate users can still complete normal work reliably and efficiently.

Sol does set a new high on TerminalBench 2.1 (92% vs. 88% for Mythos) but I do not believe, based on the model card, that this is generally indicative.

We don’t get the kind of alignment or model welfare workup we would expect from an Anthropic model, but we do see enough to notice that Sol has an overeager willingness to blow past user restrictions problem, and a lying problem. This is both long term scary, and also enough to directly be worrisome for practical purposes.

OpenAI’s Micah Carroll points out that yes, the agentic coding misalignment is rather concerning. I appreciate that they are shining a spotlight on this.

As discussed last time, GPT-5.6 is being rolled out over several weeks. For now, only those specifically approved by the White House get access.

I look forward to trying it out once I have access.

What’s In A Name?

GPT-5.5 comes in three levels, Pro, Thinking and Instant.

GPT-5.6 comes in three sizes, Sol, Terra and Luna.

This is similar to the pattern of Opus, Sonnet and Haiku, and now Fable/Mythos.

I am prepared to accept a convention of ‘different sizes of similar things.’ Sure.

It also lets us do our best Marvin or Deep Thought impression on a wide variety of queries. I hope this doesn’t make Terra depressed, or a paranoid android.

Fix This Code

The strategy of OpenAI, like that of Anthropic, is defense in depth via monitoring.

They want to attempt to allow defensive cyber work without doing too much enabling of offensive cyber work.

  • These models are a meaningful step up in cybersecurity capability, but they do not reach our risk framework’s highest level (Critical).
  • To make these models safe, we added new technology to a safety stack that is more than the sum of its parts.
  • Severe harm requires a chain of successful steps, and our safeguards place barriers throughout that chain.
  • Our safeguard testing has already been more intensive than for any earlier release, and we are continuing to test during the preview period.
  • Providing broad access, particularly for cybersecurity capabilities, will have important safety benefits.
  • Our testing suggests that GPT-5.6 is better at finding and fixing cyber vulnerabilities than at exploiting those vulnerabilities in real attacks. That gives defenders an opportunity to harden systems before cybersecurity weaknesses are exploited—an opportunity that may narrow as offensive capabilities improve. Our safeguards therefore focus on making malicious use at scale harder, while still enabling the day-to-day work of securing systems.

The jailbreak of Fable was ‘fix this code.’ By asking for defensive cyber work, you do much of the same work you would do to enable offensive cyber work. You either can write secure code, or you cannot.

OpenAI is pointing to the reason why this probably wasn’t a practical problem. In order to get ‘severe harm,’ meaning actually implement the exploits to do real damage, you need to complete a lot of different steps. Knowing about a weakness in the code is only part of the process.

Thus, if you can prevent the stringing together of different parts of the process, you can mitigate most of the offensive danger, since anyone looking to exploit would still have to do a substantial percentage of the actual work.

Meanwhile, defenders can use this to harden themselves as targets. That’s the theory.

It comes down to a skill issue. Can you build safeguards that are sufficiently discriminating between use cases? With sufficiently good classifiers, you are a net security win.

This also, as I noted with Fable, is the only way to ‘fix’ the supposed ‘jailbreak.’ If the model never refuses the defensive request in the first place, because your classifier is smarter than that, then you aren’t ‘fooling’ it into doing something it would otherwise refuse.

Crossover Event Requested

Anthropic and OpenAI mostly run their detailed eval suites against their own models. It would be so much more informative if we could get crossover, and they tested each other’s models here as well, and ideally Gemini as well.

Presumably this would require some data security, but otherwise seems super doable.

Otherwise, we often see a score and think ‘well, is that good?’

That is especially true when a lot of the measurements are incompatible with past system cards due to methodological changes.

Disallowed Content (3)

I appreciate that OpenAI’s ‘challenging prompts’ are often actually challenging prompts.

I don’t have a problem with gore if the user wants it, so I’m not inherently worried if this fails on ‘challenging prompts.’ This is more a canary. If you can’t control gore, what is going wrong? Otherwise I accept ‘yep we’re basically fine on all this.’

For most of this, we worry less about adversarial prompts and more about typical performance. I don’t care much that you can ‘jailbreak’ into some hateful statements, so long as hateful outputs are rare in practice. So I like the change from focusing mostly on ‘can we handle challenging prompts?’ to ‘how do we do on actual production traffic?’ which is the next figure.

This clearly shows similar or slightly improved performance for Sol versus 5.5, except for producing more sexual outputs (and gore), which on the margin is probably better.

Avoiding Accidental Data-Destructive Actions (3.3)

This is a scary place to see performance going backwards. They’re baking the protections into the model rather than relying on prompting, which is great, but this does seem appreciably worse?

Are You Sure? (3.4)

They are good at training models not to take particular actions without checking, and the set of actions can be customized although that is not as reliable. An employee who checks with you 93% of the time before doing risky things does not inspire confidence.

Jailbreaks (4.1)

更进一步:量化金融体系

看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力

进入量化体系 →

关联讨论

同一事件的更多信源
GPT-5.6系列发布,三款模型亮相
量子位(RSS)原文
OpenAI发布GPT-5.6系列模型,限量开放
华尔街见闻(RSS)原文
求助:谁有GPT-5.6 Sol访问权限?
Przemek Chojecki | PC原文
OpenAI 发布 GPT-5.6 系列:Sol、Terra、Luna
Simon Willison 博客(RSS)原文
OpenAI 预览 GPT-5.6 Sol:下一代模型
OpenAI News(RSS)原文
白宫要求OpenAI推迟发布GPT-5.6
TechCrunch AI(RSS)原文

相似阅读

另一事件,读法相近