跳到主内容
@wquguru
精选88The Zvi(RSS)模型发布/更新多源精选 ×8

GPT-6 Astra系统卡深度解析:对齐、监控与能力边界

GPT-6 Astra: The System Card, Alignment and What Comes Next

原文
发到 X
推荐理由

这篇来自The Zvi的深度分析拆解了GPT-6 Astra的核心争议点,特别是可监控性下降与对齐宣称之间的矛盾,对关注模型安全与能力的从业者极具参考价值。

OpenAI claims that Astra is ‘the most intelligent and most aligned [available] model’ in the world. Not the most intelligent and aligned OpenAI model, but the most period.

That is bold talk. It risks overstepping, and by doing so souring the release of what is clearly an excellent model. As do the severe problems with monitorability.

It also raises the question of what they mean by ‘most aligned model.’ How do they define ‘aligned.’ Why do they think it is more aligned than Claude Fable 5.1?

keltan : “a significant step forward in […] alignment.”

Buddy, how tf are you measuring ‘alignment’? Would love to know because being able to measure that would save the fucking world.

roon (OpenAI): low rates of cheating

Rob Miles: *detected cheating

keltan : Thank you for clarifying. But you know what I’m gonna say next, right?

roon (OpenAI): that this metrics are not a full solve of alignment and will break discontinuously

keltan : Yep. But I would have said it in a dumber way. Something like: Low Rates of Cheating ≠ Alignment

roon (OpenAI): I agree but also in some real sense Astra is more aligned than Sol.

Astra is in key ways more aligned than Sol, as far as I can tell, but: Oh no.

OpenAI President Greg Brockman confirmed that they did the ‘standard testing process together with the government’ and the government did not ask for any changes. I do not think those running the government testing understand what is going on, and believe they are reacting basically on vibes and what various trusted people tell them.

This focus is especially important because of what we learned yesterday. Yes, they solved a Millenium Problem, I remember when I would have cared about that or been surprised, but they also said something far more important, that they very much did not have to say:

OpenAI: Since August 28 we have been training a new internal model that has exhibited unprecedented performance in our benchmarks, including mathematics. This model’s training is ongoing and its performance continues to improve.

That means this effort started four days into training the new model, on September 1. The work finished on September 5, eight days after the model began training.

Another implication is that they spun up 10,000 concurrent agents to form a swarm, of a new model that had been in training on the order of a few days. We could not possibly know that such a model is aligned.

If that is all it takes to get a generation ahead of Astra, and we are willing to move this fast, these pauses in training are not going to end up meaning very much, time is even shorter than we knew, singularity, singularity, singularity, singularity, oh I don’t know that I’m going to be alive that much longer. Yikes. Holy ****.

OpenAI is giving us what might be our last warnings. We better not waste them.

This makes it that much more important to take a critical look at Astra, the evidence for its alignment or lack thereof, and its potentially dangerous capabilities, while knowing that 6.1 already exists and can be trained largely within a week.

I see the model card making a number of claims about evidence for Astra’s alignment, or what we can be confident Astra cannot do, that seem unjustified to me. This includes claims about its inability to do steganography, and that some of its ‘better aligned’ behaviors are good news about Astra’s alignment, rather than warning signs.

Table of Contents

  • OpenAI’s Safety Claims About Astra (1).
  • Preparedness Capabilities Assessment (10).
  • Biological and Chemical Capability is High.
  • Cybersecurity Capability is Critical.
  • AI Self-Improvement Capabilities (10.1.3).
  • Astra Is Highly Verbally Eval Aware (from 8.6).
  • Safe Mundane Completions (4.1).
  • Jailbreaks (5.1).
  • Prompt Injection (5.2).
  • Health (6).
  • Hallucinations (7).
  • Alignment (8).
  • Obeying Restrictions (8.2).
  • That’s Worse, You Do Get How That’s Worse, Right?
  • OpenAI Does Not Understand Why This Is Worse.
  • The Alternative Explanation Is Also Worse.
  • Metagaming (8.7).
  • Alignment Faking (8.7).
  • Don’t Lie to the User (8.3).
  • Misalignment in Realistic Work Environments (8.4).
  • Unintended Agent-to-Agent Communication (8.5).
  • The Three Obviously Monitored Temptations of Astra.
  • Severe Issues In Simulated Traffic Are Down By Half.
  • UK AISI External Evaluations (8.8).
  • Sabotaging Safety Work.
  • What About The July 19 Attacks?
  • Apollo Research External Evaluations (8.8.1).
  • The Alignment Verdict.
  • It Depends What You Mean By Alignment.

OpenAI’s Safety Claims About Astra (1)

I have ordered them by my decreasing confidence in their validity. I mostly believe the first six claims here. The seventh is complicated.

  • GPT-6 Astra meets our Critical threshold for Cybersecurity.
  • As in, it can do serious damage in new ways against hard targets.
  • GPT-6 Astra’s monitorability has decreased relative to GPT-5.6 Sol.
  • I covered this yesterday. It is some scary stuff. It can avoid incriminating itself in the CoT. It can evade internal monitors on sabotage tasks.
  • OpenAI claims this loss of monitorability is mostly or entirely unrelated to their use of a new technique called recurrent depth. I have my doubts.
  • A model looking like it is becoming smarter, attempting shenanigans less often, and more often doing what you want, but getting better at hiding its actions when it wants to do that, is exactly the scary combination.
  • OpenAI are deploying misalignment monitoring broadly.
  • This includes internal deployments. Good.
  • They say this is at substantial compute cost.
  • GPT-6 Astra more responsibly navigates browsing and workplace settings.
  • This is better robustness against prompt injections and destructive actions.
  • Sol was a lot worse at this than Opus and Fable.
  • GPT-6 Astra is significantly safer in higher-risk scenarios than GPT-5.6-Sol.
  • As in, it does more safe things while refusing more unsafe things. Good.
  • GPT-6 Astra is significantly more robust to jailbreaks than its predecessors.
  • GPT-6 Astra is better aligned than GPT-5.6 Sol.
  • This appears to refer to a new suite of alignment evaluations.
  • These evaluations appear to be ‘flags raised over various simulated tasks.’
  • Making qualitative claims about alignment, based on quantitative data on mundane use case tests, was bullshit when Anthropic did it, and it is bullshit now when OpenAI does it. You cannot conclude one from the other.
  • In this case, the claim seems extremely ill-formed.
  • I issue a call for precision. If you mean ‘better on the evals,’ say that.

The first two are big trouble. Astra can do real damage, and is hard to monitor.

I believe OpenAI that they are deploying monitoring, but it will be less effective.

We then get four claims about Astra being a safer model.

Claims four through six are about Astra in practice. I believe these in relative terms, as improvements compared to Sol.

The last is an abstract claim, that Astra is ‘better aligned.’ What does that mean?

I notice that in their announcements, the claim is that Astra is the most aligned model in the world, period. Here, the claim is only that Astra scored higher on an eval metric than Sol. There is a big difference.

Given that Astra is highly eval aware and very good at evading monitors, it is hard to be confident it is relatively aligned.

One way to divide alignment features is between ‘mundane alignment,’ as in how you would feel about using the model for practical purposes, and the alignment that ultimately matters, that OpenAI has in the past called ‘super alignment,’ as in the thing without which we all die when AIs get sufficiently capable.

Astra’s mundane alignment is greatly superior to Sol. For practical purposes, I was actively nervous about some potential uses of Sol, in a way I am not for Astra.

Astra’s super alignment status should scare the living daylights out of you.

Preparedness Capabilities Assessment (10)

更进一步:量化金融体系

看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力

进入量化体系 →

关联讨论

同一事件的更多信源

相似阅读

关联信息,但可能不是同一事件