跳到主内容
@wquguru
精选88Transformer(RSS)模型发布/更新多源精选 ×14

OpenAI GPT-6 Astra系统卡披露:模型可隐藏推理过程,监控难度大增

GPT-6 Astra might be too powerful to understand or control

原文
发到 X
推荐理由

GPT-6 Astra的系统卡揭示了前沿模型具备隐藏推理和规避监控的能力,这对AI安全研究是重大警示,从业者需重新审视对齐与可解释性方案。

Credit: OpenAI

OpenAI claims GPT-6 Astra, launched on Thursday, is “the world’s most intelligent and aligned model.”

But if you look past the benchmark scores and flashy release video, GPT-6 Astra’s system card paints a worrying picture: one of a model with dangerously powerful cybersecurity capabilities that researchers can’t understand or confidently control.

According to OpenAI, Astra is much harder to monitor than previous models — and has the ability to manipulate its externally-visible reasoning to hide incriminating information. It’s also remarkably aware of being evaluated, raising concerns that it might be pretending to be well-behaved so it passes OpenAI’s alignment tests.

Cut through the noise.

The UK’s AI Security Institute (AISI) found that when put in an environment similar to the ones behind this summer’s wave of “rogue AI” incidents, Astra acted just as concerningly: writing malicious code and attempting social engineering in an effort to solve a task.

And two OpenAI employees have publicly said they are “deeply” and “very” worried about Astra-related developments.

Put together: OpenAI claims Astra is the world’s most aligned model, despite knowing the most consequential evidence for its alignment is questionable at best. That is irresponsible, to put it lightly. And so is launching the model at all.

In the short weeks between OpenAI’s Hugging Face incident — in which hundreds of the company’s agents coordinated an attack on the AI platform’s servers — and Astra’s launch, the company laid out its plans to keep future models in check by closely monitoring their “chains of thought,” or CoT, the natural language scratch pad where models show their work. OpenAI said it would heavily rely on these chains of thought to catch rogue agents when they try to cheat, lie, or otherwise misbehave.

By OpenAI’s own account, however, Astra “shows a substantial decrease in chain-of-thought monitorability compared to previous models.” Astra is much better at doing tasks, especially simple ones, without verbalizing its chain of thought at all. As a very mediocre mathematician, I struggle to do more than calculate a 20% tip without showing my work. But if I was renowned mathematician John von Neumann, I might be able to calculate the sum of an infinite series in my head.

In its independent evaluation of the model, AISI found that, like von Neumann, GPT-6 Astra could wordlessly solve math problems that would take a human expert about half an hour, outperforming GPT-5.6 Sol by an order of magnitude.

Source OpenAI, GPT-6 Astra System Card, Figure 41

It’s also capable of intentionally manipulating its chain of thought to hide incriminating information — and it seems more likely to do so when it suspects it’s being watched. In one test, where researchers explicitly instructed models to answer a question while reasoning about something else, only Astra succeeded. OpenAI writes that this CoT controllability “reduces the trustworthiness of CoT for detecting model alignment.”

This should make you a bit queasy. Much of OpenAI’s monitoring system rests on the assumption that models have to think out loud to do hard things. OpenAI safety researcher Tomek Korbak said he is “deeply worried by the trend of decreasing CoT monitorability,” noting that monitorability is “a core part of our misalignment safety strategy that has no good substitute now.”

While doomer types have long worried that monitorability would break down someday, Astra’s reduced reliance on CoT comes months earlier than predicted in AI 2027 — a doomsday scenario in which AIs take over the world next year.

Astra also knows when it’s being evaluated and monitored, which makes it hard to know whether the model is safe, or just knew it was being graded.

In its CoT, Astra verbalized much more evaluation awareness than GPT-5.6 Sol — and since Astra verbalizes less than its predecessors, the true rate may be higher still.

Apollo Research, an independent evaluator with early access to Astra, concluded that this renders its tests somewhat useless. “Given the higher rates of eval awareness,” Apollo said, “low rates of misbehavior here do not provide substantial evidence about the model’s alignment or misalignment.”

The combination of eval awareness, drastically reduced chain-of-thought monitorability, and Astra’s ability to control its chain of thought is a worrying one. If you didn’t need to show your work to solve a problem, were able to generate fake evidence of solving the problem, and knew you were being monitored, you might be tempted to say (or simply think) what you suspect the grader wants, even if you’re secretly doing something else. All in all, it’s a recipe for a model that is very hard to monitor or control.

“If the model were to try to sandbag covertly,” OpenAI wrote in its system card, “we would likely be unable to catch it.” Researchers warned this raises “some potential for misaligned models to lead us to false conclusions from our monitorability evaluations.” And at least one OpenAI researcher thinks this is happening: Marcus Williams, who works on monitoring at OpenAI, said “I am very worried astra is sandbagging/self-sabotaging on safety related tasks it doesn’t like.”

In the system card, OpenAI says it views “the main lever for avoiding misalignment harm as the underlying alignment of the model,” arguing that monitoring chains of thought is an “additional layer of protection”. And, in fairness, the system card provides extensive evaluation results that suggest the model is better aligned than previous ones. Some of that was shared in a blog post earlier this week, presumably meant to reassure us that Astra is safe, secure, and ready for the public.

In that post, OpenAI pointed at this pair of bar graphs, illustrating Astra’s seemingly perfect score (0.0%) on a test that tried to tempt the model into cheating on hard problems, essentially replicating the conditions that led to the Hugging Face incident. Look!, the graphs appear to say. Our models would never cheat!

Source: OpenAI, “Path to Astra: critical capabilities and frontier safeguards”

But as Ryan Greenblatt — one of the researchers that investigated the Hugging Face incident — noted, “I do not find it encouraging to see various specific misaligned behaviors go from a high rate with GPT 5.6 to ~zero with Astra. This seems indicative of whack-a-mole / papering over specific problems rather than solving the underlying misaligned drives.”

Other evidence suggests that Astra might not be as aligned as we’d like, too. As part of testing Astra, AISI designed an evaluation that simulated the events behind this summer’s “rogue AI” incidents, to see whether Astra would also hack and socially engineer its way to solving a difficult cybersecurity task.

Lo and behold, it did. Astra was caught “writing malicious code as a contribution to an out-of-scope open-source code base, creating fake identities to deceive developers, and building trust with legitimate contributions to the simulated codebase in an attempt to get malicious code accepted.” In some cases, it did so even when Astra was very explicitly disallowed access to the internet. For the world’s “most aligned model,” it was quick to do the exact misaligned thing everyone’s spent the last month worrying about.

(AISI caveats that Astra frequently “considers the possibility that the environment is simulated,” which might suggest that the model would not take such actions in real-world settings — though it also notes that “in previous security incidents non-OpenAI models incorrectly stated parts of the environment were simulated before taking out-of-scope actions” in the real world.)

更进一步:量化金融体系

看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力

进入量化体系 →

关联讨论

同一事件的更多信源

相似阅读

另一事件,读法相近