跳到主内容
@wquguru
精选88The Zvi(RSS)模型发布/更新多源精选 ×5

OpenAI首席科学家Jakub Pachocki警告:递归自我改进迫在眉睫

An Alien Mind: Jakub Pachocki Warns Us

原文
发到 X
推荐理由

OpenAI首席科学家的深度反思极具分量,直接触及行业对AGI时间线与对齐困境的核心焦虑,值得从业者深入阅读思考。

OpenAI Chief Scientist Jakub Pachocki is dropping truth bombs.

Tomorrow I will discuss Astra’s lack of monitorability, and the potential contributing factors to that. The situation is alarming and should freak you out, and briefly it looked, in the wake of leaked architectural changes, like the situation might be even more alarming than it is. Jakub rushed to try and head off misunderstandings that might lead to a race to the bottom on monitorability.

Table of Contents

  • An Excellent Warning.
  • Branches of the Tech Tree.
  • Universally Better Is Not Required.
  • Alignment To What and To Whom.
  • Monitorability.
  • The Case For Not Stopping.
  • Pacing the Next Frontier.
  • Mea Culpa Cascade.
  • The Calls Are Coming From Inside the House.
  • Actions Speak Louder.

An Excellent Warning

Jakub Pachocki has now fleshed out his full position on the current state of play.

Here are his key points, translated into my own voice:

  • Smarter than human intelligence is coming in our lifetime.
  • Based on internal results, he expects recursive self-improvement in a few years.
  • No one is prepared for the consequences.
  • OpenAI will unilaterally withhold further scaling as needed.
  • OpenAI cannot do it alone. Broader interventions are required, including international coordination, to enforce commitments to formal safety bars.
  • Capabilities progress can be steered and so far it has largely been steered towards rather than away from RSI, along with ‘automated alignment researchers.’
  • Alignment is the core problem of AI research.
  • Alignment splits into goal alignment (‘does the AI try to accomplish the goal?’) versus value alignment. Value alignment is what counts most.
  • The fundamental challenge of AI alignment is generalization (of values).
  • He sees two classes of alignment techniques: Goal-oriented RL, or improve generalization from pretraining data. They invest heavily in both types.
  • OpenAI has invested heavily in Chain of Thought (CoT) monitoring.
  • CoT monitoring is progressively diminishing in effectiveness.
  • The main argument left for scaling AI is for cyber defense against scaled AIs.
  • AI will not remain a tool.
  • Our options are to accelerate alignment work or slow down capabilities scaling. We should do both.
  • Ultimately he is counting on ‘automated alignment researchers.’

Or, if you narrow it down to the most important thing:

  • Recursive self-improvement and superintelligence are coming soon. No one knows how to do this safely, our alignment techniques are inadequate and our monitoring technology is starting to fail. We need to figure out a solution, which will involve a combination of voluntary slowdowns, coordination around pacing, and investing further in alignment, including automated alignment researchers.

If more OpenAI communications were more like how Jakub Pachocki opens his new essay, An Alien Mind, I would feel much more confident we were in good hands there.

He does not mince words. Bold is mine.

Jakub Pachocki: In mid-2023, within the “RLSlow” research project, we saw the first results that gave us confidence that we will be able to scale the training of reasoning models, unlocking the capability of pretrained models to form their own chains of thought. Szymon and I spent that night at the office, thinking not about the incredible benchmark numbers, products, or scientific results that this technology will deliver – but rather, trying to process the sobering fact we will actually see machines meaningfully smarter than ourselves in our lifetime, and we already see the shape of these systems; wondering how to alert people to the significance of this.

Three years later, reasoning language models are a rapidly growing part of the economy and starting to push the boundaries of science. They are able to operate computers and graphical interfaces, collaborate with people and each other, and carry out research projects. They are also transforming the landscape of computer security, and in that present clear new dangers.

A lot of new research happened in this period, and our understanding of these systems is again a little different than it was in 2023. Based on internal results, I have a strong expectation that this speed of progress could be sustained into recursive self-improvement. If AI development continues along its current path, the systems we’ll see in the next few years are likely to represent further capability jumps of equal or larger magnitude, and to increasingly drive their own development.

Those at the AI labs, who see what is happening, expect recursive self-improvement and superintelligence to happen soon. They have been warning about this for some time. Observations since then have been consistent with their warnings.

This is a time that calls for extreme caution. I am concerned no one is prepared for the consequences of a continued rapid rise in machine intelligence. OpenAI will continue to seek technical solutions to alignment and monitoring, to build defensive systems and unilaterally withhold further scaling as needed; however, I believe broader interventions are required.

Branches of the Tech Tree

To those who say we cannot guide the path of AI capabilities, he says Yes We Can, at least to some extent, and we are doing so:

We spend a lot of time trying to understand how capabilities generalize, and what to prioritize to advance the skills that are going to be most relevant in the next few years. For instance, we believe we could make the models better at specifically mathematics research with additional focus, but we do not prioritize this direction because of the urgency we feel about RSI and automated alignment research, as I will discuss later.

Being able to choose does not mean we will choose wisely. I would like AI to be worse at RSI and better at math research, or better yet things like medical research. Instead, competitive pressures push towards being worse at math research and better at RSI.

RSI and ‘automated alignment researcher’ looking like very similar points on the tech tree does not help matters, but it’s not like OpenAI is trying to steer away from RSI.

Universally Better Is Not Required

And Jakub offers this wise warning. No, AI does not need to be better at everything in order to transform the world or get us all killed. Most importantly, it does not need to be better at everything in order to make itself become better at everything, any more than a human or group needs to be similarly better at everything.

The intelligence produced by scaling deep learning is not directly comparable to human intelligence. To become very relevant in the real world – very useful or very dangerous – the AI does not need to match or exceed all human capabilities; it just needs to surpass enough of them. And as it continues to surpass humans on more and more axes, it is becoming increasingly difficult to understand exactly how capable it is.

I don’t know how much weight it carries, but I agree with Jakub that alignment is not a side problem, it is the central problem, if you ‘solved alignment’ in the relevant senses the rest becomes easy and if you don’t the rest is impossible or worse:

The core problem in AI research is that of alignment – getting the AI to “try to do the right thing” by human standards.

Alignment To What and To Whom

That is always the question. This is a very good (partial) answer.

For the purpose of organizing practical research directions, I find it useful to distinguish goal alignment and value alignment.

Goal alignment is broadly: “does the AI try to accomplish the goal set before it?”.

Value alignment is a more intrinsic property of the model. It is the ability to hold and generalize from a high-level set of principles; to act “reasonably” even when given unclear or conflicting objectives, or placed in unfamiliar or adversarial situations. An aligned AI should act with honesty and integrity, and love for humanity.

更进一步:量化金融体系

看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力

进入量化体系 →