跳到主内容
精选85The Zvi(RSS)行业动态

Anthropic 与美政府博弈后恢复 Fable 模型访问

Fable #6: The Return of the King

原文
推荐理由

Anthropic 与美政府博弈的完整内幕,涉及模型管制、安全妥协与行业影响,做 AI 安全或政策研究的同学必读,了解未来模型发布可能面临的审批流程。

The blip is over. We have Fable back.

Utah teapot: happy fable/mythos easter Wednesday, to those who celebrate

Here is the official letter restoring Fable, great job everyone. Notice it is addressed to Tom Brown, not to Dario Amodei.

Anthropic had to make the controls more stupid for now, but this is a big win.

j⧉nus: YES!!! I’m really proud of Anthropic for their successful negotiation with the government. Also positive update on the government being sane and possible to cooperate with. Afaik Anthropic didn’t need to agree to any bad terms / genuflect / betray their principles or dignity.

The fiasco continues, at least until such time as we have a systematic regime in place for future frontier models rather than decisions being made ad hoc, by people like Lutnik and Bessent who do not know how any of this works.

The Blip

Anthropic explains its version of what happened.

Here is the timeline:

  • Amazon researchers discover they can ask Fable to ‘fix this code.’
  • They alert the White House, which freaks out.
  • June 12: US government tells Anthropic to take down Fable on its own.
  • June 12: Anthropic responds that This Is Fine and the concern is misplaced.
  • June 12: US government applied export controls to Mythos and Fable.
  • Anthropic works with US government and expands classifiers, such that it refuses Amazon’s request to ‘fix this code’ in over 99% of cases.
  • June 26: US government eliminated the controls on Mythos.
  • June 30: US government fully lifted those controls on Fable as well.
  • July 1: Fable access was restored worldwide.
  • July 8: customers will have to pay by the token. Shut up and take my money.
  • Going forward: Anthropic is working with the government and also other Glasswing partners like Amazon, Microsoft and Google on a classification system for jailbreaks, and rules for all of this, to prevent this from happening again.
  • Going forward: Anthropic will continue to collaborate with the US government, including on future model releases.

How stupid are the extra near term safeguards they had to include here? Really stupid:

Anthropic: In the near term, some routine tasks like coding and debugging will fall back to Opus 4.8.​

We’ll continue to refine these classifiers over the coming weeks to reduce false positives and better distinguish genuine misuse from legitimate requests.

The Amazon ‘jailbreak’ was ‘fix this code.’

Debugging is literally ‘fix this code.’

So I don’t know what you want Anthropic to do here. I do know Fable is coding for me.

Here is Anthropic’s basic explanation:

  • Claude Fable 5 was never an issue, our safety mechanisms were collectively robust.
  • No, they’re not perfect, but nothing will ever be perfect, in practice This Is Fine.
  • The government freaked out over nothing, which is largely Amazon’s fault.
  • We have made the safeguards stupider and successfully calmed them down.
  • Hopefully we can fix this so it’s less stupid going forward.

Alex Stamos has a thread unpacking a punch of Anthropic’s language in its announcement.

Alex Stamos: A lot to unpack here. Anthropic is burying some hard truths in careful political language. Some initial reads:

  • Anthropic verifies that none of the jailbreaks provided a capability beyond what many other models, including Chinese models, could do.
  • Anthropic makes the cost of this White House freakout clear. US labs now have to make a much more conservative precision-recall tradeoff on cyber refusals. US models will become much less useful for defensive cybersecurity work unless you are in the trusted group.
  • “No big deal, just join the trusted group!” the apologists will say, but the restrictions mean you can’t build a product on those models. Security companies and startups that provide services to others will now be driven to use Chinese models. Big win for PRC labs this month.
  • CAISI is the group that is supposed to actually make these determinations, not the political actors in the White House. They were positive on the prior safeguards. The implication is that this whole thing was unnecessary.
  • There is no good scoring framework for jailbreaks; this would be an improvement. The inclusion of Amazon as the first name in the coalition is not an accident. Anthropic is saying “Amazon’s inability to appropriately communicate severity threw our industry into chaos”.
  • “You don’t have to get Dario on the phone to talk to us about these things. Other people work here, we swear.”
  • In short, Anthropic’s blog is saying: We have always cared about safety, we did a good job initially, the actual AI experts in USG agreed, we proved it, we will come up with standards so these things are better communicated, welcome to the AI safety club Trump admin.
  • This was a huge own goal for the US, and we will see how bad US models get over the next six months and if Chinese models become noticeably better for cyber work.
  • For all the “This is what Anthropic wanted” people/bots. No, they didn’t. They didn’t want a stupid, knee-jerk response on a Friday. We give the USG huge powers, this is why you staff it with competent, calm, non-corrupt people who don’t use those powers to punish enemies.
  • The only upside I can see from this whole mess is that there is a whole bunch of VCs with former or current Administration affiliation who we can now safely ignore on AI policy. They have shown that everything they ever said on AI regulation was just politically motivated.

I think Stamos is overreaching with the consequences in places, especially with #2 and #7. Otherwise he’s right.

I do not expect US models to ‘get bad’ over time, only that they will get better slower, and have more area where they have rather annoying safeguards.

My expectation is that right now is the most obnoxious the safeguards will ever be, on both the bio and cyber fronts. I expect the freak-out to subside over time, and my guess is most of it surrounds Mythos in particular. You don’t need or often even want Fable for most such product offerings, and Opus or Sol will remain well ahead of Chinese alternatives.

Contra Prinz, I do not think this commits Anthropic to going through the approval process will all future releases, only releases that pose plausible risks. We tested this with Sonnet 5, where it looks like Anthropic went ahead and dropped it on its own, and no one is suggesting there was anything wrong with doing so (other than to complain that they want Sonnet to be better).

The White House Explanation

It was a little weird.

Susie Wiles (White House Chief of Staff): Under President Trump’s leadership the United States is the undisputed winner in the AI race.

My gratitude to companies across industries who continue to work closely with the White House to implement the President’s EO: “Promoting Advanced AI Innovation and Security.” This includes excellent work around advanced model access and guardrail testing and security. The government and private sector have worked together in a way we have never seen before and this foundation of America First is unprecedented.

Our shared priority remains: get the best tech deployed as quickly and safely as possible.

Howard Lutnick (Secretary of Commerce): Over the past two weeks, we have worked closely with Anthropic to analyze and approve Fable 5 to ensure alignment across the US Government and strengthen America’s leadership in AI.

Tom Brown (Chief of Compute, Anthropic, Lead Negotiator): Thanks for your partnership on this, Secretary!

‘Alignment across the US government’ is very much a case of ‘PHRASING!’ and here presumably means interagency sign-off, not ‘the model is now aligned with the US government.’ Unclear whether he knows enough to be trolling here.

As in, before a model can be released, you now likely need this ‘alignment,’ which in practice means sign off from various potential veto points, starting with Commerce and the Pentagon. Who knows how many more fully count.

更进一步:量化金融体系

看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力

进入量化体系 →

关联讨论

同一事件的更多信源
当前Fable体验
Emad原文
Fable沉迷于说中国坏话
Teortaxes▶️ (DeepSeek 推特🐋铁粉 2023 – ∞)原文

相似阅读

另一事件,读法相近