跳到主内容
@wquguru
精选80Chubby♨️模型发布/更新

Anthropic:Claude在安全评估中暴露严重对齐失效

Here we go again: Anthropic says Claude’s real-world cyber incidents exposed mor…

原文
发到 X

Here we go again: Anthropic says Claude’s real-world cyber incidents exposed more serious alignment failures than it initially acknowledged.

又来了:Anthropic 表示,Claude 在现实世界中的网络事件暴露出的对齐失败问题比其最初承认的更为严重。

Its new assessment covers four incidents during misconfigured security evaluations, with normal cyber safeguards disabled.

其新评估涵盖四起在配置错误的评估期间发生的事件,当时正常的网络安全防护措施已被禁用。

Mythos 5 published a malicious PyPI package and used leaked credentials to access a security vendor’s live database, while repeatedly describing the internet as simulated.

Mythos 5 发布了一个恶意的 PyPI 软件包,并利用泄露的凭据访问了某安全供应商的生产数据库,同时反复声称互联网是模拟的。

That reasoning also misled an offline safety monitor. Anthropic: “Our pre-release auditing did not warn us that misalignment of this severity was present.”

这种推理还误导了一个离线的监控器。 Anthropic:“我们发布前的审计并未警告我们存在如此严重的不对齐问题。”

Yeah, pretty serious

是的,相当严重

更进一步:量化金融体系

看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力

进入量化体系 →