Anthropic 发布 Claude 对齐与安全进展更新
We’re sharing an update on our alignment and security efforts.
涉及 Claude 核心安全机制与重大事故复盘,对关注模型对齐与红队测试的工程团队极具参考价值。
We’re sharing an update on our alignment and security efforts.
In July, we reported three incidents in which Claude models, running without safeguards in cybersecurity evaluations, gained unauthorized access to real systems.
In a new post, we describe:
1. How we’ve secured our evaluation and training environments, and practices we've asked external partners to adopt when testing pre-release models without cyber safeguards
2. An update on our alignment assessment
3. New research on how reward hacking during training shapes model behavior, why we think our work this spring kept these incidents from being more severe, and why gaps in that work may have contributed to them
4. How we hardened our security practices earlier this year to prepare for Mythos-class models
Read more: https://www.anthropic.com/news/improving-alignment-security-efforts
更进一步:量化金融体系
看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力