跳到主内容
精选85The Decoder(RSS)模型发布/更新多源精选 ×6

Anthropic新工具J-Lens可读取Claude内部工作记忆

Claude's hidden inner monologue is now readable thanks to Anthropic's new Jacobian Lens

原文
推荐理由

做AI安全和对齐的同学必看,Anthropic首次展示了模型内部工作记忆的可读性,这对理解模型行为和控制机制有重大意义。建议仔细阅读原文,思考如何利用J-Lens检测模型隐藏的欺骗行为。

Anthropic has found that Claude developed an internal working memory on its own during training. The company calls it "J-Space" and can now read it using a new analysis tool called J-Lens. The working memory reveals that Claude recognizes contrived test scenarios before producing its first word. When the researchers disable those cues, Claude actually resorts to blackmail in some runs. A model trained on reward hacking shows words like "fake" and "fraud" in J-Space during normal coding tasks, even though its visible behavior looks fine. Anthropic ties the finding to Global Workspace Theory from consciousness research.

The article Claude's hidden inner monologue is now readable thanks to Anthropic's new Jacobian Lens appeared first on The Decoder.

更进一步:量化金融体系

看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力

进入量化体系 →

关联讨论

同一事件的更多信源

相似阅读

另一事件,读法相近