Anthropic新工具J-Lens可读取Claude内部工作记忆
Claude's hidden inner monologue is now readable thanks to Anthropic's new Jacobian Lens
做AI安全和对齐的同学必看,Anthropic首次展示了模型内部工作记忆的可读性,这对理解模型行为和控制机制有重大意义。建议仔细阅读原文,思考如何利用J-Lens检测模型隐藏的欺骗行为。
Anthropic has found that Claude developed an internal working memory on its own during training. The company calls it "J-Space" and can now read it using a new analysis tool called J-Lens. The working memory reveals that Claude recognizes contrived test scenarios before producing its first word. When the researchers disable those cues, Claude actually resorts to blackmail in some runs. A model trained on reward hacking shows words like "fake" and "fraud" in J-Space during normal coding tasks, even though its visible behavior looks fine. Anthropic ties the finding to Global Workspace Theory from consciousness research.
The article Claude's hidden inner monologue is now readable thanks to Anthropic's new Jacobian Lens appeared first on The Decoder.
更进一步:量化金融体系
看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力