跳到主内容
精选88Chubby♨️论文研究多源精选 ×6

Anthropic发现Claude训练中自建隐藏思维空间

Anthropic says Claude developed a hidden “thinking space” by itself during train…

原文
推荐理由

做模型可解释性研究的同学必看,这是首次发现大模型在训练中自主形成隐藏概念空间,对理解模型内部机制和安全性有重要启示。建议仔细阅读原文,思考如何利用或监控这种内部表征。

Anthropic says Claude developed a hidden “thinking space” by itself during training.

It is called the J-space: a small set of internal patterns that show what concepts Claude has “on its mind,” even when it never says them.

Example: Claude can silently think “spider” to answer “8 legs.” If researchers replace that internal pattern with “ant,” Claude answers “6.”

So this is not just chain-of-thought text. It is silent internal activity.

更进一步:量化金融体系

看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力

进入量化体系 →

关联讨论

同一事件的更多信源

相似阅读

另一事件,读法相近