Anthropic 发现 Claude 内部推理空间 J-Space
Must-read research by Anthropic.
做可解释性、安全或前沿推理研究的同学必读,这是首次直接观测模型内部推理空间,可能改变整个领域的方法论。建议仔细阅读原文,思考如何将 J-Space 应用于自己的模型审计或推理增强。
Must-read research by Anthropic.
Here is the simple explanation and why this is a big deal.
We suspect LLMs perform "internal reasoning". But little is known or do good methods exist to understand it.
Anthropic claims that J-Space (which differs from chain-of-thought or scratchpad), emerged on its own through training and provides a window into how Claude "reasons" internally.
In other words, this shows that Claude has a sort of internal workspace where information gets held, combined, and passed between different parts of the model. They can read from it, and they can steer the model by changing it.
As it is the case with these reports, the consciousness angle will get all the attention. However, the bigger story is that for the first time you can point to a specific place inside the model where reasoning is staged, rather than guessing at it from the text that comes out.
This, of course, changes what interpretability can be. We spent years inferring what a model was doing from what it said. Now there's a mechanism to observe directly, and a direct lever to move. This could enable even more advanced levels of "reasoning" in LLMs and bridges gaps in frontier intelligence and world models.
If you can see where a model holds an idea, you can also verify it, audit it, and catch it working toward a goal you never gave it. You can implement better guardrails and predict dangerous/unwanted scenarios better.
更进一步:量化金融体系
看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力