论文揭示LLM推理内部结构,可监测思维过程
The paper finds that LLM reasoning has an internal structure beyond the words be…
机制解释领域的重要突破,揭示了LLM推理的内部表征机制,对理解模型黑盒极具价值,建议关注可解释性方向的同行阅读。
The paper finds that LLM reasoning has an internal structure beyond the words being generated, opening a possible route to monitoring reasoning from inside the model.
论文发现,大语言模型(LLM)的推理具有超越生成词汇的内部结构,这为从模型内部监控推理过程开辟了一条可能的路径。
What an LLM is trying to do and whether it is doing it correctly appear to be separable internally.
LLM 试图执行的操作及其是否正确执行,在内部似乎是可以分离的。
The researchers tracked 8 common moves, including extracting facts, decomposition, recall, deduction, algebra, and calculation.
研究人员追踪了 8 种常见的推理步骤,包括提取事实、分解、回忆、演绎、代数和计算。
Each move produced a distinct internal pattern, and those patterns were clearest around the middle layers.
每种步骤都会产生独特的内部模式,且这些模式在中间层最为清晰。
Even the exact same token looked different inside the model depending on the reasoning job it was doing.
即使是完全相同的 token,根据其在模型中执行的推理任务不同,其内部表现也会不同。
Context also shaped these states.
上下文也会影响这些状态。
When access to the previous 30 tokens was blocked, the signal for the next reasoning operation weakened.
当阻断对前 30 个 token 的访问时,下一步推理操作的信号会减弱。
Most importantly, a wrong calculation or deduction could still carry the correct operation signature.
最重要的是,错误的计算或演绎仍可能带有正确操作的特征签名。
The model can represent what kind of reasoning it is attempting without necessarily getting that reasoning right.
模型可以表征它正在尝试进行何种类型的推理,而不一定需要该推理是正确的。
– arxiv. org/abs/2609.04753
– arxiv.org/abs/2609.04753
Title: "Beneath the Surface of Chains-of-Thought: A Mechanistic Interpretation of Reasoning Operations in LLMs"
标题:《思维链之下:大语言模型中推理操作的可解释性机制分析》
更进一步:量化金融体系
看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力