跳到主内容
@wquguru
精选90r/LLMDevs(Reddit)技巧与观点

Qwen与Mistral复现模型原生数值记忆提取机制

I stored model-native numerical memory outside a frozen LLM and retrieved the associated memory through its own attention — Qwen → Mistral replication, 127/128 Top-1

原文
发到 X
推荐理由

提供了一套完整的、可复现的模型原生数值记忆提取与检索工作流,包含具体的层头定位与维度细节,对Agent长期记忆研究极具参考价值。

Two different 7B transformers. Two different internal coordinates. The same memory mechanism.

两个不同的 7B Transformer。两种不同的内部坐标。相同的记忆机制。

I started AKBASCORE MAM on Qwen2.5-7B-Instruct. I have now independently localized and replicated the mechanism on Mistral-7B-Instruct-v0.3.

我在 Qwen2.5-7B-Instruct 上启动了 AKBASCORE MAM。目前,我已在 Mistral-7B-Instruct-v0.3 上独立定位并复现了该机制。

Final Mistral result: 127/128 correct memories at Top-1 — 99.22%.

Mistral 最终结果:Top-1 正确记忆数为 127/128 —— 准确率为 99.22%。

Counterfactual retrieval: 125/128 — 97.66%.

反事实检索:125/128 —— 准确率为 97.66%。

Shifted-pointer control: 0/128.

移位指针控制:0/128。

The model weights were not changed. No fine-tuning. No LoRA. No optimizer. No learned router. No gold B-memory ID is supplied to the retriever. At query time, none of the 128 candidate B memories is forwarded through Mistral.

模型权重未发生任何改变。没有微调。没有 LoRA。没有优化器。没有学习到的路由器。检索器不会接收任何黄金 B 记忆 ID。在查询时,128 个候选 B 记忆中没有任何一个被送入 Mistral。

But I don't think 99.22% is the most interesting part of this experiment.

但我认为 99.22% 并不是这个实验最有趣的部分。

The more important question is: what exactly is the model reading?

更重要的问题是:模型究竟在读什么?

Because this system is not searching through 128 text documents in the conventional sense. (Fig. 1)

因为该系统并非以传统方式在 128 份文本文档中进行搜索。(图 1)

  • The text enters the model once
  • 文本仅进入模型一次

When a memory is formed, its source text is processed through the frozen transformer.

当记忆形成时,其源文本会经过冻结的 Transformer 进行处理。

The transformer already produces K and V states for its own attention computation. I do not treat selected parts of those states merely as temporary computational residue. I extract them as model-native numerical memory.

Transformer 已经为其自身的注意力计算产生了 K 和 V 状态。我不将那些状态的选定部分仅仅视为临时的计算残留物。我将它们提取为模型原生的数值记忆。

So instead of only having a sentence such as:

因此,我们不再只有类似这样的句子:

“Instrument Zyrhyn carries seal PCI.”

“仪器 Zyrhyn 携带密封 PCI。”

we now have a numerical memory structure derived directly from the model's own internal computation.

我们现在拥有的是直接源自模型自身内部计算的数值记忆结构。

  • That numerical structure can exist outside the model
  • 这种数值结构可以存在于模型之外

The memory does not have to remain inside the model weights.

记忆不必保留在模型权重内部。

The numerical structure can be held in RAM, serialized, written to a file, and therefore moved to persistent storage.

该数值结构可以保存在 RAM 中、进行序列化、写入文件,从而转移到持久化存储中。

This distinction matters.

这一区别至关重要。

When people say “persistent LLM memory,” the usual architecture often means storing text, documents, embeddings, database records, or summaries and later retrieving them so the model can read them again.

当人们提到“持久化 LLM 记忆”时,通常的架构往往意味着存储文本、文档、嵌入、数据库记录或摘要,并在稍后检索它们,以便模型再次读取。

The path I am investigating is different:

我所探索的路径有所不同:

instead of storing human-readable information so the model can reread it, store a machine-native numerical state derived from the model itself.

不是存储人类可读的信息供模型重读,而是存储源自模型本身的机器原生数值状态。

The model weights remain frozen.

模型权重保持冻结。

  • Then how does the model know which numerical memory to retrieve?
  • 那么,模型如何知道要检索哪个数值记忆呢?

This is the part I find most interesting.

这是我觉得最有趣的部分。

A new question arrives:

一个新问题出现了:

“What seal does instrument Zyrhyn carry?”

“仪器 Zyrhyn 携带什么印章?”

The active numerical A memory is installed into the transformer cache.

活动的数值 A 内存被安装到变压器缓存中。

Mistral processes the new question.

Mistral 处理这个新问题。

A particular channel of the frozen model's own attention mechanism then produces a question-conditioned distribution over the A-memory positions.

随后,冻结模型自身注意力机制的特定通道产生了一个关于 A 内存位置的问题条件分布。

For Mistral, that channel is:

对于 Mistral,该通道是:

L28H00.

L28H00。

I call this endogenous attention pointer ÇAĞRIİZ.

我将此内生注意力指针称为 ÇAĞRIİZ。

It is not an external neural router.

它不是外部神经网络路由器。

It is not a trained classifier.

它不是经过训练的分类器。

It is not a gold memory ID.

它不是黄金内存 ID。

It comes from the transformer's own native:

它来自变压器自身的原生:

Q · K

attention computation.

注意力计算。

  • The pointer is then applied to V space
  • 随后将该指针应用于 V 空间

This is where K and V take different operational roles.

这正是 K 和 V 承担不同操作角色的地方。

K helps answer:

K 帮助回答:

Where should I look?

我应该看哪里?

V helps answer:

V 帮助回答:

What numerical address should I construct from what I found?

我应该根据我找到的内容构建什么样的数值地址?

For Mistral, the independently localized address space is:

对于 Mistral,独立定位的地址空间是:

L00-V, uncentered.

L00-V,未居中。

8 KV heads × 128 dimensions:

8 个 KV 头 × 128 维:

1024 dimensions.

1024 维。

Applying the ÇAĞRIİZ distribution to those V states produces a question-conditioned 1024-dimensional address.

将 ÇAĞRIİZ 分布应用于这些 V 状态,会生成一个由问题条件决定的 1024 维地址。

At this point there is no filename telling the retriever which B memory to choose.

此时没有文件名告诉检索器应选择哪个 B 记忆。

There is no B-memory ID supplied by the experiment.

实验未提供 B 记忆 ID。

There is a 1024-dimensional retrieval address derived from the frozen model's own internal computation. (Fig. 3)

存在一个源自冻结模型自身内部计算的 1024 维检索地址。(图 3)

  • Then all 128 B memories compete
  • 随后所有 128 个 B 记忆展开竞争

Every B memory has already been independently converted into its own numerical structure.

每个 B 记忆都已被独立转换为其自身的数值结构。

Each retains token-level V rows in the localized address space.

每个 B 记忆在局部化地址空间中均保留了 token 级别的 V 行。

The live 1024D address is compared against every B memory.

实时的 1024D 地址将与每个 B 记忆进行比较。

For each B candidate, the system computes the maximum cosine similarity between the live address and that candidate's numerical rows.

对于每个 B 候选项,系统计算实时地址与该候选项数值行之间的最大余弦相似度。

Then:

接着:

128 candidates → 128 scores → argmax → one associated memory. (Fig. 2)

128 个候选项 → 128 个得分 → argmax(取最大值索引)→ 一个关联的记忆。(图 2)

In the real Item 001 demo run:

在实际的 Item 001 演示运行中:

Expected:

预期结果:

B#001 / PCI

Returned:

返回结果:

B#001 / PCI

Rank:

排名:

1/128

Top-1 cosine:

Top-1 余弦相似度:

0.999649

Second candidate:

第二候选项:

0.660254

Margin:

差距(Margin):

+0.339395

And the number of Mistral model forwards required for that retrieval was:

以及完成该次检索所需的 Mistral 模型前向传播次数为:

The model processed:

模型处理了:

27 question tokens over a preinstalled 32-slot A cache.

在预安装的 32 槽位 A 缓存上处理的 27 个问题 token。

Model forwards through the 128 candidate B memories:

模型对 128 个候选 B 记忆进行前向传播:

That distinction is important.

这一区分很重要。

After receiving the question, the system did not send 128 source texts back through Mistral to find the answer.

收到问题后,系统并未将 128 篇源文本通过 Mistral 回传以寻找答案。

  • Then I changed the memory
  • 然后我更改了记忆

I changed the association for the same Zyrhyn record:

我更改了同一 Zyrhyn 记录的关联:

PCI → COL

The model remained the same.

模型保持不变。

The weights remained the same.

权重保持不变。

The question remained the same.

问题保持不变。

The B-memory bank remained the same.

B 记忆库保持不变。

The changed A memory was re-forged and the same frozen retrieval mechanism was run again.

更改后的 A 记忆被重新锻造,并再次运行相同的冻结检索机制。

This time the system selected:

这一次系统选择了:

B#054 / COL

Rank:

排名:

1/128.

So changing the numerical A association redirected the same frozen mechanism to a different B memory. (Fig. 4)

因此,更改数值型 A 关联将相同的冻结机制重定向到了不同的 B 记忆。(图 4)

  • This was not evaluated on only one showcase item
  • 这并非仅在一个展示项上进行评估

For the final Mistral experiment, TEST528, the mechanism was frozen before the final evaluation panel.

对于最终的 Mistral 实验 TEST528,在最终评估面板之前,该机制已被冻结。

The result was:

结果为:

Primary retrieval: 127/128 — 99.22%

主要检索:127/128 — 99.22%

Counterfactual retrieval: 125/128 — 97.66%

反事实检索:125/128 — 97.66%

Controls:

对照组:

Shifted pointer: 0/128

偏移指针:0/128

NO-A: 1/128

无 A(NO-A):1/128

The single NO-A “success” has a simple implementation-level explanation: a zero address produces equal zero scores for every candidate, so the deterministic tie-break selects B#001. I therefore do not interpret it as meaningful memoryless retrieval. (Fig. 5)

唯一的 NO-A “成功”有一个简单的实现层面解释:零地址会导致每个候选项产生相等的零分,因此确定性平局打破机制会选择 B#001。因此,我不将其解释为有意义的无记忆检索。(图 5)

The failed cases are preserved in the experimental record.

失败案例保存在实验记录中。

  • Qwen did not use the same coordinates
  • Qwen 未使用相同的坐标

This is probably the most important result of the cross-model experiment.

这大概是跨模型实验最重要的结果。

Qwen2.5-7B-Instruct:

Qwen2.5-7B-Instruct:

L23H12 → L02-V → 512D

Mistral-7B-Instruct-v0.3:

Mistral-7B-Instruct-v0.3:

L28H00 → L00-V → 1024D

I did not copy Qwen's layer/head coordinates into Mistral.

我没有将 Qwen 的层/头坐标复制到 Mistral 中。

Mistral's pointer and address regions were independently localized and then frozen before the final evaluation.

Mistral 的指针和地址区域是独立定位的,然后在最终评估前被冻结。

So the strongest statement I think the current evidence supports is:

因此,我认为当前证据所能支持的最强结论是:

The mechanism transferred across model families; the internal coordinates did not.

机制在不同模型家族间实现了迁移;内部坐标则没有。

Two models do not establish universality.

仅两个模型无法确立普遍性。

But this is also no longer a result tied to one accidental coordinate in one transformer.

但这也不再是一个与某个 Transformer 中偶然出现的单一坐标绑定的结果。

So where is the “persistent machine memory” part?

那么,“持久化机器记忆”部分在哪里?

I want to be precise about this because it is easy to overstate what the current prototype demonstrates.

我想对此保持精确,因为很容易夸大当前原型所展示的内容。

This experiment does not demonstrate 150 million memories.

本实验并未展示一亿五千万条记忆。

It does not yet demonstrate a production-scale persistent memory database.

它尚未展示生产规模的持久化记忆数据库。

It does not yet demonstrate the model discovering the first A memory automatically from an enormous inactive bank.

它尚未展示模型从庞大的非活跃存储库中自动发现第一条 A 记忆的能力。

What it does establish is a smaller but, in my view, more fundamental chain:

它所确立的是较短但在我看来更基础的链条:

Model-derived numerical memory can be extracted.

模型衍生的数值记忆可以被提取。

Numerical memory can exist outside the model weights.

数值记忆可以存在于模型权重之外。

Active numerical memory can be connected back to transformer computation.

活跃的数值记忆可以重新连接到 Transformer 计算中。

A natural-language question can produce a pointer through the model's own attention.

自然语言问题可以通过模型自身的注意力机制产生指针。

That pointer can become a model-native V-space address.

该指针可以成为模型原生的 V 空间地址。

That address can select an associated numerical memory from an external bank.

该地址可以从外部存储库中选择关联的数值内存。

Once these numerical structures are serialized, whether the storage medium is RAM, SSD, or another persistent storage layer becomes primarily a systems-engineering question.

一旦这些数值结构被序列化,存储介质是 RAM、SSD 还是其他持久化存储层,主要就变成了一个系统工程问题。

The harder problem is not writing an array of numbers to a hard disk.

更难的问题不是将数字数组写入硬盘。

The harder problem is:

更难的问题是:

How does the frozen model later determine which machine-native numerical memory it wants back?

冻结后的模型后来如何确定它想要找回哪种机器原生数值内存?

That is the connection I am testing here.

这就是我在这里测试的连接点。

And this is where I prefer to stop the speculation and let the architecture speak for itself.

这也是我倾向于停止推测,让架构本身说话的地方。

If model-native memory states can exist outside model weights, persist in storage, and later be addressed by signals generated endogenously from natural language inside the transformer, then is long-term machine memory necessarily limited to:

如果模型原生内存状态可以存在于模型权重之外,在存储中持久化,并且随后能够由 Transformer 内部自然语言生成的内生信号寻址,那么长期机器内存是否必然仅限于:

find old human-readable information → put it back into context → make the model read it again?

查找旧的人类可读信息 → 将其放回上下文 → 让模型再次读取它?

更进一步:量化金融体系

看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力

进入量化体系 →

相似阅读

关联信息,但可能不是同一事件