跳到主内容
@wquguru
精选88r/LLMDevs(Reddit)技巧与观点

AKBASCORE NIRVANA:在Qwen与Mistral上实现可拆卸数值

AKBASCORE NIRVANA — I Built Removable Numerical Memory Cartridges for Two Different Frozen 7B LLMs. Qwen and Mistral Both Work. Now I’m Scaling the Memory Bank.

原文
发到 X
推荐理由

提供了一套完整的非微调、非RAG的数值记忆工程方案,含具体参数与跨模型验证,对Agent长期记忆构建极具参考价值。

Zenodo permanent records:

Zenodo 永久记录:

Qwen2.5-7B-Instruct:

Qwen2.5-7B-Instruct:

https://doi.org/10.5281/zenodo.23127434

Mistral-7B-Instruct-v0.3:

Mistral-7B-Instruct-v0.3:

https://doi.org/10.5281/zenodo.23143605

I want to start with the simplest possible explanation of what I have been building.

我想从对我一直在构建的内容的最简单解释开始。

Imagine taking a piece of information, letting a language model process it once, and then throwing the original text away.

想象一下,取一段信息,让语言模型处理一次,然后把原始文本扔掉。

No sentence stored in a database.

没有句子存储在数据库中。

No paragraph hidden somewhere.

没有段落隐藏在某个地方。

No readable summary.

没有可读的摘要。

No RAG system fetching the original document.

没有 RAG 系统去获取原始文档。

No fine-tuning.

没有微调。

No LoRA.

没有 LoRA。

No weight update.

没有权重更新。

What remains is numerical transformer memory derived from the model's own internal computation. I package that numerical memory into what I call a Cognitive Cartridge. Later, I can install that cartridge back into the frozen model and ask questions about the information that produced it — without putting the original source text back into the readout prompt.

剩下的是源自模型自身内部计算的数值型 Transformer 记忆。我将这种数值记忆打包成我所谓的“认知卡带”(Cognitive Cartridge)。稍后,我可以将该卡带重新安装到冻结的模型中,并询问关于生成该卡带的信息的问题——而无需将原始源文本放回读取提示中。

That was the first result. The new result is more important:

那是第一个结果。新的结果更为重要:

I have now reproduced the Cognitive Cartridge architecture on two different 7B transformer model families.

我现在已经在两个不同的 7B Transformer 模型系列上复现了认知卡带架构。

Qwen2.5-7B-Instruct.

Qwen2.5-7B-Instruct。

And now Mistral-7B-Instruct-v0.3.

以及现在的 Mistral-7B-Instruct-v0.3。

The implementations are not numerically identical. The architectures are different, the layer counts are different, the KV structures are different, and the working cartridge configurations are different. But the central mechanism survived the move.

这些实现并非在数值上完全相同。架构不同,层数不同,KV 结构不同,工作卡带配置也不同。但核心机制在迁移过程中幸存了下来。

That is the reason I am publishing this second record.

这就是我发布第二条记录的原因。

The question is no longer only:

问题不再仅仅是:

“Can I make this happen once on Qwen?”

“我能在 Qwen 上让它发生一次吗?”

Now there is a second implementation on Mistral. And the Mistral result is the cleanest version so far.

现在在 Mistral 上有第二个实现。而且 Mistral 的结果是迄今为止最干净的一个版本。

What is actually inside a Cognitive Cartridge?

认知卡带(Cognitive Cartridge)内部实际上是什么?

This is probably the most important thing to understand.

这可能是最需要理解的一点。

Suppose the source record says:

假设源记录如下所示:

Object: amber sextant

对象:琥珀色六分仪

Container: RQ-415

容器:RQ-415

Location: elm lodge

位置:榆树小屋

That text exists during the forging stage. The frozen transformer processes it. NIRVANA takes source-derived internal transformer K/V states and represents their content numerically using a fixed, source-independent codebook.

该文本存在于锻造阶段。冻结的 Transformer 对其进行处理。NIRVANA 提取源自源数据的内部 Transformer K/V 状态,并使用固定的、与源无关的词表将其内容表示为数值形式。

In the released Mistral implementation, both K and V are compressed to D120 while a source-specific OWN component is preserved. After forging, the source record is not supplied to the readout prompt.

在已发布的 Mistral 实现中,K 和 V 均被压缩至 D120,同时保留了特定于源的 OWN 组件。锻造完成后,源记录不再提供给读取提示(readout prompt)。

So the conceptual transformation is:

因此,概念上的转换过程为:

human language

人类语言

→ frozen transformer computation

→ 冻结的 Transformer 计算

→ internal K/V states

→ 内部 K/V 状态

→ compressed numerical Cognitive Cartridge

→ 压缩后的数值型认知卡带

Then later:

随后:

numerical Cognitive Cartridge

数值型认知卡带

→ reconstructed transformer K/V memory

→ 重构的 Transformer K/V 记忆

→ frozen transformer

→ 冻结的 Transformer

→ language

→ 语言

Or, in the shortest form:

或者,用最简短的形式表示:

language → internal numerical memory → language

语言 → 内部数值记忆 → 语言

The middle is no longer human-readable source text. That distinction matters.

中间部分不再是人类可读的源文本。这一区别至关重要。

A cartridge is not a text file with a different name.

卡带并非只是换个名称的文本文件。

It is not a vector database containing the original sentence.

它也不是包含原始句子的向量数据库。

It is not a prompt template.

它也不是提示模板。

It is not conventional RAG returning the source paragraph.

它不是返回源段落的传统 RAG。

It is not a fine-tuned model.

它不是一个经过微调的模型。

The model weights remain frozen. The information is carried by a numerical representation derived from transformer K/V memory.

模型权重保持冻结状态。信息通过从 Transformer K/V 记忆派生的数值表示来承载。

Why call it a cartridge?

为什么称之为“卡带”(cartridge)?

Think less about a document and more about an interchangeable machine-readable memory module. The base model stays where it is. Knowledge packages can be forged separately. Those packages can remain separate. They can be installed and queried without retraining the base model.

不要把它想成文档,而应视为一种可互换的机器可读存储模块。基础模型保持不变。知识包可以单独制作。这些包可以保持独立。它们可以在不重新训练基础模型的情况下被安装和查询。

This becomes much more interesting when there is more than one cartridge.

当存在多个卡带时,情况会变得有趣得多。

The new Mistral experiment uses 16 independently forged cartridges. Each one contains a numerical representation derived from a separate source record. They are not concatenated into one giant text prompt. They remain independent memories.

新的 Mistral 实验使用了 16 个独立制作的卡带。每个卡带包含源自单独源记录的数值表示。它们没有被拼接成一个巨大的文本提示。它们作为独立的记忆保留。

For example, in human-readable form, imagine one cartridge represents:

例如,以人类可读的形式想象,一个卡带代表:

amber sextant → RQ-415 → elm lodge

琥珀色六分仪 → RQ-415 → 榆树小屋

Now ask the cartridge bank:

现在向卡带库提问:

Which container is associated with the amber sextant?

哪个容器与琥珀色六分仪相关联?

The query is evaluated against the independent cartridge memories.

查询是针对独立的卡带记忆进行评估的。

The relevant cartridge returns:

相关的卡带返回:

RQ-415

The unrelated cartridges return:

不相关的卡带返回:

NONE

无

There is no learned router secretly selecting the correct cartridge before this happens.

在此发生之前,并没有一个学习到的路由器在暗中选择正确的卡带。

Then the recovered identifier can be used for a second lookup:

然后,恢复的标识符可用于第二次查找:

Where is RQ-415?

RQ-415 在哪里?

The bank is queried again.

再次查询该库。

The relevant memory returns:

相关的记忆返回:

elm lodge

榆树小屋

The unrelated memories return:

不相关的记忆返回:

NONE

无

So the complete retrieval becomes:

因此,完整的检索变为:

amber sextant

琥珀色六分仪

→ RQ-415

→ elm lodge

→ 榆树小屋

The important point is that the model did not reread the original source record to answer either stage. It operated from reconstructed numerical transformer memory.

关键要点在于,模型在回答这两个阶段时都没有重新读取原始源记录。它基于重构的数值转换器记忆进行操作。

The Mistral result

Mistral 的结果

The final public Mistral run used:

最终的公开 Mistral 运行使用了:

Mistral-7B-Instruct-v0.3

32 transformer layers

32 层 Transformer

hidden size 4096

隐藏层大小 4096

32 attention heads

32 个注意力头

8 KV heads

8 个 KV 头

BF16 / SDPA

K = D120

V = D120

OWN preserved

OWN 保留

16 independent Cognitive Cartridges

16 个独立的认知卡带

frozen model weights

冻结的模型权重

greedy decoding

贪婪解码

There was:

存在以下情况:

no fine-tuning

无微调

no LoRA

无 LoRA

no optimizer

无优化器

no learned router

无学习到的路由器

no model-weight update

无模型权重更新

The recorded final run produced:

记录的最终运行产生了:

Object → container ID: 16/16

对象 → 容器 ID: 16/16

Container ID → location: 16/16

容器 ID → 位置: 16/16

Complete two-stage retrieval: 16/16

完成两阶段检索: 16/16

Missing-object controls: 8/8

缺失对象控制: 8/8

Absent-ID controls: 8/8

缺失 ID 控制: 8/8

NOMEM controls: 8/8

无内存 (NOMEM) 控制: 8/8

But there is another result I think is just as important.

但还有另一个我认为同样重要的结果。

For every target query, there are 15 unrelated cartridges.

对于每个目标查询,都有 15 个不相关的弹药筒。

16 queries × 15 unrelated cartridges = 240 unrelated cartridge reads.

16 次查询 × 15 个不相关弹药筒 = 240 次不相关弹药筒读取。

Stage 1 unrelated-cartridge rejection:

第一阶段不相关弹药筒拒绝:

240/240 NONE

240/240 无

Stage 2 unrelated-cartridge rejection:

第二阶段不相关弹药筒拒绝:

240/240 NONE

240/240 无

That means the result is not simply:

这意味着结果不仅仅是:

“The correct memory can say something.”

“正确的记忆可以陈述某些内容。”

The system also demonstrated, in this controlled panel:

该系统还在受控面板中展示了:

“The memories that do not contain the requested relation can refuse to claim that they do.”

“不包含所请求关系的记忆能够拒绝声称它们包含该关系。”

For a modular memory system, I think this distinction is fundamental. A memory bank that can retrieve information but cannot distinguish relevance from irrelevance becomes increasingly dangerous as it grows.

对于模块化记忆系统而言,我认为这一区分是根本性的。一个能够检索信息却无法区分相关性与非相关性的存储库,随着规模扩大将变得日益危险。

The interesting problem is not only remembering. It is also knowing which memory does not answer the question.

有趣的问题不仅在于记忆,还在于知道哪些记忆无法回答该问题。

The no-memory control matters for the same reason. When the cartridge memory was removed, the target relations were not recovered.

无内存 (NOMEM) 控制的重要性源于相同原因。当移除弹药筒记忆时,目标关系未被恢复。

NOMEM:

NOMEM:

8/8 controls passed.

8/8 控制通过。

The source-removal audit passed.

源数据移除审计已通过。

The frozen-weight sentinel passed.

冻结权重哨兵检测已通过。

The model remained frozen.

模型保持冻结状态。

This is why I consider the Mistral result an important step beyond the first demonstration.

这就是为什么我认为 Mistral 的结果是继首次演示之后的一个重要进步。

The first Qwen release established the architecture publicly. The Mistral release gives us something else:

首个 Qwen 版本公开了架构。Mistral 的发布则为我们提供了其他内容:

cross-model evidence.

跨模型证据。

Qwen and Mistral are not the same transformer.

Qwen 和 Mistral 并非相同的 Transformer。

Qwen2.5-7B-Instruct uses 28 transformer layers.

Qwen2.5-7B-Instruct 使用了 28 层 Transformer。

Mistral-7B-Instruct-v0.3 uses 32.

Mistral-7B-Instruct-v0.3 使用了 32 层。

Their internal configurations differ. The working compression configurations also differ.

它们的内部配置不同。工作压缩配置也不同。

The Qwen public implementation used:

Qwen 的公开实现使用了:

K120 / V128 / OWN

The Mistral implementation uses:

Mistral 的实现使用了:

K120 / V120 / OWN

I did not simply copy a cache from one model into another. Each model builds and reconstructs its own source-derived internal memory.

我并没有简单地将一个模型的缓存复制到另一个模型中。每个模型都构建并重构其源自源数据的内部记忆。

What transferred was the architecture:

转移的是架构:

source

源数据

→ internal transformer memory

→ 内部 Transformer 记忆

→ numerical compression

→ 数值压缩

→ independent cartridge

→ 独立卡带(cartridge)

→ source removed

→ 源数据已移除

→ reconstructed K/V

→ 重建 K/V

→ frozen-model readout

→ 冻结模型读取输出

That is the bridge between the two releases.

这就是两个版本之间的桥梁。

I would not call two models proof of universal compatibility with every transformer architecture. That would be scientifically too strong.

我不会将两个模型作为与所有Transformer架构普遍兼容的证明。那在科学上过于强断了。

But it is now evidence that Cognitive Cartridge is not merely one accidental Qwen-specific behavior. The same broader architecture has been implemented and publicly reproduced on a second transformer family.

但如今有证据表明,Cognitive Cartridge 不仅仅是一种偶然的、针对 Qwen 特有的行为。同一更广泛的架构已在另一个 Transformer 家族中实现并公开复现。

That changes the research question.

这改变了研究问题。

The first question was:

第一个问题是:

Can this mechanism exist at all?

这种机制是否真的可能存在?

The next question became:

接下来的问题变成了:

Can multiple independent memories coexist?

多个独立记忆能否共存?

Then:

然后是:

Can one retrieved result lead to another retrieval?

一次检索结果能否引发另一次检索?

Then:

接着是:

Can unrelated memories reject a query instead of contaminating the answer?

无关的记忆能否拒绝查询,而不是污染答案?

And now:

而现在:

Does the architecture survive a move to another model family?

该架构在迁移到另一个模型家族后是否依然稳健?

We now have experimental answers to each of those questions.

我们现在对这些问题都有了实验性的答案。

So I am moving to the next problem:

因此,我转向下一个问题:

scale.

规模(scale)。

更进一步:量化金融体系

看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力

进入量化体系 →

相似阅读

关联信息,但可能不是同一事件