跳到主内容
@wquguru
精选88r/LLMDevs(Reddit)模型发布/更新

AkbasCore MAM:无需重读文本,通过记忆插件实现冻结模型增量记忆

AkbasCore MAM: The end of the "one big context window" era? A frozen LLM that remembers by plugging in memory cartridges instead of re-reading text (72/72, open-source demo)

原文
发到 X
推荐理由

这篇技术解读非常硬核,展示了绕过Transformer传统注意力机制、通过注入数值卡带实现模型增量记忆的全新架构思路,对长上下文和Agent记忆有重要启发,值得研究底层原理的同学细读。

Until now, there were basically two ways to teach a large language model something new: retrain it (fine-tuning) or search a database and paste the text in front of it every single time (RAG).

迄今为止,教大语言模型新东西的方法基本上有两种:重新训练它(微调),或者每次都在提示词前搜索数据库并粘贴文本(RAG)。

AKBASCORE MAM is a third way. It gives a completely frozen Mistral-7B a persistent, incremental memory by appending numerical cartridges directly into the model's internal layers. The source text is not in the prompt. There is no training, no weight is touched, and there is no retrieval or router. The model is never asked to read the information again.

AKBASCORE MAM 是第三种方式。它通过直接将数值型“卡带”附加到模型的内部层中,为完全冻结的 Mistral-7B 赋予持久且可增量更新的记忆。源文本不在提示词中。没有训练,不触碰任何权重,也没有检索或路由机制。模型从未被要求再次阅读这些信息。

In the sealed benchmark, the append-only memory answers 72 out of 72 questions without seeing the source text. That is exactly what the same model scores when it reads the full text. Naively stacked memories score 13/72.

在封闭基准测试中,这种仅追加的记忆在不查看源文本的情况下回答了全部 72 个问题。这与同一模型阅读完整文本时获得的分数完全相同。简单堆叠的记忆只能得 13/72。

Important: this is not a finished product. It is the first public proof that this can be done. The full-capacity, scaled version is what I am working on now. I'm sharing the proof today because the mechanism works, it is open, and anyone can run it.

重要提示:这并非成品。这是证明该方案可行的首个公开案例。全功能、可扩展的版本是我目前正在开发的内容。我今天分享这一成果,是因为该机制有效、开源,且任何人都可以运行它。

The hidden assumption inside every Transformer

Transformer 内部的隐藏假设

The original Transformer ("Attention Is All You Need", 2017) was designed around one silent assumption: everything is written and seen at the same time, inside one closed room.

原始的 Transformer(《Attention Is All You Need》,2017)围绕一个沉默的假设设计:所有内容在同一时间、在一个封闭的空间内被写入和看到。

That made sense back then. The goal was machine translation, or processing one paragraph from start to finish in a single pass. Memory was never designed as separate modules, independent cartridges or files that can be added at different times. All the weight was put on one giant context window.

这在当时是合理的。目标是机器翻译,或者一次性从头到尾处理一个段落。记忆从未被设计为独立的模块、独立的卡带或可以在不同时间添加的文件。所有的权重都押注在一个巨大的上下文窗口上。

The result is the problem we all live with today. A model cannot connect pieces of knowledge that arrive separately unless you put all the raw text back into the window and let it re-read everything together.

结果就是今天我们共同面临的问题。除非你将所有原始文本放回窗口并让模型一起重新阅读,否则模型无法连接单独到达的知识片段。

AKBASCORE MAM questions that assumption at the architecture level. The question I asked was: what if memory were not text in a window, but numerical cartridges you can plug in, one after another, and the model connects them by itself?

AKBASCORE MAM 从架构层面质疑了这一假设。我提出的问题是:如果记忆不是窗口中的文本,而是你可以逐个插入的数值型卡带,并且模型能够自行将它们连接起来,会怎样?

What is a Cognitive Cartridge, and how is it made?

什么是认知卡带,它是如何制作的?

A cognitive cartridge is a fact, written once into the model's own numerical language and stored as tensors.

认知卡带是一个事实,以模型自身的数值语言编写一次,并以张量形式存储。

How it's produced:

生产流程:

The fact (e.g. a short sentence) goes through the frozen model once, using a fixed template.

该事实(例如一个短句)通过冻结的模型一次,使用固定模板进行处理。

We stop and save two things:

我们停止并保存两样东西:

the hidden state at the output of layer 6 (we call it H6), and

第 6 层输出的隐藏状态(我们称之为 H6),以及

the attention keys and values of layers 0–6.

第 0–6 层的注意力键(keys)和值(values)。

That's the cartridge. The text can now be thrown away; the cartridge holds the knowledge.

这就是卡带。现在可以将文本丢弃;卡带承载着知识。

Each cartridge is written independently: in its own pass, at its own time, without knowing which other cartridges will exist.

每个缓存片(cartridge)独立写入:在各自的遍次中、各自的时间点完成,且不知道其他缓存片是否存在。

It is also compact: about 36 KiB per token (H6 8 KiB + layers 0–6 KV 28 KiB, BF16). A full KV cache needs about 128 KiB per token.

它也很紧凑:每 token 约 36 KiB(H6 占 8 KiB + 第 0–6 层 KV 占 28 KiB,BF16 精度)。完整的 KV 缓存每 token 需要约 128 KiB。

The belief this breaks

这一发现打破了以下信念

The common wisdom says: if you encode facts separately and just glue their caches together, the model can't connect them. Facts must "see each other" during encoding, so you have to re-read the text together.

普遍观点认为:如果你分别编码事实并将它们的缓存简单拼接,模型无法将它们关联起来。事实必须在编码过程中“互相看见”,因此你必须重新一起读取文本。

And that is true if you glue them naively. Five independent cartridges stacked side by side give only 13/72 correct answers. The facts stay strangers to each other.

如果你粗暴地拼接它们,确实如此。五个独立缓存片并排堆叠仅能给出 13/72 的正确答案。这些事实彼此仍是陌生的。

The discovery behind MAM is where that connection actually happens inside the network. Facts start binding to each other in the middle layers (roughly 7–14), not in the lowest layers. Layers 0–6 can stay completely independent per cartridge without losing anything. Only the layers above need to let cartridges "meet."

MAM 背后的发现在于这种连接实际上发生在网络内部。事实开始在中层(大约第 7–14 层)相互绑定,而不是在最底层。第 0–6 层可以保持每个缓存片完全独立而不会损失任何信息。只有上层需要让缓存片“相遇”。

So we don't need the text back. We only need to let the upper part of the model consolidate the cartridges.

所以我们不需要重新读取文本。我们只需要让模型的上半部分对缓存片进行整合。

How we bypass the Transformer's front door (DC6)

如何绕过 Transformer 的前门(DC6)

Normally a Transformer has one way in: text → tokens → embeddings → layer 0 → … → layer 31.

通常 Transformer 只有一条入口路径:文本 → token → 嵌入向量 → 第 0 层 → … → 第 31 层。

MAM changes this flow. We call the method DC6 (consolidation at cut layer 6):

MAM 改变了这一流程。我们将该方法称为 DC6(在第 6 层切割处进行整合):

Layers 0–6 are not computed again. Their keys/values come straight from the cartridges.

第 0–6 层不再重新计算。它们的键/值直接来自缓存片。

The model receives zero input embeddings. There is no text at all at the entrance.

模型接收零输入嵌入向量。入口处完全没有文本。

At the output of layer 6, a hook replaces the internal state with the cartridge's stored H6.

在第 6 层的输出端,一个钩子(hook)将内部状态替换为缓存片中存储的 H6。

Layers 7–31 then run normally on top of that state. This is where the cartridges are connected to each other and to what is already in memory.

随后第 7–31 层基于该状态正常运行。正是在这里,缓存片彼此之间以及与内存中已有的内容建立了连接。

In plain words: we skip the model's "reading" stage and inject the knowledge directly into the point where it starts thinking.

用通俗的话说:我们跳过了模型的“阅读”阶段,直接将知识注入到它开始思考的位置。

What happens mechanically when a cartridge is loaded back in

当缓存片被重新加载时机械层面发生了什么

Adding a new cartridge to an existing memory works like this:

向现有内存中添加新缓存片的工作方式如下:

New positions: the new cartridge is placed at the end of the memory, at the next free positions.

新位置:新缓存片被放置在内存末尾的下一个可用位置。

Position correction (RoPE re-phasing): the cartridge's layer 0–6 keys were written at their original positions. Mistral encodes position as a rotation (RoPE), so we rotate the stored keys to their new place: K_new = K·cos(Δθ) + rotate_half(K)·sin(Δθ) Because RoPE is a pure rotation, R(p+Δ) = R(Δ)·R(p). Moving a cartridge is an exact rotation, not an approximation.

位置校正(RoPE 重相位):缓存片的第 0–6 层键是在其原始位置写入的。Mistral 将位置编码为旋转(RoPE),因此我们将存储的键旋转到新位置:K_new = K·cos(Δθ) + rotate_half(K)·sin(Δθ)。由于 RoPE 是纯旋转操作,R(p+Δ) = R(Δ)·R(p)。移动缓存片是一个精确的旋转,而非近似。

Consolidation: layers 7–31 compute the new cartridge while attending to everything already in memory (causal attention). This is where the new fact gets connected to the old ones.

整合:第7至31层在计算新墨盒的同时,关注内存中已有的所有内容(因果注意力)。正是在这里,新事实与旧事实建立连接。

Append-only: the old memory is not recomputed. Earlier rows are checked bit by bit, and they are identical before and after.

只追加:旧内存不会被重新计算。逐位检查前面的行,它们在追加前后完全相同。

Ask: the question is run against this numerical memory only. The source text is never in the input.

提问:问题仅针对此数值内存运行。源文本永远不会出现在输入中。

The memory grows like a stack of plates: you add on top and never rebuild what's underneath.

内存像一叠盘子一样增长:你在顶部添加,而从不重建下面的部分。

Results (TEST560, sealed)

结果(TEST560,已密封)

Panel: 24 synthetic worlds × 4 relation types (current, former, near, role). Each question has a memory of 5 cartridges: 1 target fact + 4 distractor facts from other worlds. The target is tested at the first, middle and last position. That gives 72 cases.

面板:24个合成世界 × 4种关系类型(当前、前任、邻近、角色)。每个问题拥有包含5个墨盒的内存:1个目标事实 + 来自其他世界的4个干扰事实。目标事实分别在第一个、中间和最后一个位置进行测试。这构成了72个案例。

Condition | Source text in input? | First | Middle | Last | Total MAM, append-only (INCR_DC6) | No | 24 | 24 | 24 | 72/72 MAM, batch (BATCH_DC6) | No | 24 | 24 | 24 | 72/72 Model reads full text (JOINT) | Yes | 24 | 24 | 24 | 72/72 Naively stacked cartridges (INDEP) | No | 4 | 3 | 6 | 13/72

条件 | 输入中包含源文本? | 第一个 | 中间 | 最后一个 | MAM总计,只追加 (INCR_DC6) | 否 | 24 | 24 | 24 | 72/72 MAM,批处理 (BATCH_DC6) | 否 | 24 | 24 | 24 | 72/72 模型读取完整文本 (JOINT) | 是 | 24 | 24 | 24 | 72/72 简单堆叠的墨盒 (INDEP) | 否 | 4 | 3 | 6 | 13/72

+59 correct answers over naive stacking, 0 lost. The answers are identical to the batch version in 72/72 cases and to the full-text model in 71/72 cases.

相比简单堆叠多答对59题,无丢失。在72/72个案例中,答案与批处理版本完全一致;在71/72个案例中,与完整文本模型一致。

What we proved, and how it is verified

我们证明了什么,以及如何验证

This isn't "trust me". Every claim is checked by the code itself:

这不是“相信我”。每一项声明都由代码本身进行检查:

Weights never change. Every model parameter is hashed (SHA-256) at startup, before the run and after it. All three hashes are identical.

权重永不改变。每次启动时、运行前和运行后,都会对所有模型参数进行哈希计算(SHA-256)。所有三个哈希值均相同。

The old memory is never rebuilt. After each append, earlier memory rows are compared with torch.equal. They are bit-exact identical.

旧内存永远不会被重建。每次追加后,使用 torch.equal 比较之前的内存行。它们在比特级别上完全相同。

The source text is never shown. A recorder logs every forward pass. During the question, the input is only the question tokens, and the memory is pure numbers.

源文本从未显示。记录器会记录每一次前向传播。在提问期间,输入仅包含问题标记,而内存仅为纯数字。

Moving memory is exact math. Position correction is an exact rotation identity.

移动内存是精确的数学运算。位置校正是一个精确的旋转恒等式。

Accuracy equals full reading. Without the source, the append-only memory matches the model reading the full text (72/72 = 72/72).

准确性等同于全文阅读。在没有源文本的情况下,只追加内存与读取完整文本的模型表现一致(72/72 = 72/72)。

Everything is sealed with SHA-256 (engine, panel, results), and the release has a DOI.

所有内容均通过 SHA-256 密封(引擎、面板、结果),且发布版本附有 DOI。

What you will see when you run the demo

运行演示时将看到的内容

Open the code in Google Colab with an A100, then run the 3 cells in order (or the single full file). You get a web interface where you:

在 Google Colab 中使用 A100 打开代码,然后按顺序运行3个单元格(或单个完整文件)。你将获得一个 Web 界面,你可以在其中:

Pick a case and choose where the target fact sits: FIRST / MIDDLE / LAST.

选择一个案例,并选择目标事实的位置:FIRST / MIDDLE / LAST。

Watch 5 cartridges being written one by one and appended into memory. In the live run on 8 October 2026 the memory grew 41 → 55 → 68 → 80 → 93 tokens. Every step was verified bit-exact.

观察5个墨盒逐个写入并追加到内存中。在2026年10月8日的实时运行中,内存从41增长到55 → 68 → 80 → 93个标记。每一步都经过比特级验证。

See the question asked with no source text (23 tokens of pure question). In the live run the model answered "Melket", which is correct.

查看无源文本的提问(23 个标记的纯问题)。在实时运行中,模型回答为“Melket”,这是正确的。

Get a report showing that the weight hash is identical before and after. It also shows the memory map, the answer and the verification checks.

获取一份报告,显示权重哈希在前后保持一致。该报告还展示内存映射、答案以及验证检查。

Optionally replay all 72 cases live and compare them with the sealed reference.

可选地,重新实时回放全部 72 个案例,并将其与密封参考结果进行比较。

Download the full sealed package: JSON logs, figures and SHA manifest.

下载完整的密封包:JSON 日志、图表和 SHA 清单。

更进一步:量化金融体系

看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力

进入量化体系 →

相似阅读

关联信息,但可能不是同一事件