LLM推理调度循环详解:从请求入队到连续批处理
The LLM scheduling loop, clearly explained:
做LLM部署和推理优化的同学必看,把复杂的调度循环拆解得清清楚楚,直接指导工程实践。
The LLM scheduling loop, clearly explained:
大语言模型调度循环,清晰解析:
(bookmark this)
(收藏此内容)
Whenever a user sends a prompt to an LLM, the resulting inference request does not go straight from the API endpoint to the GPU.
每当用户向大语言模型发送提示时,生成的推理请求并不会直接从 API 端点发送到 GPU。
Instead, it first enters the serving engine, where a scheduler decides when and how its tokens will be processed.
相反,它首先进入服务引擎,在那里由调度器决定何时以及如何对其 token 进行处理。
Here is the complete lifecycle:
以下是完整的生命周期:
1) The request enters the waiting queue
1) 请求进入等待队列
Each request arrives with a prompt, sampling settings, and a maximum output length.
每个请求都附带提示、采样设置和最大输出长度。
Prompt lengths vary, and the scheduler cannot know the final generation length in advance. It therefore manages requests one iteration at a time.
提示长度各不相同,且调度器无法提前知道最终的生成长度。因此,它按每次迭代管理请求。
2) The scheduler builds the next batch
2) 调度器构建下一个批次
Before every model step, the scheduler checks two main constraints:
在每次模型步骤之前,调度器会检查两个主要约束条件:
- How many tokens can be processed in this iteration - How much KV-cache space remains
- 本次迭代可以处理多少 token - KV 缓存空间还剩多少
It then applies a scheduling policy, often prioritizing running decode requests before admitting new work.
然后应用调度策略,通常优先处理正在运行的解码请求,然后再接纳新任务。
3) New requests enter prefill
3) 新请求进入预填充阶段
During prefill, the model processes the prompt tokens and creates the K and V tensors needed by attention.
在预填充阶段,模型处理提示 token 并创建注意力机制所需的 K 和 V 张量。
This stage can process many prompt tokens in parallel. Long prompts may be split into chunks so they do not block active generations for too long.
此阶段可以并行处理大量提示 token。长提示可能会被拆分成块,以免阻塞活跃生成过程过久。
4) Running requests enter decode
4) 运行中的请求进入解码阶段
Once prefill finishes, the request moves to decode.
预填充完成后,请求转入解码阶段。
The model now generates the next token using the KV cache created for all previous tokens. Standard autoregressive decoding usually adds one new token per active request during a model step.
模型现在使用为所有先前 token 创建的 KV 缓存来生成下一个 token。标准的自回归解码通常在每个模型步骤中为每个活跃请求添加一个新 token。
5) The GPU executes the selected work
5) GPU 执行选定的工作
Modern serving engines can place prefill and decode work in the same GPU batch.
现代服务引擎可以将预填充和解码工作放在同一个 GPU 批次中。
For example, request A may be processing its prompt while requests B and C generate their next tokens.
例如,请求 A 可能正在处理其提示,而请求 B 和 C 正在生成它们的下一个 token。
The batch is therefore not a fixed group that runs until every request finishes. Its composition can change after every model step.
因此,批次并不是一个固定组直到每个请求完成才运行。其组成可以在每次模型步骤后发生变化。
6) The engine updates request state
6) 引擎更新请求状态
After execution, generated tokens are sampled and appended to their requests.
执行后,生成的 token 被采样并追加到各自的请求中。
A finished request releases its KV-cache blocks and leaves the batch. An unfinished request keeps its state and becomes eligible for the next scheduling iteration.
一个完成的请求会释放其 KV-cache 块并离开批次。一个未完成的请求会保留其状态,并在下一次调度迭代中变得可被调度。
The scheduler can then use the freed token budget and cache space to admit another waiting request.
然后,调度器可以使用释放出来的 token 预算和缓存空间来接纳另一个等待中的请求。
In the visual below, C finishes after iteration 1, so D enters during iteration 2. By iteration 3, the active batch contains only A and D.
在下图中,C 在迭代 1 后完成,因此 D 在迭代 2 期间进入。到迭代 3 时,活跃批次仅包含 A 和 D。
This iteration-level replacement is continuous batching. It keeps the GPU working while requests with different prompt and output lengths progress independently.
这种迭代级别的替换就是连续批处理(continuous batching)。它让 GPU 保持工作状态,同时使具有不同提示词和输出长度的请求独立推进。
If you want to dive deeper, I wrote a detailed article covering the full LLM inference pipeline, including tokenization, prefill, decode, KV caching, and the latency metrics each stage controls.
如果你想深入了解,我写了一篇详细文章,涵盖了完整的 LLM 推理流水线,包括分词、预填充(prefill)、解码(decode)、KV 缓存以及每个阶段控制的延迟指标。
Read it below.
请在下方阅读。
更进一步:量化金融体系
看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力