跳到主内容
@wquguru
精选88PostHog 博客(RSS)产品与增长

Context Engineering:PostHog Wizard

WTF is context engineering? (with real examples)

原文
发到 X
推荐理由

给出了具体的 Context Engineering 架构方案(context-mill 流水线),包含数据来源、打包格式和分发协议(MCP),读者可直接参考其自动化构建与交付逻辑来优化自己的 AI Agent 产品。

Say you're deploying an AI assistant that processes online order returns. For it to work, it would need access to your store's purchase policy, item prices, and order history.

假设你正在部署一个处理在线订单退货的 AI 助手。为了使其正常工作,它需要访问你商店的购买政策、商品价格和订单历史。

The problem is that LLMs can only hold a finite amount of information, so you can't give it the whole database of purchases made by every user since you launched. Instead, you might provide details on one specific order, which you can only get after the customer provides their order number.

问题在于,大语言模型(LLM)只能容纳有限量的信息,因此你不能把自你推出以来每位用户的所有购买数据库都交给它。相反,你可能只提供特定订单的详细信息,而这只有在客户提供订单号后才能获取。

Building the systems to deliver that information is what context engineering is about. You're designing what an agent should read – instructions, tools, data, examples, memory – and when. Andrej Karpathy, one of the co-founders of OpenAI, described it well:

构建提供这些信息的系统就是上下文工程的核心。你在设计智能体应该读取什么——指令、工具、数据、示例、记忆——以及何时读取。OpenAI 的联合创始人之一 Andrej Karpathy 对此描述得很到位:

"Context engineering is the delicate art and science of filling the context window with just the right information for the next step."

"上下文工程是一门精细的艺术和科学,旨在为下一步操作填充恰到好处的信息到上下文窗口中。"

You might think that as models and agents improve, context engineering's relevance would decrease over time, but it's actually the opposite.

你可能会认为,随着模型和智能体的进步,上下文工程的相关性会随着时间推移而降低,但事实恰恰相反。

Better models just give agents the potential to tackle a challenge, but context engineering is what gets them to actually succeed at scale.

更好的模型只是赋予智能体应对挑战的潜力,而上下文工程才是让它们真正大规模成功的关键。

What does context engineering actually look like?

上下文工程实际上是什么样子的?

To understand what context engineering is more concretely, take a look at the system behind the PostHog Wizard, our custom onboarding agent.

为了更具体地理解什么是上下文工程,让我们看看 PostHog Wizard(我们的自定义入职引导智能体)背后的系统。

The PostHog Wizard is an AI assistant that automatically installs PostHog in an existing codebase. This used to take developers at least 2 hours of manually reading docs, pasting code snippets, and testing integrations. Now, users can accomplish this in 8 minutes – all with a simple npx @posthog/wizard@latest command. And it's been wildly successful since launch, 5x'ing our paid conversion rates and 2x'ing activation speeds.

PostHog Wizard 是一个 AI 助手,能够自动在现有代码库中安装 PostHog。过去,开发人员至少需要花费 2 小时手动阅读文档、粘贴代码片段并测试集成。现在,用户可以通过一个简单的 `npx @posthog/wizard@latest` 命令在 8 分钟内完成此操作。自发布以来,它取得了巨大成功,将付费转化率提高了 5 倍,将激活速度提高了 2 倍。

Building it, however, was not simple.

然而,构建它并不简单。

When we started on it back in 2025, language models were already good enough at generalized coding tasks. The challenge was in making them good at PostHog-specific code. It's like how the world's greatest rocket scientist won't know how to install a toilet or cook a mean lasagna if they've never tried or been taught how. No matter how super-intelligent models are, general knowledge can only take you so far in highly specific scenarios.

当我们于 2025 年开始着手该项目时,语言模型在通用编码任务方面已经足够出色。挑战在于让它们擅长处理 PostHog 特定的代码。这就像世界上最伟大的火箭科学家如果从未尝试或学习过,也不会知道如何安装马桶或制作美味的千层面一样。无论模型多么超级智能,通用知识在高度特定的场景中所能发挥的作用是有限的。

Then, when you consider that PostHog has 20+ products, 17+ SDKs, and 25+ frameworks that are constantly getting updated, you'll start to see the moving combinatorial explosion we were dealing with. On top of that, LLMs are diabolically good at sounding correct, so there were many wrong solutions that slipped and passed our review in the early stages.

然后,当你考虑到 PostHog 拥有 20 多种产品、17 多种 SDK 和 25 多种框架,并且它们都在不断更新时,你就会开始意识到我们当时所面临的动态组合爆炸问题。此外,大语言模型(LLM)在听起来正确方面表现得极其出色,因此在早期阶段,许多错误的解决方案溜走并通过了我们的审查。

(We won't go into the details of how we solved those problems in this blog. If you're interested in that story and want to see the lessons we learned while building the PostHog Wizard, check out our newsletter.)

(我们将不会在本博客中深入探讨我们如何解决这些问题的细节。如果你对那个故事感兴趣,并希望了解我们在构建 PostHog Wizard 过程中吸取的教训,请查看我们的新闻通讯。)

It took several months of hacking and experimenting until we finally landed on the design that's still in use today: the context-mill.

经过几个月的黑客式开发和实验,我们最终确定了至今仍在使用的设计:context-mill(上下文工厂)。

It's an automated pipeline that independently gathers PostHog installation context and serves it to downstream wizard client agents. It works by first grabbing the latest info about PostHog from three main sources:

它是一个自动化管道,独立收集 PostHog 安装上下文并将其提供给下游的向导客户端代理。它的工作原理是首先从三个主要来源获取有关 PostHog 的最新信息:

  • Docs. For example, the Next.js integration skill pulls the Next.js library docs, plus the shared .md guide about how to identify users for analytics.
  • Example apps. These are 40+ few-shot example apps with PostHog installed across frameworks like Django, Swift, Rust and more. The wizard is instructed to match these as closely as possible.
  • Hand-written instructions. This contains technical "gotchas" to avoid, step-by-step installation guides, and other prompts for runtime.
  • 文档。例如,Next.js 集成技能会抓取 Next.js 库的文档,以及关于如何为分析识别用户的共享 .md 指南。
  • 示例应用。这些是安装了 PostHog 的 40 多个少样本示例应用,涵盖 Django、Swift、Rust 等框架。向导被指示尽可能匹配这些应用。
  • 手写说明。这包含了需要避免的技术“陷阱”、逐步的安装指南以及其他运行时提示。

Once it's gathered all of the source information, the context-mill assembles the material into a collection of zip files and cuts a versioned GitHub release. It then packages the context as downloadable skills and registers them with MCP servers as a publicly accessible resource. That way, downstream agents can install PostHog with the latest version from anywhere, anytime.

一旦收集了所有源信息,context-mill 就会将材料组装成一系列 zip 文件,并发布一个带版本号的 GitHub Release。然后,它将上下文打包为可下载的技能,并在 MCP 服务器上注册为公开可访问的资源。这样,下游代理就可以随时随地从任何地方安装最新版本的 PostHog。

This is just one example of context engineering in practice. Anyone in the field will tell you that no two context engines will ever look the same since each one is typically built to solve a specific, unique, contextual problem.

这只是上下文工程在实际应用中的一个例子。该领域的任何人都会告诉你,没有两个上下文引擎看起来会是相同的,因为每个引擎通常都是为了解决特定的、独特的、情境化的问题而构建的。

Why is everyone talking about context engineering now?

为什么现在大家都在谈论上下文工程?

Context engineering first started trending around 2024, but it exploded in popularity in 2025 for a few reasons.

上下文工程大约在 2024 年开始流行,但在 2025 年因以下几个原因爆红。

1. Prompt engineering turned into context engineering

1. 提示工程转变为上下文工程

In 2025, most of us were still using ChatGPT in the browser as a chatbot and marveling at Cursor's auto-complete.

在 2025 年,大多数人仍然在浏览器中使用 ChatGPT 作为聊天机器人,并对 Cursor 的自动补全功能感到惊叹。

At the time, the only real surface in which engineers could improve their AI systems was just in their prompts. This was the era when prompt engineering was hot; people would work on mastering the craft of writing the best combination of text for one-shot classification or generation tasks. Companies were hiring "prompt engineers", communities like r/PromptEngineering thrived, and Google published a 68-page whitepaper on advanced prompting techniques.

当时,工程师唯一能真正改进其 AI 系统的途径仅限于提示词(prompts)。那是提示工程(prompt engineering)风靡的时代;人们致力于掌握为一-shot 分类或生成任务编写最佳文本组合的技巧。公司开始招聘“提示工程师”,r/PromptEngineering 等社区蓬勃发展,Google 还发布了一份关于高级提示技术的 68 页白皮书。

And then agents happened. Agents changed the meta from chatbots that primarily work in single-turn interactions to systems that can operate over multiple turns of inference:

随后,智能体(agents)出现了。智能体改变了范式,从主要在单轮交互中工作的聊天机器人,转变为能够在多轮推理中运行的系统:

This opened up way more possibilities for AI engineering.1 Developers could now start designing strategies for the entire context state including system instructions, tools, message history, and more.

这为 AI 工程打开了更多的可能性。1 开发者现在可以开始设计涵盖整个上下文状态的策略,包括系统指令、工具、消息历史等。

2. The models just kept getting better

2. 模型持续变得更好

With every major model release, there's a matching rise in interest in context engineering since no two generations are the same. Boris Cherny, the creator of Claude Code, described:

随着每次重大模型版本的发布,由于每一代模型都不尽相同,对上下文工程(context engineering)的兴趣也随之上升。Claude Code 的创作者 Boris Cherny 描述道:

"Every model generation behaves differently. It has a slightly different personality, and you have to take the time to get to know it and then adjust."

“每一代模型的行为方式都不同。它们具有略微不同的个性,你需要花时间了解它们,然后进行调整。”

One of the most dramatic advancements in model capabilities of 2026 was Anthropic's Fable 5. As part of the release, Anthropic reported that they removed 80% of Claude Code's system prompt because the new models needed fewer rules and repetition. Instead, they said that allowing more space for LLM judgment was the better approach. The new rules for context engineering now favor leaner prompts and thinner context.

2026 年模型能力最显著的进步之一是 Anthropic 发布的 Fable 5。作为发布的一部分,Anthropic 报告称他们移除了 Claude Code 系统提示词的 80%,因为新模型需要的规则和重复内容更少。相反,他们认为允许 LLM 拥有更多判断空间是更好的方法。如今上下文工程的新规则更倾向于精简的提示词和更薄的上下文。

For example, in older versions of the PostHog Wizard, the installation would often land on the wrong project in monorepos because our scripts pointed to root by default. Once models got smart enough to reliably infer repo structure on their own, we updated it so headless runs take advantage of those capabilities.

例如,在旧版本的 PostHog Wizard 中,安装过程经常会在 monorepos(单体仓库)中定位到错误的项目,因为我们的脚本默认指向根目录。一旦模型足够聪明,能够可靠地自行推断仓库结构,我们就进行了更新,使无头运行(headless runs)能够利用这些能力。

This is an ongoing cycle in context engineering. As models improve, your context needs to be reshaped and often cut down in response. Otherwise, you risk "hobbling" the model – A.K.A., overconstraining it by giving too much information, rather than letting it use built-in reasoning.

这是上下文工程中一个持续的循环。随着模型的改进,你的上下文需要随之重塑,并且通常需要缩减。否则,你会有“束缚”模型的风险——即通过提供过多信息来过度约束它,而不是让它利用内置的推理能力。

3. Agentic systems keep getting more complex

3. 智能体系统持续变得更加复杂

Alongside model improvements, harnesses, infrastructure, and tasks keep getting more complex. In turn, agents have more complicated needs for context, which is why interest in context engineering has grown.

随着模型的改进, harnesses(框架/工具链)、基础设施和任务也持续变得更加复杂。反过来,智能体对上下文的需求也变得更加复杂,这也是为什么人们对上下文工程的兴趣日益增长的原因。

Consider that in early 2025, most people were working with just a handful of agents at a time, usually on a single task. Now, we have things like:

考虑到在2025年初,大多数人一次只使用少数几个智能体,通常仅处理单一任务。而现在,我们有了以下情况:

  • Multi-agent orchestration: Agents leading other agents to break up, dispatch, and execute work in parallel
  • Long-horizon tasks: Agents working autonomously for hours on multi-step projects with complete verification
  • Scheduled agents: Agents that run on a schedule or trigger on events, without a human starting the session
  • Multiplayer AI: Systems where people and agents work in the same shared workspace and context, like in Slack
  • Sandboxing: Isolated environments that limit what agents can touch, which lets them run often with even higher autonomy
  • 多智能体编排:由智能体领导其他智能体,以并行方式分解、分发和执行工作
  • 长周期任务:智能体在具有完整验证的多步骤项目中自主工作数小时
  • 定时智能体:按预定计划运行或在事件触发时运行的智能体,无需人类启动会话
  • 多人协作AI:人与智能体在同一共享工作空间和上下文中协同工作的系统,例如在Slack中
  • 沙箱隔离:限制智能体可访问范围的隔离环境,使其能够在更高程度的自主性下运行

更进一步:量化金融体系

看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力

进入量化体系 →

相似阅读

关联信息,但可能不是同一事件