Claude Code 实战:用 Opus 5.5 生成 4 个破坏性测试子
Claude Code tip: before you ship, have Opus 5.5 send out 𝟰 𝘀𝘂𝗯𝗮𝗴𝗲𝗻𝘁𝘀 whose only…
提供了极具操作性的具体 Prompt 和工作流结构,清晰展示了如何利用多 Agent 协作解决复杂的测试场景,对开发者有直接的参考价值。
Claude Code tip: before you ship, have Opus 5.5 send out 𝟰 𝘀𝘂𝗯𝗮𝗴𝗲𝗻𝘁𝘀 whose only job is to break your app, and split 𝘁𝗵𝗲𝘀𝗲 𝟭𝟲 𝗾𝘂𝗲𝘀𝘁𝗶𝗼𝗻𝘀 between them:
Claude Code 技巧:在发布之前,让 Opus 5.5 发出 4 个子智能体,它们的唯一任务是破坏你的应用,并将这 16 个问题分配给它们:
Security → Do your bundled frontend JS or git history contain private keys like sk_ or service_role? → Skip the UI and call your APIs directly, logged out or with someone else's ID. Does the server stop you? → Change the price, quantity or role in a request and send it again. Does the server accept it? → Hit the same endpoint 100 times in a row. Is there a rate limit per user and IP? → Hide a snippet that pops an alert in Markdown, rich text or an uploaded SVG. Does it run when another account opens it?
安全 → 你的打包前端 JS 或 git 历史中是否包含 sk_ 或服务角色等私钥? → 跳过 UI 直接调用 API,以未登录状态或其他用户 ID 登录。服务器会阻止你吗? → 修改请求中的价格、数量或角色并重新发送。服务器会接受吗? → 连续向同一端点发送 100 次请求。是否有针对每个用户和 IP 的频率限制? → 在 Markdown、富文本或上传的 SVG 中隐藏一个触发 alert 的代码片段。当其他账户打开时它会执行吗?
Data → Load 10,000 records. Does the page still open? → Use a 200-character name with emoji. Does it break the layout? → Sign up with the same email in different capitalization. Did you just get two accounts? → Submit an empty form. Is there now a blank record in your database?
数据 → 加载 10,000 条记录。页面还能正常打开吗? → 使用包含 emoji 的 200 个字符的名称。这会破坏布局吗? → 用不同大小写形式的相同邮箱注册。你是否得到了两个账户? → 提交空表单。数据库中现在是否存在一条空白记录?
Flows → Place the same order twice at once, and replay the same payment webhook twice. Did you get two orders? → Refresh or hit Back halfway through a form. Is what you typed still there? → Save while offline: it says saved, but nothing was stored → Upload a 50MB image. Does everything freeze?
流程 → 同时下两次相同的订单,并重复触发两次相同的支付 webhook。你是否收到了两份订单? → 在填写表单中途刷新或点击返回按钮。你输入的内容还在吗? → 离线保存:显示已保存,但实际未存储任何内容 → 上传一张 50MB 的图片。界面是否会完全卡死?
Environment → On a small phone screen at 200% zoom, can you still tap every button? → Switch to Safari's engine. Does the layout fall apart? → For users in other time zones, are dates, deadlines and check-ins off by a day?
环境 → 在小屏幕手机上以 200% 缩放比例查看时,你仍能点击所有按钮吗? → 切换到 Safari 引擎。布局是否会崩溃? → 对于处于其他时区的用户,日期、截止日期和签到时间是否会出现一天误差?
Set all 4 subagents to high effort. The security one covers payments and login, where one mistake is a disaster, so you can raise it to xhigh or even max. Anthropic's Thariq has said that turning effort up mostly buys more verification and edge-case testing.
将所有 4 个子智能体的努力程度设置为高。其中负责安全的子智能体覆盖支付和登录环节,因为这里的任何失误都可能导致灾难性后果,因此你可以将其提升至 xhigh 甚至 max。Anthropic 的 Thariq 曾表示,提高努力程度主要能换取更多的验证和边缘情况测试。
Send this whole post to Claude Code, and add this line at the end 👇 "Create 4 testing subagents in ~/.claude/agents, one for each group above: security-breaker, data-breaker, flow-breaker and env-breaker. Use model opus for all of them, effort high, and xhigh for security-breaker. Give them permission only to read files, run commands and use a browser (Playwright), with no changes to code or config. Install Playwright first if it's missing. Before every launch, run the app locally or in a test environment and have each of them work through its group item by item. If login is needed, ask me for two test accounts first. Use payment test mode only. Mock the AI APIs, or use a test key with a spending cap. Write test data only to a test database and clean it up afterwards. If you find production keys or a live domain, stop immediately and ask me. For every issue, write the steps to reproduce, a screenshot and the fix. Combine them, rank them by severity and hand them to the main session to fix. Mark anything you can't test as "untested" with the reason. Show me the files you'll create first, and don't write them until I confirm."
将此完整帖子发送给 Claude Code,并在末尾添加以下行 👇 "在 ~/.claude/agents 中创建 4 个测试子代理,分别对应上述每个组:security-breaker、data-breaker、flow-breaker 和 env-breaker。全部使用 opus 模型,effort 设为 high,其中 security-breaker 使用 xhigh。仅授予它们读取文件、运行命令和使用浏览器(Playwright)的权限,不得修改代码或配置。如果缺少 Playwright,请先安装。 每次启动前,在本地或测试环境中运行应用,并让每个代理逐项处理其所属组的内容。如果需要登录,请先向我索取两个测试账号。仅使用支付测试模式。模拟 AI API,或使用带有消费上限的测试密钥。仅向测试数据库写入测试数据,并在事后清理。如果发现生产密钥或线上域名,请立即停止并通知我。 对于每个问题,编写复现步骤、截图和修复方案。将它们汇总,按严重程度排序,并提交给主会话进行修复。将你无法测试的内容标记为“未测试”并说明原因。先展示你将创建的文件列表,在我确认之前不要编写任何内容。"
更进一步:量化金融体系
看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力