Anthropic发布指南:用Claude Code优化API调用
Anthropic just published a guide where they had Claude tune an AI feature round…
提供了具体的Claude Code指令模板和可量化的优化数据(成本与准确率变化),结构清晰,对开发者有直接参考价值。
Anthropic just published a guide where they had Claude tune an AI feature round by round, and the same job ended up costing 𝗮𝗯𝗼𝘂𝘁 𝟭/𝟱 𝗮𝘀 𝗺𝘂𝗰𝗵, with higher accuracy. You can have Claude Code follow it and put any feature in your product that calls the Claude API through the same process, with two commands.
Anthropic 刚刚发布了一份指南,其中展示了他们让 Claude 逐轮优化一个 AI 功能的过程,最终相同工作的成本降至原来的约五分之一,且准确率更高。你可以让 Claude Code 遵循该流程,通过同样的步骤将任何调用 Claude API 的功能集成到你的产品中,只需两条命令即可。
It takes three steps. First, cut the extra steps and contradictory rules from the prompt. Then step down to cheaper models and lower effort, one level at a time. Finally, add the rules that were missing.
整个过程分为三步。首先,从提示词中剔除多余的步骤和相互矛盾的规则。然后逐步降级到更便宜的模型并降低投入程度,每次只调整一个层级。最后,补上之前缺失的规则。
Every step is checked against a set of cases it has never seen, and any change that lowers accuracy gets rolled back. In the end, accuracy went from 𝟳𝟴.𝟲% 𝘁𝗼 𝟵𝟬.𝟱%. The new Sonnet 5.5 is close to Opus 5.5 in capability but cheaper and faster, so if you want to use it, measure it the same way first.
每一步都会用一组它从未见过的案例进行检查,任何导致准确率下降的更改都会被回滚。最终,准确率从 78.6% 提升至 90.5%。新的 Sonnet 5.5 在能力上已接近 Opus 5.5,但成本更低、速度更快,因此如果你想使用它,请先以相同的方式进行评估。
Send this prompt to Claude Code 👇
将此提示词发送给 Claude Code 👇
"Read this guide: https://claude.dev/blog/automating-eval-design-and-hillclimbing/
"阅读这份指南:https://claude.dev/blog/automating-eval-design-and-hillclimbing/
Then list every place in this project that calls the Claude API, estimate which one costs the most, and let me pick one. Use build-eval from the claude-api skill to build a test set for it, and tell me roughly what it will cost in API fees before you run anything.
然后列出该项目中所有调用 Claude API 的地方,估算哪一部分成本最高,并让我选择其中一个。使用 claude-api 技能中的 build-eval 为其构建测试集,并在运行任何操作前告诉我大致需要多少 API 费用。
Once I've approved the test cases and the grading, run hillclimb on a new branch. The goal is to keep accuracy and cut cost. You can try lower effort and cheaper models, including Sonnet 5.5. The report should spell out three things: 1. What changed at each step, and how accuracy and cost per call moved. 2. Which cheaper setups got things wrong, and where. 3. Which sentences were removed from or added to the prompt.
在我批准测试用例和评分标准后,在新分支上运行 hillclimb(爬山优化)。目标是保持准确率的同时降低成本。你可以尝试降低投入程度并使用更便宜的模型,包括 Sonnet 5.5。报告应明确说明以下三点: 1. 每一步发生了哪些变化,以及准确率和每次调用的成本如何变动。 2. 哪些更便宜的配置出现了错误,具体错在哪里。 3. 提示词中删除或新增了哪些句子。
Show me the report first. Don't merge anything until I confirm."
先向我展示报告。在我确认之前,不要合并任何内容。"
更进一步:量化金融体系
看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力