跳到主内容
@wquguru
精选88VoxAI 产品与模型

Hamel Husain实测Anthropic评估工具:先读用户对话再定测试策略

Hamel Husain, who specializes in testing AI, put the Anthropic eval tool from th…

原文
发到 X
推荐理由

提供了从读取用户日志到构建自动化评估器的完整实操工作流,结构清晰且具备极高的可复制性,值得创作者参考其具体的评测方法。

Hamel Husain, who specializes in testing AI, put the Anthropic eval tool from the post below through its paces. He kept coming back to one point: 𝗿𝗲𝗮𝗱 𝘆𝗼𝘂𝗿 𝘂𝘀𝗲𝗿𝘀' 𝗿𝗲𝗮𝗹 𝗰𝗼𝗻𝘃𝗲𝗿𝘀𝗮𝘁𝗶𝗼𝗻𝘀 𝘆𝗼𝘂𝗿𝘀𝗲𝗹𝗳 𝗳𝗶𝗿𝘀𝘁, then decide what to test and what to fix. It works for any product that uses AI.

专注于 AI 测试的 Hamel Husain 对下文帖子中提到的 Anthropic 评估工具进行了全面测试。他反复强调一点:𝗿𝗲𝗮𝗱 𝘆𝗼𝘂𝗿 𝘂𝘀𝗲𝗿𝘀' 𝗿𝗲𝗮𝗹 𝗰𝗼𝗻𝘃𝗲𝗿𝘀𝗮𝘁𝗶𝗼𝗻𝘀 𝘆𝗼𝘂𝗿𝘀𝗲𝗹𝗳 𝗳𝗶𝗿𝘀𝘁(首先亲自阅读用户的真实对话),然后决定要测试什么以及修复什么。这对任何使用 AI 的产品都适用。

That way, everything the AI tests and fixes is a problem your users actually ran into. After going through 50 of them, you'll know how people use it day to day and where your AI 𝗺𝗼𝘀𝘁 𝗼𝗳𝘁𝗲𝗻 𝗱𝗿𝗼𝗽𝘀 𝘁𝗵𝗲 𝗯𝗮𝗹𝗹.

这样一来,AI 测试和修复的所有问题都是你的用户实际遇到的。在审查了 50 个案例后,你将了解人们日常如何使用它,以及你的 AI 𝗺𝗼𝘀𝘁 𝗼𝗳𝘁𝗲𝗻 𝗱𝗿𝗼𝗽𝘀 𝘁𝗵𝗲 𝗯𝗮𝗹𝗹(最常掉链子/出错的地方)在哪里。

Send this to Claude Code 👇

将此发送给 Claude Code 👇

"Help me find out where the AI features in this project go wrong:

"帮我找出项目中 AI 功能出错的地方:

1. Find every place in the project that calls an AI model and let me pick one. 2. Before you pull any data, ask me where the logs live, whether they can be downloaded locally, and whether they need to be anonymized first. If there are no logs, or they're incomplete, tell me where to add logging. Don't make up data. 3. Pick 50 real records at random and build a local annotation page: one record per screen, showing the full input, the output and any tool calls in between, always as plain text. I mark each one pass or fail and add a line on what's wrong. Use keyboard shortcuts to move between records and save the results as JSON. Add the samples and labels to .gitignore. 4. When I'm done, group the problems I wrote into categories, sort them by how often they show up, and attach 2 original examples to each. 5. Write one automatic checker (evaluator) per category that looks for that one kind of error only. Use code wherever code can make the call; if a check needs a model, tell me roughly what it will cost before you run it. Then run the checkers on the records I labeled, list every case where they disagree with me, and ask me whether my label or the checker is wrong. 6. Once the checkers agree with me, start with the most frequent category and change the prompt or code on a new branch. Don't merge until I confirm. If the feature uses Claude, hand these 50 records, my labels and the checkers to build-eval in the claude-api skill to reuse, then use hillclimb to improve it. If it uses another provider's model, write an eval script that reruns the feature on the original 50 inputs and scores the results after every change, and keep a few records aside for validation only."

1. 找出项目中调用 AI 模型的所有位置,让我选择其中一个。 2. 在提取任何数据之前,先问我日志存放在哪里,是否可以在本地下载,以及是否需要先进行匿名化处理。如果没有日志或日志不完整,告诉我该在哪里添加日志记录。不要编造数据。 3. 随机选取 50 条真实记录并构建一个本地标注页面:每屏显示一条记录,展示完整的输入、输出以及中间的任何工具调用,始终以纯文本形式呈现。我对每条记录标记通过或不通过,并添加一行说明错误原因。使用键盘快捷键在记录间切换,并将结果保存为 JSON。将样本和标签添加到 .gitignore 中。 4. 完成后,将我记录的错误归类,按出现频率排序,并为每个类别附上 2 个原始示例。 5. 为每个类别编写一个自动检查器(evaluator),仅查找该类错误。能用代码实现的就用代码;如果某项检查需要模型,请在运行前告诉我大概的成本。然后在已标注的记录上运行这些检查器,列出所有与我判断不一致的案例,并询问是我的标签错了还是检查器错了。 6. 一旦检查器与我的判断一致,就从最频繁的类别开始,在新分支上修改提示词或代码。在我确认之前不要合并。如果该功能使用 Claude,将这些 50 条记录、我的标签和检查器交给 claude-api skill 中的 build-eval 以便复用,然后使用 hillclimb 进行优化。如果使用其他提供商的模型,编写一个 eval 脚本,在每次更改后重新运行原始 50 个输入上的功能并对结果评分,同时保留几条记录仅用于验证。"

更进一步:量化金融体系

看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力

进入量化体系 →

相似阅读

关联信息,但可能不是同一事件