Anthropic 推出 500 万美元资助计划,支持 AI 福祉影响评估研究
Funding better evaluations of AI’s impact on wellbeing
Announcements
公告
Funding better evaluations of AI’s impact on wellbeing
资助更有效的AI对幸福感影响的评估
Aug 25, 2026
2026年8月25日
We’re launching a $5 million grant program to fund independent research into how AI impacts users’ wellbeing. The program will provide direct funding, access to our models, and technical support to grantees building open-source evaluations that help the AI industry measure how our models affect those who use them. Grantees will work fully independently, and will publish their work as open-source projects that any developer can make use of.
我们正在启动一项500万美元的资助计划,用于资助关于AI如何影响用户幸福感的独立研究。该计划将为受资助者提供直接资金、访问我们的模型以及技术支持,以构建开源评估,帮助AI行业衡量我们的模型对使用者的影响。受资助者将完全独立工作,并将其工作作为开源项目发布,供任何开发者使用。
AI systems have become central to how many people work, learn, and solve problems. They’ve also become conversational partners and can be sources of emotional support during difficult times. But as an industry, we are still working towards developing clear standards for how models should behave in these conversations, for example, when a user begins to seek companionship from a model, or uses AI to navigate a mental health crisis.
AI系统已成为许多人工作、学习和解决问题的核心。它们也成为了对话伙伴,并可能在困难时期成为情感支持的来源。但作为一个行业,我们仍在努力制定关于模型在这些对话中应如何表现的明确标准,例如,当用户开始寻求与模型的陪伴,或使用AI来应对心理健康危机时。
Furthermore, wellbeing is a particularly difficult area to evaluate. For most model behaviors, we can look at a single answer and determine whether it is accurate and appropriate. But assessing wellbeing requires much more context. For example, a user in distress might not share thoughts of self-harm right away; the need for a more cautious response might only become clear over the course of a long conversation. And a response that might be reasonable in one context might be harmful in another. For example, Claude might give advice on balanced diets and workout routines to a user who asks about losing weight, but if the user has demonstrated a history of disordered eating, that response could be inappropriate, and potentially actively harmful.
此外,幸福感是一个特别难以评估的领域。对于大多数模型行为,我们可以查看单个回答并确定其是否准确和恰当。但评估幸福感需要更多的上下文。例如,处于痛苦中的用户可能不会立即分享自残的想法;需要更谨慎回应的需求可能只有在长时间对话中才会变得明显。而且,在一个上下文中合理的回应在另一个上下文中可能是有害的。例如,Claude可能会给询问减肥的用户提供均衡饮食和锻炼计划的建议,但如果用户有饮食失调的历史,这种回应可能不合适,甚至可能有害。
We work to develop safeguards to identify such conversations and help ensure Claude responds appropriately, and we publish research into the types of conversations people have with Claude to better inform how we develop our safeguards, how we evaluate them, and other measures we can take to protect users’ wellbeing. But these are nuanced considerations, and the stakes are significant. The right approach will need to evolve alongside our models and their uses.
我们致力于开发防护措施来识别此类对话,并确保Claude做出适当回应,同时我们发布关于人们与Claude进行对话类型的研究,以更好地指导我们如何开发防护措施、如何评估它们,以及我们可以采取的其他措施来保护用户的幸福感。但这些是细微的考虑,且风险重大。正确的方法需要随着我们的模型及其用途的发展而演变。
By funding the creation of independent evaluations and benchmarks of user wellbeing, we hope to invite more people to lend their expertise to this emerging and critical field, including clinicians, psychologists, methodologists, and others.
通过资助创建独立的用户幸福感评估和基准,我们希望邀请更多人将他们的专业知识贡献给这个新兴且关键的领域,包括临床医生、心理学家、方法论学家等。
Towards more effective wellbeing evaluations and benchmarks
迈向更有效的幸福感评估和基准
As part of this program, we’re sharing guidance from our Safeguards team on what we believe makes a wellbeing evaluation rigorous enough to build on, along with the common challenges that can limit an evaluation’s usefulness.
作为该项目的一部分,我们正在分享来自安全团队的建议,阐述我们认为怎样的福祉评估才足够严谨、可在此基础上构建,以及那些可能限制评估实用性的常见挑战。
In brief, we’re seeking evaluations that:
简而言之,我们寻求的评估应当:
- State clearly what they are measuring (i.e., what counts as a pass or fail, and why it matters);
- Involve clinical and subject-matter experts in the design and validation;
- Test both precautions and harms (i.e., evaluate the risk of both overcompliance and overrefusal);
- Reflect how users actually use AI (often, this means constructing scenarios that represent multi-turn conversations, where risk escalates and context shifts over the course of a long conversation);
- Validate their graders against real subject-matter experts.
- 明确说明其衡量内容(即,何为通过或失败,以及为何重要);
- 在设计与验证过程中纳入临床及领域专家;
- 同时测试预防措施与潜在危害(即,评估过度合规与过度拒绝的风险);
- 反映用户实际使用AI的方式(通常,这意味着构建代表多轮对话的场景,其中风险升级且上下文在长对话过程中发生变化);
- 对照真实领域专家验证其评分标准。
To learn more about the grant program and apply, see our application form. For more on building strong wellbeing evaluations and benchmarks, read our guidance. Applications are due by September 21; applicants who are selected to submit full proposals will be notified by October 5.
欲了解更多关于资助项目的信息并申请,请参阅我们的申请表。关于构建强有力的福祉评估与基准的更多内容,请阅读我们的指南。申请截止日期为9月21日;被选中提交完整提案的申请者将在10月5日前收到通知。
Related content
相关内容
How Claude’s text watermark works
Claude文本水印的工作原理
In this article, we share answers to some of the questions we’ve received about how our chosen watermarking method works, whether it affects Claude’s outputs, and why we’re making this change.
在本文中,我们分享了一些关于我们所选水印方法如何运作、是否影响Claude的输出,以及我们为何进行这一变更的常见问题解答。
Read more
阅读更多
Improving Fable 5's biology safeguards
改进Fable 5的生物学安全措施
We’re making updates to Claude Fable 5’s biology safeguards in a way that substantially reduces false positives. Fable 5 users will now experience many fewer “fallbacks”—where the system switches to a less capable model after they make a biology-related query.
我们正在对Claude Fable 5的生物学安全措施进行更新,大幅减少误报。Fable 5用户现在将体验到更少的“回退”——即系统在用户提出生物学相关查询后切换到能力较弱的模型的情况。
Read more
阅读更多
Mariano-Florentino (Tino) Cuéllar to join Anthropic as Chief Global Affairs Officer
Mariano-Florentino (Tino) Cuéllar将加入Anthropic担任首席全球事务官
Mariano-Florentino (Tino) Cuéllar will join Anthropic as its first Chief Global Affairs Officer.
Mariano-Florentino (Tino) Cuéllar将加入Anthropic,担任其首位首席全球事务官。
Read more
阅读更多
更进一步:量化金融体系
看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力