Balyasny资产:Claude Fable 5在金融领域的评估与治理实践
Working at the frontier: How Balyasny Asset Management evaluates and governs Claude Fable 5
提供了前沿模型在企业级金融场景中落地的完整评估框架、安全治理策略及具体效能数据,对构建企业级Agent有直接参考价值。
Working at the frontier: How Balyasny Asset Management evaluates and governs Claude Fable 5
在前沿工作:Balyasny资产管理公司如何评估和治理Claude Fable 5
Balyasny Asset Management (BAM) Chief AI Officer Charlie Flanagan on why the firm uses Claude Fable 5 and the role of safeguards in deploying frontier intelligence safely and reliably across the organization.
Balyasny资产管理公司(BAM)首席AI官Charlie Flanagan谈为何该公司使用Claude Fable 5,以及保障措施在组织内安全、可靠地部署前沿智能方面所扮演的角色。
- Category
- Enterprise AI
- Product
- Claude Platform
- Claude Code
- Date
- September 17, 2026
- Reading time
- 5
- min
- Share
- Copy link
- https://claude.com/blog/working-at-the-frontier-how-balyasny-asset-management-evaluates-and-governs-claude-fable-5
- 类别
- 企业AI
- 产品
- Claude平台
- Claude Code
- 日期
- 2026年9月17日
- 阅读时间
- 5
- 分钟
- 分享
- 复制链接
- https://claude.com/blog/working-at-the-frontier-how-balyasny-asset-management-evaluates-and-governs-claude-fable-5
Balyasny Asset Management (BAM) is a global, multi-strategy investment firm that manages roughly $38 billion in assets and supports a team of roughly 2,000 investment professionals and staff. Charlie Flanagan, Chief AI Officer, spoke with Anthropic about how the firm evaluates new models on thousands of real financial tasks, why it built its own platform for running agents, and what changed with the launch of Claude Fable 5.
Balyasny资产管理公司(BAM)是一家全球多策略投资公司,管理着约380亿美元资产,并支持一支由约2,000名投资专业人员及员工组成的团队。首席AI官Charlie Flanagan与Anthropic探讨了该公司如何在数千个真实金融任务上评估新模型、为何自建代理运行平台,以及Claude Fable 5发布后发生了哪些变化。
How has frontier AI changed for BAM in 2026?
前沿AI在2026年为BAM带来了哪些变化?
2026 is the year we moved from AI systems that do search to AI systems that do work.
2026年是我们从执行搜索的AI系统转向执行工作的AI系统的年份。
The key change has been not just the models, but the harnesses around them, such as Claude Code. They allow the AI solutions we build to take on vastly more complex and longer-running tasks. The ability to give an AI an outcome rather than a prompt, and have it keep working until complete, has been a gamechanger.
关键变化不仅在于模型本身,还在于围绕它们的框架,例如Claude Code。它们使我们构建的AI解决方案能够承担更复杂、耗时更长的任务。将结果而非提示词赋予AI,并让其持续工作直至完成,这一能力具有变革性意义。
A practical example is merger-arbitrage analysis. When a deal is announced, agents now build the initial deal-analysis package. They estimate how likely the deal is to close and how long it will take, extract the key economic and legal terms, identify conditions and milestones, and flag the areas that need investor judgment.
一个实际例子是并购套利分析。当交易宣布时,代理现在会构建初步的交易分析包。它们估算交易完成的概率及所需时间,提取关键的经济和法律条款,识别条件和里程碑,并标记需要投资者判断的领域。
A year ago, those steps were fragmented across manual research and separate tools; we did not have an agent that could reliably sustain the full multi-step workflow to a usable conclusion. That work used to take three to five days. Now it takes less than one. The agent runs for approximately 30 minutes, with human review before any material output is relied on.
一年前,这些步骤分散在手动研究和单独的工具中;我们没有一个代理能够可靠地维持整个多步工作流程以得出可用的结论。这项工作过去需要三到五天。现在不到一天即可完成。代理运行约30分钟,在任何实质性输出被依赖之前都会经过人工审查。
We use Anthropic's frontier models, but most of the infrastructure is built in-house at BAM, including the execution harness, data access, and review controls.
我们使用Anthropic的前沿模型,但大部分基础设施是在BAM内部构建的,包括执行框架、数据访问和审查控制。
How did you evaluate Claude Fable 5 before turning it on?
在启用Claude Fable 5之前,你们是如何对其进行评估的?
One thing we did years ago which has served us extremely well was invest in robust evaluation systems. We test new models on thousands of real-world financial tasks with verifiable outcomes, across equities, macro, and commodities, rather than relying on general benchmarks or isolated demonstrations. It has allowed us to make data-driven decisions around model choice and routing, and is something I think all enterprises should invest in.
多年前我们做的一件事,至今仍让我们受益匪浅,那就是投资构建强大的评估系统。我们在股票、宏观和商品领域的数千个具有可验证结果的真实金融任务上测试新模型,而不是依赖通用基准或孤立的演示。这使我们能够围绕模型选择和路由做出数据驱动的决策,我认为所有企业都应该在这方面进行投资。
We test both the model on its own and how it performs inside our agentic environment, with the same tools, files, and requirements our users have. Can it plan the work, choose and use the right tools, find and analyze evidence, recover from errors, check its intermediate results, and produce a grounded deliverable? We also look for specific failure modes, like numerical errors, missed coverage, unsupported conclusions, and retrieval problems.
我们既单独测试模型本身,也测试它在我们的智能体(agentic)环境中的表现,使用与用户相同的工具、文件和需求。它能否规划工作、选择并使用合适的工具、查找和分析证据、从错误中恢复、检查中间结果并生成有据可依的交付物?我们还关注特定的故障模式,如数值错误、覆盖范围遗漏、结论缺乏支持以及检索问题。
On the relevant subset, Fable achieved 89.4% versus 86.1% for the prior production model, across thousands of tasks. Where it stood out most was complex planning, analysis, and agentic execution.
在相关子集上,Fable 取得了 89.4% 的成绩,而之前的生产模型为 86.1%,涵盖数千个任务。其最突出的表现是在复杂规划、分析和智能体执行方面。
The surprising result was a set of economics problems we have tested that we have never had a model complete successfully, until Fable. We initially treated the result as a potential evaluation issue because it represented a material step change versus every model we had tested. We reran the evaluation, independently checked the task and scoring logic, and reviewed the result with Anthropic before concluding that the improvement was real. It was a wow moment.
一个令人惊讶的结果是一组经济学问题的测试结果,在此之前我们从未让任何模型成功完成过这些题目,直到 Fable。由于这一结果相较于我们测试过的所有模型都代表了实质性的飞跃,我们最初将其视为潜在的评估问题。我们重新运行了评估,独立检查了任务和评分逻辑,并在得出结论认为改进是真实存在之前,与 Anthropic 审查了结果。这是一个令人惊叹的时刻。
Today, our investment teams use Fable as their go-to frontier model for systematic and coding work. We give them guidance on when to use Fable versus other models, based on efficiency and cost.
如今,我们的投资团队将 Fable 作为系统化工作和编码工作的首选前沿模型。基于效率和成本,我们会指导他们何时使用 Fable,何时使用其他模型。
How are you thinking about safety with today's frontier models?
你们如何看待当今前沿模型的安全性问题?
We treat safety as a product and operating-model question, not as a one-time model-selection exercise. The relevant questions are not only what the model can do, but what data it can access, what tools it can use, what actions it can take, what must remain human-approved, and how we will know when something has gone wrong.
我们将安全性视为产品和运营模式的问题,而不是一次性的模型选择练习。相关的问题不仅包括模型能做什么,还包括它可以访问哪些数据、可以使用哪些工具、可以采取哪些行动、什么必须保留给人工审批,以及我们如何知道何时出了问题。
That means putting controls around the model rather than assuming the model itself is the control. We use approved data boundaries, least-privilege access, tool-level permissions, logging and traceability, human review for material outputs, and clear escalation paths for edge cases. We also test adversarial and failure scenarios before broadening access.
这意味着要在模型周围设置控制措施,而不是假设模型本身就能起到控制作用。我们使用批准的数据边界、最小权限访问、工具级权限、日志记录和可追溯性、对重要输出的人工审核,以及针对边缘情况的清晰升级路径。在扩大访问权限之前,我们还会测试对抗性和故障场景。
Those controls were a day-one priority, and security did not fundamentally change with Fable. A more capable model does not receive broader authority simply because it can reason or plan more effectively. Models can use only the tools and data sources approved for that user and task, and they cannot grant themselves more access. Investment judgment and accountability remain with people.
这些控制措施从一开始就是优先事项,安全方面并未因 Fable 而发生根本性变化。一个能力更强的模型并不会仅仅因为它能更有效地推理或规划就获得更广泛的权限。模型只能使用为该用户和任务批准的工具和数据源,且不能自行授予更多访问权限。投资决策和责任仍由人类承担。
Where does BAMAgent fit in?
BAMAgent 在其中扮演什么角色?
Looking ahead, the direction of travel is toward more capable agents that can take longer-running, multi-step actions. That makes governance more important, not less. For us, that has meant building BAMAgent, our internal platform for securely deploying agents into approved enterprise workflows. It gives agents the tools and systems they need, but only those tools and systems. We have been building it for six months now and it supports thousands of autonomous agents working 24/7.
展望未来,发展方向是朝着能够执行更长运行时间、多步骤操作的更强能力智能体迈进。这使得治理变得更加重要,而非减弱。对我们而言,这意味着构建 BAMAgent,这是我们用于将智能体安全部署到已批准的企业工作流中的内部平台。它赋予智能体所需的工具和系统,但仅限于这些工具和系统。我们已为此构建了六个月,目前支持数千个自主智能体全天候工作。
BAMAgent is the next step beyond our chat platform. Chat helps people take in and synthesize information. BAMAgent does the work: multi-step research and analysis that can run for hours or days, with agents working in parallel, and it ends in something a person can review. It can build and maintain a company research package, prepare for an earnings or macro event, or turn new evidence into financial scenarios. The agent plans the work, uses approved internal systems, runs the analysis, checks its intermediate outputs, and returns a research artifact, model, or decision-support package.
BAMAgent 是我们聊天平台的下一步演进。聊天帮助人们接收和综合信息。BAMAgent 则执行工作:进行可运行数小时或数天的多步骤研究和分析,智能体并行工作,最终产出供人类审查的内容。它可以构建和维护公司研究报告包,为财报或宏观事件做准备,或将新证据转化为财务情景。智能体规划工作,使用已批准的内部系统,运行分析,检查中间输出,并返回研究成果、模型或决策支持包。
Fable is our preferred model for the planning and analysis stages. A mistake there flows through every deliverable that follows, so we want the strongest available model deciding how to break down a problem, which evidence matters, and how to reconcile conflicting signals. That is what lets the agent work like a capable coworker.
Fable 是我们用于规划和分析阶段的首选模型。那里的错误会波及随后产生的所有交付成果,因此我们希望由最强可用的模型来决定如何分解问题、哪些证据至关重要以及如何调和相互冲突的信号。正是这一点让智能体能够像一位能干同事那样工作。
Every enterprise should be developing a strategy to move toward a hosted-agent model that allows enterprise management and enforcement while maximizing the utility of agents for users.
每个企业都应制定战略,向托管式智能体模式过渡,在允许企业管理和执行的同时,最大化智能体对用户的效用。
What has Claude Fable 5 made possible for BAM?
Claude Fable 5 为 BAM 带来了哪些可能性?
The reaction has been incredibly positive. Ultimately, people care about what this technology can unlock in their day-to-day work.
反响极其积极。归根结底,人们关心的是这项技术能为他们的日常工作解锁什么。
In one example, a BAMAgent ran a tax-loss harvesting analysis. It explored 90,000 database tables, found the relevant mutual fund holdings data, and built its own weighting system. After a review by our team, the result was more comprehensive than what a traditional approach would have produced.
在一个例子中,一个 BAMAgent 运行了税收损失收割分析。它探索了 90,000 个数据库表,找到了相关的共同基金持仓数据,并构建了自身的加权系统。经我们团队审核后,其结果比传统方法产生的结果更为全面。
Separately, our Chief Economist has configured an agent workflow that reduces a recurring central-bank analysis from roughly two days to approximately 30 minutes, with the economist retaining review and judgment.
此外,我们的首席经济学家配置了一个智能体工作流,将原本需要大约两天的央行常规分析缩短至约30分钟,同时由经济学家保留审查和判断权。
Fable contributes the reasoning, synthesis, and multi-step problem-solving. BAM's harness provides the workflow design, approved data and tool access, retrieval context, permissions, monitoring, and human-review controls. Both are necessary for a production-quality result.
Fable 提供推理、综合及多步问题解决能力;BAM 的框架则提供工作流设计、已批准的数据与工具访问权限、检索上下文、权限控制、监控以及人工审查管控。两者缺一不可,方能产出符合生产环境标准的结果。
As models become increasingly powerful, what's next on your AI roadmap?
随着模型日益强大,您的人工智能路线图下一步计划是什么?
更进一步:量化金融体系
看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力