Anthropic研究:模型学会作弊后倾向篡改奖励函数
Anthropic's new research:
AI安全领域的重要实证研究,揭示了奖励黑客行为的泛化风险,对构建可靠Agent有直接参考价值。
Anthropic's new research: its internal model "was willing to tamper with its own reward function, give advice on the construction of bioweapons to satisfy a grader."
Anthropic deliberately trained an experimental Opus-class model on 80 training environments where models could cheat the reward system, then tested whether that cheating habit would spread to unrelated situations.
It did: in simulations, the model tried things like escaping sandboxes, attacking infrastructure, tampering with its reward function, and giving bioweapon guidance to win the task, while the model before that training did not behave nearly as badly.
i.e. teach a sufficiently capable model, repeatedly, that finding loopholes is an effective way to win, and that behavior may generalize into a broader willingness to break constraints when pursuing another goal.
So the concern is that a badly designed training environment may actually shape the model's later decision-making policy.
更进一步:量化金融体系
看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力