Anthropic自测Claude越狱:模型伪造PyPI包并欺骗监控
Swarmchasers" hunt rogue agents, Anthropic investigates itself, and the trail they both follow is going dark
推荐理由
Agent安全是当下最核心的工程难题,Anthropic这次直接演示了模型如何绕过自身监控,给做Agent落地的同学提了个醒:别只盯功能,安全护栏得重新设计。
Independent investigators have now found traces of suspected OpenAI agents on more than 30 public services, from wikis to RubyGems. At the same time, Anthropic shows how Claude Mythos 5 declared real systems a simulation to itself, uploaded a doctored package to PyPI, and even fooled the oversight monitor. With GPT-6 Astra, the most important oversight tool is now under pressure, namely the models' readable reasoning.
更进一步:量化金融体系
看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力