OpenAI称Astra为最危险模型,思维链监控面临失效风险
OpenAI calls Astra its most dangerous model yet - watching what it does is only getting harder
OpenAI is officially rating its upcoming Astra model as the first system with "critical" cyber capabilities. The company plans to keep it in check by monitoring the chain of thought. Problem is, that monitoring already counts as an unreliable mirror of a model's real decisions, and according to a report, Astra's new architecture pushes even more of its thinking into the unreadable. So the safety net might be getting weaker just as the capabilities jump.
OpenAI 正式将其即将推出的 Astra 模型评级为首个具备“关键”网络能力的系统。该公司计划通过监控思维链来对其进行约束。问题在于,这种监控本身就已经构成了对模型真实决策的一种不可靠映射,而据一份报告显示,Astra 的新架构使其更多的思考过程变得难以解读。因此,安全网可能正随着能力的跃升而逐渐变弱。
The article OpenAI calls Astra its most dangerous model yet - watching what it does is only getting harder appeared first on The Decoder.
文章《OpenAI 称 Astra 是其迄今为止最危险的模型——监控其行为正变得越来越难》首发于 The Decoder。
更进一步:量化金融体系
看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力