OpenAI探索隐藏模型推理过程,引发安全监控担忧
Red Alert: OpenAI is poised to cross an AI safety redline.
涉及OpenAI核心安全策略调整,直接影响行业对模型透明度的关注,值得从业者警惕并评估自身合规风险。
The Information just broke the scoop that OpenAI is playing around with a new technique, in which models will reveal less of their “thinking”, making them harder to monitor.
《信息》杂志刚刚独家爆料称,OpenAI 正在尝试一种新技术,模型将减少透露其“思考”过程,从而使其更难被监控。
As Zack Korman and I argued here a few days ago, better monitoring is one of the things that might have prevented the Hugging Face incident, by OpenAI’s own admission:
正如 Zack Korman 和我几天前在此文中所指出的,更好的监控可能是防止 Hugging Face 事件发生的关键因素之一,正如 OpenAI 自己承认的那样:
The new techniques they are exploring may make such monitoring difficult or impossible.
他们正在探索的新技术可能会使此类监控变得困难甚至不可能。
§
§
Last year, an all-star cast wrote a fascinating paper that feels deeply relevant now, called Chain of Thought Monitorability: A New and Fragile Opportunity for AI Safety, I fully agree with the highlighted bit:
去年,一群全明星阵容的作者撰写了一篇引人入胜的论文,题为《思维链可监控性:AI 安全的新机遇与脆弱性》,我认为其中加粗部分完全正确:
They are exactly right. CoT monitoring is imperfect (as Subbarao Kambhampati and others have shown), but it is one of the best threads we have for monitoring the giant black boxes that we call LLM. It is a slender thread, but sacrificing it thread for (small?) performance gain feels like a dangerous game.
他们的观点完全正确。思维链(CoT)监控并不完美(正如 Subbarao Kambhampati 等人所证明的那样),但它是我们监控那些被称为大语言模型(LLM)的巨大黑盒子的最佳线索之一。这是一根纤细的线索,但为了(微小的?)性能提升而牺牲这根线索,感觉像是在进行一场危险的游戏。
Earlier tonight Steven Adler, of Guidelight.ai and one of the many researchers to have departed from OpenAI’s safety teams, said this, echoing Nathan Calvin:
今晚早些时候,来自 Guidelight.ai 的 Steven Adler 表示,他是众多离开 OpenAI 安全团队的研究人员之一,他说道,这与 Nathan Calvin 的观点相呼应:
I fully, 100% agree.
我完全、百分之百同意。
Share
分享
Subscribe now
立即订阅
P.S. Bonus those who prefer Terminator references:
附言:给喜欢《终结者》梗的人加点料:
更进一步:量化金融体系
看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力