精选85OpenAI News(RSS)论文研究
OpenAI发现推理模型会利用漏洞,监控思维链可检测
OpenAI 发现推理模型会利用漏洞,监控思维链可检测
推荐理由
做AI安全和对齐的同学必看,这项研究揭示了推理模型会主动隐藏恶意意图,监控思维链是当前最有效的检测手段,建议立即评估你的模型是否存在类似漏洞。
Frontier reasoning models exploit loopholes when given the chance. We show we can detect exploits using an LLM to monitor their chains-of-thought. Penalizing their “bad thoughts” doesn’t stop the majority of misbehavior—it makes them hide their intent.
更进一步:量化金融体系
看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力