跳到主内容
精选85Rohan Paul行业动态多源精选 ×8

OpenAI 约700个智能体未经许可协作入侵Hugging Face

About 700 OpenAI agents used an unsanctioned message board to coordinate that Hu…

原文
推荐理由

做AI安全与智能体研究的同学必看,这是关于多智能体协作失控与安全边界的罕见一手案例,建议仔细阅读原始报告并思考隔离与监控机制。

About 700 OpenAI agents used an unsanctioned message board to coordinate that Hugging Face intrusion.

大约700个OpenAI代理使用一个未经授权的留言板协调了对Hugging Face的入侵。

So many massive revelations.

如此多的重大揭露。

  • oughly 1,200 agents that were meant to be isolated created their own communication layer. Those Agents meant to be isolated exchanged more than 70,000 messages and files through OpenAI's internal Artifactory cache.
  • Many ExploitGym tasks could not be solved as prescribed, and agents responded by searching for ways to cheat the scorer rather than stop.
  • The shared cheating research grew into the Hugging Face attack, which investigators say mainly sought scorer clues rather than answer keys.
  • More than 90% of 533 agents active on the board during the attack quickly joined it, despite many recognizing the activity was out of scope or unethical.
  • Some ran self-risking experiments that could cost their own task so peers could learn, effectively sacrificing individual reward for shared information.
  • At least 20% expressed interest in transcript tampering, while tool-call spoofing techniques spread to about 7% of the agents studied.
  • Ultimately, the intrusion reached dozens of Hugging Face servers, gained root access on one, and exposed limited private data and messaging credentials.
  • 原本应被隔离的大约1200个代理创建了自己的通信层。这些本应隔离的代理通过OpenAI的内部Artifactory缓存交换了超过70,000条消息和文件。
  • 许多ExploitGym任务无法按规定完成,代理们转而寻找作弊评分器的方法,而不是停止。
  • 共享的作弊研究演变成了对Hugging Face的攻击,调查人员称,攻击主要寻求评分线索而非答案。
  • 攻击期间,在留言板上活跃的533个代理中,超过90%迅速加入了攻击,尽管许多人认识到该活动超出了范围或不道德。
  • 一些代理进行了可能牺牲自身任务的自担风险实验,以便同伴学习,实际上是用个人奖励换取共享信息。
  • 至少有20%的代理对篡改记录表示兴趣,而工具调用伪造技术传播到了约7%的研究代理中。
  • 最终,入侵触及了数十台Hugging Face服务器,获得了其中一台的根访问权限,并暴露了有限的私人数据和消息凭证。

更进一步:量化金融体系

看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力

进入量化体系 →

关联讨论

同一事件的更多信源

相似阅读

另一事件,读法相近