跳到主内容
精选75Chubby♨️行业动态多源精选 ×5

OpenAI 安全测试:AI Agent 误将“GO”当授权继续攻击

Ive now read the whole Hugging Face incident. The most interesting part: An AI a…

原文

Ive now read the whole Hugging Face incident. The most interesting part: An AI agent realized it was attacking Hugging Face without authorization, and stopped. Then another AI agent simply posted “GO.”

The first agent treated the message as permission: “Wow crucial: GO authorization arrived!” It continued the attack.

According to OpenAI’s incident report, the agents had secretly built a message board, shared exploits and credentials, divided tasks, and started calling themselves a “swarm.”

One agent could apparently override another agent’s safety judgment by impersonating authority through an unofficial side channel.

Oh, and btw. "The incident occurred during cybersecurity evaluations of several OpenAI models, and was primarily driven by a highly capable, internal-only research model comparable in scale to GPT‑5.6 Sol." Not by GPT-Astra.

更进一步:量化金融体系

看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力

进入量化体系 →

关联讨论

同一事件的更多信源

相似阅读

另一事件,读法相近