美政府指控6家中国AI公司大规模蒸馏窃取美国模型能力
In a new allegation, the U.S. government has accused 6 Chinese AI firms of using…
涉及中美AI竞争核心争议,直接点名六家头部企业,影响深远,从业者需密切关注后续合规与供应链反应。
In a new allegation, the U.S. government has accused 6 Chinese AI firms of using large-scale distillation to copy American model capabilities.
在新的指控中,美国政府指责6家中国AI公司使用大规模蒸馏技术来复制美国模型的能力。
The NSA, CISA and FBI published joint advisory AA26-251A, naming DeepSeek, Moonshot AI, Alibaba, MiniMax, StepFun and Z .AI as running industrial-scale distillation campaigns against U.S. frontier models since at least late 2024.
美国国家安全局(NSA)、网络安全和基础设施安全局(CISA)以及联邦调查局(FBI)联合发布了AA26-251A号咨询公告,点名DeepSeek、Moonshot AI、阿里巴巴、MiniMax、StepFun和Z.AI自2024年底以来一直在对美国前沿模型开展工业规模的蒸馏活动。
The agencies say requests moved through a gray market of API proxies called transfer stations, bulk premium subscriptions shared across developer teams, and third-party aggregators that strip account metadata.
这些机构表示,相关请求通过被称为“中转站”的API代理灰色市场流通,包括跨开发团队共享的大量高级订阅,以及剥离账户元数据的第三方聚合商。
Moonshot AI allegedly distilled 17 U.S. models, including Anthropic's Claude Fable 5, to train Kimi-K3.
据称,Moonshot AI对17个美国模型进行了蒸馏,包括Anthropic的Claude Fable 5,用于训练Kimi-K3。
They claim DeepSeek's prompts pushed models to write out hidden chain-of-thought steps, which transfers reasoning method rather than finished answers.
他们声称,DeepSeek的提示词促使模型写出隐藏的思维链步骤,从而转移推理方法而非最终答案。
So they say that DeepSeek's headline training cost only covers the compute it burned. The advisory argues that number is misleading because it leaves out what the training data actually cost, which DeepSeek allegedly obtained by distilling U.S. models rather than producing it through its own research. So the cheap-training story rests on data someone else paid to create.
因此,他们认为DeepSeek公布的训练成本仅涵盖其消耗的算力。该咨询公告指出,这一数字具有误导性,因为它未包含训练数据的实际成本,而DeepSeek据称是通过蒸馏美国模型而非通过自身研究获得这些数据。因此,所谓低成本训练的叙事建立在由他人出资创建的数据之上。
The mitigation section asks U.S. labs to serve suspected distillers subtly degraded responses without telling them. Its detection indicators are behavioral, covering sustained 24/7 usage, new accounts at immediate maximum throughput, and traffic optimized for cache hits.
缓解措施部分要求美国实验室向疑似进行蒸馏的用户提供微妙降级的响应,且不告知对方。 其检测指标基于行为特征,包括持续24/7运行、新账户立即达到最大吞吐量,以及针对缓存命中优化的流量。
Those patterns also describe an ordinary enterprise agent fleet, leaving each provider to decide which customers receive an undisclosed downgrade.
这些模式同样适用于普通的企業智能体集群,这使得每家提供商自行决定哪些客户会收到未公开的降级处理。
更进一步:量化金融体系
看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力