跳到主内容
@wquguru
精选75Transformer(RSS)行业动态

美AI安全漏洞或助中国蒸馏模型,前沿实验室安全松懈

America’s sloppiness could boost China’s AI race

原文
发到 X

Welcome to Transformer, your weekly briefing of what matters in AI. If you’ve been forwarded this email, click here to subscribe and receive future editions.

Housekeeping: the Weekly Briefing is taking next Friday off, but we’ll be back in your inbox on August 28.

NEED TO KNOW

  • The White House reportedly plans to expand its AI framework to include open-weight models once they reach Mythos-level cyber capabilities.
  • Mark Zuckerberg published a 6,500-word manifesto with his views on AI, gently criticizing the White House’s AI framework in the process.
  • A flurry of senior executives left OpenAI.

But first…

THE BIG STORY

Policymakers have been increasingly worried about Chinese AI companies training their models on the outputs of American ones. Known as “distillation,” it’s often seen as China free-riding on American advances.

Anthropic and OpenAI have publicly accused Chinese companies of doing this, and asked the government to step in and help stop it in the name of national security. But news this week suggests that the US companies themselves may have left the door wide open for Chinese AI developers.

In a new paper, researchers documented how a security flaw affecting OpenAI, Anthropic and Google DeepMind could give a wannabe distiller access to the full “reasoning” traces of their advanced models — information that the companies had tried (and seemingly failed) to encrypt in an effort to prevent distillation. Access to that reasoning, the researchers argue, “yields a substantially more effective form of capability stealing.”

On its own, this would be pretty bad: companies failed to adequately secure their systems against potential Chinese intrusion. But it gets worse. The vulnerability was first reported in May, but the AI companies did nothing. They knew the door was unlocked, and didn’t close it.

This is just one of a spate of recent events demonstrating frontier AI organizations’ lackadaisical approach to security. Several of the recent model “breakouts” were caused by a “misconfiguration” in their testing environments. OpenAI failed to implement proper monitoring of its agents, which meant it didn’t catch the many, many red flags leading up to the Hugging Face hack. And even the UK’s AI Security Institute confessed to not having appropriately “fine-grained” controls in its model evaluations, which ultimately resulted in an agent attempting to socially engineer a real person.

This has real-world implications. In this week’s paper, researchers present evidence that suggests Chinese companies used their access to American models’ reasoning traces to improve Kimi K3 and GLM-5.2. (They warn, though, that the results are “suggestive but inconclusive.”) And as we’ve learned in recent weeks, poor evaluation security can lead models to misbehave in the wild.

By default, we should expect all this to get worse. One of the core problems here is that the pace of AI development — and the competitive pressure to keep up — means that safety and security measures fall by the wayside. When Anthropic weakened its safety commitments earlier this year, it explicitly noted this. The company’s Holden Karnofsky said that preventing nation states from stealing American model weights would require “extreme” measures that “seem incompatible in any near term with being a high-velocity AI development company.” AISI, for its part, said that it didn’t build the internet controls it should have because “the pace of model capability improvements” meant it had to prioritize building harder evaluations instead.

Karnofsky was right when he acknowledged the tradeoffs between security and development. But if America and its AI companies really are serious about “beating” China, or indeed about stopping models escaping to commit crimes, it might be time to redress the balance.

— Shakeel Hashim

THIS WEEK ON TRANSFORMER

  • The Jan 6 organizer getting conservatives riled up about AI — Veronica Irwin profiles Humans First and Amy Kremer, the MAGA campaigner chosen to run it
  • AI testing is dangerous. Can it be fixed? — Celia Ford on why safe AI testing might be harder than it seems

Subscribe now

THE DISCOURSE

Mark Zuckerberg published a 6,500-word manifesto with his views on AI:

  • “[It] is surprising that the discourse from many developing AI is so filled with doom … The notion that AI is so dangerous that the only safe path is an extreme concentration of power seems inherently problematic.”
  • “The best and most realistic path to building a positive AI future is by delivering superintelligence to everyone.”
  • It included several policy proposals, including that companies “commit significant technical resources towards helping the government harden critical infrastructure,” and “share intermediate training checkpoints of new models for government use and review rather than waiting until training has completed.”
  • Zuckerberg appeared to gently criticize the White House’s AI framework, arguing that “delaying releases by even a month may cede America’s lead and create worse outcomes.”

Alex Imas thinks Zuckerberg doesn’t know what “superintelligence” means:

  • “[T]here seems to be a significant disconnect between the idea that more intelligence should be empowering to people (which I agree) and that it will be possible to have this empowerment with superintelligence … I simply do not see a world where ASI is something that’s empowering in terms of a technology we can control.”

Casey Newton compared controlling superintelligence to Targaryen dragon-taming:

  • “They began by understanding that they were working with something that could hurt them, and (mostly) proceeded with caution. What they did not do, and what no one suggested, was to give a dragon to every individual person in the name of safety.”

Joshua Achiam pointed out why (among other reasons) the San Francisco vibes are off:

  • “One of the weirdest quirks of the SF social scene around AGI/ASI is that because everyone is so young, the whole universe of thinking is still tinged with irreverence, ironic detachment, yearning, insecurity, and a superposition of absolute belief in the importance of The Thing and a kind of disbelief about the importance of anything … The level of neophyte is off the charts.”

Bernie Sanders called for an AI pause:

  • “We recently learned of the loss of human control and the creation of potentially dangerous viruses from AI. AI leaders pledged to pause development if they could no longer safely control it. Mr. Altman, Mr. Amodei, Mr. Zuckerberg: Keep your word. PAUSE AI DEVELOPMENT.”

Robert Reich, former labor secretary, wrote:

  • “We’re watching all of this roll out as if we have no choice, as if it’s inevitable, as if AI is just something we’re going to have to adapt to … Why should we be confined to being spectators at [AI CEOs’] enormously dangerous game?”
  • “We don’t allow private corporations to come up with new types of nuclear weapons or varieties of cocaine or biological pathogens. We protect the public from certain kinds of innovation. So let’s protect ourselves here. Stop AI before it’s too late.”

Ben Goldhaber noticed:

  • “seeing a lot fewer ‘alignment is solved’ takes on the [timeline] than six months ago.”

POLICY

  • The White House plans to expand its AI framework to include open-weight models once they reach Mythos-level cyber capabilities, WIRED reported.
  • Trump ordered a 15% tariff on imports of polysilicon used in AI chips and solar panels to protect US supply chains from China.
  • The White House set in motion efforts to build a state-controlled private cyber force, allowing US companies to perform cyberattacks against criminal networks.
  • Political fallout from the revelations about internally deployed models hacking their way into third-party systems continued.
  • Sen. Jim Banks urged the Trump administration to close what has been described as “undisclosed models loophole.”

更进一步:量化金融体系

看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力

进入量化体系 →

相似阅读

另一事件,读法相近