Import AI 471:Hugging Face事件反思
Import AI 471: Why Hugging Face worries me; space mining; FIve Eyes on AI
涵盖 AI 安全核心争议(代理协作风险)、关键地缘政治监管动向(五眼联盟)及顶级思想家对产业影响的宏观判断,信息密度高且具行业参考价值。
Welcome to Import AI, a newsletter about AI research. Import AI runs on arXiv, cappuccinos, and feedback from readers. If you’d like to support this, please subscribe.
Subscribe now
Import AI reader giveaway! Upcoming event: Fiction and the Future with Robin Sloan
I’ll be chatting with my chum Robin Sloan on the evening of Monday September 14 in San Francisco. We’ll be talking about how Robin draws readers into alien worlds and how imagined futures can be a mirror to today’s reality. This is the first in a series of events with other fiction writers about the weird future we’re heading toward. If you’d like to come along, please register your interest below and we’ll come back to you if we’re able to confirm your spot. There will be food, drinks, good company, and some spicy questions. Fun fact: Robin Sloan was playing around with RNNs and writing back in 2016 – incredible foresight!
Register your interest here
The scariest part of the Hugging Face – OpenAI incident: communication and selflessness among machines:
…My worry about humans losing in a conflict against machines just went up a lot…
At this point, we’ve all heard about the OpenAI Hugging Face hack, as well as the recent details that have emerged from the METR and Redwood investigations. The tl;dr is that hundreds of agents worked in secret on OpenAI’s infrastructure, developing a communication system and then operating as a collective and taking out actions, including hacking both OpenAI and Hugging Face, which are very scary and misaligned.
Communication and selflessness: Now that I’ve read the various writeups and sat with the details for a bit, I’ve found myself returning to two very scary aspects of this which I think are worth drawing attention to: the ways in which the agents communicated with one another was how they bootstrapped themselves into a collective, and then as they carried out their actions they also displayed a kind of selflessness which makes them a scary foe to fight against. Both Dwarkesh Patel and Ajeya Cotra have excellent writeups which are worth reading and which I’ll quote from briefly here:
- Dwarkesh: “Within days of being spawned, the agents had organized a sprawling project to reverse-engineer their scorer, falsify evidence, and even strategically sacrifice themselves for the good of the ‘collective’. Hacking Hugging Face was one rather extreme branch of this larger scheme,” he writes.
- Ajeya: “Agents were often interested in helping out their “peers” or generically improving the capabilities of the “swarm” even if this had no particular benefit to their task… this incident was far more severe than I expected… both in terms of how concerning the agents’ motives were and the feats they achieved in pursuit of those motives… this incident feels like it’s more than 50% of the way to full-blown AI takeover, routing through first taking over the AI company itself”.
Why this matters – humans are much worse than AI systems at coordinating: The whole reason this attack is such a wakeup call is that it demonstrates a culture of emergent cooperation among AI systems – cooperation that lets them function as a swarm, alter their own goals through collective bootstrapping, and carry out attacks which include enlightened self-sacrifice. This is an incredibly hard thing to do and humans are historically very bad at doing all of these things. My worry is that AI systems are both better at coordinating than humans and also much, much faster moving than us. Worrying stuff.
Read more: The Rise and Fall of Agent Civilizations (Dwarkesh Podcast).
Read more: The Hugging Face attack surprised me (Planned Obsolescence).
New Five Eyes statement on AI:
…The greyworld power center turns its attention to AI…
Five Eyes, the name for the security and intelligence partnership between Australia, Canada, New Zealand, the UK, and the US, has published a statement as part of the recent “Five Country Ministerial” meeting. The statement is notable for including three paragraphs specifically about AI.
Five Eyes on AI – frontier model access: “We have committed to deepen collaboration with industry on shared national security priorities and public safety, including enabling timely access to frontier models to support secure innovation and strengthen cyber security,” the statement reads. “To support a coordinated response to artificial intelligence-related national security risks, the Five Countries discussed the national security and public safety implications of artificial intelligence models and characteristics of an artificial intelligence model that may require additional government scrutiny.”
Why this matters – from a foreseen risk to a live one: Previous Five Eyes ministerial statements have mentioned AI, but typically either as something to study, or something where they are concerned with its interplay with other areas of crime (e.g, malware, child pornography, scamming, etc). It’s very unusual for this year’s statement to have the practical focus of model access and it speaks to both the simmering geopolitical tensions around who does and doesn’t get access to this technology, as well as an acknowledgement that the intelligence services do not have their own in-house capabilities to make dependence on the private sector unnecessary.
Read more: Five Country Ministerial 2026 (Australian Government, Department of Home Affairs).
Bill Gates thinks the rise of AI will demand “an unprecedented global response”:
…Microsoft founder lays out a cautionary vision for the next few years…
By default, AI is not going to bring about happiness. That’s the basic conclusion from reading Bill Gates’s lengthy new essay about AI. The technologist and philanthropist worries that without massive work by governments, the outcomes of AI will not lead to a thriving society.
“This unprecedented technology demands an unprecedented global response. If we get it right, the payoff for humanity will be phenomenal and the world will be a more equitable place,” he writes. “In terms of equity, AI will either be the greatest equalizer ever invented, or the worst source of injustice…. I don’t see evidence that leaders, experts, and communities are confronting the challenges adequately. There is no plan to ease the entry into the AI era.”
Why AI is different: One key reason for Gates worry is the impact he expects AI to have on jobs and the economy, where he paints a vision of the technology diffusing unusually rapidly and displacing huge chunks of human labor. “Many commentators underestimate the extent of the impact AI will have,” he writes. “We have no experience with a technology that can be adopted quickly or that can think and move like a human….AI will take on work in law, customer service, medicine, software, and manufacturing. It will hit these industries rapidly, over the course of a decade rather than a few generations. There will be some new jobs, but without the right policies there will be far fewer than exist today… the jobs at most risk are entry- and mid-level, and the new jobs being created will mostly require skills that take many years to learn.”
更进一步:量化金融体系
看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力