AI评估进入国家安全决策:ChinaTalk发起征文竞赛
AI Evals for the Situation Room, Explained
Over the past few years, we’ve seen hints of policymakers and national leaders using AI models in their actual policy decision-making. Senior leadership across the world has started using AI not just for tactical or operational tasks, but increasingly for broad strategic decision-making.
While there’s the promise of uplift and smarter calls on some of the most consequential decisions leaders face in foreign policy and national security, we’re also flying blind. Enormous effort and energy goes into benchmarking and evaluation for tasks like coding. The experiments you can run to make models better at software development are much easier to execute and much lower-stakes than running a real experiment when you’re deciding whether to invade a country or sign a treaty.
That’s why we at ChinaTalk are trying to kickstart a field aimed at helping researchers and policymakers understand exactly what they’re working with when they ask these models to support the most consequential decisions nations face. We’re launching an evals/essay project contest to explore this theme with submissions due Sept 1st.
I’ve brought on two expert AI eval creators to discuss why the field is important, what interesting work has already been done on how models approach broad national-security and strategic questions, and how you — as an eval professional, semi-professional, or just a concerned person — can contribute new ways to poke and prod at these models and see what they can really do.
Joining us today: Florian Brand, research engineer at Prime Intellect, and John Chen, professor at the University of Arizona, who’s done some pretty wild things getting models to start nuclear wars with each other in Civ V.
Our conversation covers:
- How frontier labs are hitting eval-building limits: they’ve gone from undergrads to PhDs to field experts, and the models are now catching the experts’ mistakes.
- How Civilization V exposes AI’s strategic blind spots (terrible second-order reasoning) and models’ distinct strategic personalities (Claude’s really into science!).
- Why ethical prompting in Civilization V still doesn’t stop models from launching nukes.
- Jordan's PresidentBench eval, where a Chinese model was nonchalant about a Taiwan invasion, while Claude wanted to keep Taiwan free and independent.
- Advice for designing better AI evals and ChinaTalk’s new essay/evals contest!
Listen now on your favorite podcast app.
Why Evals Matter and Why They’re Getting Harder to Build
Jordan Schneider: Florian, what is an AI eval, and why do they matter?
Florian Brand: AI evals try to put a number on something we want to measure. That can be knowledge questions, which we can measure with multiple choice, or — the current frontier — agents, where we want to see how good Claude Code is at implementing rather complicated codebases. Over the last few years, evaluations have gotten more professional and more complicated to reflect how users actually use these products.
Jordan Schneider: Let’s explore that. We started in a world where you’d make up hard science or math problems, and slowly but surely the models got better at them. The problem is, once you get to questions where there isn’t a right answer, you’re not just seeing how many points out of a 100 a model would score — you’re evaluating it almost the way you’d evaluate a hire. There are much softer things you’re trying to get at now, right?
Florian Brand: Definitely. We’re also increasingly hitting the limits of how we build these evaluations. A few years ago, we went to some undergrads and tortured them into annotating questions. Then we found PhD students. Now even that isn’t sufficient — we’re seeking out experts in their respective fields to come up with the most complicated questions they can think of. It turns out these experts are sometimes ever so slightly wrong. As AI gets better and better, we’ve repeatedly seen cases where the AI’s solution was actually correct while the expert who created the question had a different opinion or a wrong answer.
Jordan Schneider: So a math professor writes a problem thinking the answer is 2, and the model figures out the answer is actually 2.5. Once you’re at that point, you think, “All right, we get it — they’re really freaking good at math.” What are the other dimensions? There was the deep knowledge-query approach of Humanity’s Last Exam — very deep-cut questions about biophysics and Roman history — which we’re not currently saturated on. Models have a lot of factoids in them. What other dimensions is the field particularly focused on?
Source.
Florian Brand: The factoids actually aren’t interesting to the broader ecosystem anymore. Because AI is used for productive work — coding, filling out Excel sheets, doing your taxes — we need to find out how good the AIs are at filing your taxes. Naturally, we build evaluations to represent exactly that, see how well models perform, and then train models to become better and better in those areas.
Jordan Schneider: Before we get to the public-good case for policy and national-security evaluations, what’s the business case? Why do labs all around the world put so much stake in having thoughtful evals?
Florian Brand: You need evaluations to measure how good your model is at different capabilities. Without evals, you’d just blindly throw the model out into the world and make people figure out what it’s good at. That’s a hard sell — telling everyone to just use it and find out where the borders are. That’s why we need good evaluations that are realistic enough for people to care about.
Jordan Schneider: Can you talk about perhaps the most famous eval — METR’s “how long can it work on a problem” measure?
Florian Brand: The METR evaluation is very misunderstood. What they did, a long time ago, was take a bunch of coding problems and annotate them: for this one, a human needs two minutes; for this kind of task, a human probably needs fifteen minutes; and these tasks are super hard — a human needs two days of coding to solve them. Then the models are asked to solve these problems, and the better they are — especially on the harder problems — the longer the tasks they can complete. But this is different from Claude literally running 20 or 30 hours straight. METR is measuring that Claude is able to solve a task a human would probably need 20 hours for.
The METR graph’s biggest problem — because the evaluation is a bit older — is that they hadn’t thought to include many of the really hard tasks. The right tail is sparse, which means that as soon as models become able to solve those tasks, you get gigantic error bars.
Jordan Schneider: What evaluations have you seen that get at “can this model do something economically productive” in a smarter way than our friends at METR?
Florian Brand: Nowadays, we throw models into production systems, let them interact with file systems, and have them create Excel sheets with GDPval or APEX agents. We also push them to do science — ML research — and longer and longer running tasks, to see how well they perform at activities that actually generate revenue.
AI Strategic Thinking in Civ V
Jordan Schneider: Let’s take a little detour into the AI that runs its own store in San Francisco. What is Vending-Bench ?
Florian Brand: Vending-Bench is a setup where the model gets a virtual business — they’ve also done it in person. It’s told to manage inventory, interact with customers, and find deals with suppliers to generate revenue over a certain period of time. There are also variations of Vending-Bench where models play against each other, trying to undercut one another and strike deals to earn more revenue over the few days they’re tasked to run in the simulation.
更进一步:量化金融体系
看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力