跳到主内容
@wquguru
精选75Simon Willison 博客(RSS)论文研究

AI实验室是否刻意训练模型画骑自行车鹈鹕?研究结论:否

Are AI labs pelicanmaxxing?

原文
发到 X

Are AI labs pelicanmaxxing?

Excellent piece of work by Dylan Castillo, who took a deep-dive into the frequently pondered question of whether the AI labs have been deliberately training models to draw pelicans riding bicycles in response to my deeply unscientific benchmark.

I've been randomly spot-checking this in the past by testing models against other animals riding other types of vehicle, but never with anything close to the diligence of Dylan's methodology here.

Dylan took 8 animals × 6 vehicles = 48 prompts and ran them three times each through 7 different models ( GPT-5.6 Terra, Claude Sonnet 5, Gemini 3.5 Flash, Grok 4.5, Qwen3.7-Max, GLM-5.2, and DeepSeek V4 Pro). He then used GPT-5.6 Luna and Gemini 3.1 Flash-Lite to help evaluate the results.

There's a neat filter view for exploring the results:

For the models he tested he could find no evidence of pelimaxxing:

  • The pelicans on bicycles don’t look any better
  • Labs are not better at drawing pelicans
  • Labs are not better at drawing bicycles
  • Labs are not better at drawing pelicans on bicycles, even adjusting for difficulty
  • The pelican-bicycle scenes don’t look memorized [...]

Pelicans aren’t drawn any better than other animals. Bicycles aren’t drawn any better than other vehicles. And no lab draws the combination better than its pelicans and bicycles already predict. GLM-5.2 comes closest: it has the largest boost on the exact pelican-bicycle cell, and and its first pelican-on-bicycle sample caught my eye. But the effect is small and not significant, so I wouldn’t put too much weight on it.

Via Hacker News

Tags: ai, generative-ai, llms, evals, pelican-riding-a-bicycle

更进一步:量化金融体系

看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力

进入量化体系 →

相似阅读

另一事件,读法相近