跳到主内容
@wquguru
精选80Exponential View(RSS)模型发布/更新多源精选 ×5

OpenAI GPT-6 Astra基准领先,ARC-AGI展现符号推理能力

🔮 Astra outruns visibility EV#600

原文
发到 X
推荐理由

GPT-6 Astra展示了从试错到符号推理的能力跃迁,且伴随关于安全对齐的争议讨论,值得从业者关注其技术路径与安全性评估。

Hi,

Welcome to our milestone 600th Sunday edition of Exponential View. Eleven years of analysis and writing about AI, every week. I’ll be in the comments for a 600th‑edition AMA. Members can post their questions on AI or the future of the economy, and I’ll do my best to answer.

Leave a comment

🎁 6️⃣0️⃣0️⃣

To celebrate 600 editions, we’re offering a limited-time discount on your annual membership: 60% off your first year. This is the biggest discount we’ll give and the lowest price you will ever get for Exponential View as we review our prices this fall. The offer is open for 24 hours, so make sure you take advantage of it.

Get 60% off for 1 year

An astral leap

OpenAI’s GPT-6 Astra leads Claude Fable 5.1 and other leading models on several benchmarks. My own experience of Astra concurs: it is a fantastic model. Right now it’s crunching away tidying the 5,932 files I had stashed in my Desktop and Download folders. (Don’t ask.) Fable 5.1 is no slouch either. It’s now speed-running useful analysis that previously took several steps and occasional intervention. One extract below:

Extract of an analysis I ran with Fable 5.1

But Astra really is very good—and mostly cheaper than the Anthropic alternative. On difficult math problems, Astra’s time horizon is 30.9 minutes vs 3.6 minutes for GPT 5.6 Sol. Mathematician Bartosz Naskręcki says: “For a mathematician it feels like finally we arrived in the era where we can focus entirely on the ideation and exploration”.

Real-world demos show a capability jump on technical and design tasks (two of my favorites are this simulated world inhabited by agents communicating and working together and 3D modeling of Zillow listings).

Astra’s performance on ARC-AGI-3 is quite interesting. Dropped into an abstract game it had never seen, it used fewer actions than the human median on 96% of the levels it completed, averaging 51.7% fewer actions per level. This goes against researchers’ original expectation that even when an AI solves an environment, it might fumble around and be less efficient than humans. But Astra invented a symbolic model to hold an entire environment in a compact notation system. In a way, it replaced trial-and-error, an enormously expensive part of discovery, with reasoning.

Astra is highly controversial. Researchers don’t seem to trust OpenAI’s claim that this is their “most-aligned model.” AI safety researcher Ryan Greenblatt, who investigated the Hugging Face incident, noted: “I do not find it encouraging to see various specific misaligned behaviors go from a high rate with GPT 5.6 to ~zero with Astra. This seems indicative of whack-a-mole / papering over specific problems rather than solving the underlying misaligned drives.”

More for paying members this week:

  • Smarter AI, fewer clues. Why the latest models leave us guessing about how they think.
  • Cancer vaccines meet the factory floor. Breakthroughs are coming. Who will supply them?
  • A bigger pie, a smaller slice. Will workers be better off under advanced AI?
  • The problem with life after work. The Versailles Court’s sobering glimpse of what happens when status becomes your job.

Upgrade to read the full analysis.

Read more

更进一步:量化金融体系

看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力

进入量化体系 →