跳到主内容
精选70Przemek Chojecki | PC模型发布/更新多源精选 ×9

GPT-5.6 Sol Ultra 测试:局部推理显著提升,全局连贯性进步缓慢

What stands out from my tests with GPT-5.6 Sol Ultra is how good it is at concis…

原文

What stands out from my tests with GPT-5.6 Sol Ultra is how good it is at concise elegant constructions and finding ways to expand and enhance existing arguments.

It feels different from GPT-5.5 Pro which excelled at finding counterexamples.

A shared weakness of both models is writing longer papers over longer human-AI interactions. It usually becomes a mess: either by plastering new results on top of the previous ones making a paper into a brain dump or cutting too many details so that the essence is gone.

However the coherence limit has moved. Previously anything over 20-pages long would automatically contain too many errors to be useful. Right now the model can write coherent 30-40 pages of mathematics - at least when it comes to some domains/arguments and over multiple prompts. But when it comes to heavy analytical computations and bookkeeping, the limit is still around 12-15 pages.

So in a way GPT-5.6 Sol has gotten significantly smarter at local-scale thinking, but its progress on global-scale thinking is much slower.

更进一步:量化金融体系

看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力

进入量化体系 →

关联讨论

同一事件的更多信源
GPT-5.6 在 Bedrock 上正式可用
Greg Brockman原文
GPT-5.6 完成近千行复杂编码任务
OpenAI Developers原文
OpenAI 澄清 ChatGPT Work 云与桌面数据隔离
Simon Willison 博客(RSS)原文

相似阅读

另一事件,读法相近