GPT-6 Astra在ErdosBench数学基准测试中排名第一
GPT-6 Astra xhigh scored 1st on our ErdosBench, a benchmark of 226 open math pro…
GPT-6系列新模型发布并展示能力突破,对关注前沿模型进展的从业者有参考价值。
GPT-6 Astra xhigh scored 1st on our ErdosBench, a benchmark of 226 open math problems inspired by Erdos problems.
GPT-6 Astra xhigh 在 ErdosBench 上排名第一,该基准测试包含 226 道受埃尔德什问题启发的开放数学问题。
Interestingly, it's not that big of a leap compared to GPT-5.6 Sol xhigh. Astra had a couple more problems settled than Sol with stronger scientific writing, less overclaiming (even underclaiming sometimes). A solid 5%-10% gain on various math-research skills tested.
有趣的是,与 GPT-5.6 Sol xhigh 相比,提升幅度并不大。Astra 解决的题目略多于 Sol,科学写作能力更强,夸大其词的情况更少(有时甚至有所保守)。在各种数学研究技能测试中取得了 5%-10% 的稳健提升。
Also great to see, that ErdosBench is far from saturated. Bring on GPT-7.
同样令人欣喜的是,ErdosBench 远未饱和。期待 GPT-7 的到来。
更进一步:量化金融体系
看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力