跳到主内容
@wquguru
精选70The Decoder(RSS)模型发布/更新

新基准显示多模态AI视觉感知能力仍不足

New benchmark confirms AI models still perform poorly at visual perception

原文
发到 X

Moonshot AI's PerceptionBench tests how well multimodal AI models can actually "see," separate from logical reasoning. No frontier model reaches 60 percent accuracy, and GPT-5.6 Sol leads by a narrow margin. Many supposed reasoning errors actually happen as early as the image-reading stage.

Moonshot AI的PerceptionBench测试多模态AI模型实际“看见”的能力,与逻辑推理分开。没有前沿模型达到60%的准确率,GPT-5.6 Sol以微弱优势领先。许多所谓的推理错误实际上早在图像读取阶段就已发生。

The article New benchmark confirms AI models still perform poorly at visual perception appeared first on The Decoder.

文章《新基准证实AI模型在视觉感知方面仍表现不佳》首次出现在The Decoder上。

更进一步:量化金融体系

看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力

进入量化体系 →

相似阅读

另一事件,读法相近