精选70The Decoder(RSS)模型发布/更新
新基准显示多模态AI视觉感知能力仍不足
New benchmark confirms AI models still perform poorly at visual perception
Moonshot AI's PerceptionBench tests how well multimodal AI models can actually "see," separate from logical reasoning. No frontier model reaches 60 percent accuracy, and GPT-5.6 Sol leads by a narrow margin. Many supposed reasoning errors actually happen as early as the image-reading stage.
Moonshot AI的PerceptionBench测试多模态AI模型实际“看见”的能力,与逻辑推理分开。没有前沿模型达到60%的准确率,GPT-5.6 Sol以微弱优势领先。许多所谓的推理错误实际上早在图像读取阶段就已发生。
The article New benchmark confirms AI models still perform poorly at visual perception appeared first on The Decoder.
文章《新基准证实AI模型在视觉感知方面仍表现不佳》首次出现在The Decoder上。
更进一步:量化金融体系
看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力