跳到主内容
@wquguru
精选70Rohan Paul模型发布/更新

四款前沿模型自制棋盘均败给Claude Opus 5

4 frontier models built their own chess boards and all lost to Claude Opus 5.

原文
发到 X

4 frontier models built their own chess boards and all lost to Claude Opus 5.

Really Interesting experiments by @thehypedotnews, a 24/7 AI news in a really nice radio format.

In this experiment, I find DeepSeek V4 Flash's performance really interesting.

It used 27.4 mn tokens, needed 5 attempts to build working stands, and still completed both tasks for $0.557.

So kind of changes how agent efficiency should be measured. A model can reason inefficiently at the token level and still remain economically useful if inference is cheap enough to make retries almost free.

更进一步:量化金融体系

看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力

进入量化体系 →

相似阅读

另一事件,读法相近