跳到主内容
@wquguru
精选70Chubby♨️模型发布/更新

模型进步,评测基准需同步进化

As models improve, the benchmarks we use to evaluate them have to evolve too.

原文
发到 X

As models improve, the benchmarks we use to evaluate them have to evolve too.

“Create an SVG of a pelican riding a bicycle” was once surprisingly helpful. Today, frontier models routinely produce convincing results, so the test tells us increasingly little about where the actual frontier is.

Karpathys experiment is a fascinating attempt to push the benchmark forward: give Opus 5 the opening of The Lord of the Rings, a massive token budget, and two hours to turn it into an interactive Three.js world. And voila!

This tests far more than one-shot generation: The model has to maintain coherence over thousands of lines of code, translate prose into a spatial system, coordinate objects and animations, and continuously inspect its own work.

Really love where the testing / benchmarks are moving to!

Love to see GPT-5.6 and the new DeepSeek flash on this one.

更进一步:量化金融体系

看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力

进入量化体系 →

相似阅读

另一事件,读法相近