模型进步,评测基准需同步进化
As models improve, the benchmarks we use to evaluate them have to evolve too.
As models improve, the benchmarks we use to evaluate them have to evolve too.
“Create an SVG of a pelican riding a bicycle” was once surprisingly helpful. Today, frontier models routinely produce convincing results, so the test tells us increasingly little about where the actual frontier is.
Karpathys experiment is a fascinating attempt to push the benchmark forward: give Opus 5 the opening of The Lord of the Rings, a massive token budget, and two hours to turn it into an interactive Three.js world. And voila!
This tests far more than one-shot generation: The model has to maintain coherence over thousands of lines of code, translate prose into a spatial system, coordinate objects and animations, and continuously inspect its own work.
Really love where the testing / benchmarks are moving to!
Love to see GPT-5.6 and the new DeepSeek flash on this one.
更进一步:量化金融体系
看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力