3D编码测试:DeepSeek-V4-Pro 令牌消耗为 Muse Spark
This 3D coding test by @thehypedotnews found DeepSeek-V4-Pro-0813 used 48X more…
This 3D coding test by @thehypedotnews found DeepSeek-V4-Pro-0813 used 48X more tokens than Muse Spark 1.2.
这个由@thehypedotnews发布的3D编码测试发现,DeepSeek-V4-Pro-0813使用的令牌数是Muse Spark 1.2的48倍。
thats why for agentic coding, the model's ability to reach an answer matters almost as much as the answer itself. A model that keeps reopening files, reconsidering decisions, and resending context can turn a small build into a huge inference loop.
这就是为什么在代理式编码中,模型达成答案的能力与答案本身几乎同等重要。一个不断重新打开文件、重新考虑决策并重新发送上下文的模型,可能将一个小型构建变成巨大的推理循环。
The test used Nous Research's Hermes Agent CLI through OpenRouter, identical prompts, and the same Three.js constraints across three voxel-city scenes.
该测试通过OpenRouter使用了Nous Research的Hermes Agent CLI,在三个体素城市场景中使用了相同的提示和相同的Three.js约束。
Prompt caching can suppress billing without fixing the latency and retry burden created by a call-heavy agent loop.
提示缓存可以抑制计费,但无法解决由调用密集的代理循环造成的延迟和重试负担。
- total cost #1 muse spark 1.2 – $0.53 #2 gemini 3.7 flash – $0.56 #3 deepseek v4 pro – $4.57
- 总成本 #1 muse spark 1.2 – $0.53 #2 gemini 3.7 flash – $0.56 #3 deepseek v4 pro – $4.57
更进一步:量化金融体系
看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力