3B模型在可验证推理任务上达到竞品级水平
Crazy: A 3B model is now reaching highly competitive results on verifiable reaso…
做推理模型或小模型方向的同学注意了,3B 模型在可验证推理上逼近前沿水平,说明推理能力可以大幅压缩。建议关注其训练方法,可能改变小模型部署策略。
Crazy: A 3B model is now reaching highly competitive results on verifiable reasoning tasks.
VibeThinker-3B scores 94.3 on AIME26, 80.2 Pass@1 on LiveCodeBench v6, and 96.1% on unseen LeetCode contests.
The gains appear to come primarily from post-training on top of Qwen2.5-Coder: curriculum SFT, multi-domain RL, offline self-distillation, and a final RL-based instruct stage.
The core implication: certain forms of verifiable reasoning may be highly compressible into small dense models.
Frontier-scale models still matter for broad knowledge and general-purpose capability, but compact reasoning models are becoming a serious complementary path.
Love to see it!
更进一步:量化金融体系
看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力