蒸馏研究:教师模型推理习惯比分数更重要
Interesting study on model distillation.
Interesting study on model distillation.
关于模型蒸馏的有趣研究。
A distilled student model copies how its teacher reasons rather than what it knows, so teacher lineage matters far more than teacher benchmark scores.
蒸馏出的学生模型复制的是教师模型的推理方式,而非其知识内容,因此教师模型的来源比其基准分数重要得多。
A weaker teacher model built from your student's own base model will take it further than a stronger teacher from a different family.
基于学生模型自身基础模型构建的较弱教师模型,比来自不同家族的更强教师模型更能推动学生进步。
This paper finds, the data barely matters. Train on problems the teacher always solves, never solves, or a random mix, and you land in the same place. Even grade-school math gets you over 80% of the benefit.
这篇论文发现,数据几乎无关紧要。无论训练用的是教师总能解决的问题、从未解决的问题,还是随机混合的问题,结果都殊途同归。即便是小学水平的数学题,也能带来超过80%的收益。
A teacher can teach a problem it can't solve itself, because what's moving across is a habit of reasoning.
教师模型可以教授它自己无法解决的问题,因为传递的是推理的习惯。
更进一步:量化金融体系
看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力