Netflix 揭秘推荐理由生成:AI 写、AI 评、人工周审
Netflix has explained the system behind those short "because you watched" lines:…
Netflix has explained the system behind those short "because you watched" lines: an AI writes them, another AI grades them, and humans audit weekly.
Netflix已解释了那些简短的“因为您观看了”提示背后的系统:一个AI撰写它们,另一个AI评分,人类每周进行审核。
An AI judge is usually validated once and then trusted forever.
AI评判者通常经过一次验证后就被永久信任。
New Netflix paper argues a judge running in production has a lifecycle and needs maintaining like any other model.
Netflix的新论文认为,在生产环境中运行的评判者有其生命周期,需要像其他模型一样进行维护。
At Netflix, one model writes the short lines telling members why a title was recommended, and another scores every one before it is shown.
在Netflix,一个模型撰写简短说明,告诉会员为何推荐某个标题,另一个模型则在每条说明展示前对其进行评分。
Netflix splits that work into four stages: building labelled examples, tuning the judge, running it as a gate, and watching it for drift.
Netflix将这项工作分为四个阶段:构建标注示例、调整评判者、将其作为门控运行、以及监控其漂移。
Checking whether the judge agrees with human labels is not enough.
仅检查评判者是否与人类标注一致是不够的。
It also has to reject a bad explanation for the same reason a person would, so tuning runs on written rationales rather than pass-fail marks.
它还必须以人类会拒绝的理由拒绝糟糕的解释,因此调整基于书面理由而非通过/失败标记进行。
A weekly human review then sets the bar by how far the raters disagree among themselves.
每周的人工审核随后根据评分者之间的分歧程度设定标准。
In a 5-week test against no explanation at all, members shifted slightly toward titles they had not watched and more often ended a browse by playing something.
在为期5周的测试中,与完全没有解释相比,会员略微倾向于观看未看过的标题,并且更频繁地以播放某内容结束浏览。
– arxiv. org/abs/2608.18300
– arxiv.org/abs/2608.18300
Title: "The Lifecycle of LLM-as-a-Judge for Large-Scale Recommendation Explanations"
标题:“大规模推荐解释中LLM作为评判者的生命周期”
更进一步:量化金融体系
看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力