跳到主内容
@wquguru
精选75Rohan Paul论文研究

Netflix 揭秘推荐理由生成:AI 写、AI 评、人工周审

Netflix has explained the system behind those short "because you watched" lines:…

原文
发到 X

Netflix has explained the system behind those short "because you watched" lines: an AI writes them, another AI grades them, and humans audit weekly.

Netflix已解释了那些简短的“因为您观看了”提示背后的系统:一个AI撰写它们,另一个AI评分,人类每周进行审核。

An AI judge is usually validated once and then trusted forever.

AI评判者通常经过一次验证后就被永久信任。

New Netflix paper argues a judge running in production has a lifecycle and needs maintaining like any other model.

Netflix的新论文认为,在生产环境中运行的评判者有其生命周期,需要像其他模型一样进行维护。

At Netflix, one model writes the short lines telling members why a title was recommended, and another scores every one before it is shown.

在Netflix,一个模型撰写简短说明,告诉会员为何推荐某个标题,另一个模型则在每条说明展示前对其进行评分。

Netflix splits that work into four stages: building labelled examples, tuning the judge, running it as a gate, and watching it for drift.

Netflix将这项工作分为四个阶段:构建标注示例、调整评判者、将其作为门控运行、以及监控其漂移。

Checking whether the judge agrees with human labels is not enough.

仅检查评判者是否与人类标注一致是不够的。

It also has to reject a bad explanation for the same reason a person would, so tuning runs on written rationales rather than pass-fail marks.

它还必须以人类会拒绝的理由拒绝糟糕的解释,因此调整基于书面理由而非通过/失败标记进行。

A weekly human review then sets the bar by how far the raters disagree among themselves.

每周的人工审核随后根据评分者之间的分歧程度设定标准。

In a 5-week test against no explanation at all, members shifted slightly toward titles they had not watched and more often ended a browse by playing something.

在为期5周的测试中,与完全没有解释相比,会员略微倾向于观看未看过的标题,并且更频繁地以播放某内容结束浏览。

– arxiv. org/abs/2608.18300

– arxiv.org/abs/2608.18300

Title: "The Lifecycle of LLM-as-a-Judge for Large-Scale Recommendation Explanations"

标题:“大规模推荐解释中LLM作为评判者的生命周期”

更进一步:量化金融体系

看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力

进入量化体系 →

相似阅读

另一事件,读法相近