跳到主内容
@wquguru
精选75Rohan Paul行业动态

wikiHow 起诉 OpenAI 未经许可抓取逾 1.1 万篇文章训练模型

Reuters: wikiHow sued OpenAI, alleging it scraped more than 11,000 articles to t…

原文
发到 X

Reuters: wikiHow sued OpenAI, alleging it scraped more than 11,000 articles to train GPT models without permission.

路透社:wikiHow 起诉 OpenAI,指控其未经许可抓取超过11,000篇文章用于训练GPT模型。

The complaint says that copying infringed at least 1,200 registered copyrights and fed models that can reproduce wikiHow text.

诉状称,这种复制行为侵犯了至少1,200项已注册版权,并喂养了能够复现wikiHow文本的模型。

This broadens the dispute beyond training because wikiHow says ChatGPT answers can substitute for the how-to pages that supplied the material.

这使争议超出了训练范畴,因为wikiHow表示ChatGPT的答案可以替代提供素材的指南页面。

OpenAI says its models use publicly available data under fair use, which also considers whether unlicensed use harms the original work's market.

OpenAI表示,其模型在合理使用原则下使用公开数据,该原则也考虑未经许可的使用是否损害原作品的市场。

The case was filed Aug 21 in the Southern District of New York and remains at the complaint stage, with no ruling on wikiHow's allegations.

该案于8月21日在纽约南区法院提起,目前仍处于诉状阶段,尚未对wikiHow的指控作出裁决。

By alleging both ingestion and substitution, wikiHow is connecting training provenance directly to the market-effect question courts consider under fair use.

通过同时指控数据摄取和内容替代,wikiHow将训练来源直接与法院在合理使用原则下考虑的市场影响问题联系起来。

The nature of wikiHow’s content gives OpenAI a meaningful defense even if ChatGPT answers the same questions. Copyright protects original expression, but U.S. law does not protect the underlying procedure, process, method, or fact described in an article. A model can therefore explain how to perform a task without automatically infringing the article containing similar instructions.

wikiHow内容的性质为OpenAI提供了有力的辩护,即使ChatGPT回答相同的问题。版权保护原创表达,但美国法律不保护文章中描述的基础程序、过程、方法或事实。因此,模型可以解释如何执行任务,而不自动侵犯包含类似说明的文章。

更进一步:量化金融体系

看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力

进入量化体系 →

关联讨论

同一事件的更多信源

相似阅读

另一事件,读法相近