Publishers Argue AI Firms Should Pay for Content as Copyright Case Advances

News Room
6 Min Read

A host of copyright infringement lawsuits filed by news organizations against the ChatGPT developer OpenAI are heading toward a new phase, and the Trump administration’s intervention is raising the question of just how costly it would be for AI companies to pay to license the content their models are trained on.

The cases include those filed by The New York Times and by Ziff Davis, which owns CNET. The parties filed briefings this month asking a federal court to rule on a handful of issues, namely whether OpenAI’s use of the published material constituted fair use under US copyright law. 

But a statement of interest from the federal government, saying that licensing barriers would burden AI companies, drew a rebuke out of court from Ziff Davis CEO Vivek Shah. Writing in Fortune, Shah compared the total value of royalties and licensing payments in the US music industry — under $20 billion per year — with the $750 billion that OpenAI has told investors it expects to spend on computing infrastructure by 2030. 

Something like $20 billion a year for licensing to news publishers would be “a rounding error” for AI firms but would have significant meaning for publishers, Shah wrote. “A sensible fee structure creates a flywheel of quality inputs and quality outputs, benefiting the AI consumer and the public good.”

A representative for OpenAI pointed to the company’s blog post on the litigation, which notes its partnerships with journalism organizations. “AI systems are a force for good in journalism and everyday Americans,” the blog post says. 

We’re All Copyright Owners. Welcome to the Mess That AI Has Created

The op-ed comes as a lengthy discovery process wraps up in the copyright cases. For the purpose of discovery, a variety of separate suits against OpenAI and Microsoft were consolidated before one judge. Now the parties have taken the evidence gathered over the past few years and presented initial arguments to the court, which could rule on certain key questions before sending the individual cases back to their respective districts for potential trials.

In a joint brief (PDF), the news publishers — including Ziff Davis, The New York Times and the New York Daily News — argued that the use of their published content to train ChatGPT has undermined web traffic and diluted the market with AI-generated “pink slime” content.

“There is no doubt about ‘what would happen’ if Defendants’ conduct were to become widespread and unrestricted — because it already is happening,” the publishers argued. “One effect of ChatGPT’s release in November 2022 was to spark other companies like Google to fast-track their own GenAI products, including, most significantly, Google’s AI Overviews.”

Read more: Google’s AI Overviews ‘Misconduct’ Undermines Publishers, Lawsuit Says

OpenAI claimed in its filing (PDF) that its use of data scraped from the internet constituted fair use because it was “highly transformative, the works are highly factual, and pretraining has caused no cognizable harm to the market for or value of those works.”

As for licensing, OpenAI argued that its access to publishers’ context was implicitly licensed because the publishers didn’t, at the time, block the company’s scrapers using robots.txt (which tells bots whether they can access a page) or other means. 

“For decades, search engines like Google — which have built web indexes by crawling (i.e., copying) virtually every webpage on the Internet — and other web-based products and services have relied on scraping or crawling the Internet,” OpenAI argued. “Indeed, News Plaintiffs use web scrapers of their own. This copying has long been understood to be impliedly licensed unless the website owner opts out via mechanisms like the robots.txt protocol.”

Many publishers now block bots from OpenAI and other AI firms from accessing their sites. In his op-ed, Shah wrote that those kinds of barriers and paywalls will grow more onerous, limiting access to information and hindering the open web. “All because we allow the false assumption that licensing is hard and expensive to go unchallenged.”

Read the full article here

Share This Article
Leave a comment

Leave a Reply

Your email address will not be published. Required fields are marked *