The papers, The New York Times (NYT) and New York Daily News among them, accuse OpenAI of lying to a federal court about its ability to search its systems for evidence that it misused millions of their news stories to train its AI model, Reuters reported Thursday (July 9).
“For over two years, OpenAI lied to The Times, The Daily News Plaintiffs, the public, and the court,” the New York Times’ lead attorney Ian Crosby said in a statement to Reuters. “It claimed searching ChatGPT outputs for copies of The Times’ and the Daily News Plaintiffs’ content was infeasible, burdensome, and invasive of users’ privacy – while at the same time concealing that it had already done such searches.”
The filing also alleges that OpenAI deleted billions of relevant ChatGPT conversations or made them unsearchable. The plaintiffs are seeking sanctions including attorneys fees, as well as a finding that the company’s chat logs showed OpenAI misused the papers’ copyrighted material.
Reached for comment by PYMNTS, a spokesperson for OpenAI shared a statement rejecting the plaintiffs’ allegations.
“As the Times’ case weakens and they’ve been forced to drop claims against us, they’re persisting with their efforts to invade the privacy of people who have nothing to do with this case, including by making these blatantly false allegations,” the company said. “We’ll continue defending our users’ privacy and the long-established principles of fair use.”
The suit was filed by the NYT in 2023, accusing OpenAI and its benefactor Microsoft of using millions of its articles without consent to train OpenAI’s ChatGPT. The NYT has since dropped some of its claims, per a recent Bloomberg Law report.
As Reuters noted, this is one of many suits filed by copyright owners against AI companies for allegedly misusing their books, songs and news articles to train their systems.
So far, PYMNTS wrote last month, the judicial record on cases like these shows a divide. For example, U.S. District Judge William Alsup in San Francisco called AI training “quintessentially transformative” and said copyright law “seeks to advance original works of authorship, not to protect authors against competition.”
But U.S. District Judge Vince Chhabria, also in San Francisco, came to a different conclusion two days later, warning that widespread AI training could impede the economic incentives that drive human creative work.
Daryl Lim, H. Laddie Montague Jr. Chair in Law at Penn State Dickinson Law School, told PYMNTS that only a handful of companies can train frontier models at scale, since these firms control compute, data, cloud infrastructure, and distribution all at once.
“When you train frontier models, you need to ingest vast repositories of works that may include copyrighted works,” Lim said.