The New York Times and a group of news publishers filed a sanctions motion on July 9 accusing OpenAI of a deliberate, two-year effort to hide evidence in their copyright lawsuit. The case, originally filed in late 2023, accuses OpenAI and Microsoft of using millions of news articles to train ChatGPT without permission or payment.
The central allegation is straightforward. OpenAI told the court it lacked the tools to search its training data and chat logs for the publishers' copyrighted content. Then, during a court-ordered deposition in April, an OpenAI data engineer testified that the company had already conducted internal searches of that same data and had quietly assembled a database of roughly 78 million anonymized conversations it was using to measure how much it was reproducing others' work.
That contradiction is serious. Under court rules, a party in a lawsuit must preserve and hand over evidence. If a company tells a judge it cannot do something and then a witness reveals it could, the judge has broad power to punish that company, including by assuming the hidden evidence would have hurt them.
The publishers are asking for exactly that. They want the judge to declare, as an established fact, that the chat logs would have shown ChatGPT reproducing their journalism at scale. They also want OpenAI barred from using the 20 million logs it did eventually supply, arguing those were so heavily redacted the court already called them "unusable." The publishers also claim OpenAI deleted billions of conversations after the lawsuit was filed, in direct violation of a preservation order.
OpenAI denies the accusations. Its spokesperson called the allegations "blatantly false" and framed the fight as an attempt by the Times to access private user conversations as its case weakens.
The underlying legal question for the whole AI industry is whether using copyrighted content to train an AI tool counts as "fair use," a legal concept that allows limited copying without payment for purposes like research or commentary. OpenAI's position is that training an AI is transformative: the tool learns patterns, it does not store copies. The Times argues the opposite, pointing to cases where ChatGPT reproduced its articles nearly word for word when prompted.
This matters well beyond the media industry. More than 50 copyright lawsuits are currently pending against AI companies in the US, covering books, music, visual art, and software. A ruling that forces OpenAI to pay licensing fees for training data would set a cost baseline that every AI company would eventually face. Anthropic already settled a similar case brought by book authors for $1.5 billion.
The sanctions motion is not the final verdict, but it is a signal of how the case is going. Courts rarely reach the sanctions stage unless one side has genuinely exhausted the judge's patience. If the judge agrees with the publishers, OpenAI will enter the trial with its key evidence disqualified and a formal finding of misconduct already on record. That is a very difficult position from which to win.
For any business that uses AI tools built on publicly scraped content, this case is worth following. The outcome will determine whether licensing fees become a structural cost of every AI product, or whether the current free-training model holds.