Publishers have filed a class action lawsuit against Google in a federal court in New York, accusing the company of copying millions of books to build its Gemini AI without authorisation or payment. The named plaintiffs are Hachette Book Group, Cengage Learning, Elsevier, and bestselling author Scott Turow.
The core argument is straightforward and pointed. Publishers gave Google access to their books years ago, for specific, agreed purposes: to let users search for books and see short extracts on Google Books, to sell digital copies on Google Play, and to support academic research through Google Scholar. The publishers say Google then took those same books and used them to train Gemini, a completely different and commercially competing product. That, they argue, was never part of the deal.
What makes this suit different from others is the internal evidence. According to court documents, Google's own files warned that using books from Google Play Books for AI training was "highly problematic" and could expose the company to fines of between $10 billion and $100 billion. Google proceeded anyway. The lawsuit also claims Google removed or altered copyright information on those books, which is a separate legal violation under US law.
The legal context is not straightforward. In June 2025, two California judges ruled, in separate cases against Anthropic and Meta, that training an AI on copyrighted books can be considered "fair use," the legal concept that allows someone to use protected work without permission in certain circumstances. Both rulings, however, were narrow. The judges said their conclusions rested heavily on the specific evidence presented by the plaintiffs, and both noted that stronger evidence of market harm could have led to different outcomes. Neither ruling binds courts in New York, where this case will be decided.
There is also a live benchmark for what losing looks like. In September 2025, Anthropic settled a copyright case for $1.5 billion, the largest copyright payout in US history, covering roughly 500,000 books at about $3,000 each. Anthropic paid that sum specifically because it had used books downloaded from pirate websites, not books obtained through legitimate channels. Google's situation is arguably worse: it obtained its books through legitimate partnerships, then allegedly used them for something beyond what those partnerships allowed.
And while other AI companies have been signing licensing deals with publishers, Google has largely refused to. OpenAI, Meta, Microsoft, Amazon, and Anthropic have all reached licensing agreements with major publishers or news organisations over the past year. Google signed a deal with the Associated Press for news content, but has struck no broad licensing agreements with book publishers. Meanwhile, Google is reportedly pressuring news publishers to hand over AI training rights as a condition of staying in its existing paid partnership programmes.
Publishers are asking the court for three things: financial damages, an order permanently stopping Google from continuing the alleged copying, and a requirement that Google destroy all unauthorised copies of their works currently in its systems. That last point matters: it would force Google to potentially retrain parts of Gemini from scratch, at enormous cost.
For anyone whose business produces, licenses, or relies on written content, the direction of travel is clear. AI companies need vast amounts of text to build and improve their systems. The question of who owns that text and who gets paid for it is now being answered one lawsuit at a time. The Anthropic settlement set a floor. This case against Google may set the ceiling.