A group of prominent publishers and an author have filed a class action lawsuit against Google, accusing the technology company of illegally using millions of copyrighted books to train its Gemini artificial intelligence models. The plaintiffs, including Hachette Book Group, Cengage Learning, and Elsevier, along with best-selling author Scott Turow, contend that Google copied these works without obtaining the necessary permissions or providing compensation. The lawsuit describes the alleged actions as "one of the most prolific infringements of copyrighted materials in history".
The legal action, filed in the U.S. District Court for the Southern District of New York, centers on Google's use of books that were provided for specific services such as Google Books, Google Play Books, and Google Scholar. The publishers argue that these agreements allowed for limited uses, such as displaying search snippets or selling e-books, but did not grant Google the right to copy the entire works for training commercial AI products. The complaint states that Google "reproduced millions of copyrighted works without permission, without providing any compensation to authors or publishers, and with full knowledge that its conduct violated copyright law".
The lawsuit also claims that Google removed or altered copyright management information from the works to obscure their use in AI training. According to internal Google documents cited in the filing, the company was aware of the legal risks involved. One document reportedly flagged that using "publisher provided copyrighted books" from Google Play Books in connection with AI development was "highly problematic" and warned of potential fines ranging from $10 billion to $100 billion. Other internal analyses identified issues such as "Restrictive licenses," "Publishers are sensitive about training on their data," and a "Heightened risk around fair use defenses".
The plaintiffs assert that Google's actions have harmed authors and the publishing industry. They argue that AI-generated content can directly compete with and negatively impact book sales, with Gemini capable of producing works that substitute for original copyrighted material. The lawsuit highlights that Gemini can generate content such as a 100-page murder mystery quickly and at a low cost, a capability that stems from Google's alleged use of copied works for training.
This case adds to a growing number of legal disputes between content creators and artificial intelligence companies over copyright infringement. Authors and publishers have filed similar lawsuits against other major AI developers, including OpenAI, Anthropic, and Meta. In one instance, Anthropic reached a settlement of $1.5 billion with authors over allegations of using pirated books to train its AI chatbot Claude. However, previous court rulings in California have sided with AI companies, determining that training on copyrighted materials could qualify as "fair use" under U.S. law.
The plaintiffs are seeking statutory damages, a permanent injunction to prevent further infringement, and a court order for Google to destroy any unauthorized copies of their works used in AI training. Google had not publicly responded to the allegations at the time of reporting.
