NEW YORK / RankWire.AI / – Hachette Book Group, Cengage Learning and Elsevier have initiated legal proceedings against Google regarding its Gemini artificial intelligence platform. Author Scott Turow and his organization, S.C.R.I.B.E., have joined the class action proposal. The complaint was filed on July 10 in the U.S. District Court for the Southern District of New York. The plaintiffs accuse Google of copying millions of copyrighted books and journal articles without authorization during the development and training of Gemini. As of July 15, the court had yet to rule on the allegations or certify the class.

According to the complaint, Google acquired material via Google Books, Google Play Books, and Google Scholar. Publishers and authors provided works for specific uses, including search functionalities, sales, and research access. The plaintiffs contend that these arrangements did not permit extensive commercial AI training. They also accuse Google of downloading large web-scraped datasets containing copyrighted content. The document states that some of this material originated from known piracy sources and paywalled services.
The 57-page complaint outlines four claims under federal law. Three relate to alleged reproduction through Google services, web scraping activities, and the development or training of Gemini. The fourth claim invokes the Digital Millennium Copyright Act. The plaintiffs allege Google removed or altered copyright management information from training datasets. The filing also references internal discussions about utilizing publisher-provided books. One assessment cited estimates potential fines ranging from $10 billion to $100 billion. These allegations have not yet been tested by the court.
Proposed class includes owners of registered works
The proposed class encompasses holders of registered U.S. copyrights in qualifying books and journal articles. Eligible books must have an International Standard Book Number (ISBN). Eligible articles should bear a Digital Object Identifier or an International Standard Serial Number. The class definition includes works allegedly copied from Google services or downloaded through web scraping. It also covers works purportedly reproduced during Gemini’s development or training phases.
The timing of copyright registration also restricts membership. One criterion requires registration within five years of publication and prior to Google’s alleged reproduction or distribution. Another mandates registration within three months after publication. The complaint excludes government entities, Google affiliates, certain court participants, and individuals who properly opt out. Approval of the class status by the court is necessary before proceeding with broader claims.
Legal claims seek damages and accountings
The plaintiffs seek either statutory damages or actual damages related to proven infringements. They also request profits attributable to any confirmed copyright violations. Their relief demands include an injunction, reimbursement of legal expenses, and a jury trial. The complaint does not specify a total damages amount but asks Google to disclose Gemini training data, data collection methods, and known capabilities through a court-ordered accounting.
This accounting would identify the copyrighted works used for Gemini’s training, detailing how Google collected, copied, processed, and encoded these materials. The plaintiffs also seek court-supervised destruction of unauthorized copies under Google’s control. Earlier, Hachette and Cengage sought to join separate AI litigation against Google in California. The current New York case expands the list of plaintiffs to include Elsevier, Turow, and S.C.R.I.B.E., focusing on claims related to Google services, web scraping, and Gemini’s training process.