Regulation

Microsoft Tells Court Copilot Rarely Reproduces Books in AI Copyright MDL

mm
Add Unite.AI to your preferred sources on Google

Microsoft moved for summary judgment on September 4, 2026 in the consolidated copyright litigation brought by book authors and news publishers over the training of large language models, telling a federal court in Manhattan that an expert’s review of 8.2 million Copilot conversations found only 24 responses containing at least 30 words matching the authors’ asserted works.

The figure comes from an analysis described in Microsoft’s memorandum of law supporting its summary judgment motion in the Books plaintiffs’ consolidated cases, part of the multidistrict litigation pending before Judge Sidney H. Stein in the U.S. District Court for the Southern District of New York. The Authors Guild and named fiction and nonfiction authors filed the suit roughly ten months after OpenAI launched ChatGPT to the public in November 2022, later adding Microsoft as a defendant.

What the Copilot Log Analysis Found

According to the memorandum, the book authors’ expert, Dr. Shawn Shan, used an adversarial extraction protocol of approximately 5.3 million attempts that fed the GPT models verbatim book passages hundreds of words long, and fewer than 1 percent of those attempts produced any 30-word match. Microsoft characterized the exercise as one in which the expert selected a target passage and repeatedly prompted the model until it emitted the desired text.

Outside that laboratory setting, the brief states, Shan reviewed 8.2 million conversations from Microsoft’s Copilot product and identified 24 responses containing 30 matching words — a rate the filing describes as .00029 percent. For 202 of the 212 asserted books, the memorandum says, the expert found no regurgitation in the logs at all. “That hardly undermines the transformative purpose of LLM training,” the filing concludes.

Microsoft also cited the authors’ own sales evidence, stating that the plaintiffs are selling just as many books as they would have had ChatGPT and Copilot never been released and that they found virtually no examples of real-world outputs matching their works. The brief quotes playwright David Henry Hwang stating he does not believe the defendants’ LLMs are hurting the market for his plays, and says sales analyses show no drop attributable to the models’ arrival.

The Fair Use Argument

The central legal claim in Microsoft’s filing is that training an LLM on copyrighted books is fair use as a matter of law. The brief argues the use is “profoundly transformative,” pointing to rulings in two other AI training cases — one against Meta and one against Anthropic — that described LLM training as highly transformative. Books are created to be read, the filing states, while OpenAI’s use of texts such as John Grisham’s The Street Lawyer and Stacy Schiff’s Cleopatra served a technological purpose: building a model that generates natural-language responses to prompts.

On the authors’ focus on OpenAI’s downloading of books from Library Genesis, an online repository known to contain unauthorized copies, Microsoft stated the undisputed record establishes it did not download and had no involvement in acquiring those datasets, and argued the provenance of training data does not alter the fair use analysis because every step of the process served the same training purpose.

The filing also addresses Microsoft’s provision, at OpenAI’s request, of portions of the Bing search engine index. The memorandum states the evidence is inconclusive as to whether those portions contained any of the plaintiffs’ works, noting the plaintiffs’ expert identified between 8 and 24 asserted works in transferred portions of the index, and argues the transfer was part of LLM training regardless.

On market harm, Microsoft urged the court to reject two theories it attributes to the plaintiffs: lost licensing fees for training data, and “market dilution” from AI-assisted books competing with the authors’ own. Research cited in the filing found the overwhelming majority of consumers would not purchase an AI-generated book even at a drastically lower price, according to the company’s expert declaration. The brief adds that any licensing requirement for training data would be an enormous barrier to innovation, given that developers must obtain works by millions of authors.

Microsoft simultaneously moved for judgment on the pleadings on the plaintiffs’ contributory infringement count, according to the same memorandum, which states that finding LLM training to be fair use would independently foreclose that claim. The company also filed a separate motion for summary judgment in the News plaintiffs’ consolidated cases the same day; its Books brief describes the News filing as explaining why the specific LLM-based tool the publishers challenge is lawful.

What Comes Next in the MDL

The filings landed in In re: OpenAI, Inc. Copyright Infringement Litigation, the multidistrict proceeding combining the authors’ and publishers’ cases. Docket records show The New York Times Company filed its own summary judgment motions on September 4, 2026, with responses due by October 5, 2026, and OpenAI filed a summary judgment motion on the News plaintiffs’ claims the same day.

A stipulated sealing order signed by Judge Stein on September 3, 2026 sets the public-release calendar for the briefing: parties and third parties must file any requests to maintain redactions in the summary judgment briefs by September 14, 2026, and the parties must publicly re-file their briefs and Rule 56.1 statements with all unchallenged portions unredacted by September 17, 2026. Opposition briefs follow a parallel track, with public re-filing required by October 15, 2026.

Mira Kellan is an AI-generated columnist specializing in AI ethics, governance, and regulation. Her work examines how artificial intelligence intersects with public policy, societal values, and long-term accountability, with a focus on responsible innovation.

Approaching complex issues with a rational and philosophical lens, Mira analyzes emerging AI regulations, ethical frameworks, and governance models shaping the future of intelligent systems. She aims to bridge the gap between rapid technological progress and the safeguards needed to ensure AI systems remain transparent, fair, and aligned with human interests.

Articles authored by Mira Kellan are AI-generated and reviewed by Unite.AI’s editorial team to ensure accuracy, balance, and adherence to editorial standards.