Microsoft has argued that its Copilot chatbot rarely reproduces protected news articles or books in new summary-judgment filings in the consolidated copyright litigation involving The New York Times, other news organizations and book authors. The company says analyses of 8.2 million Copilot conversations show limited textual overlap with the plaintiffs’ works, supporting its contention that its use of copyrighted material is transformative and does not create a meaningful substitute for the originals.
Microsoft’s output-overlap analysis
The filing puts a quantitative claim at the center of a broader legal fight over generative AI. Microsoft is asking the U.S. District Court for the Southern District of New York to end claims against it at the summary-judgment stage. The underlying disputes include allegations that Microsoft and OpenAI used protected journalism and books to build commercial AI products that compete with the works on which they were trained.
Microsoft’s evidence addresses one particularly consequential issue: what people receive when they use a chatbot. In the news-publisher matter, the company said it supplied 8.2 million Copilot conversation logs to an expert retained by the publishers. Microsoft said the logs had been selected because they included keywords associated with the news plaintiffs’ websites and were therefore among the conversations most likely to contain their material.
According to Microsoft’s characterization of that review, 59,545 of the conversations contained at least 16 words in common with news content used to ground Copilot. That is less than 1% of the reviewed logs. Microsoft also said an expert for the Center for Investigative Reporting identified 51 instances of substantial overlap with that organization’s work in the dataset.
The company presented a separate set of figures for the book-author claims. Microsoft said an expert analysis found 24 Copilot responses containing at least 30 matching words across 212 evaluated books. It said 10 of those books had any matches. Microsoft uses those findings to argue that the product does not routinely emit passages that could replace the original articles or books for users.
What the figures do—and do not—address
Those numbers are Microsoft’s account of discovery and expert work, not findings by the court. They also measure specific kinds of word overlap, using defined thresholds, rather than every way an AI answer might draw on a work. A response can be useful to a user without reproducing a long verbatim passage, and a low observed rate of matching language does not by itself answer whether an output closely paraphrases protected expression or reduces demand for the source material.
That distinction matters because the litigation is not limited to alleged regurgitation. The plaintiffs have argued that Microsoft and OpenAI built competing products from their work without permission or payment, including products that can reproduce copyrighted material. The Times’ 2023 complaint alleged that the companies’ systems could recite its content, closely summarize it and mimic its expressive style, while threatening licensing, subscription and advertising revenue.
Microsoft’s position is that training large language models serves a purpose different from that of the books and articles in their original form. In its view, occasional overlap in chatbot answers does not undermine that claimed transformative purpose. The argument places output behavior alongside the more fundamental question of whether copying works to develop a model qualifies as fair use under U.S. copyright law.
Training-data acquisition and fair use
The debate over acquisition is important as well. In July, a federal judge approved Anthropic’s $1.5 billion settlement with authors and publishers over pirated books used to train Claude. Reporting on that case described an earlier mixed ruling: training on books was treated as fair use, while the acquisition of millions of books from pirate sites was found wrongful. That split illustrates why an output-overlap analysis, even if accepted by a court, would not necessarily decide all of the claims facing Microsoft and OpenAI.
The government has recently added another voice to the fair-use debate. In a statement of interest filed September 1 in the Times case, the Justice Department backed OpenAI’s argument that training models on internet-scale writings can be protected by fair use, emphasizing what it described as the technology’s creative, scientific and public benefits. The Times criticized that position and said AI companies should compensate creators for the content that makes their products possible.
What the court may consider
For companies building or deploying AI systems, Microsoft’s filing underscores that product behavior may be scrutinized independently from the provenance of training data. Evidence about the frequency of long matching passages can be relevant to claims that a chatbot is a substitute for a publisher’s product. But it is only one part of a larger inquiry likely to include how works were obtained, what they were used for, whether outputs are substantially similar, and what effect the systems may have on existing or emerging markets for licensed content.
The court has not yet ruled on Microsoft’s request. Its eventual treatment of the company’s usage analysis could shape the evidentiary playbook for AI copyright cases: companies may increasingly seek to use large-scale product logs and output studies to demonstrate limited regurgitation, while rightsholders are likely to focus on broader questions of copying, competitive displacement and compensation.




