· via The Verge
Microsoft filing says Copilot rarely reproduces NYT and book authors' work
Microsoft told a court that only a tiny fraction of 8.2 million Copilot chats showed overlap with publishers' content, arguing the figures support its fair use defense in the NYT-led copyright case.

Microsoft has told a federal court that its Copilot chatbot almost never reproduces meaningful portions of news articles or books, submitting new figures as part of its push to end copyright lawsuits from The New York Times, other news publishers and book authors without a full trial.
The numbers Microsoft is relying on
According to The Verge, Microsoft turned over 8.2 million Copilot chat logs during discovery to an expert retained by the news publishers. The company says the logs were deliberately selected because they matched keywords tied to the publishers' websites, making them the conversations most likely to contain protected material.
Even so, the analysis Microsoft cites found that 59,545 of those exchanges contained at least 16 words in common with the news content used to ground the model — under one percent of the total. An expert for the Center for Investigative Reporting identified 51 instances of "substantial overlap" with that organization's work in the same dataset, Microsoft says, while an expert in the authors' case found only 24 responses across all 8.2 million conversations that contained at least 30 matching words. Of the 212 books evaluated, just 10 produced any matches at all.
The Verge reports that The New York Times, the Center for Investigative Reporting and the Authors Guild did not immediately respond to requests for comment.
The fair use argument
Microsoft contends the figures support its broader position that assembling AI training datasets from copyrighted material qualifies as fair use. While systems like Copilot do depend on copyrighted works, the company argues the resulting products serve significantly different purposes than the originals, and that occasional reproduction of text, in Microsoft's words, "hardly undermines the transformative purpose of LLM training."
Where the case stands
The filing is the latest move in litigation brought by publishers and authors against Microsoft and OpenAI. The plaintiffs argue the two companies built commercial products on top of their work and now compete with them directly, partly by regurgitating copyrighted content. The news publishers' and book authors' claims were consolidated under a single judge to streamline the proceedings, over the plaintiffs' objections.
Microsoft submitted its filing on Friday as part of a request for summary judgment, which would resolve the case at an early stage. If the judge declines to grant it and instead sides with the publishers and authors, the dispute continues toward trial. The Verge also notes that the Trump administration filed a statement of interest in the New York Times case this week, backing OpenAI.
Why it matters
The filing puts a rare quantified data point at the center of one of the most consequential AI copyright disputes. Plaintiffs in these cases lean heavily on examples of models reproducing protected text as evidence of harm, and Microsoft's numbers, if accepted, undercut that narrative by suggesting verbatim output is vanishingly rare even in conversations primed to find it. But the argument cuts both ways: the core legal question is whether ingesting copyrighted material for training is permissible in the first place, not merely what a model emits afterwards. How the judge weighs output statistics against training-based claims in a summary judgment ruling could shape the legal playbook for AI developers and content owners well beyond this case.
- #ai
- #copyright
- #microsoft
- #fair-use
- #openai