Microsoft: Few Users Grabbed NYT Articles via Chatbot

▼ Summary
– Microsoft submitted legal filings arguing that its Copilot AI rarely reproduces substantial copyrighted text, using this data to support a fair use defense.
– The company provided 8.2 million chat logs to experts, who found only minor overlaps with news articles and book content in the responses.
– Microsoft contends that using copyrighted material for training large language models is transformative and serves different purposes than the original works.
– Publishers and authors are suing Microsoft and OpenAI, claiming the companies built competing products by regurgitating their copyrighted content.
– The Trump administration filed a statement supporting OpenAI as Microsoft seeks summary judgment to potentially end the consolidated lawsuit early.
Microsoft has submitted new legal arguments asserting that its Copilot chatbot rarely reproduces full sentences or substantial passages from copyrighted news articles and books. These filings are part of the company’s ongoing defense against copyright infringement lawsuits brought by major publishers, including The New York Times, as well as various book authors. The tech giant contends that the limited textual overlap found in user interactions does not negate the legality of using such content for training artificial intelligence models.
To support its position, Microsoft provided an expert hired by news publishers with 8.2 million Copilot chat logs. The company stated these records were selected specifically because they contained keywords associated with the plaintiffs’ websites, making them the most probable sources to contain relevant copyrighted works. According to Microsoft, an analysis of this dataset revealed that only 59,545 conversations included at least 16 words in common with the news content used to ground the AI model. Furthermore, an expert representing the Center for Investigative Reporting identified just 51 instances of what was described as “substantial overlap” with the organization’s work. In the separate suit filed by book authors, a corresponding expert found that among the 8.2 million conversations, only 24 responses contained at least 30 matching words. Additionally, Microsoft noted that out of 212 books evaluated, only 10 showed any matches at all.
Representatives for The Times, the Center for Investigative Reporting, and the Authors Guild did not immediately provide comment on these figures.
These statistics are central to Microsoft’s argument that utilizing copyrighted material for AI training constitutes fair use. The company maintains that while systems like Copilot depend on vast amounts of existing data, their ultimate function serves a significantly different purpose than the original source material. Microsoft argues that the occasional reproduction of text snippets “hardly undermines the transformative purpose of LLM training.” This perspective is designed to demonstrate that the AI tools do not merely act as substitutes for the original creative works but rather offer new value through synthesis and generation.
The legal proceedings involve consolidated cases against both Microsoft and OpenAI, with plaintiffs alleging that these companies built products that directly compete with their own works by regurgitating copyrighted content. Although the publishers and authors objected to the consolidation, a single judge was assigned to streamline the process. Microsoft is currently seeking a summary judgment, which would resolve the case early without a full trial. Meanwhile, the Trump administration recently filed a statement of interest in the New York Times case, expressing support for OpenAI. If the judge rejects Microsoft’s motion for summary judgment, the litigation will proceed to court, potentially setting a critical precedent for the future of AI development and intellectual property rights.
(Source: The Verge)



