AI Training on Copyrighted Books: Is It Legal?

▼ Summary
– AI models like ChatGPT and Claude are trained on vast amounts of published works, raising legal concerns among authors about copyright infringement.
– Judge William Alsup recently ordered Anthropic to pay a $1.5 billion settlement for pirating books, while ruling that the AI training process itself was lawful.
– Legal experts argue the ruling favors AI companies by distinguishing between copying copyrighted material and the transformative use involved in reading or training models.
– The case highlights the complexity of applying outdated 1976 copyright laws to modern artificial intelligence technologies.
– Current legal uncertainty stems from debates over whether AI training constitutes fair use through transformative processes rather than direct replication.
The development of generative AI systems like ChatGPT, Gemini, and Claude relies on massive datasets comprising hundreds of millions of published works. These databases include books, academic papers, and online articles, meaning that countless authors have inadvertently contributed to the creation of tools that could potentially disrupt their careers. While this dynamic appears to be a clear-cut violation of intellectual property rights, the legal reality is far more nuanced.
Cathy Gellis, an attorney specializing in intellectual property, copyright, and technology, told TechCrunch: “I think one of the issues with this entire area of law and this entire area of technology is there’s a lot going on.” She noted that the situation is highly complex and driven by strong emotions from all sides.
A pivotal moment occurred last year when Judge William Alsup ordered Anthropic to pay a $1.5 billion settlement to writers whose works were used to train the company’s models. Although this appeared to be a victory for authors, the judge ultimately ruled that the AI training process itself was lawful. The penalty was specifically imposed because Anthropic had obtained these texts from illegal shadow libraries rather than through authorized channels.
Judge Alsup drew a parallel between machine learning and human creativity, stating: “Like any reader aspiring to be a writer, Anthropic’s LLMs trained upon works not to race ahead and replicate or supplant them , but to turn a hard corner and create something different.” He compared the ingestion of trillions of words by a large language model to a writer studying literature to inform their own unique output.
Legal experts suggest this ruling favors tech companies. Given that some firms project revenues nearing $200 billion annually by 2028, a $1.5 billion fine is manageable. Gellis explained: “I think it is generally good news for AI training that he looked at what was going on and really sort of thought it analogous to reading a copyrighted work as opposed to copying a copyrighted work.” She emphasized that copyright law prohibits unauthorized copying but does not restrict the act of reading, consuming, or using a work.
The core challenge lies in applying copyright statutes established in 1976 to modern technological advancements. Jason Henderson, Senior Attorney and Founder of the IP & Media Practice at JWL International, highlighted the uncertainty facing the industry. He told TechCrunch: “Everybody is very worried right now because the law is all over the place, and it’s because of this question.” He added that while companies know their models are trained on vast amounts of data, the legal framework has not kept pace with these developments.
Central to these disputes is the doctrine of fair use, which permits limited use of copyrighted material without permission for purposes such as criticism, education, and parody. Courts evaluate whether a use is “transformative” by considering factors like the purpose of the use, the amount taken, and the impact on the original work’s market.
Henderson pointed out that courts often look at competitive intent. He stated: “Copyright is always about protecting and growing the market.” He observed that if an AI company trains its model to directly compete with the original creator, courts tend to view this unfavorably. Conversely, if the use does not serve as a market substitute, judges are more likely to find it permissible.
This distinction was evident in the case where Thomson Reuters sued Ross Intelligence for building a competing legal research platform. In that instance, Judge Stephanos Bibas ruled against fair use, writing: “Ross’s use is not transformative because it does not have a ‘further purpose or different character’ than Thomson Reuters’s.” Unlike chatbots, which generate new content, Ross aimed to replace Thomson Reuters’s existing service, leading to a finding of infringement.
The legal landscape also extends to the copyrightability of AI-generated content itself. In Thaler v. Perlmutter, the court determined that works created entirely by AI cannot be copyrighted. This raises difficult questions about how to verify AI involvement in creative works. Gellis offered an analogy to clarify the distinction between training and generation: “If you write your novel in [Microsoft] Word and run spell check, we kind of feel comfortable with the idea of saying that Word does not own your novel.” She warned that AI is forcing society to reconsider decisions previously ignored.
Litigation remains ongoing, meaning definitive resolutions are unlikely in the near future. Gellis noted: “What you are seeing is that the initial opening volleys are being influential, and that influence itself could be undone if other courts decide different things, and it’ll take later states of litigation to figure out which one will prevail.” She concluded that despite the uncertainty, it would be unwise for AI companies to ignore these evolving legal precedents, as they continue to shape the industry’s trajectory.
(Source: TechCrunch)




