· via TechCrunch
Training AI on copyrighted books remains legally unsettled as courts diverge
According to TechCrunch, early court rulings split on whether training AI on copyrighted books is legal: US copyright law dates to 1976, and judges disagree on when training counts as fair use.

No simple answer to a basic question
Whether AI companies may lawfully train their models on copyrighted books still has no clean answer, and the first court decisions to address the issue point in different directions, according to a TechCrunch explainer on the subject. Models behind chatbots such as ChatGPT, Gemini and Claude are trained on enormous collections of books, articles and papers, and most published authors contributed without knowing or consenting — yet that alone does not make the practice illegal.
Cathy Gellis, an attorney specializing in intellectual property, copyright and technology, told TechCrunch the area is "very complex" with strong feelings on both sides.
A billion-dollar penalty, but training ruled lawful
The most consequential ruling so far involved Anthropic. Judge William Alsup ordered the company to pay $1.5 billion to a group of writers whose works were used to train its models — but, as TechCrunch notes, the judge actually found the training itself lawful. What Alsup penalized was Anthropic obtaining the books from illegal online shadow libraries.
Alsup compared a large language model ingesting trillions of words to a writer studying literature, writing that Anthropic's models trained on works not to replicate or supplant them, but to create something different.
Gellis sees the ruling as largely good news for AI companies, because the judge treated training as analogous to reading a copyrighted work rather than copying one. Copyright law, she explained, hinges on copying — not on using, experiencing or consuming a work. She also questioned how much deterrence a $1.5 billion fine provides to a company projecting roughly $200 billion in annual revenue by 2028.
A 1976 law confronting 2020s technology
US copyright law has not been updated since 1976, which means judges are interpreting half-century-old guidelines to resolve questions that could shape the entire AI industry. Jason Henderson, senior attorney and founder of the IP & Media Practice at JWL International, told TechCrunch that "the law has not really caught up" to models trained on vast amounts of material.
These cases largely turn on fair use — the carve-out in copyright law that permits use of copyrighted material without permission, protecting criticism, parody, education and similar activities. Judges weigh factors including the purpose and nature of the use, how much of the work was used, and the impact on the market for the original.
Competition is the dividing line
According to Henderson, courts have reasoned inconsistently in AI cases, but a pattern is emerging: when training is done to directly compete with the source material, courts frown on it; when it is not, they tend to find ways to allow it. "Copyright is always about protecting and growing the market," he told TechCrunch.
He pointed to Thomson Reuters' lawsuit against research firm Ross Intelligence, which copied Reuters content to build a competing AI-based legal platform. Judge Stephanos Bibas ruled that Ross's use was not transformative because it lacked a "further purpose or different character" than Thomson Reuters' own offering.
Authors could try to argue that chatbots compete with them by generating new, synthetic books from their works — but as TechCrunch reports, that argument has not yet prevailed in court.
AI-generated output is a separate question
Gellis draws a useful distinction between copyright in training inputs and copyright in AI-generated output. In Thaler v. Perlmutter, a court ruled that a work that is entirely AI-generated cannot be copyrighted at all, which opens further problems: how to prove whether something was AI-generated, and to what degree. Her analogy: if you write a novel in Microsoft Word and run spell check, most people accept that Word does not own your novel. AI, she argues, is forcing courts to revisit assumptions that were long left unexamined.
Why it matters
Most major AI companies remain tied up in pending litigation, so no definitive resolution is coming soon. Gellis told TechCrunch that the early rulings are already influential, but that their influence "could be undone if other courts decide different things," with later stages of litigation determining which approach prevails. In her words, "it would be kind of foolish for the AI companies to ignore them."
The stakes run in both directions. For AI developers, the emerging case law shapes how they source data, whether they pursue licensing deals, and how much legal and financial risk attaches to existing models. For authors and publishers, the rulings so far suggest training itself may survive fair-use scrutiny, leaving liability concentrated on how books were acquired — and leaving open the harder question of whether AI-generated content that competes with original works crosses a legal line.
- #ai
- #copyright
- #fair-use
- #legal
- #large-language-models