ai-tools

Is it legal to train AI models on copyrighted books? It's complicated

vybecodingBy vybecoding.ai Editorial
August 24, 20266 min readOfficial
A $1.5 billion copyright settlement against Anthropic, handed down in 2025 by U.S. District Judge William Alsup, sounds like a decisive win for authors — until you read what the judge actually ruled.

A $1.5 billion copyright settlement against Anthropic, handed down in 2025 by U.S. District Judge William Alsup, sounds like a decisive win for authors — until you read what the judge actually ruled. Training an AI on copyrighted books, Alsup concluded, is lawful. The fine was for how Anthropic obtained the books, not for using them.

What Changed

The case, Bartz v. Anthropic PBC, centered on a group of authors who alleged that Anthropic had ingested millions of copyrighted works — including pirated copies pulled from illegal online shadow libraries — to train its Claude models. Judge Alsup, ruling out of California's Northern District in June 2025, applied the fair use doctrine and found that training itself qualifies: every factor in the four-part copyright test favored Anthropic except the nature of the copyrighted work. In Alsup's framing, the LLM's ingestion of text to produce something new is "exceedingly transformative" — comparing it, in the ruling's own language, to a writer studying literature before turning a corner to create something original, not a replica.

The $1.5 billion penalty landed not on the training act but on the piracy that preceded it. Sourcing books from shadow libraries when legitimately purchased copies existed, the judge found, cannot be excused by a downstream transformative purpose. Multiple legal analyses published after the ruling — including a detailed breakdown from Thomas Horstemeyer IP Law — note that Alsup drew a hard line between two acts: using lawfully acquired books for training (fair use) versus building a pirated library to do so (not fair use, even if the copies are discarded immediately afterward). The distinction sounds technical, but it has enormous commercial consequences.

A parallel ruling in Kadrey v. Meta Platforms, issued two days later, reinforced the same split. Judge Vince Chhabria similarly declined to grant Meta an automatic win because its training data came from shadow libraries, but he also rejected the idea that downloading and training should be evaluated as completely separate acts — the source of the data colors everything that follows. Together, the two rulings signal that federal courts are willing to protect AI training under fair use doctrine, while treating the piracy question as live legal exposure that proceeds to trial on liability and damages.

What makes the Anthropic fine potentially a bargain rather than a deterrent: one intellectual property attorney quoted by TechCrunch noted that $1.5 billion is a significant number in isolation but modest for a company projecting roughly $200 billion in annual revenue by 2028.

How It Works

Fair use in U.S. copyright law isn't a blanket exemption — it runs through four factors that courts weigh together: the purpose and character of the use, the nature of the copyrighted work, how much of it was used, and what effect it has on the market for the original. The key question in AI training cases has been whether "transformative use" covers mass ingestion for statistical pattern learning.

Alsup's reasoning leans on the transformative-use thread: an LLM doesn't reproduce books, it builds internal representations of language patterns across trillions of tokens. The model isn't a copy of any work; it's a probability engine shaped by exposure to many works. This is where "reading versus copying" becomes the operative distinction — copyright law attaches to reproduction, not to the experience of a work influencing a reader (or, by this logic, a model). A separate legal analysis from Thomas Horstemeyer's IP practice confirms this framing, noting that courts are applying a "reasonably necessary" test: did the training actually require that specific sourcing method, or were lawful alternatives available?

The gap the 1976 Copyright Act leaves here is real and growing. The statute was written decades before the internet, let alone large language models, and judges are applying general doctrine to facts the legislators never imagined. The result is that outcomes vary considerably based on the specifics of sourcing, competitive intent, and what the model actually produces — not on a settled rule.

What It Means for Developers

The practical upshot for anyone building with or on top of AI systems divides into two distinct risk zones. The first is training-data sourcing: if you or a vendor you depend on has assembled training sets from pirated or unlicensed corpora, the Alsup and Chhabria rulings make clear that the training act being transformative won't protect you. The shadow library question goes to trial — damages are live. Our read is that this is the sharper near-term risk for AI labs; it's also the one developers rarely have visibility into, since training provenance is almost never disclosed publicly.

The second risk zone is output ownership. Thaler v. Perlmutter — a separate federal case — established that works generated entirely by AI, without meaningful human authorship, cannot be copyrighted. This matters concretely for developers generating code, text, images, or other assets with AI tools and assuming those outputs are automatically protectable. They are not, unless a human's creative decisions are documentable. That distinction becomes commercially significant the moment you're trying to defend exclusive rights to an AI-assisted product.

There's also a competitive-use wrinkle that the Thomson Reuters v. Ross case introduced: using copyrighted material to train a model that directly competes with the source of that material fails the fair use test. Training a legal research AI on Westlaw content to undercut Westlaw is different from training a general-purpose model on the same content. Developers building narrow vertical models in markets with dominant incumbent data providers should treat this as a specific flag, not a general clearance.

The honest assessment heading into late 2026 is that there is no stable answer here yet. Two significant rulings in the same district, issued two days apart, reached compatible but not identical conclusions — and both leave damages questions unresolved. Legislation hasn't moved. The 1976 Copyright Act will keep getting applied to facts it wasn't designed for until either Congress acts or enough appellate rulings solidify a real precedent. For now, sourcing transparency and documented human involvement in creative outputs are the two levers developers actually control.

Sources

techcrunch.com Is it legal to train AI models on copyrighted books? It's complicated Training AI Using Copyrighted Materials: Is It Fair Use? - Thomas Horstemeyer Federal Judge Rules AI Training on Copyrighted Books Is Fair Use—With Key Limitations MSN

Based on

https://techcrunch.com/2026/08/23/is-it-legal-to-train-ai-models-on-copyrighted-books-its-complicated/techcrunch.com

This article is an original, AI-assisted summary and analysis. Credit for the underlying reporting or footage belongs to the source above.

vybecoding

Written by the vybecoding.ai editorial team

Published on August 24, 2026

TOPICS

#ai#news
Is it legal to train AI models on copyrighted books? It's complicated