$1.5 billion for pirating books to train Claude: AI's fair use comes out unscathed, but not free

🕒 Published on Zendoric: July 22, 2026 · 01:59
A federal judge has given final approval to the largest copyright settlement in U.S. history: Anthropic will pay $1.5 billion to authors and publishers for using pirated copies of their books to train Claude. The ruling confirms something more important than the figure: training on protected text remains legal; stealing it does not.
By Android Headlines · July 21, 2026. Federal judge Araceli Martínez-Olguín gave final approval on Monday to the $1.5 billion settlement between Anthropic and a class action of authors and publishers, closing the largest AI copyright case ever resolved in the United States. The lawsuit, filed in 2024 by novelist Andrea Bartz and other writers, accused the company—backed by Amazon and Alphabet—of using unauthorized copies of their works to train the chatbot Claude.
The distribution is concrete: each rights holder will receive around $3,000 per work, from a fund covering more than 500,000 books. According to Anthropic's own legal team (as reported by Reuters), more than 91% of the affected authors and publishers had already filed their claim before final approval, a level of participation that explains why barely 350 people, out of several hundred thousand class members, chose to opt out of the settlement to litigate on their own—and the judge blocked attempts to do so past the deadline.
The settlement did not arise in a vacuum: it comes after last year's ruling by then-judge William Alsup, who drew a line that still shapes the rest of the AI litigation underway. Alsup determined that training generative models on copyrighted text constitutes fair use (legitimate use protected under the U.S. copyright fair use doctrine); but he also determined that downloading millions of books from pirate repositories—in this case, Library Genesis—to build that training base is infringement, regardless of what is done with the text afterward. Buying and scanning physical books, legal; pirating them at scale, not. Anthropic chose to settle rather than risk a trial over statutory damages for piracy that, according to the article itself, could have run to figures far higher than the $1.5 billion ultimately paid.
Approval was not a formality. Several authors objected that the $3,000 per work was less than they could have obtained by litigating individually, and there was friction over the fees for the plaintiffs' lawyers: Judge Martínez-Olguín trimmed the initial request of $187 million to about $101 million—roughly 7% of the total fund—redirecting the difference to the plaintiff class. The settlement also includes an unusual clause: Anthropic must destroy the central library of pirated works it built for training, a gesture aimed at closing the door to future uses of that same material.
In general, this case has become the de facto reference point for the rest of the industry: OpenAI, Meta, Google, and Microsoft face similar lawsuits over the origin of their training data, and all share the same underlying legal dilemma—where research on protected text ends and the plundering of pirate repositories begins. That the fair use of training itself was confirmed (and not challenged in this settlement) is, for AI companies, the good half of the news; that pirating data now has a de facto market price—on the order of $3,000 per work, multiplied by 500,000 works—is the serious warning for anyone who built their training corpus with the same kind of shortcuts.
Our reading is that this settlement does not halt generative AI: it ratifies it, but with conditions. The principle that training on human knowledge is legitimate—the legal basis that allows these models to exist—still stands, and it is a necessary precondition for the technology to keep advancing toward that horizon of abundance we defend in the long term: better, cheaper models trained on more data are, ultimately, engines of scientific and medical progress. But how that knowledge is obtained now has real, quantified economic consequences, not just moral ones. In the short term, this makes access to quality training data more expensive and slower—AI companies will have to license content instead of simply downloading it—which will probably accelerate commercial deals between AI labs and publishers, something already emerging in the sector with content licensing agreements. In the long term, it is a healthy signal: an industry that pays for its inputs, just as it pays for energy or talent, is a more sustainable industry with less social friction, not one that erodes public trust by building its advantage on a fragile legal foundation. The price has been high—the largest copyright settlement in U.S. history—but it has also been the price of buying legal certainty, something this industry, still young in its relationship with the law, badly needed.
🔗 Related on Zendoric
- Anthropic will pay 1.5 billion for pirating books, but AI training is shielded as fair use · 2026-07-23
- $1.5 billion: judge approves the largest copyright settlement of the AI era with Anthropic · 2026-07-21
- Judge gives final green light to Anthropic's $1.5 billion settlement, but cuts lawyers' fees from $375 million to $101 million · 2026-07-24


