Training AI on books is fair use — piracy is the crime
Copyright law officially turned 50 years old with zero updates, and the AI industry's biggest unresolved legal question still hinges on a statute written in an era of photocopiers, not neural networks. TechCrunch's explainer this weekend crystallizes where the fight actually stands after a year of landmark rulings: courts keep telling AI companies that training on copyrighted books is legal, even as they hand authors record payouts. The result is a legal landscape that rewards one thing and punishes another, and that split is the story nobody should flatten into a simple "AI lost to writers" headline.
The win everyone misread
The anchor is the Anthropic settlement. Last year Judge William Alsup, in the Northern District of California, issued one of the first substantive rulings on LLM training — and it was widely reported as a $1.5 billion win for authors. The nuance mattered more than the number. Alsup ruled that training Claude on copyrighted books was "quintessentially transformative" and protected as fair use. What he actually penalized Anthropic for was a separate, narrower sin: the company had pulled more than seven million books from illegal shadow libraries like Library Genesis and stashed them in a "central library" it kept filed away. That copying wasn't fair use. Federal Judge Araceli Martinez-Olguin gave final approval to the $1.5 billion settlement in July — the largest copyright settlement in U.S. history — but the ruling that unlocked it had effectively said the core AI-training practice is lawful.
Read closely, and the message Alsup sent was a green light with a ransom note attached. Training on legitimate sources is fine. Piracy — regardless of the model it feeds — is not, and seven million pirated books is expensive. Anthropic settled because defending the piracy claim alone could have exposed it to statutory damages running into the tens of billions, not because it lost the fair-use argument. Lawyers now openly describe a "shadow library strategy": skip the debate over whether training is fair use and go straight for the piracy angle, which loses the industry the moral ground even when it wins the law.
The dividing line: do you compete?
The consistent principle emerging across the federal courts is blunt and economic. Jason Henderson, an IP attorney, put it plainly: courts tend to bless training that doesn't directly compete with the copyright owner and frown on training whose purpose is to build a rival product. The cautionary counter-example is Thomson Reuters v. Ross Intelligence, where training on legal content to build a competing legal-research platform was ruled not transformative, under exactly that logic.
That test is what keeps publishers nervous. A novelist can argue a chatbot trained on their books is gunning for their market — the model can generate synthetic books in their style. Yet that argument has not prevailed in court. The 1976 Copyright Act gives judges four factors to weigh — purpose, nature, amount used, and market impact — and so far each judge's reading of "transformative" has done most of the work. It is a fragile stability. As Cathy Gellis, a copyright attorney, told TechCrunch, the law is "all over the place," and every early ruling is shaping the next round of litigation even though none is binding precedent for the others.
Who wins, who loses, and what's actually still open
Right now the clear winners are the large labs, which now have a moderately friendly default: train on licensed or publicly circulating books, destroy the pirated archives, and pay settlements sized to the piracy rather than the training. The losers are the small players and the open-weights ecosystem, where clean, trustworthy training data costs real money — a dynamic we saw play out in Apple's pay-per-use news licensing deals, which reset how content owners price AI access. For a startup without Anthropic's cash flow — the company projects roughly $200 billion in annual revenue by 2028 — even a "win" on fair use carries the risk of a ruinous discovery fight over where training data came from. The uncertainty is itself a moat that favors the incumbents who can lawyer their way through it.
The second front — not covered by the training rulings — is the output side. In Thaler v. Perlmutter the D.C. Circuit confirmed that a 100% machine-generated work cannot be copyrighted, because authorship requires a human. That leaves a messy open question: how much human involvement is enough to make an AI-assisted work copyrightable, and how would anyone prove it? It is the reason AI watermarking suddenly matters as more than a provenance nicety — it is becoming the evidence mechanism for a rule nobody has cleanly defined.
What to watch next
The tells to track are the appeals and the second wave of lawsuits. Several authors and publishers opted out of the Anthropic settlement and are pursuing their own cases, which could test whether that "transformative" reasoning survives a different judge. The publishing houses suing Google, xAI, and OpenAI over training data will put more flesh on the competing-product doctrine. And watch for any movement out of Congress — fifty years is a long time for a statute to carry this much weight, and the courts have made clear they are applying the word as written and leaving the policy call to lawmakers. If a new copyright act ever lands, it will sweep away the case law that this entire industry's legal strategy is currently betting on.
Do you think the "it's only illegal if you directly compete" reading of fair use is sustainable — or should training itself be licensed? Tell us in the comments.
Sources: TechCrunch · Reuters: Anthropic settlement approval · Wolters Kluwer Copyright Blog · Skadden on Thaler v. Perlmutter · Reuters: AI copyright cases hit a pivotal year