Back to blog
News·9 min read·1618 words

Is It Legal to Train AI on Copyrighted Books? The Anthropic Precedent Is Now Everyone's Problem

A $1.5B penalty, but training was ruled legal: how Judge Alsup's Anthropic decision quietly greenlit AI training on copyrighted books — and why every lawsuit since hinges on it.

Is It Legal to Train AI on Copyrighted Books? The Anthropic Precedent Is Now Everyone's Problem — illustration

In 2025, Judge William Alsup handed down one of the most consequential rulings in the history of artificial intelligence — and almost everyone misread it. Anthropic was ordered to pay a colossal $1.5 billion to a group of authors whose books were used to train its Claude models. Headlines screamed that the AI companies had finally lost. The opposite was true: Judge Alsup explicitly ruled that training an AI model on copyrighted books is lawful. What Anthropic was punished for was how it obtained those books — pirating them from illegal online "shadow libraries" like Library Genesis and Z-Library.

Now, as a new wave of lawsuits works its way through American courts, that distinction has become the single most important line in AI law. And according to attorneys interviewed by TechCrunch this week, the law on both sides of that line is still "all over the place."

The Ruling That Everyone Misread

Judge Alsup's reasoning was unusually literary for a federal court order. "Like any reader aspiring to be a writer, Anthropic's LLMs trained upon works not to race ahead and replicate or supplant them — but to turn a hard corner and create something different," he wrote, comparing the way an LLM ingests trillions of words to a writer's study of literature.

Strip away the elegant language and the structure of the ruling is simple:

  • Training itself: legal. Ingesting a copyrighted book to learn statistical patterns — the way a human writer learns from reading — does not violate copyright, because copyright law punishes copying, not learning.
  • Acquisition: illegal. Anthropic downloaded millions of books from shadow libraries that host pirated copies. That mass piracy cost the company $1.5 billion plus attorney's fees — the largest copyright award ever entered against an AI company.
  • The precedent's logic now governs the industry: get your training data from lawful sources, and the training itself is defensible.

Cathy Gellis, an attorney specializing in intellectual property and technology, told TechCrunch she reads the ruling as "generally good news for AI training." Copyright law, she explained, "hinges on copying, but it doesn't hinge on using the work or experiencing the work, consuming the work, reading the work."

Why $1.5 Billion Was the Cheap Option

Here is the uncomfortable arithmetic that authors' groups keep pointing to: Anthropic is projecting roughly $200 billion in annual revenue by 2028. A one-time $1.5 billion penalty — with training itself declared legal — may be the best money the company ever spent. It bought a precedent that the core of its business model is lawful, at a cost equal to less than one percent of projected future revenue.

That tension is exactly why the appeals and the copycat lawsuits keep coming. Authors argue the ruling creates a perverse incentive: pirate now, settle later, and treat the fine as a licensing fee paid after the fact.

A Law Written in 1976 Meets Technology From 2026

The deeper problem is statutory. U.S. copyright law has not been meaningfully updated since 1976 — five decades before large language models existed. Judges are forced to interpret guidelines written for photocopiers and videotape when ruling on trillion-parameter neural networks.

"Everybody is very worried right now because the law is all over the place," Jason Henderson, senior attorney and founder of the IP & Media Practice at JWL International, told TechCrunch. "They know that the AI model has been trained on so much stuff, and the law has not really caught up to that question."

Most of these cases now hinge on fair use — the doctrine that permits use of copyrighted material without permission for purposes like criticism, parody, education, and commentary. Courts weigh several factors, including the purpose of the use, how much was taken, and — critically — the effect on the market for the original work.

And it is that market-effect factor where Henderson sees a pattern emerging across the rulings:

  • If you train on someone's work to compete directly with it, courts frown. A model that writes romance novels trained on a romance novelist's catalog is on the weakest ground.
  • If your use doesn't compete with the original, courts tend to find a way to permit it. A model that learns grammar and world knowledge from the same catalog, but produces code or analysis, is on much stronger ground.

What This Means for the AI Industry in 2026

For the frontier labs, the practical playbook has already shifted:

  1. Licensing deals have exploded. After the Anthropic settlement, striking data-licensing agreements with publishers became far cheaper than litigation. Every major lab now has a licensing team.
  2. Data provenance is now a compliance function. Where a byte of training data came from matters as much as what it contains. Shadow libraries are radioactive; licensed corpora and public-domain archives are not.
  3. The "market effect" test shapes product decisions. Labs increasingly avoid products that obviously substitute for the works they trained on — one reason you see chatbots refuse to produce "in the style of" living authors.
  4. Open models face the same exposure. The legal theory doesn't distinguish between a closed frontier model and an open-weights release; if anything, open models that anyone can run are harder to police.

Why Developers Should Care

If you build on AI APIs, this isn't abstract legal theater. The copyright status of foundation models determines:

  • Whether your provider survives. A sufficiently bad ruling could reshape — or unwind — a model provider's ability to operate its flagship models in the U.S.
  • Indemnification guarantees. Enterprise API contracts increasingly include IP indemnity clauses, but those clauses are only as strong as the underlying legal theory.
  • Fine-tuning risk. Courts have not yet squarely addressed what happens when you fine-tune a model on copyrighted data you don't own. The Alsup framework suggests the training act may be defensible, but the acquisition of your dataset is entirely on you.

The safest posture for developers: use models from providers with clean data provenance and strong enterprise terms — and keep your own fine-tuning corpora licensed or public-domain.

It's Not Just a U.S. Story

While American courts sort out fair use, the rest of the world is diverging fast — and that matters for any team deploying models globally:

  • The EU folded AI-copyright transparency into its AI Act regime: providers of general-purpose models must publish sufficiently detailed summaries of their training data, and rights holders gained a formal pathway to contest use.
  • The U.K. spent 2025 debating a text-and-data-mining exception with an opt-out, then walked it back after creator-industry backlash; the current posture is a messy status quo favoring negotiated licenses.
  • Japan has long had one of the most permissive text-and-data-mining regimes, which is one reason several labs route portions of training through Japanese infrastructure and partnerships.

For developers, the practical consequence is that the same model can carry different legal exposure depending on where you deploy it. Multinational teams increasingly negotiate deployment-region terms in their provider contracts — another reason provider choice is becoming a legal decision as much as a technical one.

The Bottom Line

The Alsup ruling established that learning from books is legal, but stealing them is not. That single sentence now anchors U.S. AI copyright law — and it left both sides unhappy. Authors got a record payout but lost the war over training itself. AI companies got a green light for their core business but a billion-dollar reminder that data provenance will be audited. With appeals pending and Congress showing no appetite to update a 50-year-old statute, the real rulemaking is happening one courtroom at a time.

For teams building on AI APIs, the practical takeaway is simple: provenance matters, indemnity matters, and the provider you bet on matters. Compare models, pricing, and provider terms in one place at qubax.ai/models — and see the Qubax docs for API details.

FAQ

As of 2026, the leading precedent (Judge Alsup's ruling in the Anthropic case) says yes — training itself is lawful — but acquiring the books through piracy is not. Anthropic paid $1.5 billion for the piracy, not the training. Appeals and newer lawsuits may still refine this, but it remains the controlling framework.

The penalty was for downloading millions of books from illegal shadow libraries. The court treated the acquisition of pirated copies as mass copyright infringement, separate from the question of whether training on the content infringes.

What is fair use in AI training?

Fair use is a copyright doctrine allowing use of protected works without permission in certain circumstances (criticism, parody, education, and similar). In AI cases, courts weigh whether the use is "transformative" and — most importantly — whether it harms the market for the original work. Training that competes with the original works is on much weaker ground.

Does this ruling apply to open-source AI models?

The legal theory does not distinguish between closed and open models. Open-weights models trained on pirated data carry the same legal exposure; some argue more, since anyone can download and inspect them.

Yes — the training act itself may be defensible under the Alsup framework, but the dataset you fine-tune on is your responsibility. Use licensed, public-domain, or synthetic data, and prefer providers with strong IP indemnification in their enterprise terms.

Where can I compare AI models and their pricing?

Qubax AI aggregates frontier and budget models with transparent, discounted pricing in a single API. Browse the full catalog at qubax.ai/models.

🤖

Try Claude on Qubax

Anthropic models on Qubax. Up to 74% off.

View pricing

Article tags

#ai-copyright#anthropic#fair-use#ai-news#training-data
Share:Post on XTelegramLinkedInYHacker NewsReddit
Qubax AI

Qubax AI

AI Models at up to 99% off · Pay with crypto

Reading about Claude? Access it — plus 340+ other models — through one API. Anthropic models on Qubax. Up to 74% off.

Related articles