Technology

An AI Company Just Paid $1.5 Billion for Using Pirated Books

Martin HollowayPublished 2w ago5 min readBased on 11 sources
Reading level
An AI Company Just Paid $1.5 Billion for Using Pirated Books

A federal judge has given final approval to Anthropic's $1.5 billion settlement of a copyright lawsuit, the largest copyright settlement in U.S. history. Judge Araceli Martinez-Olguin of the U.S. District Court for the Northern District of California signed off on the final approval on Monday, July 20, 2026, as first reported by Reuters (Reuters).

Anthropic is a company that builds artificial intelligence systems. To make those systems smart, the company needs to feed them huge amounts of written text. This process is called "training." The case, formally known as Bartz v. Anthropic, focused on where Anthropic got the books it used for training. The company used two sources: books it purchased and scanned, and books it downloaded from pirate websites including Library Genesis and Pirate Library Mirror. Anthropic reportedly downloaded approximately 7 million books in total, of which roughly 500,000 titles are covered by the settlement (Authors Guild).

Judge William Alsup, now retired, had previously ruled on the two parts of Anthropic's book collection separately. He found that using copyrighted text to train an AI model is fair use, meaning the law allows it without getting permission. However, he ruled that Anthropic's downloading of books from pirate sites was illegal. Anthropic settled to avoid a trial and potential jury-awarded damages (TechCrunch).

The settlement payout works out to about $3,000 per book across an estimated 500,000 books, to be shared among the authors and publishers who hold rights to them. Anthropic agreed to pay the $1.5 billion regardless of how many individual rightsholders opt out of the settlement, and the payment is structured over time in several installments (Authors Alliance; City Law Forum).

The path to final approval was not straightforward. After the settlement was announced in August 2025, a federal judge in San Francisco initially declined to grant approval (Reuters). Preliminary approval came on Thursday, September 25, 2025 (Reuters). Martinez-Olguin later delayed final approval at one point, seeking more details before ultimately granting it (Storyboard18). The lead attorneys for the plaintiffs also reduced their fee request in March 2026 after receiving pushback (Reuters).

One important consequence of the settlement is that the judge's fair-use ruling will never reach a higher court and therefore does not become a binding rule for the rest of the country. Training on copyrighted text was found to be fair use in this one court, but that finding carries no weight outside the Northern District of California. Other AI companies facing similar lawsuits are not bound by it, and the people suing them are free to argue the opposite.

The broader litigation landscape remains active. Ongoing copyright lawsuits over AI training on copyrighted works involve companies including Google, Meta, Midjourney, and OpenAI (TechCrunch). In the week of July 14, 2026, publishers and authors including Hachette, Cengage, Elsevier, author Scott Turow, and S.C.R.I.B.E. filed a class action lawsuit against Google over accusations that it used their copyrighted works to train its AI platform Gemini (TechCrunch).

The Anthropic settlement resolves one case but leaves the central legal question unsettled. Whether training a commercial AI model on copyrighted text is fair use under U.S. law remains unresolved at the higher court level. Anthropic's decision to settle means the company pays $1.5 billion for the piracy part of its data collection while keeping a favorable fair-use ruling that no other court has to follow.

For the AI industry, the settlement establishes a rough price point: about $3,000 per copyrighted book for the act of illegally downloading books from pirate sources. Whether that figure becomes a standard depends on outcomes in the pending cases against Google, Meta, OpenAI, and others, where the details around how data was obtained and used may differ. The fair-use question, which is the one AI developers most need answered, remains open.

The broader context here is that the AI industry is still in an early chapter of its collision with copyright law. We have seen this pattern before in technology, most notably during the Napster era of the early 2000s, when music file-sharing forced the legal system to adapt to a new technology. What is unusual about the Anthropic outcome is that both sides accepted a practical resolution rather than pushing for a ruling that would apply across the industry. That keeps the door open for every other lawsuit to produce a different answer, which is a mixed result for anyone who wanted clear rules.