Technology

Music Publishers Sue Anthropic Over Alleged Piracy in AI Training Data

Martin HollowayPublished 18h ago6 min readBased on 7 sources
Reading level
Music Publishers Sue Anthropic Over Alleged Piracy in AI Training Data
Photo by Euronewsweek Media on Unsplash

Sony Music Publishing, Warner Chappell, and other music publishers filed a copyright-infringement lawsuit against Anthropic and its co-founders Dario Amodei and Benjamin Mann on Friday in the U.S. District Court for the Northern District of California. The suit, first reported by Music Business Worldwide and covered by TechCrunch, accuses the AI lab of running a "brazen campaign of illegally torrenting, scraping, and downloading copyrighted works."

The complaint alleges that Anthropic used thousands of copyrighted works to train its Claude model family, calling the conduct "blatant theft" and "flagrant piracy." Specifically, the publishers claim Anthropic used torrenting (a peer-to-peer file-sharing method) to obtain millions of copies of books, including ones containing song lyrics and sheet music.

This is not the first time Anthropic has faced copyright litigation over its training data practices, and the new filing arrives in a legal landscape that has already shifted against the company. In Bartz v. Anthropic PBC (docketed as 4:24-cv-05417), a group of authors accused Anthropic of using copyrighted works to train products including Claude. That case resulted in a $1.5 billion judgment against Anthropic after a judge ruled that acquiring copyrighted content through piracy was not legal. The Bartz filings reference Sony Music Entertainment v. Cox Communications, 93 F.4th 222, a precedent that may influence how courts handle repeat infringement claims.

The music publishers' legal campaign against Anthropic also has an earlier chapter. In October 2023, Reuters reported that publishers alleged Anthropic violated their rights by using lyrics from at least 500 songs, including the Beach Boys' "God Only Knows." That original action was filed as Concord Music Group, Inc. v. Anthropic PBC, docketed as 3:23-cv-01092 in the Northern District of California. A second federal case under the same caption, docketed as 5:24-cv-03811, followed in June 2024. Most recently, in July 2026, the publishers filed an amended complaint against Anthropic over alleged copying of copyrighted song lyrics by the Claude chatbot, according to Music Business Worldwide.

The latest lawsuit broadens the scope considerably. Where the earlier actions focused on lyric reproduction (Claude outputting song lyrics when prompted), the new complaint extends to the mechanics of data acquisition itself, alleging that the torrenting and scraping infrastructure used to build Claude's training corpus involved millions of infringing copies. Some of the lawyers behind the new Sony Music Publishing and Warner Chappell filing also represented Concord Music Group and Universal Music Group in the January case, indicating a coordinated and deepening legal strategy across the publishing sector.

The named defendants include Anthropic as a corporate entity as well as co-founders Amodei and Mann individually. That choice targets the company's leadership directly alongside the organization. For AI practitioners and legal teams, the inclusion of individual co-founders as defendants in a copyright action centered on training-data acquisition is a detail worth noting. It signals that plaintiffs are pursuing a theory of direct individual accountability for decisions about how training data was sourced, not merely corporate liability for what a trained model later produces.

The cumulative legal pressure on Anthropic is substantial. A $1.5 billion judgment in Bartz has already established that courts are willing to treat large-scale acquisition of copyrighted material via piracy as infringing conduct, not as a fair-use-protected activity incidental to model development. Fair use is the legal doctrine that permits limited use of copyrighted material without permission, typically for purposes like criticism or research; AI labs have argued that training models on copyrighted text qualifies. The music publishers' new complaint, if it proceeds on similar grounds, would extend that liability framework into the specific domain of musical compositions, lyrics, and sheet music contained within torrented texts.

The broader context here is that the filing adds another data point to a now well-established pattern: rights holders are systematically pursuing AI labs through copyright litigation, and they are winning. The Bartz ruling's logic, that obtaining copyrighted works through torrenting and piracy cannot be laundered into lawful training data, cuts against a defense strategy that has relied heavily on fair use and transformative purpose arguments. Music publishers, author collectives, and other rights holders appear to be operating from a shared legal playbook, building on precedents from one case to strengthen the next.

There is an open question worth flagging: whether the industry's training-data pipeline can be retrofitted to exclude pirated content without degrading model quality. The technical feasibility of removing specific copyrighted works from a trained model's weights (the internal parameters a neural network learns during training) is not settled, and the legal system is now moving faster than the research on that problem. For AI labs still relying on large-scale scraped corpora, the gap between legal exposure and technical remediation is widening.