Sony Music and Warner Chappell Sue Anthropic Over Copyrighted Training Data

Sony Music and Warner Chappell filed a copyright infringement lawsuit against Anthropic in the US District Court for the Northern District of California, seeking damages for what the complaint calls "one of the largest and most blatant ongoing thefts of intellectual property in history" The Verge.
The suit covers "tens of thousands" of copyrighted works and seeks up to $150,000 per work, plus up to $25,000 for each instance where copyright management information — the metadata that identifies who owns a work — was stripped from the file. If the court awards the maximum on both counts, the total could reach several billion dollars.
Anthropic co-founders Dario Amodei and Benjamin Mann are named as individual defendants, not just the company itself. The complaint alleges Mann personally used BitTorrent (a peer-to-peer file-sharing protocol) to download over five million pirated books, and that Anthropic employees downloaded at least two million additional pirated books from Pirate Library Mirror, a shadow library. According to the filing, the defendants conducted a campaign of illegally torrenting, scraping, and downloading copyrighted works to develop and profit from Anthropic's Claude series of AI models.
On the music side, the complaint alleges Anthropic scraped lyrics from licensed services including MusixMatch and LyricFind, both of which paid to license content from the labels. Specific songs the plaintiffs say they found in Anthropic's training data include Marvin Gaye and Tammi Terrell's "Ain't No Mountain High Enough," Bon Jovi's "Livin' On a Prayer," Earth, Wind & Fire's "September," Leonard Cohen's "Hallelujah," and Taylor Swift's "Paper Rings."
Naming Amodei and Mann individually rather than solely the corporate entity means the plaintiffs are pursuing personal liability for decisions made at the founding stage of the company. In most corporate litigation, the company itself is the defendant; piercing the corporate veil to attach personal liability to executives requires showing that individuals directed or participated in the alleged wrongdoing with specific intent. Whether the specific factual allegations about Mann's torrenting activity meet that bar is a question the court will have to resolve, but the legal strategy itself signals the publishers' appetite for a fight that reaches beyond a settlement check.
The inclusion of both music and book piracy allegations in a single complaint from music publishers is also notable. Sony Music and Warner Chappell are music-side entities, yet the complaint foregrounds book piracy as evidence of a broader pattern of deliberate infringement. This builds a narrative: that Anthropic's approach to training data was systematic and top-down, not incidental. Whether a jury accepts that framing will turn on internal communications, download logs, and any evidence linking specific individuals to specific decisions about data acquisition.
The damages architecture is designed to scale. At $150,000 per work across tens of thousands of works, the statutory exposure alone is enormous. The additional $25,000-per-violation claim for stripped copyright management information under Section 1202 of the DMCA adds a separate count that compounds with the infringement damages. This is a structure that makes early settlement more attractive to defendants and makes a full trial more expensive and risky.
This lawsuit joins a growing body of litigation in which rights holders are testing whether AI training constitutes fair use (the legal doctrine that permits limited use of copyrighted material without permission), whether scraping licensed platforms for underlying content bypasses those licenses, and whether individual executives can be held personally liable for data acquisition decisions. The Northern District of California has become a central venue for these disputes.
What separates this filing from some earlier AI copyright suits is the specificity of the alleged conduct. The complaint does not merely assert that copyrighted material ended up in training data; it describes particular mechanisms — BitTorrent downloads, Pirate Library Mirror, scraping of MusixMatch and LyricFind — and attaches them to named individuals. If these allegations are supported by discovery (the pre-trial process where each side must hand over relevant evidence), they move the conversation from "did infringement occur" to "who directed it and how."
The broader context here is about what this means for the AI industry at large. If courts accept the argument that scraping licensed platforms for their underlying content constitutes infringement, the data acquisition pipeline for lyric and text-based training becomes legally hazardous. If executives can be held personally liable for data sourcing decisions, the risk calculus for founders and CTOs changes in ways that go beyond corporate indemnification. And if the damages model here holds, the financial exposure for training on copyrighted works at scale becomes a number that even well-capitalized companies cannot absorb without material impact.


