Sony, UMG Sue Suno Again, Allege v6 Was Distilled From Infringing Models

Sony Music and Universal Music Group have sued Suno over its v6 generative music model, alleging copyright infringement. The complaint is the labels' second action against the company and targets v6 specifically, rather than Suno's earlier models.
The suit was filed in the U.S. District Court for the District of Massachusetts. The complaint runs to 45 pages, according to case details reported on Sept. 18 Variety. Trade coverage has described it as a second lawsuit following the labels' original 2024 action The Hollywood Reporter.
At the center of the new filing is how v6 was built. Sony and UMG allege v6 was trained on user outputs generated by previous Suno models. Those earlier models, the labels allege, were themselves trained on unlicensed music ripped from YouTube and other sources The Verge. No license exists between the labels and Suno. That point is explicit in the complaint.
The labels frame that training lineage as "model laundering." Their argument, as pleaded, is that training a new model on outputs of an infringing model does not eliminate infringement. Sony further alleges Suno used distillation to train v6 to replicate the results of its previous teacher models.
Distillation, in this pleading, refers to a student model learning to match the output distribution of teacher models. The labels treat that process as continuity of infringement, not as a clean break. For practitioners, the claim targets a common workflow: using synthetic outputs and preference data from a deployed system to bootstrap its successor.
Suno disputes that characterization. Jack Brody of Suno said v6 was trained from the ground up with a new set of data including user data The Verge. The company has not licensed repertoire from Sony and UMG, but its stated position separates v6 training data from the prior models at issue in the earlier litigation.
The history matters here. On June 24, 2024, Sony Music, Universal Music Group and Warner Records sued Suno and Udio for copyright infringement Reuters. That case established the labels' baseline theory that training on copyrighted recordings without authorization infringes. The v6 suit extends the theory to a second generation trained in part on synthetic data derived from the first.
Udio has taken a different path. Universal Music Group and Udio settled copyright infringement litigation and announced strategic agreements for a new licensed AI music creation platform Universal Music Group. The companies plan to launch that platform in 2026. Suno has no equivalent agreement with Sony or UMG.
The broader context here is the industry's unresolved question around training-data provenance across model generations. If a v1 model ingested unlicensed audio, do v2 outputs trained on v1 generations inherit that liability. The labels answer yes. Suno's answer, at least publicly, is that retraining from the ground up with a new dataset resets the chain.
In my view, the technical stakes extend beyond music. Distillation, synthetic-data loops, and user-interaction logs are now standard inputs for iterative model improvement across modalities. A ruling that synthetic outputs carry forward the copyright status of the teacher model's training corpus would force closer tracking of lineage, filtering, and consent at each retraining step. It would also give licensors leverage over any successor model, not only the model that directly ingested the original works.
Worth flagging for builders is the licensing split now emerging. One frontier lab settles and builds a licensed platform with UMG. Another litigates a second complaint over whether retraining cleans the dataset. For enterprise buyers and API consumers, that divergence affects indemnity, data governance, and product roadmaps. Licensed training corpora cost more and constrain coverage. Unlicensed or user-data-trained systems may offer broader generation, with higher legal uncertainty.
The case will turn on specifics not yet in the public record. What audio was in the teacher training set. What proportion of v6 training data came from user outputs. How distillation was implemented. Whether filtering or deduplication altered memorization. Those details decide infringement, not abstract debate about synthetic data.


