Sony Music and Universal Music Group (UMG) have filed a new copyright lawsuit against artificial intelligence music creation service Suno. This time, record companies are focusing on the training process of Suno’s v6 model and claim that the new model benefits from user content produced by previous versions. According to the plaintiffs, since the old models in question were trained with copyrighted music obtained without permission from YouTube and other sources, using their output in v6 training does not eliminate the copyright issue. Sony and UMG are among the major music companies that have not signed a licensing agreement with Suno. The new case also raises the debate on how generative artificial intelligence models will evaluate not only the training data they directly use, but also the information they inherit from previous models, in terms of copyright law.
In the lawsuit shared by The Verge, Sony and UMG describe Suno’s method as “model laundering”, roughly speaking “model laundering”. The companies argue that training another model with the printouts of a model alleged to have infringed copyright does not eliminate the infringement at the first stage. According to the petition, the value taken from the plaintiffs’ works was first transferred to the old models, then to the printouts of these models, and finally to v6. Sony and UMG therefore argue that v6 cannot be considered a fresh start created with a completely independent and clean dataset. The main disagreement here centers on how the use of synthetic data created by another artificial intelligence system, as well as direct source files, in artificial intelligence training will be handled in terms of copyright.
Training data of the Suno v6 model is at the center of the case
When Suno v6 launched, the company’s Jack Brody told The Verge that the model was trained “from scratch” and using a new dataset. While Brody stated that training data also included user data, he did not provide detailed information about its scope. Suno spokesperson Rachel Racusen later confirmed that v6 was trained using content licensed from the company’s partners, as well as interactions such as work and preference signals created by community members, and the knowledge of the Suno team. However, it was not disclosed whether the audio files uploaded by users to the system or the outputs produced using these files were included in the training data. Therefore, Sony and UMG’s new complaint focuses on whether Suno’s definition of “new dataset” includes content produced by older models.
Sony also claims that Suno used the model distillation method known as “distillation” when developing the v6. In this approach, a new model can be trained by using the results produced by a larger or previously trained “teacher” model and some capabilities of the previous system can be transferred to the new model. Sony alleges that Suno trained the v6 to mimic the results of its previous models, and that the teacher models in question were created using copyright-infringing data. Plaintiffs therefore argue that the mere fact that v6 was not directly trained with Sony or UMG recordings is not enough. In the petition, it is claimed that the new model indirectly obtained information and benefited from unauthorized copies held by Suno.
If the record companies’ approach is accepted, the debate may not be limited to which audio recordings Suno uses directly as training data. Using the outputs of one model to train a later model may become a separate copyright issue, especially if the legal status of the initial model’s training data is disputed. On the other hand, since synthetic data and model distillation techniques can be used for different purposes in the artificial intelligence sector, legal evaluations in this case will need to be made on the concrete training process and materials used. Sony and UMG’s claims at this stage reflect the views of the parties in the petition, and it is not yet clear how the court will evaluate them. The dispute around Suno v6 thus shows that the source of information transferred between models, as well as licensing in music production with artificial intelligence, may be the subject of copyright lawsuits.