
When AI Learns to 'Create,' Who Should Copyright Law Protect?
On the day the Munich District Court verdict came down, I was chatting with a junior lab mate working on music generation about his experimental setup. He said he used Meta's MusicGen for fine-tuning, mixing some song snippets scraped from the internet into the training set, and the results sounded like "a cat stepped on its tail." I told him to be careful; someone in Germany has already faced lawsuits over this. He froze for a second and said, "But I'm just doing academic research."
Academic research—these four words have never been a get-out-of-jail-free card in copyright law. The Suno case is basically a lab-level experiment placed under the magnifying glass of commercialization. The court ruled that Suno had no right to process works represented by GEMA (Germany's state-authorized music copyright management organization) artists. This conclusion itself isn't surprising. What's surprising is that the court required Suno to disclose illegal profits—which effectively exposes the AI model's training data sources and revenue sharing to the sunlight.
The "Black Box" of Training Data and the Reproducibility Dilemma
Those of us in computer vision often encounter similar problems. ImageNet datasets contain facial photos, and trained models might identify personal features, but lack portrait rights authorization. Last year, a CVPR paper was retracted directly because it used unauthorized medical image data. There's an unwritten rule in academia: If you use public datasets, you bear the responsibility; if you build your own dataset, you must prove the data source is legal.
The Suno case pushes this rule to the commercial level. GEMA, as a German collective music copyright management organization, represents the collective interests of authors. The core logic of the court's ruling is: Suno had no right to "process" these works, even merely as training data. The definition of "processing" is very broad—including scraping, parsing, feature extraction, and parameter updates. In other words, breaking a song into spectral fragments and feeding it to a Transformer also counts as infringement.
This raises an academic issue: reproducibility of experimental results. If the training data itself is a black box, how can other researchers reproduce it? I checked Suno's technical report; they only stated they used "large amounts of public music data," with vague descriptions regarding specific sources, proportions, and cleaning methods. This is a fatal flaw in academic paper writing. Reviewers seeing the phrase "proprietary dataset" will likely reject the paper outright. Because without transparency in data sources, experimental results lose the foundation for reproducibility.
"Copyright Compliance" Operational Guide for Experimental Design
From a technical perspective, AI music generation is essentially a conditional probability modeling problem. Given a melody, the model predicts the next note; given lyrics, the model predicts corresponding harmony. But training this model requires massive paired data (melody + lyrics + arrangement). If you use copyrighted works as supervision signals, it's equivalent to fitting parameters on protected data.
I advise peers, especially those working on music generation, image generation, and text generation, to focus on the following operational steps to significantly reduce legal risks:
1. Clarify Data Source Agreements: Create a LICENSE.md file in the data/ directory, listing the license agreement for each part of the data item by item. E.g., CC0, CC BY-SA, or commercial licenses. Don't use vague terms like "publicly available internet data." Courts only care if you obtained authorization, not whether you scraped it.
2. Use "Clean" Datasets as Baselines: Look for music datasets labeled "no copyright" or "public domain" on https://huggingface.co/datasets, such as MusicNet or MAESTRO. Although smaller in capacity, they at least allow experiments to run. If budget allows, buy commercial licensed databases, like extended packages of AudioSet.
3. Add Copyright Filtering Modules in Training Scripts: Write a filter_copyright.py that calls APIs from acrcloud or audd for each audio file, skipping protected works upon identification. This step loses about 20% of the data, but allows you to clearly state the compliance process in the paper's "Ethics Statement," leaving reviewers no grounds for complaint.
4. Disclose Statistical Characteristics of Training Data: In the paper appendix, report the duration distribution, genre distribution, and country distribution of the dataset, but do not provide specific filenames. This protects privacy while meeting the minimum requirements for reproducibility. Many top conferences now require submission of "Data Cards," which serve this purpose.
The Game Between Tight Lab Budgets and Copyright Costs
To be honest, our lab's funding is chronically tight. Doing AI music generation, buying a set of commercial music databases costs roughly 50,000 to 100,000 RMB. Using public datasets costs almost nothing. But the Suno case tells us that zero cost does not equal zero risk. If students scrape works managed by GEMA via crawlers and get sued, compensation could exceed the lab's annual budget.
The German court ordered Suno to pay damages; the specific amount hasn't been announced yet, but referencing GEMA's previous case with Google, compensation per song ranges from 0.5 to 1 Euro. Suno likely used millions of songs, totaling millions of Euros. This figure is enough to bankrupt a startup. For academic labs, although litigation risk is relatively lower, once involved, the university's legal department will directly demand deletion of all data, papers may be retracted, and experiments wasted.
Last month, a PhD student asked me if he could use Spotify's API to download user playlists for song recommendation research. I told him to read Section 7.2 of the Spotify Developer Agreement first, which states "data shall not be used to train machine learning models." He didn't believe it, saying "everyone does this." I showed him the Suno verdict, and he went silent.
The bottom line of academic research is not "what everyone else does," but "what the law permits." When AI learns to "create," copyright law protects not the creator's emotions, but the economic foundation upon which creators survive.
Physix Frontier