Regulation2 min read

Over 21 Million Songs Were Taken to Train AI

June 16, 2026Synthesized from 1 source: Engadget

The Atlantic published searchable databases showing more than 21 million songs were fed into AI music tools without artist consent, and the legal fights that followed are now reshaping how the entire music industry handles copyright.

The Atlantic has published four searchable databases revealing the scale of music used to train AI tools without permission. The total across all four databases runs to more than 21 million tracks. Songs from some of the world's most recognizable artists sit in these lists alongside millions of lesser-known recordings.

This was not a surprise to the music industry. The Recording Industry Association of America sued two AI music companies, Suno and Udio, back in June 2024 on behalf of Universal Music Group, Sony Music, and Warner Music Group. Both Suno and Udio largely admitted to using copyrighted recordings in their training but argued the practice was legal under fair use, the part of copyright law that allows limited, unlicensed use of protected work in certain contexts.

That defense has not held up well. Expert testimony in the cases showed that both platforms could reproduce recognizable fragments of the songs they trained on, which weakens the fair use argument considerably. Courts tend to take a harder line when the tool being built competes directly with the original work it was trained on.

The cases have since split. Universal Music settled with Udio in October 2025, and Warner Music settled with Suno in November 2025. Under the Warner deal, Udio agreed to abandon its existing model and build a subscription platform where artists are properly licensed, credited, and paid. These settlements are sometimes described as a resolution, but they are closer to a pivot forced by legal pressure.

Sony Music is still litigating. A hearing on the core fair use question is scheduled for July 2026, and the outcome matters well beyond music. If the court rules that AI training on copyrighted material without permission is illegal, every AI company that built tools on scraped creative work faces serious exposure. If the court rules in favor of the AI companies, the labels lose most of their negotiating leverage overnight.

The settlements themselves have not satisfied artists. Independent musicians filed separate class-action lawsuits in late 2025 because the major-label deals do not cover them. The musicians' union has also sued Universal and Warner, arguing the labels settled without protecting the actual performers whose work was used. Artists whose recordings appeared in training data are reportedly receiving nothing from the settlement agreements so far.

The Atlantic databases are significant because they give artists, lawyers, and courts something concrete to work with. Previously, AI companies were vague about exactly what they trained on. Now there is a searchable record. That kind of documentation is what drove a $1.5 billion initial settlement in a related publishing case, where piracy claims proved more compelling to a judge than copyright infringement arguments alone.

For businesses that use AI-generated audio in advertising, training videos, branded content, or customer-facing applications, the legal status of the tools matters. Content produced by a platform later found to have trained on stolen material could carry its own risk. The safest position right now is to ask vendors directly whether their training data was properly licensed, and to keep records of that answer.

Stay informed

Get AI intelligence like this delivered to your inbox.


You May Also Find Valuable