SZA, Kenneth Blume demand answers after AI scraped their music

A new analysis tool has put fresh scrutiny on how generative music models are being trained, revealing millions of tracks — including songs by major stars and independent creators — were swept into commercial datasets. The disclosure has prompted outraged responses from artists, renewed legal pressure and fresh questions about how AI firms gather and use musical works.

What the detection tool revealed

Researcher Alex Reisner built an AI detection tool published last week by The Atlantic that checks whether an artist’s work appears in datasets used to train music generators. The probe found collections containing roughly 21 million songs, drawing on material ranging from mainstream catalogues to obscure uploads.

Among the names identifiable in the sample were globally known performers as well as smaller independent musicians, raising alarms about both scale and scope: the data sets reportedly included whole albums, single tracks and some material that artists say was never officially released.

Artists voice anger and suspicion

Several musicians reacted strongly after using the tool.

A high-profile singer said the search returned hundreds of her tracks and accused the AI industry of taking advantage of creators, including claiming that Black artists are disproportionately targeted in scraping practices. A producer with a large social following criticized specific companies — singling out Suno — and described working for services that train on scraped music as morally indefensible.

Other artists described how automated training outputs made their catalogs sound like the very “AI” imitations critics had accused them of producing, only for them to discover the datasets actually contained dozens of their own songs. One electronic producer offered a different take, arguing that the entertainment and tech sectors have historically tolerated unfair practices such as unlicensed sampling or unauthorised releases.

Read also  Live Nation, DOJ agree to settle antitrust suit: what it means for ticket prices

How the data appears to have been collected

Reisner’s write-up points to a mix of sources. Three of the featured datasets included automated links to tracks hosted on streaming platforms, and the researcher says developers often employ scraping tools that can bypass login walls and advertising — actions that would conflict with the platforms’ terms of service.

The fourth dataset in the investigation was built from content on the Free Music Archive, a repository of openly licensed and often freely available tracks.

Unclear chain of use; some firms admit involvement

While companies such as Google and Stability have acknowledged training models on large music collections, it remains uncertain which commercial music generators drew directly from the specific datasets flagged by the tool. The lack of transparency about training sources is at the center of the dispute.

Legal activity has already begun to catch up. Industry-facing music generators like Suno and Udio have faced lawsuits from major record labels in recent months; one label previously involved in litigation later agreed a license with Suno. This month, the American Federation of Musicians filed suit against two major record companies over alleged AI use of their members’ recordings.

  • Artist control: Creators worry about losing control over how their work is used and reproduced.
  • Economic impact: If models can replicate an artist’s sound, that could erode royalties and future income streams.
  • Disparate effects: Several musicians say Black and independent artists may face disproportionate exploitation.
  • Legal uncertainty: Ongoing lawsuits and shifting licensing deals mean the regulatory landscape is unsettled.
  • Platform responsibility: Questions linger about how streaming services and archives police automated scraping.

For listeners, the controversy raises questions about what “original” music will sound like when AI models can mimic specific voices and styles. For creators, the immediate stakes are practical — copyright, consent and compensation.

The story is likely to evolve quickly. Expect more artists to test their catalogs against detection tools, additional legal actions, and increased pressure on AI developers to disclose training sources or negotiate licensing. Until clearer rules or agreements emerge, many musicians say they will push for stronger protections and transparency around the datasets powering music-generating AI.

Similar Posts

Rate this post

Leave a Comment

Share to...