The Conflict Between the Music Industry and Generative AI Enters a New Phase
The copyright issues surrounding generative AI are once again on the verge of escalating into a major courtroom battle.
On August 28, 2026, Sony Music Publishing, Warner Chappell Music, and their affiliated publishers filed a lawsuit against Anthropic, the developer of the generative AI "Claude," in the U.S. District Court for the Northern District of California. The defendants include not only Anthropic but also its co-founder and CEO, Dario Amodei, and co-founder Benjamin Mann, who are named as individuals.
What is noteworthy this time is that the focus is not on record companies like "Sony Music" or "Warner Records," but on music publishers that manage the rights to the songs themselves, such as lyrics and compositions.
The dispute is not only over the master rights of sound recordings. A major focus is on how lyrics, songs, and sheet music, referred to as "musical works," were obtained and used as data to create Claude.
The publishers argue in their complaint that Anthropic's actions were not merely incidental use of copyrighted works but involved systematic actions with large-scale illegal downloads and scraping.
Of course, this is the plaintiffs' claim at this stage, and the court has not yet made any factual determinations regarding this lawsuit. This distinction is very important.
Anthropic has also stated that it disagrees with the publishers' claims and intends to strongly contest them in court.
"Tens of Thousands of Songs" Could Be Involved, Potential Damages Could Reach Billions of Dollars
The complaint alleges that the songs claimed to have been used without permission in the development of Claude number "tens of thousands."
Examples include globally recognized songs such as "Ain’t No Mountain High Enough," "All I Want for Christmas Is You," "Eye of the Tiger," and "Paper Rings."
The statutory damages sought by the publishers could reach up to $150,000 per work if deemed willful copyright infringement.
Additionally, if it is determined that Copyright Management Information (CMI), such as the copyright holder's name or copyright notice, was illegally removed or altered, damages of up to $25,000 per violation are also being sought.
If the scope expands to tens of thousands of songs, the simple calculation could reach billions of dollars.
However, it is premature to conclude that "Anthropic will have to pay billions of dollars."
The complaint does not present a finalized total claim amount but is seeking the statutory maximum per work. The actual damages will depend on which works are found to have been infringed, whether willfulness is recognized, and the outcome of the trial.
The Main Issue Is Not "AI Learning" Itself, but How the Data Was Obtained
What is particularly important in this lawsuit is that it goes beyond the conventional debate of "is it legal to let AI learn copyrighted works?"
The publishers are strongly questioning the "route" through which Anthropic obtained its learning data.
According to the complaint, in 2021, Benjamin Mann allegedly obtained at least about 5 million books using BitTorrent from an online library called LibGen.
Furthermore, in 2022, it is claimed that Anthropic employees additionally obtained about 2 million books from the Pirate Library Mirror, known as PiLiMi, which did not overlap with LibGen.
The issue for the publishers is that these books allegedly included songbooks containing lyrics and sheet music.
In other words, it is not a simple matter of "Claude is not a music-generating AI, so it has little to do with music publishers."
Large language models learn from text. Naturally, this text includes lyrics, song explanations, and text accompanying sheet music.
If copyrighted works managed by music publishers were included in books that were massively acquired, then music publishers have a direct interest in the matter.
Scraping from Lyrics Sites Also a Point of Contention
The publishers are not only concerned about pirate libraries but also about data collection from the internet.
The complaint alleges that Anthropic collected lyrics data from websites, including services like Musixmatch and LyricFind, which display lyrics under proper licenses.
There is an interesting twist here.
Lyrics sites contract with rights holders and pay fees to display content. However, when AI companies collect large amounts of text from these sites to use as training data for creating other commercial services, how should the original licensing relationships be handled?
Is "being freely viewable on the internet" the same as "being free to replicate for commercial AI training"?
This is a very significant issue in the era of generative AI.
The publishers are also raising issues about acquiring copyrighted works through methods such as purchasing large quantities of used books to scan pages, and through large datasets like Common Crawl, The Pile, and Books3.
Were Copyright Notices Also Removed?
Another aspect that distinguishes this lawsuit from a mere "AI learning data issue" is the claims regarding CMI.
CMI refers to information that accompanies a copyrighted work, such as the author's name, rights holder's name, and copyright notice, indicating the rights relationship.
The publishers claim that Anthropic used processes that removed such information during the formatting of learning texts.
In AI development, "cleaning" is commonly done to extract only the main text from web pages, removing unnecessary parts like navigation, ads, and footers.
However, if it is determined that this process intentionally removed the author's name or copyright notice, a different legal responsibility could arise, separate from the issue of "learning a large amount of text."
This is why the publishers are seeking statutory damages related to CMI of up to $25,000 per instance.
Claude's Output of Lyrics Also Under Scrutiny
The lawsuit is not limited to input data.
The publishers claim that Claude sometimes outputs existing song lyrics as-is or in a very similar form upon user instruction.
This is a very troublesome issue for generative AI companies.
"Analyzing learning data by the model" and "re-providing learned copyrighted works to users" have different legal and business implications.
AI companies usually explain that models do not store and search text like a database but learn patterns from large amounts of data to generate new responses.
However, the issue of models "remembering" and reproducing long portions of very famous lyrics or texts has been pointed out before.
The publishers argue that such outputs compete with the existing lyrics licensing market.
If users can obtain lyrics that are normally provided for a fee just by asking an AI, the claim that "AI services are replacing existing content markets" could become stronger.
In future trials, the actual reproducibility and the extent to which Anthropic's safeguards are functioning will also be important.
Why a Ruling that "AI Learning is Fair Use" Alone Cannot Protect Anthropic
One reason this lawsuit is tough for Anthropic is that there is already a significant judicial decision in a copyright lawsuit involving books.
In the 2025 Bartz v. Anthropic lawsuit, the court ruled that using copyrighted works for training large language models could be considered fair use under certain conditions.
This was a major victory for the AI industry.
However, the court did not allow "obtaining data from anywhere."
Particularly, obtaining pirate books and holding them as a central library was treated as a separate issue from AI training.
This difference is extremely significant.
In other words,
"AI learning from legally owned books"
and
"Obtaining pirate copies for AI learning"
are not the same legal issue.
Furthermore, Anthropic agreed to a $1.5 billion settlement in a class-action lawsuit with those authors, which received final court approval in July 2026.
The Sony and Warner lawsuit is precisely targeting this weakness of "how data was obtained."
Why Are the Two Founders Being Sued Individually?
Another unusual aspect is that not only the corporation Anthropic but also Amodei and Mann are named as individual defendants.
The publishers claim that Mann's large-scale acquisition from LibGen was conducted from the early days of the company's establishment, and Amodei was aware of and approved the data acquisition.
In other words, the plaintiffs are trying to position it not as a scenario where "some employees downloaded copyrighted works on their own," but as a decision close to the core of management and development.
Whether this claim is recognized will be a major focus of the trial.
When AI companies face copyright issues, the "responsibility as a company" is usually discussed.
However, if cases increase where management is recognized to have specifically approved illegal acquisition methods, the way learning data is procured could become a personal legal risk for AI company CEOs and technical officers in the future.
On Social Media, "AI Companies Should Be Sued" Clashes with "Major Publishers Can't Be Trusted"
This lawsuit is also sparking significant debate on social media and online communities.
On Reddit's technology-related community, a post sharing a TechCrunch article gathered thousands of upvotes shortly after being published.
A prominent view is that "AI companies may have factored in litigation risks from the start."
For AI companies with vast resources, if copyright lawsuits and settlements become part of the research and development costs, they may not serve as a deterrent against rights holders.
Users claiming to be musicians also express the opinion that while they use AI daily, if pirate copies were used, strict responsibility should be enforced.
Moreover, there are many opinions that "not only Anthropic but other AI companies should be scrutinized under the same standards."
Since foundational models of generative AI require vast amounts of learning data, there is a question of whether this issue should be resolved as a problem of just one company.
Meanwhile, Skepticism Toward Sony and Warner
Interestingly, users critical of AI companies do not necessarily fully support Sony or Warner.
On social media, there are repeated observations that "just because major music companies are filing lawsuits doesn't necessarily mean it benefits artists."
In the traditional music industry, conflicts over revenue sharing between artists and record companies or publishers have long persisted.
Therefore,
"While opposing unauthorized use by AI companies, should we equate large rights holders receiving huge compensation with 'protecting creators'?"
This complex reaction is emerging.
Especially for individual creators, they lack the financial power to fight massive lawsuits over several years like large corporations.
Even if rules favorable to copyright holders are formed through this trial, whether the benefits will reach small musicians, writers, and illustrators is another matter.
The Counterargument: "How Is AI Learning Different from Humans Listening to Music?"
Users relatively positive about AI are also presenting long-standing counterarguments.
Human musicians listen to a lot of past music and create new songs under its influence. So why is it called copyright infringement when AI learns patterns from a large number of works?
However, this comparison alone cannot resolve the issues in this lawsuit.
Even if the "learning" of AI models is similar to human learning, if works were copied from pirate sites in the preliminary stage, it becomes a different issue.
Just as reading books in a library is not the same as copying millions of books from pirate sites to create a personal database, AI also needs to separate "the nature of learning" from "the method of data acquisition."
This lawsuit is likely to question precisely that boundary.
The Real Shift: From "Allowing AI to Use Copyrighted Works" to "How to Source Them"
The copyright debate over generative AI has tended to focus on the single question of "Is AI learning fair use?"
However, the situation is changing.
Even if courts recognize certain AI learning as fair use, that does not make all data use by AI companies legal.
Where was the data obtained from?
Were the terms of use followed?
Was it not recognized as pirated?
Was the copyright notice not removed?
Were sufficient measures taken to prevent the model from reproducing the original works?##HTML_TAG_277