On August 17, 2026, Round Hill Music filed two copyright lawsuits in California federal court — one against Anthropic, another against Suno. The complaints alleged that both companies had used hundreds of the publisher's songs without permission to train their AI models. The potential damages, according to Round Hill, could "conceivably exceed $1 billion."
Round Hill is not alone. Universal Music Group and Sony Music Entertainment are still litigating their own copyright case against Suno. Book authors, newspapers, visual artists, movie studios, and record labels have all sued AI labs over training data. The legal theory is consistent: you cannot build a multi-billion-dollar business on other people's creative work and pay them nothing.
The U.S. AI industry is drowning in data litigation.
Meanwhile, in China, the legal picture looks different. Chinese AI companies face a different set of constraints — not copyright lawsuits over training data, but a regulatory framework that makes real-world data collection expensive and legally fraught. The response to that constraint has been unexpected: China has built the world's largest synthetic data industry.
The American Data Crisis
The Round Hill lawsuit is just the latest front in a war that has been escalating for years.
In the Anthropic case, the complaint preliminarily identifies 500 musical compositions — including "Iris" by the Goo Goo Dolls, "Total Eclipse of the Heart" by Bonnie Tyler, and "I Got You (I Feel Good)" by James Brown — with plans to expand to "ten thousand or more" works. The suit against Suno also names Bright Data, a scraping company, as a defendant.
Round Hill CEO Josh Gruss made the company's position clear: "We are not against artificial intelligence. We are against the idea that you can build a business worth billions on top of other people's creative work and pay the creators nothing". The company has stated it intends to take the cases to trial rather than settle.
The fundamental problem is structural. U.S. AI companies built their models by scraping the public internet — often without permission and often in defiance of copyright law. They argued that training AI on publicly available information is comparable to human learning. That argument is now being tested in courts across the country.
Every new lawsuit increases the cost and uncertainty of building AI in America. Every new plaintiff adds another layer of legal risk. The result: AI companies now spend more on legal defense than on data acquisition, and the data they need to train next-generation models is increasingly locked behind litigation.
The Chinese Data Problem — and Its Unintended Solution
China faced a similar problem, but for a different reason.
In November 2021, China's Personal Information Protection Law (PIPL) went into effect. It is one of Asia's most consequential data localization laws — stricter than Indonesia's PDP Law and comparable to Europe's GDPR, but with tighter cross-border transfer rules. Data generated from China-based systems must be processed and stored within China. Processing China-origin data outside the mainland requires an explicit legal basis — and for AI training purposes, that basis is increasingly difficult to establish.
For Chinese AI companies, the problem was immediate: they could not simply scrape the open internet and train models the way American companies did. PIPL requires consent for using personal information in processing activities. Violations carry significant fines — up to 50 million yuan or 5% of a company's annual revenue from the previous fiscal year.
So Chinese companies did something else. They built synthetic data.
Academic research has identified that under PIPL and the Data Security Law, synthetic data generation offers one of the only viable compliance architectures. The logic is straightforward: synthetic data is not real personal information. It does not require consent. It does not trigger cross-border restrictions. It is, from a regulatory perspective, clean.
Chinese AI labs began investing heavily in synthetic data infrastructure — not as an experiment, but as a necessity. The result is a large and growing synthetic data industry, built not by choice, but by legal constraint.
What Synthetic Data Actually Does
Synthetic data is not fake data. It is manufactured data — generated by AI models themselves, designed to preserve the statistical properties of real data without containing any actual personal information. By combining generative AI with differential privacy, these systems can produce datasets that are statistically equivalent to real data but free of any identifiable individual information.
In robotics, simulation data can cost roughly 50 yuan per 10,000 frames — a fraction of what real-world collection costs. In language modeling, synthetic data can be generated at marginal cost near zero.
But there is a catch. Naive recursive training on synthetic data risks model collapse — a degenerative process where repeated training on model-generated outputs erodes distributional tails and homogenizes outputs. When generative models are iteratively trained on their own outputs, performance can degrade due to narrowed coverage and accumulated bias.
Chinese labs have been working on this too. Multiple research papers from Chinese institutions have explored mitigation strategies for model collapse. Some work challenges the inevitability of collapse under certain conditions. The underlying insight is that synthetic data can work — but only if you design for it.
The result is a production-grade synthetic data industry. Chinese labs now routinely use synthetic data for pre-training, fine-tuning, and reinforcement learning — not just as a supplement, but as a primary data source.
Two Models, Two Outcomes
The contrast is stark.
In the U.S.: AI companies are locked in litigation over who owns the training data. Every new lawsuit adds uncertainty. Every new plaintiff increases the cost. The legal system is effectively rationing data, and the rationing is slowing American AI development.
In China: AI companies are locked in a regulatory framework that makes real data expensive and risky. The solution — synthetic data — has become an industry. Chinese labs now have production-scale synthetic data pipelines that many U.S. labs are only beginning to explore.
One system is fighting over the past. The other is building the future.
The Numbers Tell the Story
The American litigation machine is not slowing down. Round Hill's lawsuits are the latest in a long line. Anthropic and Suno have faced numerous such lawsuits across industries. The damages being sought are enormous — not just millions, but potentially billions.
Meanwhile, the Chinese synthetic data industry is scaling. The country's data element market is being institutionalized through the National Data Administration. In March 2026, the National Data Administration released the "Three-Year Action Plan for High-Quality Dataset Construction (2026-2028)" — China's first national-level special plan specifically for dataset construction. The plan explicitly calls for "vigorously developing synthetic data technology" to address the pain points of sample scarcity and incomplete coverage in key areas.
In June 2026, the administration followed up with the "Implementation Plan for Promoting Industry High-Quality Dataset Construction Actions" , which calls for applying simulation and synthesis technologies to expand data supply. The plan also promotes the exploration of token-based value systems for dataset trading. Data is being treated not as a legal liability, but as a tradable asset.
The U.S. is fighting over who owns the data. China is building systems that don't need to ask.
What This Means
The Round Hill lawsuits are not an anomaly. They are the natural outcome of a system where the legal framework for data ownership was written for the 20th century, and the AI industry operates in the 21st. The courts are being asked to resolve questions that no one anticipated — and the answers are coming slowly, expensively, and unpredictably.
China's solution is not without its own problems. Synthetic data can introduce bias. Model collapse is a real risk. The regulatory environment is still evolving. But the direction is clear: China has built a data infrastructure that is legally compliant, scalable, and increasingly sophisticated.
The next time someone tells you that AI is a technology race, remind them: it is also a data race. And the U.S. is fighting yesterday's war.
Sources: Billboard (August 17, 2026); Hollywood Reporter (August 17, 2026); Variety (August 17, 2026); TechRepublic (August 19, 2026); DLA Piper PIPL analysis; ejournals.eu PIPL compliance research; National Data Administration policy documents (March–June 2026); arXiv:2607.17043 "Learning from Synthetic Data without Model Collapse" (July 2026).
Disclaimer
The information provided in this article is for general informational and educational purposes only. It does not constitute legal, financial, or professional advice. The author and publisher are not responsible for any actions taken based on the content of this article. Readers should consult qualified professionals for advice specific to their situation. All trademarks and references to third-party products, services, or organizations are the property of their respective owners. The performance data and benchmarks discussed are based on specific research studies and may not generalize to all use cases or environments. As of the publication date, the AI landscape continues to evolve rapidly, and readers should verify current information independently.
Limitations
This analysis is based on reporting and public data available as of the article date; figures may be revised as sources update.
Forecasts from third-party analysts can change with market conditions.
Cost and pricing examples are point-in-time estimates; actual rates vary.
Country and company comparisons rely on public reporting, not operational data.
This sector moves fast; timelines and deal terms may be updated later.
Company deals and regulatory rulings may evolve; verify current status.
AI infrastructure is changing quickly; claims can become outdated soon.
Sources
- Billboard (August 17, 2026)
- Hollywood Reporter (August 17, 2026)
- Variety (August 17, 2026)
- TechRepublic (August 19, 2026)
- DLA Piper PIPL analysis
- ejournals.eu PIPL compliance research
- National Data Administration policy documents (March–June 2026)
- arXiv:2607.17043 "Learning from Synthetic Data without Model Collapse" (July 2026).
The information provided in this article is for general informational and educational purposes only. It does not constitute legal, financial, or professional advice. The author and publisher are not responsible for any actions taken based on the content of this article. Readers should consult qualified professionals for advice specific to their situation. All trademarks and references to third-party products, services, or organizations are the property of their respective owners. The performance data and benchmarks discussed are based on specific research studies and may not generalize to all use cases or environments. As of the publication date, the AI landscape continues to evolve rapidly, and readers should verify current information independently.
Limitations: This analysis is based on reporting and public data available as of the article date; figures may be revised as sources update.; Forecasts from third-party analysts can change with market conditions.; Cost and pricing examples are point-in-time estimates; actual rates vary.; Country and company comparisons rely on public reporting, not operational data.; This sector moves fast; timelines and deal terms may be updated later.; Company deals and regulatory rulings may evolve; verify current status.; AI infrastructure is changing quickly; claims can become outdated soon.