For years, the dominant Western narrative about data labeling has been built around a single image: rows of low-paid workers in developing countries, clicking boxes around objects in images, transcribing audio clips, earning pennies per task. It is the mental picture most Westerners have of the industry—the "AI sweatshop."
That picture is roughly eight years out of date.
In 2026, China's data labeling industry has completed a transformation that few outside the industry have noticed. The work is no longer about "clicking boxes." It is about teaching AI to reason. And the companies that dominate this space are no longer labor brokers. They are technology companies with AI-powered platforms, robotics labs, and global footprints.
The "sweatshop" narrative is inaccurate enough to mislead Western observers about who controls the data supply chain—and who will capture the value at the top of it.
What the 2026 Labeling Industry Actually Looks Like
Data labeling in 2026 is a three-tier system. According to leading Chinese data service providers, the breakdown is roughly 10% to 30% fully manual (used only for novel, long-tail tasks where no pre-trained model exists), 50% to 70% human-machine collaboration (the industry standard), with the remainder representing deeper automation in standardized tasks.
The dominant mode is human-machine collaboration, or "human-in-the-loop." The workflow works like this: machines do the bulk of the labeling work, and humans step in to handle edge cases, verify quality, and correct errors. A system might call two or three different pre-trained models to independently generate labels, cross-fuse them, and then use active learning to identify samples where the models disagree, have low confidence, or show bias. Those samples are flagged for human review. Everything else passes automatically.
As one industry analyst put it, the industry has moved from "selling labor" to "selling assets." Instead of charging per labeled data point, providers are selling API access, full-stack solutions, and even exploring token-based trading and data subscriptions. The relationship between data service providers and their clients has shifted from outsourcing to long-term strategic partnership.
The scale of the industry tells the story. According to the National Data Administration's Digital China Development Report (2025), China had over 110,000 high-quality datasets totaling more than 908 petabytes by the end of 2025, representing year-over-year growth of 61% and 142% respectively. The cumulative scale of data labeling exceeded 85 petabytes, generating 18.3 billion yuan ($2.5 billion) in related output. Seven pilot cities for data labeling have supported 260 large model development projects, cultivated 364 labeling enterprises, and employed 95,000 labeling professionals.
The global data labeling solutions and services market reached $20.4 billion in 2025, according to industry research, growing at a compound annual rate of 24.5%.
This is not a sunset industry. It is a growth industry—and it is growing in sophistication.
The Companies Driving the Shift
Haitian Ruisheng, a publicly traded Chinese data company, is a case study in the transformation. In the first quarter of 2026, the company reported a 2,161% year-over-year increase in net profit, driven by the shift toward higher-margin, platform-based services. The company's 2025 annual revenue reached 377 million yuan ($52 million) , up 59% year-over-year. Its training data application services revenue grew by 1,394%.
The company operates Beijing's first embodied intelligence data training facility, with over 100 robots generating high-quality data for physical AI systems. Its AI-powered automation platform runs 24/7 algorithmic quality checks, balancing low cost with expert-level quality. The company's business has expanded far beyond basic labeling. It now provides full-stack data services: collection, annotation, platform tools, data management, quality assurance, and talent training. It operates in over 70 countries, supporting more than 300 languages and dialects for data collection and annotation.
Testin Cloud Testing has emerged as another major player. Its "human-machine collaborative labeling" systems achieve accuracy rates of 99.99% in some configurations. The company has been recognized as one of China's top AI enterprises and a top data labeling technology provider in the country.
DataTang has developed an AI-powered platform that integrates pre-labeling and assisted labeling, enabling "human-machine interactive semi-automatic labeling."
China Mobile, the state-owned telecom giant, has built a national data labeling infrastructure. In its labeling platform, the company has implemented chain-of-thought labeling—workers are now required to write out the reasoning process behind each label, teaching AI not just what to think, but how to think. Smart labeling adoption exceeds 80%, with efficiency three times higher than purely manual labeling.
The New Frontier: Embodied AI Data
The most advanced edge of China's data labeling industry is embodied AI—training robots to interact with the physical world. This is where the "sweatshop" narrative breaks down completely.
In August 2026, Ant Group's digital data base opened in Wuhan, featuring full-scale replicas of pharmacies, retail stores, restaurant kitchens, living rooms, and bedrooms. Human trainers wear VR headsets and exoskeletons, simulating tasks like dispensing medicine, restocking shelves, preparing food, and doing housework. The system captures visual, tactile, and motion data simultaneously, generating training datasets for embodied AI systems.
The bottleneck is data scarcity. Industry estimates suggest that only about 500,000 hours of usable embodied data existed by the end of 2025, with just tens of thousands of hours actually used for model training—while future demand could reach hundreds of millions of hours. A single hour of effective teleoperation data can cost 1,000 yuan ($140) or more.
To address this, companies are developing low-cost data collection methods. Ant's AoE (Always-On Egocentric) framework uses a smartphone and a low-cost neck-mounted device to capture first-person video of real workers performing actual tasks. The cost drops from "millions of yuan in robotics hardware" to "thousands of yuan in wearable gear." The effectiveness is measurable: for one robot's computer-shutdown task, just 50 teleoperation data samples yielded a 45% success rate. Adding 200 AoE data samples pushed the success rate to 95%.
This is not data labeling as the West imagines it. This is data engineering at the frontier of physical AI.
The Economic Logic
China has built an AI data supply chain that the West has largely overlooked. It is not a low-cost labor play. It is a technology play.
The National Data Administration's 2026 Implementation Plan for Promoting Industry High-Quality Dataset Construction explicitly calls for developing intelligent annotation services including "model pre-annotation + human calibration," "human annotation + model verification," and "model pre-annotation + model verification."
The plan aims to cultivate a tiered ecosystem of data labeling innovation zones, industry leaders, unicorns, and high-growth enterprises.
In Rizhao, a dedicated data labeling industrial base has trained over 30,000 digital talent professionals, working with 58 universities and processing over 6,000 student internships annually. The base is structured as a "front shop, back factory" model: front-end offices in major cities securing contracts from companies like JD.com, ByteDance, and UBTech; back-end operations in the base executing the work.
The goal is not to keep labor costs low. The goal is to build a vertically integrated data supply chain—from collection to labeling to quality assurance to dataset sales—that can power China's AI industry at scale.
What This Means
The "AI sweatshop" narrative describes a version of the data labeling industry that existed in 2018, before AI-assisted labeling, human-in-the-loop workflows, and chain-of-thought reasoning transformed the field.
In 2026, China's data labeling industry is a high-tech ecosystem. It uses AI models to do the bulk of the labeling work, with humans handling edge cases and reasoning tasks. It produces chain-of-thought data that teaches AI how to think, not just what to output. It operates robotics training facilities and generates embodied data for physical AI systems.
The companies that dominate this space are not labor brokers. They are technology companies with global footprints, proprietary platforms, and growing market capitalizations.
Western observers who dismiss data labeling as a low-value, low-skill industry are missing the real story. The data supply chain is not a cost center. It is a strategic asset. And China is building it at scale.
Sources: China Business Journal (June 29, 2026); National Data Administration, Digital China Development Report (2025) (April–June 2026); Haitian Ruisheng investor relations disclosures (April–June 2026); Testin Cloud Data industry reports (2026); China Youth Daily (June 23, 2026); Dazhong Daily (August 25, 2026); Jimu News (August 5, 2026); Stratistics MRC global data labeling market report (2026).
Disclaimer
The information provided in this article is for general informational and educational purposes only. It does not constitute legal, financial, or professional advice. The author and publisher are not responsible for any actions taken based on the content of this article. Readers should consult qualified professionals for advice specific to their situation. All trademarks and references to third-party products, services, or organizations are the property of their respective owners. The performance data and benchmarks discussed are based on specific research studies and may not generalize to all use cases or environments. As of the publication date, the AI landscape continues to evolve rapidly, and readers should verify current information independently.
Limitations
This analysis is based on reporting and public data available as of the article date; figures may be revised as sources update.
Forecasts from third-party analysts can change with market conditions.
Cost and pricing examples are point-in-time estimates; actual rates vary.
Country and company comparisons rely on public reporting, not operational data.
This sector moves fast; timelines and deal terms may be updated later.
Company deals and regulatory rulings may evolve; verify current status.
AI infrastructure is changing quickly; claims can become outdated soon.
Sources
- China Business Journal (June 29, 2026)
- National Data Administration, Digital China Development Report (2025) (April–June 2026)
- Haitian Ruisheng investor relations disclosures (April–June 2026)
- Testin Cloud Data industry reports (2026)
- China Youth Daily (June 23, 2026)
- Dazhong Daily (August 25, 2026)
- Jimu News (August 5, 2026)
- Stratistics MRC global data labeling market report (2026).
The information provided in this article is for general informational and educational purposes only. It does not constitute legal, financial, or professional advice. The author and publisher are not responsible for any actions taken based on the content of this article. Readers should consult qualified professionals for advice specific to their situation. All trademarks and references to third-party products, services, or organizations are the property of their respective owners. The performance data and benchmarks discussed are based on specific research studies and may not generalize to all use cases or environments. As of the publication date, the AI landscape continues to evolve rapidly, and readers should verify current information independently.
Limitations: This analysis is based on reporting and public data available as of the article date; figures may be revised as sources update.; Forecasts from third-party analysts can change with market conditions.; Cost and pricing examples are point-in-time estimates; actual rates vary.; Country and company comparisons rely on public reporting, not operational data.; This sector moves fast; timelines and deal terms may be updated later.; Company deals and regulatory rulings may evolve; verify current status.; AI infrastructure is changing quickly; claims can become outdated soon.