On August 25, 2026, Figure AI came out of stealth with something that wasn't a robot. It was an app called Index.
The app pays people in 108 countries to record themselves doing everyday tasks — folding laundry, washing dishes, cooking, cleaning. In four months of testing, Index accumulated 16 million videos from more than 100 countries. Figure paid creators $15 million. The company has committed to spending more than $1 billion on data and compute over the next 12 months.
Index is not a side project. It is Figure's core bet on how to win the humanoid robotics race. And it reveals a fundamental divergence in how the U.S. and China are approaching the single biggest bottleneck in physical AI: data.
What Figure Actually Bought With $1 Billion
Figure's thesis is simple. Large language models scaled because the internet already had trillions of tokens of text. Robots don't have an equivalent source. Physical interaction data must be collected from the real world.
So Figure built a crowdsourcing machine. Anyone can download Index, record themselves doing tasks, and get paid. The platform now has over 264,000 downloads, more than 44,000 weekly active contributors, and processes roughly 30 minutes of uploaded video every second. Every 1,000 hours of collected data contains 373 unique tasks, 1,146 unique objects, and 116 unique environments.
The goal is diversity at scale. Thousands of contributors across 108 countries generate variations that would be impossible to engineer in a lab. The data feeds Figure's Helix model — a 7-billion-parameter vision-language model paired with an 80-million-parameter visuomotor network.
But there are catches.
First, data quality is hard to control. Figure acknowledges the need for extensive filtering, fraud detection, deduplication, and annotation pipelines. The company processes videos through a five-stage pipeline: filtering, fraud review, deduplication, rebalancing, and annotation. That infrastructure costs money — part of the $1 billion.
Second, embodiment mismatch. Human actions are not robotic actions. A video of a human folding laundry is not the same as a robot folding laundry. The greater the gap between human demonstrator and robot, the harder the transfer learning problem becomes.
Third, the cost structure is relentless. Figure is paying creators per video — about $0.94 per video on average. At 16 million videos, that's $15 million. To scale 100x, as the company plans, the payout scales with it. This is not a one-time investment. It is a recurring operational expense that grows with every new capability the company wants to teach.
The Chinese Alternative: Less Money, More Engineering
Chinese robotics companies are taking a different path. They are not betting billions on crowdsourced video. They are betting on a hybrid strategy: simulation, teleoperation, and targeted real-world data — each used where it is most cost-effective.
Simulation first. Synthetic data is dramatically cheaper than real-world collection. Industry estimates put simulation costs at roughly 50 yuan per 10,000 frames, while the overall synthetic data market is valued at about 500 to 600 million yuan. One analysis found that 90% of training data in some Chinese robotics pipelines relies on synthetic simulation, at costs as low as 1% of real-world data.
Chinese companies have built sophisticated simulation pipelines to generate training data at scale. Lightwheel AI, a Beijing-based physical AI simulation infrastructure company, has developed the world's first industrial-grade simulation evaluation platform, RoboFinals, designed to accelerate embodied intelligence training. WUWENAI has deployed a "world model × data factory × world simulator" integrated solution, creating a closed-loop system for virtual-real fused data.
Teleoperation for precision. When simulation isn't enough, Chinese companies use teleoperation — humans remotely controlling robots to demonstrate specific tasks. This is the "gold standard" for data quality. But it is expensive: teleoperation equipment costs over 200,000 yuan ($28,000) per unit, and labor runs about 300 yuan ($42) per day.
Companies use teleoperation strategically — only for tasks that simulation cannot handle. Zhiyuan Robotics operates a 4,000-square-meter data collection facility in Shanghai, where it generates thousands of data samples per robot per day for precision assembly and complex manipulation. But the cost is significant: one hour of effective teleoperation data can cost 1,000 yuan ($140) or more. Even with optimized workflows, professional teleoperators produce only 2 to 3 hours of usable data per 8-hour shift.
Real-world data for grounding. Finally, Chinese companies collect real-world data — but not through crowdsourcing. They deploy robots in actual factories and homes, collecting data as a byproduct of operations. Unitree Robotics, which shipped over 5,500 humanoid robots in 2025, uses its hardware footprint to generate a "data flywheel". The company has open-sourced datasets like UnifoLM-WBT, with 340 hours and 1.89 million trajectory data points collected from real G1 robots in actual home and industrial environments. It has also partnered with Hugging Face to release HIW-500, the largest open-source humanoid teleoperation dataset collected in real home environments.
Zhiyuan Robotics has built a 2,000-square-meter real-world collection facility covering 217 tasks and over 3,000 objects, generating 850 terabytes of data. NVIDIA's GR00T N1 model reportedly uses over 80% real-robot training data from Zhiyuan's open-source dataset.
The Cost Curve Comparison
The cost difference is stark.
Figure's model: Pay humans per video. Scale with volume. At $0.94 per video and 16 million videos, that's $15 million already spent. To reach the diversity needed for general-purpose robots, the cost scales linearly with every new task, object, and environment. The $1 billion commitment is not a ceiling — it is a down payment.
China's model: Pay for simulation infrastructure once, then scale marginal cost near zero. Pay for teleoperation only for high-value tasks. Collect real-world data as a byproduct of deployed robots. The cost curve is front-loaded (infrastructure) but flat thereafter.
One industry analysis estimated that a million hours of real-world teleoperation data would cost over 1 billion yuan ($140 million) to collect. The same volume of simulation data would cost a fraction of that — and can be generated in weeks, not years.
The math is unforgiving. Figure is spending billions to collect data from humans who are not robots. Chinese companies are spending millions to build systems that generate data from robots that already exist.
The Factory Density Advantage
There is one more factor that rarely appears in Western analysis: China has more robots in the real world than anyone else.
Unitree shipped over 5,500 humanoids in 2025. Zhiyuan shipped 5,168.
Together, they represent roughly 39% of global humanoid shipments.
Every one of those robots, deployed in factories, warehouses, and homes, is a data collection device. Every task it performs generates training data for the next generation.
Figure, by contrast, has shipped approximately 350 robots. Its data comes from humans with phones, not robots with sensors. The embodiment mismatch problem — human actions are not robotic actions — means Figure must eventually bridge the gap between what humans demonstrate and what robots can execute.
Chinese companies start with robots. The data is already in the right format.
What This Means
The $1 billion data gamble reveals two competing theories of how to scale physical AI.
Figure's theory: Data is the bottleneck. The company is spending aggressively to collect it from humans at global scale, building a massive, diverse dataset and trusting the model to figure out the rest.
China's theory: Data is the bottleneck, but the solution is engineering, not spending. Use simulation for scale, teleoperation for precision, and deployed robots for grounding. Keep costs low and iterate fast.
One approach is capital-intensive and linear. The other is engineering-intensive and compounding.
Figure may win the data race in the short term. But the company is spending like data will always be scarce. Chinese companies are building systems that assume data will eventually be abundant — and the winners will be those who can generate it cheapest.
The question is not whether $1 billion is enough. The question is whether spending $1 billion on human videos is the right way to solve a problem that Chinese engineers are solving for a fraction of the cost.
Sources: ARC Web (August 28, 2026); 36Kr (August 27, 2026); AgentLocker (August 27, 2026); ezone.hk (August 27, 2026); Unscarcity.ai (August 29, 2026); The Paper (April 24, 2026); EO Intelligence (July 2026); TMTPost (July 17, 2026); China Association of Mechatronics Technology Application (July 2026); The Paper (April 12, 2026); Zhidx (June 26, 2026).
Disclaimer
The information provided in this article is for general informational and educational purposes only. It does not constitute legal, financial, or professional advice. The author and publisher are not responsible for any actions taken based on the content of this article. Readers should consult qualified professionals for advice specific to their situation. All trademarks and references to third-party products, services, or organizations are the property of their respective owners. The performance data and benchmarks discussed are based on specific research studies and may not generalize to all use cases or environments. As of the publication date, the AI landscape continues to evolve rapidly, and readers should verify current information independently.
Limitations
This analysis is based on reporting and public data available as of the article date; figures may be revised as sources update.
Forecasts from third-party analysts can change with market conditions.
Cost and pricing examples are point-in-time estimates; actual rates vary.
Country and company comparisons rely on public reporting, not operational data.
This sector moves fast; timelines and deal terms may be updated later.
Company deals and regulatory rulings may evolve; verify current status.
AI infrastructure is changing quickly; claims can become outdated soon.
Sources
- ARC Web (August 28, 2026)
- 36Kr (August 27, 2026)
- AgentLocker (August 27, 2026)
- ezone.hk (August 27, 2026)
- Unscarcity.ai (August 29, 2026)
- The Paper (April 24, 2026)
- EO Intelligence (July 2026)
- TMTPost (July 17, 2026)
- China Association of Mechatronics Technology Application (July 2026)
- The Paper (April 12, 2026)
- Zhidx (June 26, 2026).
The information provided in this article is for general informational and educational purposes only. It does not constitute legal, financial, or professional advice. The author and publisher are not responsible for any actions taken based on the content of this article. Readers should consult qualified professionals for advice specific to their situation. All trademarks and references to third-party products, services, or organizations are the property of their respective owners. The performance data and benchmarks discussed are based on specific research studies and may not generalize to all use cases or environments. As of the publication date, the AI landscape continues to evolve rapidly, and readers should verify current information independently.
Limitations: This analysis is based on reporting and public data available as of the article date; figures may be revised as sources update.; Forecasts from third-party analysts can change with market conditions.; Cost and pricing examples are point-in-time estimates; actual rates vary.; Country and company comparisons rely on public reporting, not operational data.; This sector moves fast; timelines and deal terms may be updated later.; Company deals and regulatory rulings may evolve; verify current status.; AI infrastructure is changing quickly; claims can become outdated soon.