In March 2026, DoorDash launched an app that pays its couriers to strap on body cameras and film themselves washing dishes, folding laundry, and making beds. Not to improve food delivery. To manufacture training data for humanoid robots.
That is what a binding constraint looks like when capital runs into it. Roughly six billion dollars went into humanoid robotics in 2025, and it did not buy the one input the field actually needs, because that input does not exist yet and cannot be purchased. It has to be recorded, in the physical world, one interaction at a time.
This changes where the contest is decided, and almost every analysis of China’s position in embodied AI is watching the wrong scoreboard.
The bottleneck is not compute, and the gap is not close
Start with the number that should reframe the whole field.
A language model trains on a corpus of somewhere between 1.5 and 4.5 billion examples. That data already existed. It was scraped, not commissioned, and the internet produced it as a byproduct of people going about their lives.
Robots have no equivalent. A robot learning to wipe a counter needs synchronized traces of vision, force, joint position, and motor command, captured during a real physical interaction. Nothing on the internet contains that. The pixels of a YouTube video do not carry the torque a wrist applied.
So the data has to be collected, and collection is brutally slow. To train RT-1 in 2022, Google ran 13 robots for 17 months in an office kitchen and came away with 130,000 trajectories across roughly 700 tasks. Open X-Embodiment, the largest pooled dataset the field has assembled, brought together 60 datasets from 21 institutions and 34 labs across 22 robot types, and reached about one million trajectories.
One million, against a language corpus of one and a half billion. The field’s entire shared dataset, contributed by dozens of the best labs on earth, is roughly three orders of magnitude smaller than what a language model consumes as a matter of routine.
A trajectory and a text example are not the same unit, and anyone comparing them should say so. A trajectory is a rich, seconds-long, multi-sensor recording; a text example may be a single sentence. The comparison is not apples to apples and it does not need to be. A gap of three orders of magnitude survives any reasonable adjustment you want to make to the units. The shape holds even when the measurement is generous.
Compute is not the scarce thing here. GPU clusters capable of training vision-language-action models are commercially available, and have been for a while. The scarce thing is a recording of a hand doing something in the world.
This is not a wall. It is a curve, and its slope is a fleet size.
Here is where most coverage of the data problem stops, and where it goes wrong. Stating that robot data is scarce sounds like a verdict. Run the arithmetic and it becomes something more useful: a schedule.
Take Google’s own collection rate as a baseline. Thirteen robots over seventeen months is 221 robot-months, yielding 130,000 trajectories, which works out to roughly 590 trajectories per robot-month. Now ask how long it would take to reach language-model scale, 1.5 billion examples, at that rate.
With a fleet of 1,000 robots: about 212 years.
With a fleet of 100,000 robots: about 2.1 years.
With a fleet of 1,000,000 robots: about ten weeks.
These are illustrative, not forecasts. Real collection rates vary enormously with task, and nobody actually needs a full 1.5 billion trajectories to build something useful. The point is not the absolute number. The point is the shape.
The data problem is not a physical law. It is a deployment function. The same gap that looks unbridgeable at a thousand robots becomes a two-year project at a hundred thousand. Nothing about the physics changed. The fleet changed.
Which means the question that decides embodied AI is not who has the best algorithm or the most compute. It is who can get the most robots doing real work in real environments, for the most hours, soonest.
The bottleneck resource is not made in a fab
Now sit with what that implies for the way China’s position is usually assessed.
The standard analysis runs through silicon. Export controls limit access to advanced accelerators, therefore Chinese AI is constrained, therefore China is behind in embodied intelligence too. Every part of that chain is defensible except the last step, which quietly assumes that embodied AI is bottlenecked by the same input as a language model.
It is not. The binding constraint on a vision-language-action model is not the flops used to train it. It is the trajectories available to train it on. And trajectories are not manufactured in a fab. They are manufactured wherever robots are doing real work: on assembly lines, in warehouses, on factory floors.
This does not make export controls irrelevant. Compute still matters, and a constrained training budget still hurts. It relocates the decisive variable. In language, the scarce input was compute and the abundant one was data. In embodied AI, that ratio inverts. The scarce input is data, and the country with the largest manufacturing base has the largest surface on which to generate it.
Two systems have read the same scoreboard and answered it differently
China is not inferring this. It is building it.
As of January 2026, the Chinese government had funded 40 dedicated robot training centers, according to Rest of World. Inside facilities like the National and Local Co-Built Humanoid Robotics Innovation Center in Suzhou, human trainers stand beside humanoid robots and repeat motions like folding clothes and wiping tables hundreds of times a day. In Shanghai, workers have spent entire weeks in VR headsets and exoskeletons repeating a single microwave-door motion, over and over, to train the machine next to them. Cities including Beijing and Wuxi are competing to build embodied-AI data infrastructure of their own.
Set that against the American approach and the contrast tells you something about both systems. The United States is crowdsourcing the problem through the gig economy: DoorDash pays couriers piecework to film their own kitchens, Uber runs a parallel program, and startups mail out sensor gloves to volunteers. China is centralizing it as state infrastructure: dedicated facilities, funded from above, generating data as a public input the way a highway generates freight capacity.
Neither approach has been proven superior. But note what the Chinese posture reveals. When a state starts funding dataset production the way it funds semiconductor fabs, it is telling you where it thinks the binding constraint is. The institutional machine has read the same scoreboard, and it has moved.
What this says about the two billion yuan Unitree is about to spend
We have spent three issues on Unitree’s numbers, so run this against them.
The IPO earmarks roughly 2 billion yuan for the embodied-intelligence model build, more than ten times the company’s entire historical research spend. We showed that the money comes out of the hardware margin, that the margin is already falling, and that the valuation holds that margin constant while the company spends it.
Now add the constraint from this piece. That 2 billion yuan can buy compute. It can buy engineers. It can buy teleoperation rigs and simulation licenses. It cannot buy trajectories that were never recorded.
So the spend is a bet, and it is worth naming precisely what the bet is. It is a wager that embodied data can be bought, simulated, or collected fast enough to matter, by a company whose robots are largely sold to researchers and developers rather than deployed on production lines at scale. Every G1 in a university lab is generating data for somebody, and it is not necessarily generating the kind of data that trains a model to do industrial work.
The company that wins the embodied-model race will not be the one that spent the most on models. It will be the one whose robots were already out there, doing something repetitive and real, logging it.
That is a different company profile than the market is currently paying for.
Read the seam, and the seam is between the model and the floor
The honest counterweight has to be stated, because the deployment story is seductive and the obstacles are real.
Data collected on a factory floor is narrow. A robot that has palletized ten million boxes has learned to palletize boxes, not to fold a shirt. Chinese industry participants themselves concede that the sector is mired in data silos, with incompatible formats, metadata conventions, and annotation standards making the data nearly impossible to pool. One prominent view holds that hardware deployment cannot possibly supply most of the training data in the near term, and that simulation must carry the bulk of it, which reopens the question of how well simulated experience transfers to physical reality. None of that is settled, and anybody claiming the factory floor solves embodied AI outright is selling something.
But the direction of the constraint is clear enough to price. Whatever fraction of the data must come from the real world, that fraction comes from deployment, and deployment is a manufacturing question before it is an AI question.
Which brings the argument back to the machine as a whole.
Silicon is the substrate. Models are the intelligence. Robots are the body. Factories are where all of it lands as capacity, and this is the issue where the last layer stops being a destination and becomes an input. The factory floor is not merely where physical AI gets deployed. It is where physical AI gets trained. Deployment is not the end of the stack. It loops back to the top of it.
The seam between the intelligence layer and the deployment layer is where this whole industry will be decided, and it is a seam almost nobody is pricing, because pricing it requires reading a model and a production line at the same time.
Inside China’s Machine is research, not investment advice. Dataset figures for RT-1 and Open X-Embodiment, and language corpus scale, are Confirmed from published research and reporting. DoorDash’s Tasks program and courier count are Confirmed from the company’s announcement and Bloomberg’s reporting of March 2026. The count of 40 state-funded robot training centers is reported by Rest of World as of January 2026 and is Confirmed to that source. The fleet-size arithmetic is the author’s illustrative calculation from RT-1’s disclosed collection rate and is a scenario, not a forecast. Views on simulation’s share of future training data are Estimated and attributed to industry participants. Current as of July 11, 2026.


