Is Robotics Paying for its Own Data Layer?

Every software revolution is preceded by a hardware revolution. The iPhone before the App Store, twenty years of Broadcom before the iPhone, gaming GPUs training AlexNet. Looking back on most of these cases, the hardware was paid for by someone else, for a different reason. Is robotics the first one trying to pay for its own data enabling layer?

The internet gave language and video models their pretraining corpus for free. Without a learning paradigm breakthrough, robotics has to manufacture its own, one trajectory and one real-world success at a time.

SCARCE ▲ ABUNDANT ▼ T1 T1 · WEB VIDEO & TEXT T2 T2 · HUMAN EGO VIDEO T3 T3 · SIM + WORLD MODELS T4 T4 · HANDHELD CAPTURE T5 T5 · TELEOP DEMOS T6 T6 · ON-POLICY EXPERIENCE
Figure 001 — Robot training data, sorted by what it costs to make. Each tier up is roughly 10x the price.

Robotics inherits some of its stack. Motors and batteries have benefited from smartphone and EV progress (Unitree is this exactly). Video semantics come from a billion hours of already established YouTube content. (I love Taylor Swift, she is my top artist every year on Spotify Wrapped, but I still call Physical Intelligence’s ‘coke bottle to Taylor’ inflection point underwhelming.) But the action data layer in robotics has not had its moment in the past. Unfortunately, nobody in the 2000s collected millions of hours of robot-arm trajectories.

It is one reason among many why atoms are going to be a decisive factor in the AI race. Everyone already believes this for compute: hyperscalers will spend $725B on capex in 2026. But in robotics, the intelligence is made of atoms. Right now it is comically underpriced. The entire projected two-year robot-data market of ~$3B is about 0.2% of a single year of compute capex.

One tempting but mistaken reason to value the robotics data market so low is that learning against a cheap, automatic verifier often replaces expensive human data. We’ve seen this in text, math, video, and even humanoid locomotion. But it doesn’t work where the verifier is made of atoms, whether the task is grasping or DeepMind’s video-model evaluator, which had to be calibrated against 1,600+ real-world robotic trials. Today’s best sim-to-real simply isn’t good enough and the reasoning here holds until it improves dramatically.

Eventually the vertically integrated player who owns the hardware, the successes and failures, and the verifier will win because they hold the speed and flexibility that no one else outside of China has. I don’t have a humanoid folding my laundry yet, though. So until this happens, robotics competitors will pay the data collection toll in an attempt to reach the vertical integration level of scale. The rest of this essay maps where the robotics data market stands now.


Robotics Data Layer Landscape

In June 2015, the most sophisticated robots to date competing for a $2 million DARPA prize kept losing fights with a basic set of stairs. DARPA asked machines to do things your cat does without thinking. Not even MIT’s Atlas robot made it back to its feet after falling. “Fall seven times, stand up eight” apparently wasn’t a behavior hard-coded in the ~650,000-line codebase.

I would have predicted that vision would be solved before language. It’s the order that evolution did it in, and it is such a richer source of information. But the approach to physical learning has played out to be closer to the way language models learned. And this type of learning runs on data.

Today’s leading open models pretrain on tens of trillions of words but a robot’s ‘internet’ archive has to be performed. Embodied data used for training ranges from free web video at the base of the data pyramid to on-the-job learning at the top of the pyramid, with each layer costing an order of magnitude more to acquire.


Each tier in the physical data pyramid from Figure 001 asks the question: what does this data teach that the cheaper stuff below it can’t?

T1 — Web data. Pretty much free. Usually in the form of a vision-language model pretrained on internet text and images with an action decoder bolted on. In Physical Intelligence’s own ablations of its home-robot model π0.5, removing the web co-training dropped success on out-of-distribution tasks from 94% to 74%.

T2 — Egocentric human video. Two different products share this tier. Scraped ego-centric video is free. Commissioned capture using smart glasses, wrist cameras, and hand-tracking worn by paid collectors is performed for ~$20/hr. It’s the fastest-growing tier in the market. NVIDIA’s EgoScale corpus runs 20,854 hours, and claims the first scaling law for robot dexterity when task completion doubled after growing human video from 1,000 to 20,000 hours.

T3 — Synthetic data. A great example is GR00T’s generation of 780,000 simulated trajectories in 11 hours of compute. The groundedness is questionable and the sim-to-real gap remains an open problem, especially in tasks like contacting floppy objects.

T4 — Handheld capture. This tier is newer to me. Generalist AI’s GEN-0 training set, which reportedly passed 500,000 hours by mid-2026, appears to be built this way using $300 UMI-style grippers.

T5 — Teleoperation. Expensive leader-follower style action data costing $100–350/hr in the US. Generalizing on this data across tasks and settings is showing promise but still has flaws.

T6 — On-policy experience. This tier is robots in the wild gathering on-the-job experience, like Amazon’s Vulcan picker handling real production stows.

Some takeaways to focus on in the data pyramid:

1. I like to ground in real numbers, so check out just how small the apex is. The model driving Figure’s Helix is trained on about 500 hours of teleoperation, under 5% of total training data. GR00T’s real-robot layer was 88 hours. Whoever wins the race here probably becomes the Tesla/Waymo of robotics.

NVIDIA GR00T N1 (2025)88 HRS
Figure Helix (2025)~500 HRS
PI π0 (2024)10,000+ HRS
Generalist GEN-0 (2025)270,000 HRS
10¹10²10³10⁴10⁵10⁶

Log scale. 88 HOURS ≈ TWO HUMAN WORK-WEEKS.

2. Empirical results show diversity beats volume. A policy’s ability to generalize follows a power law in the number of environments and objects it saw. Demonstrations per environment saturate around 50. π0.5, for example, ran remove-one-ingredient tests and concluded variety of environments matters most, benchmarking at only 31% without multi-environment data, and 49% without data from other robot types.

MORE DEMOS, SAME PLACES DEMOS PER ENVIRONMENT-OBJECT PAIR SUCCESS % SATURATES ~50 SAME DEMOS, MORE PLACES ENVIRONMENT × OBJECT PAIRS (LOG) LIN ET AL., ICLR 2025 — 4 COLLECTORS · 1 AFTERNOON · ~1,600 DEMOS
Figure 003 — Left: piling demos into the same places flattens hard around 50. Right: adding new places keeps paying, as a power law.

3. Each tier leads to capability only when anchored in data above it. For example, Meta’s V-JEPA 2 pretrained on a million-plus hours of ordinary video needed just 62 hours of robot data to hit 65–80% pick-and-place success on a real arm, zero task-specific training.

4. Unlike an image that can be labeled in a few seconds, for a few cents, actions have to be performed, so cost scales with time and fleet size and no cleverness parallelizes it. The hardware bill does continue to improve, web and egocentric video are cheap at $0.10 to $15 per data-hour, but handheld demos are $10–40 and US teleoperation estimates are $100–350 (we are not in Shenzhen). Until cheap evals exist, I’m not sure what each of these data types is worth. The format matters, hours are not fungible, and the real question is whether a data source makes the robot perform better through model changes.

$0.1 $1 $10 $100 $1000 WEB & EGO VIDEO$0.1–15 SIMULATION$1–30 HANDHELD DEMOS$10–40 OFFSHORE TELEOP$20–60 US TELEOP$100–350 ON-POLICY EXPERIENCE

Scale led the vendor market around the data opportunity when it launched its Physical AI Data Engine in September 2025, partnering with frontier labs including Physical Intelligence, Generalist, and Dyna as named customers. XDOF is now out of stealth with $70 million and Mecka with a claimed $100 million run-rate, all chasing Bessemer’s projection of over $3 billion in industry data spend across two years. If the AV era is the template, I expect one breakout and a long tail.

Looking ahead, every lab is now racing towards robots manufacturing their own data in real deployed settings. Physical Intelligence’s π*0.6 with RECAP added reinforcement learning so the robot improves from its own successes and failures on top of autonomous experience and human corrections (I volunteer my La Marzocca!). PI reports the recipe more than doubled throughput and cut failure rates by half. 1X ships its NEO home robot with teleop fallback as the explicit data strategy. Tesla is pursuing the same deployed-learning loop with Optimus.

So here’s my bet, so you can score it in the future: by July 4, 2028, only two of Physical Intelligence, 1X, Figure, Tesla Optimus, Google DeepMind will have >50% of their new training hours come from deployed robots rather than dedicated data-collection operations. We need this progress; I hope there are more! I am tired of cleaning my house and taking packages to the post office.