RWA / THE DATA OPPORTUNITY

Experience is
the scarce part.

Robots need records of what they saw, felt and did. That experience cannot simply be scraped from the web.

~300K hours

Estimated global robot manipulation data

Bessemer estimate · April 2026

A limited supply of diverse, task-relevant experience is a bottleneck for physical AI.

Recorded robot action — conceptual robotics artwork.

Recorded robot action

Cameras, joints and contact signals tied to what the machine actually did.

Collected through physical interaction.
First-person human demonstration of holding a cup.

Human demonstrations

Human activity offers a scalable source of task knowledge and visual experience.

Human video is not robot-action data.
Simulated robot hand with a movement trajectory.

Simulation at scale

Virtual environments multiply practice across tasks, objects and conditions.

Real-world transfer still needs validation.

SUPPLY & AMBITION

Scaling the experience.

Different kinds of data are growing at different speeds.

DISCLOSED ROBOT DATA~10K h

Physical Intelligence · π0

Robot demonstrations used in the original 2024 model. A concrete historical example of directly collected experience.

HUMAN INTERACTION500K+ h

Generalist · GEN-1

Physical interaction pretraining data reported in April 2026. Generalist explicitly says this pretraining contains no robot data.

HUMAN VIDEO1M+ h

Dyna · Dyna-2

Egocentric human video reported in 2026. Million-hour collections are emerging, but are not interchangeable with robot hours.

A LONG-TERM RESEARCH VISION100M hours

Researcher Joel Jang describes a future trained on 100 million hours of human egocentric video. It is a personal research outlook, not an agreed requirement or a company collection target.

There is no audited industry-wide total or agreed “hours needed” threshold. Text tokens, video hours and robot experience measure different things; a universal 10×–200× gap cannot be inferred from them.

WHAT A USEFUL TRACE CONTAINS

One action. Many signals.

The connected evidence behind vision-language-action and world models.

  1. Cameras

    Scene & objects

  2. Joint states

    Position & motion

  3. Force & touch

    Contact & pressure

  4. Actions

    Commands & results

  5. Language

    Intent & task context

Built for human spaces.

Humanoids can operate around human tools, furniture and workspaces. Their experience can map to homes, factories and offices—the environments where the tasks happen.

Reality is the final test.

Simulation makes practice cheaper. Real contact remains difficult to reproduce: cloth, liquids and tight insertions expose the gap. Physical trajectories provide direct evidence of what actually works.

LEARNING THROUGH DEPLOYMENT

Work becomes the next lesson.

With capture and review, deployed robots can generate fresh training experience.

  1. 01

    Deploy

    Work in real environments

  2. 02

    Capture

    Record actions and outcomes

  3. 03

    Train

    Curate data and improve models

  4. 04

    Improve

    Validate and return to the field

Context & sources

The ~300K-hour estimate comes from Bessemer, not an audited census. The 500K and 1M-hour disclosures describe human interaction or human video, not a 300K–500K total of robot demonstrations. Million-hour corpora are not evidence of a common 2026 robot-data target, and 10M–100M-hour ambitions are not established requirements for general-purpose humanoids. Text corpora are measured in tokens; they cannot be converted into equivalent physical experience by a single reliable multiplier.

  1. Bessemer — Robot manipulation data estimate, April 2026
  2. Physical Intelligence — π0 training data
  3. Generalist — GEN-1 and its data mixture
  4. Dyna — Dyna-2, million-hour human-video scaling
  5. Joel Jang — 100M-hour research vision, personal views
  6. Skild AI — Data sources and deployment transfer
  7. Humanoid Everyday — Multimodal humanoid data research
  8. Goldman Sachs — Deployment, data and capability feedback