Robohouse ’26 Library
Contents

Chapter 7

The Robot Data Supply Chain

6 sections · about 7 minutes

7.1 Teleoperation vendors and human data farms

The data-labelling industry has repositioned wholesale around "Physical AI," but only some of it genuinely touches robots.

Verified robot-data operators. Scale AI runs a Physical AI Data Engine claiming 1,000+ hours of demonstration data uploaded daily, across three collection modalities: teleoperated bimanual manipulators, Scale Harness (an egocentric wearable rig for collection without a robot), and operation of customer hardware — across data factories, residential homes and industrial sites. Objectways explicitly lists teleoperation demonstrations, egocentric capture, robot trajectories and RGBD collection, with 2,200+ specialists across eight US and India sites. Encord markets itself as "the multimodal data layer for Physical AI," supporting LiDAR and sensor-fusion annotation plus teleoperation facilities; customers include Woven by Toyota and Zipline. Micro1 lists a robotics line for high-fidelity real-world robotics data.

Do not assume. Mercor — valued at $10B in October 2025 — is an expert marketplace for knowledge work; no evidence of robot data collection was found. Turing's robotics claims could not be verified.

Skeptic's view: teleoperation hours are the new token count — easy to sell, hard to audit for diversity, and structurally success-biased, because operators do not deliberately produce failures and vendors are not paid for them. Chapter 8 explains why that bias matters more than the raw number.


7.2 Egocentric and wearable human data

The bet: humans wearing sensors are the cheapest hands in the world, and about a billion times more numerous than robots.

Ego4D (February 2022) set the template — 3,670 hours of video from 923 participants across 74 locations in 9 countries, a 13-university consortium with Meta AI, and benchmarks for hand-object interaction and action forecasting.

Ego-Exo4D (December 2023) added the critical third-person view: 1,286.3 hours, 5,035 takes, 740 participants, 13 cities, 123 sites, with Aria glasses time-synchronised to four or five stationary GoPros, plus seven-microphone audio, IMU, eye gaze, 6-DoF localisation and point clouds. Cooking alone is 564 hours.

Project Aria now serves 200+ partner institutions. Aria Gen 2 adds four SLAM cameras, two eye-tracking cameras, PPG heart-rate sensing, a contact microphone, GNSS, on-device hand and eye tracking, and 6–8 hours of battery versus 1–2 for Gen 1.

EgoDex (Apple, ICLR 2026) contributes 829 hours of Apple Vision Pro egocentric video with 3D finger tracking over 194 tasks — the most robot-usable egocentric corpus because the hand pose is already estimated.

Skeptic's view: egocentric video gives you hands and intent but no forces, no torques, and no action labels. The inverse-dynamics step is where the information is lost, and it is not a solved problem. Meta's own V-JEPA 2 result — 1M+ hours of video plus 62 hours of robot data — is the honest framing: video is a prior, and you still need robot data to ground it.


7.3 Low-cost open data-collection hardware

The most consequential cost curve in robotics, and the layer where academia is unambiguously ahead of industry.

SystemOriginWhat it isKey number
UMIStanford/Columbia/TRI, RSS 2024Handheld parallel-jaw gripper + wrist GoPro111 demos/hour vs 35 for SpaceMouse teleop; ~48% of bare-hand speed; two-minute setup in a new environment
DexUMIStanford/Columbia/CMU/NVIDIA, CoRL 2025Wearable exoskeleton for multi-fingered hands + robot-hand video inpainting86% average success on XHand and Inspire Hand
DEXOPMIT Improbable AI + Adelson, 2025Passive exoskeleton mechanically coupling human to robot fingers, with full-hand tactile captureHigher task performance per unit collection time than teleoperation
ALOHA / Mobile ALOHAStanford, 2023–24Bimanual leader-follower teleop rig, then with a mobile baseUp to 90% success from 50 demos/task when co-trained with static data
TWIST2Stanford Movement Lab, Nov 2025Mocap-free whole-body humanoid capture: VR headset + $250 2-DoF robot neck~100 demos in 15 minutes
LeRobot + SO-100/SO-101Hugging FaceOpen software stack + sub-$500 armsReachy Mini at $299/$449

The intellectual move that unifies UMI, DexUMI and DEXOP is worth stating plainly, because it is the most important idea in robot data collection since teleoperation itself: rather than making the robot easier for a human to drive, make the human's natural demonstration already be in the robot's action space. DEXOP is the purest version — a mechanical linkage means the human's finger motion is the robot's finger motion, and the tactile sensors are on the same surfaces, so no retargeting or inverse-dynamics estimation is needed at all.


7.4 The large open datasets

DatasetDateScaleNotes
Open X-EmbodimentOct 20231M+ trajectories, 22 embodiments, 527 skills, 160,266 tasks, 60 datasets, 34 labsThe canonical cross-embodiment pool; RT-1-X and RT-2-X trained on it
DROID202476,000 trajectories, 350 hours, 564 scenes, 86 tasks, 50 collectors, 13 institutions, 12 monthsStandardised Franka + ZED rig; the diversity benchmark
BridgeData V2CoRL 202360,096 trajectories (50,365 teleoperated, 9,731 scripted), 24 environments, 13 skillsWidowX 250; the low-cost precedent
AgiBot World ColosseoMar 20251,001,552 trajectories, 2,976.4 hours, 217 tasks, 100+ robots, 4,000 m² facilityLargest purpose-built real-robot corpus; GO-1 reports ~30% gain over Open-X pretraining
RoboMINDRSS 2025107k trajectories, 479 tasks, 96 object classes, four embodiments, plus 5,000 labelled failuresThe failure labels are rare and valuable
Galaxea Open-World202511 real residential, retail and office sitesMobile dual-arm, genuinely in-the-wild
EgoDexICLR 2026829 hours human egocentric with 3D finger tracking, 194 tasksHuman, not robot — see §7.2

Two observations. First, the Chinese entries now dominate on raw scale, and AgiBot World specifically is the largest purpose-built corpus in existence. Second, almost nobody publishes failures. RoboMIND's 5,000 labelled failures are an outlier, and the absence of failure data is arguably the field's most under-discussed dataset problem — you cannot learn a good reward model, or a good uncertainty estimate, from a corpus of successes.


7.5 Simulation stacks

Isaac Sim is now Apache-2.0 open source on Omniverse, with Isaac Lab as the GPU-accelerated learning framework and Isaac Lab Arena for scalable in-sim policy evaluation. The consolidating development is Newton: an open-source GPU physics engine co-developed with Google DeepMind and Disney Research, governed by the Linux Foundation, now selectable in Isaac Lab alongside PhysX, Warp and MuJoCo. That is the first credible attempt at a shared physics substrate for the whole field, and its governance structure is deliberately neutral.

MuJoCo Playground wraps MJX/JAX and MuJoCo-Warp with locomotion, dexterous and vision-based environments framed explicitly around sim-to-real.

ManiSkill3 (RSS 2025, SAPIEN-based) claims 30,000+ FPS RGBD-plus-segmentation collection on a single RTX 4090 with heterogeneous parallel scenes.

Genesis unifies rigid, FEM, MPM, PBD/SPH and IPC solvers. Its original speed claims drew community scepticism and current documentation no longer foregrounds them.

Drake (TRI/MIT) remains the rigorous outlier — optimisation-first, hydroelastic contact, favoured where correctness matters more than throughput.

OmniGibson / BEHAVIOR-1K provides 1,000 household activities, 50 interactive scenes and 10,000+ objects with fluids, heat, cloth and transparency.

Synthetic-data generators. GR00T-Dreams / DreamGen fine-tunes Cosmos-Predict, generates robot video from one image plus a language instruction, then extracts actions with an inverse-dynamics model into LeRobot format. NVIDIA Cosmos is the world-foundation-model line for this. World Labs' real-to-sim-to-real engine (§5.10) and MSL's Gaussian-splatting reconstructions (§2.9) attack the same problem from the reconstruction side.

The unresolved question: simulation solves appearance diversity and kinematic diversity well and contact dynamics badly. Every honest practitioner in this book agrees on that; they disagree only on whether contact matters enough to be disqualifying.


7.6 Evaluation and benchmarking

Robot evaluation is the field's least-solved problem, and unusually, everyone knows it.

The structural difficulty: real-robot evaluation is unreproducible — different labs, lighting, object instances, reset procedures and human judgement — while simulation benchmarks saturate and correlate imperfectly with hardware.

LIBERO (Spatial/Object/Goal/Long/100 suites) has become the default VLA leaderboard and is widely regarded as near-saturated. (Note: "LIBERO is solved" is community lore, not a documented result from the project itself.)

SIMPLER (2024, UCSD/Stanford/Berkeley/DeepMind) attacks the correlation question head-on: green-screening and texture matching close the visual gap, system identification closes the control gap, and it introduces Mean Maximum Rank Violation (MMRV) to measure whether simulation ranks policies the way real hardware does — a much more useful target than absolute success rate.

RoboArena (Levine et al., 2025) abandons standardisation entirely in favour of crowd-sourced, double-blind pairwise comparisons on distributed DROID platforms — 7 institutions, 600+ real-robot pairwise trials, 7 generalist policies — arguing this ranks policies more accurately than centralised evaluation ever could.

BEHAVIOR Challenge 2026 pushes the opposite direction: 100 full-length household tasks in 7 scenes, 20,000 teleoperated demos totalling 1,950 hours, scored by BDDL partial credit, with π0.5 and GR00T N1.7 as baselines and an October 2026 deadline.

TRI's lbm_eval is the methodology, not the benchmark: blind evaluators, sufficient rollouts per pair, and quality control on the graders themselves.

What does not exist: no benchmark measures data-efficiency per dollar of collection. That is the number that would actually adjudicate the arguments in Chapter 9, and its absence is why those arguments remain unresolved.