Robohouse ’26 Library
Contents

Chapter 10

The Embodiment Question

5 sections · about 5 minutes

10.1 Can one policy control many bodies? The case for

Open X-Embodiment is the founding positive result: RT-1-X beat original per-dataset methods by 50% in small-data domains, and RT-2-X showed roughly 3× improvement on emergent skills. That is transfer, not mere co-existence.

Octo (800k OXE trajectories, 9 platforms) and CrossFormer (30 embodiments, 900k trajectories, including quadcopters and quadrupeds) showed a single transformer matching state-of-the-art across six action spaces without any action-space alignment.

HPT (Wang, Chen, Zhao, He; NeurIPS 2024) formalised the architecture — embodiment-specific "stems," a shared "trunk," task-specific "heads" — over 50 datasets and 200k+ trajectories, beating baselines by more than 20% on unseen tasks.

Physical Intelligence trains π0 across eight distinct robots; π0.5 ablations find "data from other robots... is important across all evaluation conditions"; π0.7 transferred laundry folding to a bimanual UR5e with zero task-specific data for that platform, matching expert teleoperator zero-shot rates.

Google DeepMind's motion transfer is the most striking single demonstration: Gemini Robotics 1.5 moved ALOHA-2-only tasks onto Apptronik's Apollo and a bi-arm Franka without specialisation, and Gemini Robotics 2 adapts to new embodiments "typically with less than 200 examples," even for bodies with "drastically different shapes, sensors, and degrees of freedom."

Boston Dynamics and TRI train one policy jointly across Atlas (50 DoF) and a 29-DoF manipulation test stand — a production-grade version of the same claim.

Skild's omni-bodied thesis makes morphological diversity the mechanism rather than the obstacle: train across 100,000 bodies and the model "cannot memorize the solution for one body, it must find a strategy that works across all of them," with reported zero-shot recovery from limb loss and jammed wheels.

10.2 The case against

The honest reading of the literature is that most published cross-embodiment results are positive-transfer-or-neutral by construction. CrossFormer's own design motivation is that forced observation and action-space alignment "sometimes hinders performance on complex navigation and third-person single-arm manipulation tasks" — that is, naive alignment causes interference, and the architecture is designed to route around it.

Three specific objections:

  1. What transfers is semantics and coarse kinematics, not dynamics. Torque limits, backlash, compliance, gear ratios, thermal behaviour — none of this is shared. Every cross-embodiment result to date operates in a regime where the policy outputs end-effector poses or joint targets that a separate, embodiment-specific controller executes. The controller is doing the embodiment-specific work.
  2. Levine's intersection argument (§9.7) applies with full force to human-to-robot transfer.
  3. "Fewer than 200 examples" is itself the admission. Transfer is a warm start, not a free lunch. The interesting question is not whether transfer is positive — it clearly is — but whether the constant factor shrinks toward zero with scale, and nobody has published evidence that it does.

A genuine gap in the literature: no controlled, large-scale study of negative transfer in cross-embodiment training could be located. The case against is assembled from design motivations and theory, not from a direct ablation paper. Someone should run that experiment.

10.3 The humanoid form factor: the believers

The clearest statement is Figure's master plan:

"We could have either millions of different types of robots serving unique tasks or one humanoid robot with a general interface, serving millions of tasks."

Figure sizes the opportunity as "over 10 million unsafe or undesirable jobs in the U.S. alone," in a market where manual labour is roughly 50% of global GDP.

1X's Bernt Børnich adds the data argument, which is subtler and better: the humanoid is the data-compatible body. "After years of developing our world model and making Neo's design as close to human as possible, Neo can now learn from internet-scale video and apply that knowledge directly to the physical world." On this view, anthropomorphism is not sentimentality — it is the thing that makes the world's largest video corpus usable as training data.

Elon Musk remains the loudest believer, though his 2026 statements are notably deflationary: "No, Optimus production will be extremely slow at first, as everything is new... This is not like making a car."

XPeng's "extreme anthropomorphism" position is the strongest industrial version, paired with a cross-domain VLA shared between cars, robots and aircraft.

10.4 The humanoid form factor: the skeptics

Three lines of attack, each from someone with standing.

Rodney Brooks attacks the sensor channel. In "Why Today's Humanoids Won't Learn Dexterity" (September 2025): "Collecting just visual data is not collecting the right data," because human dexterity runs on roughly 17,000 mechanoreceptors per hand, and "we as a species have not developed technologies to capture touch, to store touch, to transmit touch over distances and time." He also notes that a falling humanoid's kinetic energy scales cubically with size — a safety argument, not just a capability one.

Ken Goldberg attacks the data premise (Chapter 8): there is no path from current collection rates to a ChatGPT-scale robot corpus, so betting the form factor on a corpus that will not exist is a category error.

Russ Tedrake attacks the business case, and does so having spent twenty years building legged robots:

"I thought about legs for 20 years; that's the class I teach at MIT. There are many reasons to build a robot with legs. But the question is, what's the addressable market?... Factories already have autonomous mobile [wheeled] robots. They already have safety cases built around AMRs. You can piggyback on that."

That last clause is the underrated part. The barrier to a legged robot in a factory is not locomotion competence; it is that nobody has written the safety case, and someone already wrote one for wheels.

Melonee Wise, then Agility's Chief Product Officer, adds the demand-side version: "I don't think anyone has found an application for humanoids that would require several thousand robots per facility."

10.5 How the market is actually voting

At the margin, toward wheels. Walden Robotics launched at $1.1B building wheeled-base humanoids. Ai² Robotics raised $735M at $3B and Holiday Robotics $105M — both wheeled. Galbot deliberately dropped legs. Dexterity calls Mech an "industrial superhumanoid" with wheels. Sanctuary abandoned the form factor as a near-term strategy entirely.

Meanwhile the pure bipeds — Figure, Tesla, 1X, Apptronik, Boston Dynamics — hold the highest valuations and have the least deployed. That is not a refutation; it is a bet with a longer horizon. But the divergence is the most informative single pattern in the 2026 market.

The synthesis this book would offer: the humanoid debate conflates two separable claims. Claim A: human-like hands and arms are necessary, because the world's tools, handles, containers and fasteners are designed for them. Claim B: human-like legs are necessary, because the world has stairs. Claim A is well supported and almost uncontested — every serious manipulation company builds anthropomorphic end effectors. Claim B is contested, expensive, and the source of nearly every reliability and safety problem in Chapter 12. Most of the 2026 evidence supports A and undercuts B, which is exactly what a wheeled humanoid is.