Robohouse ’26 Library
Contents

Chapter 13

Open Problems, 2026–2030

9 sections · about 4 minutes

13.1 Dexterity beyond parallel grippers

The hardware has arrived — 1X's 25-DOF tendon-driven hands, Apptronik's 22-DOF SharpaWave tying knots and sealing ziplock bags under Gemini Robotics 2, Kepler's 96 fingertip tactile contact points. The data has not. There is no multi-fingered corpus at Open X-Embodiment scale.

Watch: whether DexUMI- and DEXOP-class devices produce one; and whether Tedrake's durability objection holds — "I have not seen a more dexterous hand that could have done the work our hand has done."

13.2 Long-horizon reliability

State of the art in 2026 is a four-minute, 61-action autonomous sequence. Industrial requirements are eight-hour shifts at 99.99% uptime. This is the field's largest single gap, and the one least addressed by better models — it is a systems, hardware and error-recovery problem.

Watch: MTBF disclosure; intervention rates in teleoperation-backed deployments; and whether Physical Intelligence's Multi-Scale Embodied Memory (up to 15 minutes of task memory) extends toward hours.

13.3 Tactile integration

Brooks' challenge stands: there is no ImageNet of touch, no standard encoding, no transmission format. Sparsh (460k tactile images) is the first general-purpose tactile encoder; Digit 360 the first high-fidelity finger; DenseTact/TensorTouch the first attempt at full stress-tensor recovery; DOT-Sim the first differentiable optical tactile simulator.

Watch: a cross-sensor tactile foundation model; and — the decisive experiment nobody has run cleanly — whether tactile-conditioned policies beat vision-only policies on the same tasks with the same data budget.

13.4 On-robot compute and latency

The System 1 / System 2 hierarchy exists largely because of silicon budgets. 1X's Redwood runs at ~5 Hz on an embedded GPU. Helix splits 7B / 80M / 10M across 7–9 Hz / 200 Hz / 1 kHz. Figure 03 added 10 Gbps mmWave offload — which is to say, it gave up and moved compute off the robot.

Watch: Jetson Thor-class silicon; one-step generative models (Kaiming He's Mean Flows line) that collapse diffusion inference; and whether the hierarchical split survives a 100× improvement in edge inference.

13.5 Evaluation

The field's most under-appreciated crisis, and the one where progress would compound fastest. TRI's methodology paper showed that much prior reporting was too noisy to support its own claims. 1X poses the target directly: "can you predict how well a robot performs before you test it in the real world?"

Watch: adoption of lbm_eval; whether RoboArena-style distributed pairwise evaluation becomes standard; whether SIMPLER-style rank-correlation metrics displace raw success rates; and whether learned world models become accepted as evaluation substrates.

13.6 Safety certification

No standard covers a legged robot that becomes a falling mass on power loss. ISO 10218-1:2025 excludes consumer and public-access service robots; ISO/TS 15066 is a decade old. Home deployment will force the question.

Watch: the first ISO working group specifically for mobile bipedal machines; liability frameworks; and, realistically, the first serious injury and its regulatory aftermath.

13.7 Whole-body loco-manipulation

2026's clearest technical trend. Helix 02 learned 1 kHz balance from 1,000+ hours of retargeted human motion. Gemini Robotics 2 controls "feet to fingertips." Boston Dynamics argues real work requires "a broadening of what we mean by physical intelligence" — using shoulders, forearms and hips, not just fingertips. Stanford's TWIST2 and Karen Liu's group supply the academic data-collection answer.

Watch: whether whole-body data collection scales, and whether wheeled bases make the whole problem moot.

13.8 Memory and continual learning

Physical Intelligence's Multi-Scale Embodied Memory pairs a short-horizon video encoder with language-based long-term memory, reaching tasks requiring up to fifteen minutes of memory while explicitly fighting causal confusion. Combined with on-robot RL (3× speedups from ~15 minutes of real data), this is the most credible route to Goldberg's flywheel: robots that improve from their own deployment rather than from a data-collection factory.

If that works, the entire data bottleneck of Chapter 8 dissolves, because the fleet becomes the collection apparatus. If it does not, the field is back to buying hours.

13.9 The three questions that will decide the decade

  1. Does human data substitute for robot data? Generalist's GEN-1 says yes at 500,000 hours with zero robot data. Levine says structurally no. This is empirically resolvable, and someone should resolve it publicly.
  2. Does the transfer constant shrink? Gemini Robotics 2 needs fewer than 200 examples for a new body. Does that go to zero with scale, or asymptote at 200? Nobody has published the curve.
  3. Does anyone reach 99.99%? Every commercial claim in Chapters 4 through 6 depends on crossing from demo reliability to industrial reliability. No published result is close, and no published result even reports the metric.

Answer those three and the rest of this book is bookkeeping.