Robohouse ’26 Library
Contents

Chapter 8

Scalability

3 sections · about 2 minutes

8.1 The arithmetic

Consider a realistic protocol for EMG-teleoperated demonstration collection, using the numbers established in §4:

  • 15 minutes electrode–skin impedance stabilisation before calibration is meaningful (Sousa et al. 2023).
  • 15–30 minutes decoder calibration (ForceBand: 15 min; DexEMG requires per-user calibration; the Nature 2025 personalisation protocol: 20 min).
  • Several minutes donning, per the HD-EMG in-home study, with a median of two recalibrations per day.
  • Session length bounded by fatigue. Vogel, Bayer & van der Smagt (2013) observed increased false grasp triggers in the second trial with SMA patients; the Wang et al. (2025) armband study explicitly tested "under fatigue."
  • Re-donning introduces 1.16 ± 0.34 cm of electrode shift, costing 7.6–20% decoder accuracy unless re-calibrated.

Against a leader-follower rig, where setup is switching on two arms, and against UMI, which is deployable in a new environment in about two minutes and yields 111 demonstrations per hour.

Even under generous assumptions, EMG teleoperation carries an overhead of roughly 30–45 minutes per session before the first useful demonstration, plus intra-session recalibration, plus a fatigue-bounded session length. For a target corpus of the size that matters — DROID's 350 hours, AgiBot's 2,976 — this is disqualifying. The overhead does not amortise, because it recurs per session and per user, and multi-day collection re-introduces the cross-session distribution shift of §4.5 which converts a single-operator corpus into a robomimic-MH-like mixture.

8.2 The lever that changes the arithmetic

Calibration-free cross-user decoding is the only development that could alter this conclusion, and it arrived in 2025. The Nature paper's zero-shot performance across thousands of held-out participants removes per-user calibration from the critical path for discrete gesture and low-dimensional continuous control.

Two caveats bound its relevance. It is not robotics — no manipulator, no 6-DoF pose, no contact. And DexEMG, the closest robotics analogue, explicitly reports that it still requires per-user calibration, which suggests the generic-decoder result has not yet transferred to dexterous hand pose at the fidelity retargeting needs.

Hypothesis H3: a generic, calibration-free EMG2Force model — the ForceBand pipeline trained at Nature-paper scale — would eliminate the 15-minute calibration and make EMG-as-annotation a genuinely scalable augmentation of egocentric video collection. This is a straightforward, well-motivated engineering programme and nobody has announced it.

8.3 Comparative cost model

MediumSetup per sessionThroughputAction-space gapForce/stiffness observableScalability verdict
Leader-follower (GELLO/ALOHA)~minutesModerateNoneNo (unless instrumented)High
VR controller~minutesModerateRetargetingNoHigh
SpaceMouse~minutes35/hrSmallNoModerate
Kinesthetic~minutesFast per demo, physically limitedNoneYes, via joint torqueLow (fatigue)
Handheld (UMI)~2 min111/hrEliminated by constructionNoVery high
Exoskeleton (DEXOP/DexUMI)~minutes> teleop per unit timeEliminatedYes (tactile)High
Egocentric video~noneUnboundedSevereNoHighest
EMG teleoperation30–45 min + recal.Low; fatigue-boundedSevere (12.2° pose error)YesLow
EMG as annotation15 min, amortisedInherits video throughputN/A — labels, not actionsYesHigh

The last two rows are the argument of this review in one place.