Robohouse ’26 Library
Contents

Chapter 10

Synthesis: Four Architectures, Three Viable

4 sections · about 6 minutes

10.1 Architecture A — EMG as force annotation for human-video demonstrations

Viable now. Highest expected value.

The human demonstrates with their own hand at full biological bandwidth; a wearable band supplies a force label the camera cannot see; the label enters the learned policy as an extra action dimension executed by a force-tracking controller.

Why it works. Latency is irrelevant because there is no control loop (§4.2 nullified). Electrode shift and session drift degrade label quality, and label noise is far more benign than state-correlated action noise (§4.8 nullified). The 15-minute calibration amortises across hours of collection (§8.1 nullified). And the contribution — a force channel — is exactly the intervention the force-conditioning literature identifies as most valuable, namely relabelling the action space (§7.6).

Evidence: ForceBand, 87% pick–squeeze–place, 0/10 for binary-gripper baselines, forces spanning 3.2–19.3 N. One paper, unrefereed, four subjects, ten hours.

Open risks. Scaling to a generic decoder is unproven for force (as opposed to gesture and pose). The 18% MAE gain from anatomy-aware electrode placement implies sensitivity to donning that a consumer band may not survive. And FELT's result — that hallucinated tactile from RGB recovers most of real tactile's benefit — raises the possibility that a sufficiently large vision model learns to predict grip force anyway, making the band redundant. That is a genuine threat to the whole architecture and should be tested directly.

10.2 Architecture B — EMG as a stiffness channel for compliant manipulation

Viable, narrow, and with an unexamined baseline.

The tele-impedance evidence is real: 83.3% versus 60.0% success on drilling; peak contact force halved under one second of delay; operator EMG effort reduced to 0.57 of bilateral control. Co-contraction is genuinely unobservable to pose sensors, so the information is real.

The problem is the baseline. Every result compares against constant impedance. The covariance-based alternative (KΣ1K \propto \Sigma^{-1}, Calinon 2010; minimal-intervention control) extracts a stiffness profile from demonstration variability at zero hardware cost, and no head-to-head exists. Worse for this architecture, Stiffness Copilot achieves tele-impedance's objective from wrist-camera RGB alone, with better NASA-TLX than either constant-stiffness condition, and it comes out of a lab that also builds sEMG wristbands — which is a signal about where they think the value is.

Verdict: worth pursuing only after H2 (§7.2) is run. If covariance-derived or vision-derived stiffness recovers most of the benefit, this architecture has no reason to exist.

10.3 Architecture C — EMG as a latent-action interface for assistive robots

Viable, but it is a different objective.

Here the goal is not to produce training data; it is to give a person with impaired motor function control of a robot. The relevant results are excellent.

Yang, Hodgson, Sun, Erickson & Weber (CMU) [arXiv:2602.02773] [C] deployed 128 electrodes per arm in a spandex sleeve on a Hello Robot Stretch 3 with two users with cervical spinal cord injury. Gesture classification 90.9 ± 5.3% (left, 5 gestures) and 98.0 ± 2.0% (right, 3 gestures); online false positives 0.1–0.5%. Multi-room energy-drink task: 738 s pure teleoperation → 517 ± 62 s with auto-alignment and room mode, a 30.0% time reduction. Learning trend −33 s/day teleoperation versus −55 s/day with autonomy; teleoperation variability 2.3× higher.

The theoretical case is stronger still. Losey, Jeon, Li, Srinivasan, Mandlekar, Garg, Bohg & Sadigh, Autonomous Robots 2021 [PDF] [A] frame the problem exactly right: "users are challenged by an inherent mismatch between low-dimensional interfaces and high-dimensional robots." Their learned latent action space, trained on at most twenty minutes of kinesthetic demonstrations, achieved 88% success (44/50) against a HARMONIC High Assist baseline, with faster completion (t(158)=2.95, p<.05), reduced joystick input magnitude (t(158)=2.49, p<.05) and shorter trajectories (t(158)=9.39, p<.001). Jeon, Losey & Sadigh (RSS 2020) showed latent actions and shared autonomy are complementary rather than substitutes in a 2×2 design.

The BCI literature confirms the principle at even lower input bandwidth. Downey et al. (2016), J. NeuroEng. Rehabil. [link] [A]: ARAT success 78% shared versus 22% BMI-only (p<0.001) for one participant with tetraplegia, 46% versus 0% for another; path length 2.44 m versus 5.00 m. Lee, Lee, Mishra et al. (2025), Nature Machine Intelligence [link] [A] report a 3.9× higher target hit rate with an AI copilot, and a participant with spinal cord injury completing a pick-and-place task they "could not do without" it.

And in prosthetics, the George lab's Nature Communications 2025 result [link] [A] is the best-quantified cognitive benefit of autonomy anywhere in this literature: fragile-object transfer 89 ± 10% versus 59 ± 23% (p<0.01); holding time for amputees 51.56 ± 45.0 s versus 7.81 ± 12.33 s (p<0.01, a 6.6× improvement); and detection-response task latency 0.62 ± 0.38 s versus 0.74 ± 0.41 s, p<0.001 — a 24% reduction in cognitive load.

The gap: nobody has combined a low-bandwidth biosignal with a learned latent action space on top of a pretrained VLA. Losey and Sadigh built latent actions from 20 minutes of kinesthetic data on a single robot. Yang and Weber put HD-EMG on a mobile manipulator with discrete gesture modes. No one has put an EMG decoder on π₀'s or GR00T's latent space. Given that On-Device 2 adapts to new embodiments with fewer than 200 examples, and that a foundation policy supplies exactly the low-dimensional, semantically meaningful action abstraction that a noisy biosignal needs, this looks like the most under-explored opportunity in the entire survey.

10.4 Architecture D — EMG as the primary pose channel for scaled data collection

Not viable, and the evidence is unambiguous.

Best-case cross-user hand-pose decoding is 12.2° mean joint error (emg2pose held-out users), against 3.5° rotational error for UMI's SLAM and effectively zero for leader-follower encoders. Add a ~50 ms physiological delay against a 100–125 ms total budget; 7.6–20% accuracy loss per 2 cm of electrode shift with 1.16 cm re-introduced at every donning; a 3.8% → 18% error increase under limb-position change, in a domain that is limb-position change; 10–12 points of cross-session degradation converting a single-operator corpus into a mixed-operator one; and 30–45 minutes of per-session overhead against UMI's two minutes and 111 demonstrations per hour.

The M4Bench numbers are the empirical confirmation: 26% slower, 58× the error rate, 2.8× the workload, against baselines that are themselves substantially worse than a leader-follower rig.

And the theory says the resulting data is worse than its task-level metrics suggest, because closed-loop operator compensation (Hwang et al. 2017) converts decoder error into action inconsistency rather than task failure — hiding the damage from teleoperation metrics while preserving it in exactly the quantity that behaviour cloning is sensitive to.

The exoskeleton community has already voted. DEXOP, DexUMI, DexEXO, ACE, AnyTeleop, Open-TeleVision, HOMIE and TWIST2 place hardware directly over the forearm and none of them records EMG, because mechanical encoding is more accurate, needs no calibration and does not drift. That unanimity across eight independent systems is the most economical summary of this section.