Robohouse ’26 Library
Contents

Chapter 5

Comparative Evaluation of Teleoperation Media

4 sections · about 4 minutes

5.1 The only direct comparison

Yoshida, Dossa, Di Vincenzo, Sujit, Douglas & Arulkumaran (2025), M4Bench, Frontiers in Robotics and AI [DOI] [A] is, as far as our search could establish, the only benchmark that places a biosignal interface head-to-head against conventional input devices on a manipulation task with a full metric suite. N = 50 participants, 7-DoF Franka Panda arms, pick-and-place of coloured blocks into matching bins, three device pairs.

MetricMouse + KeyboardGamepadEye tracker + EMG
Completion time (s)105.6 ± 4.2106.9 ± 5.8132.7 ± 22.5*
Command selection time (s)0.423 ± 0.4430.596 ± 1.1550.846 ± 1.284*
Error rate0.005 ± 0.0290.014 ± 0.0410.292 ± 0.263*
NASA-TLX overall14.7 ± 13.517.4 ± 15.340.8 ± 20.1*

* corrected p < 0.001 versus both conventional interfaces.

The biosignal pair was significantly worse on every metric: ~26% slower, ~58× the error rate, ~2.8× the workload. There was no significant difference between mouse/keyboard and gamepad.

Three caveats before this is over-read. The EMG hardware was g.tec forearm and calf electrodes in a supervisory multi-robot paradigm, not a continuous manipulation interface. Eye tracking is confounded with EMG in the same condition. And the comparison devices are themselves poor manipulation interfaces — GELLO's user study found mouse-class devices at 63% average success versus 92% for leader-follower. So M4Bench establishes that EMG is worse than mediocre baselines, which is a stronger negative result than it first appears.

5.2 A more favourable read

Lobo-Prat, Keemink, Stienen, Schouten, Veltink & Koopman (2014), J. NeuroEngineering and Rehabilitation 11:68 [DOI] [A] compared EMG, force and joystick interfaces for active arm supports with 8 healthy participants. EMG was best on tracking error and gain-margin crossover frequency; force was better on information transmission rate and effort; joystick was indistinguishable from force. Differences emerged beyond 0.9 Hz (tracking error) and 1.4 Hz (information transmission).

The bandwidth framing is the useful part. EMG leads the neural-to-mechanical delay — it precedes movement — which is why it can win on high-frequency tracking. This is a real advantage and it is why EMG persists in assistive contexts. It does not translate into an advantage for producing accurate 6-DoF pose trajectories.

5.3 The conventional interfaces, for calibration

InterfaceThroughputSuccess / accuracySource
UMI (handheld gripper)111 demos/hr (vs 35 SpaceMouse); 48% of bare-hand speed; 149/hr on dynamic tossing where SpaceMouse produced zero demos in 15 minSLAM ATE 6.1 mm / 3.5°; 70% success OODChi et al., RSS 2024 [arXiv:2402.10329] [A]
GELLO (leader-follower, ~$300)Fastest completion times across all 5 tasks92% avg (Hat 92, Mask 92, Banana 100, Towel 92, USB 83)Wu, Hoque, Mandlekar, Abbeel & Goldberg [arXiv:2309.13037] [B]
VR controller (Quest 2, ~$300)72% avg (USB 50)same
3D SpaceMouse (~$150)35 demos/hr63% avg (USB 58)same / UMI
Kinesthetic teachingFastest demos (17.4–24.0 s vs 26.1–65.3 joystick) but physically demandingHighest downstream policy performance; replay success 88.0/83.0/58.1%Vanc et al. 2026 [arXiv:2605.28033] [C]; Li, Cui & Sadigh 2025 [arXiv:2503.07017] [C]
EMG, discrete-gesture12 gestures >90% online; 81.8–101.4 s task times (best mapping strategy)Wang et al. 2025 [A]
EMG + eye tracker26% slowererror rate 0.292M4Bench [A]

Two observations. First, the ordering leader-follower > VR > space-mouse is consistent across studies, and the gap widens with task precision (USB insertion: 83 / 50 / 58). Second, and more important for this review, kinesthetic teaching produces the best downstream policies despite being the least scalable — a finding independently reached by Li, Cui & Sadigh (2025) through user study and by RINSE (2026) through smoothness metrics. Li et al. also report that mixing a small amount of kinesthetic data with additional VR-teleoperation data yields ~20% higher average performance than either alone, because the two modalities trade off action consistency against state diversity.

That last result is the template for the constructive argument in §10: heterogeneity in proficiency hurts (robomimic MH), but heterogeneity in modality can help when the modalities contribute orthogonal quality dimensions. The question for EMG is which kind of heterogeneity it introduces.

5.4 A structural gap in the comparison literature

Almost no study connects teleoperation-interface throughput to downstream policy success. GELLO and UMI measure collection; ACT and Diffusion Policy measure policies. Only Li, Cui & Sadigh (2025) and RINSE (2026) close the loop, and neither includes a biosignal condition. No paper anywhere trains a visuomotor policy on EMG-collected demonstrations and compares it against the same policy trained on conventionally-collected demonstrations of the same task. This is the central empirical hole in the field this review surveys.