Chapter 11
Research Agenda
about 3 minutes
Ordered by expected value per unit of effort. Several are cheap.
H1 — The missing measurement (cheap, high value). Collect matched demonstrations of the same task from the same operators via (a) leader-follower, (b) EMG-teleoperation, (c) kinesthetic teaching. Report SPARC, LDLJ, Belkhale action variance and state similarity for each. Then train identical ACT and diffusion policies on each and report success. Prediction: the EMG policy gap will exceed the teleoperation-level success gap, because operator compensation masks decoder error at the task level. Nobody has ever trained a visuomotor policy on EMG-collected demonstrations and compared it to a conventional control.
H2 — The decisive stiffness experiment (cheap, decisive). Compare EMG-derived stiffness, covariance-derived stiffness () from the same demonstrations, and vision-derived stiffness (Stiffness Copilot-style) on identical contact-rich tasks with identical data budgets. Prediction: covariance recovers most of EMG's benefit at zero hardware cost. If so, Architecture B is dead. If not, it is the field's most important result of 2027.
H3 — Generic EMG2Force at scale. Train the ForceBand pipeline at Nature-2025 scale (thousands of participants) and test zero-shot force regression on held-out users. Prediction: achievable, since force regression is a lower-dimensional target than 22-DoF hand pose. This would remove the last per-user calibration from Architecture A.
H4 — The redundancy test for Architecture A. Apply the FELT protocol to grip force: train a model to predict fingertip force from RGB alone, and compare a policy conditioned on predicted force against one conditioned on EMG-derived force under matched data. Prediction: vision recovers a substantial fraction. This is the strongest threat to Architecture A and should be run before large-scale collection begins.
H5 — Biosignal-driven latent actions on a foundation policy. Take π₀-class or GR00T-class latent action representations and drive them with an sEMG decoder for an assistive manipulator. Compare against discrete gesture mode-switching (Yang et al.) and against learned latent actions from kinesthetic data (Losey et al.). Prediction: the combination outperforms both, because the foundation model supplies precisely the low-dimensional abstraction the noisy channel needs. Nobody has done this.
H6 — Release the corpus. Either ForceBand's promised release or an equivalent should ship: synchronised vision, robot action logs, and biosignal provenance, with failures included and licence stated. Absent this, every claim in §7 remains single-source.
H7 — Report smoothness in EMG teleoperation papers. Every EMG-teleoperation paper we surveyed reports task success, completion time and NASA-TLX; none reports SPARC, jerk or action variance. Adding three lines of analysis to existing datasets would immediately populate the comparison this review could not make.
H8 — Control for action parameterisation. Given that delta versus absolute actions is worth 11 points on average and chunk-wise versus step-wise another 10, any force- or biosignal-conditioning ablation that does not fix and report the action parameterisation is confounded. This should become a reporting norm.
H9 — Amputee-inclusive evaluation. Ninapro DB3, DB7, DB8 and DB10 include amputee participants; the HD-EMG mobile-manipulation study included two users with SCI; the George lab study included four transradial amputees. Most EMG-robotics work uses able-bodied participants exclusively — including, explicitly, the npj Robotics HD-EMG study. Since the strongest use case for this technology is assistive, this is both an ethical and a validity problem.
H10 — Cross-session protocol standardisation. Adopt a mandatory 15-minute impedance stabilisation period, report donning shift, report per-session decoder accuracy, and treat multi-day collection as multi-operator data for analysis purposes. Without this, any EMG corpus is a mixed-quality dataset by construction and should be curated accordingly.
H11 — Failure data. Almost no manipulation dataset contains labelled failures; RoboMIND's 5,000 are a conspicuous exception. EMG interfaces fail in characteristic, diagnosable ways (misclassification, false triggers, drift), which makes them an unusually good source of labelled failure modes for training reward models and uncertainty estimators. This is a genuinely novel argument for collecting EMG data that has nothing to do with its use as a control channel.