Chapter 12
Threats to Validity
about 2 minutes
Single-source dependence. Architecture A's entire empirical case rests on one unrefereed preprint with four subjects and ten hours of data. Every quantitative claim about EMG-as-annotation should be read as provisional until replicated.
Preprint density. A substantial fraction of the 2026 literature cited here — ForceBand, DexEMG, ForceVLA2, ManipForce, TaCo, FELT, RINSE, Stiffness Copilot, the HD-EMG in-home study — is unrefereed. Self-reported baseline comparisons in this genre are systematically optimistic.
Inaccessible primaries. We could not obtain the 2012 tele-impedance paper's numbers, several electrode-shift primaries (cited second-hand through Tanaka 2025), Hahne et al.'s per-method regression values, or the Zhuang et al. Nature Machine Intelligence participant counts. The optimal-window-length result of Smith, Hargrove, Lock & Kuiken (2011) is commonly quoted as 150–250 ms; we could not verify it and do not cite the number.
Heterogeneity precludes meta-analysis. The corpus spans different robots, tasks, metrics, participant populations and policy architectures. No pooled effect size is defensible. This is itself a finding: the field has no shared protocol, which is why §11's cheap experiments have not been run.
Corrections to widely-repeated claims. Three premises we tested did not survive. Manus Robotics does not use electromyography — its wearable uses optics-based muscle activity sensing, and we found no evidence for a product named "Hemyo." The March 2026 MIT wristband that controls a robotic hand is ultrasound, not EMG (Lu, Chen, Li et al., tracking 22 DoF across 8 volunteers) — genuinely relevant as a non-EMG biosignal alternative but frequently miscited. And Faye Wu and Asada's supernumerary-finger work is glove-based, using a ShapeHand fibre-optic data glove with PLS regression (first two principal components capturing ~82% variance, >80% of predictions within ±10°); the actual EMG supernumerary-finger work is Hussain, Spagnoletti, Salvietti & Prattichizzo (2016), Frontiers in Neurorobotics 10:18, from Siena.
Publication bias toward positive results. TaCo and FELT are unusual in reporting null or deflationary findings. Their existence suggests the positive-result literature (§7.4) is somewhat inflated, and readers should weight the FoAR ablation — where naive force concatenation delivered essentially nothing — accordingly.