Robohouse ’26 Library
Contents

Chapter 7

Performance Impact and Human Intent Capture

6 sections · about 11 minutes

7.1 Does EMG capture stiffness and force better than rigid controllers? The tele-impedance evidence

The affirmative case rests on the tele-impedance literature, which originates with Ajoudani, Tsagarakis & Bicchi (2012), IJRR 31(13):1642–1656 [DOI] [B]. The idea: send the remote robot both a desired motion trajectory and an impedance profile, with an algorithm that decouples force from stiffness so that co-contraction — which changes stiffness without net torque — is separated from net force. Validated on peg-in-hole and ball catching. (We could not obtain quantitative results; the full text is paywalled and the abstract asserts "significant differences" without figures. Do not cite 2012 numbers.)

Two mechanisms recur in this lineage. A co-contraction index from an antagonist pair (typically biceps/triceps) drives a common-mode stiffness term, while arm posture drives a configuration-dependent stiffness term via the muscle Jacobian [Ajoudani, Fang, Tsagarakis & Bicchi, IROS 2015]. Alternatively a linear stiffness–activation model maps normalised activations to a pseudo-stiffness matrix via subject-specific regression.

The quantitative evidence comes from successors.

Fani, Ciotti, Catalano, Grioli, Tognetti, Valenza, Ajoudani & Bianchi (2018), IEEE RAM [IEEE] [A]. Pisa/IIT SoftHand plus KUKA LWR IV+, Delsys Trigno at 1 kHz, EDC/FDS for hand stiffness and BB/TB for arm endpoint impedance. Drilling task, 10 subjects.

ConditionSuccess
Tele-impedance83.3%
Constant high stiffness78.3%
Constant low stiffness60.0%

TI vs LS: p < 0.005, χ² = 8.04, df = 1. Adding CUFF force feedback raised success across all conditions (p < 0.005), significantly for LS but not HS. Questionnaire medians (7-point): "tele-impedance intuitive" 6 (IQR 0); "robot as body extension" 6 (IQR 1).

Laghi, Ajoudani, Catalano & Bicchi (2020), IJRR 39(4) [DOI] [A] — the best quantitative tele-impedance source we located. Two Franka Panda arms, 1 kHz master/slave threads, Myo armbands at 50 Hz and Delsys Trigno at 1 kHz, 10 subjects, round-trip delays of 0 / 500 / 1000 ms.

Maximum interaction force fzf_z (N) in a contact-recognition task:

Architecture0 ms500 ms1000 ms
4-channel bilateral8.0816.1724.31
FT28.0214.8620.36
TIFT2 (tele-impedance)6.6811.2612.33

At one second of delay, tele-impedance halves peak contact force (12.33 vs 24.31 N, p = 1.69 × 10⁻⁵). Normalised operator EMG effort at 1000 ms: TIFT2 0.57 versus 1.0 for classic bilateral. Peg-in-hole (1 mm clearance) fyf_y max: TIFT2 4.73 / 5.95 / 7.06 N versus 4C 5.75 / 9.00 / 8.96 N. Caveat the authors report: TIFT2 scored worse on peg extraction difficulty, attributed to lower compliance during withdrawal.

So the affirmative answer to RQ4's second half is: yes, EMG-derived stiffness demonstrably improves contact-rich teleoperation over constant-impedance control, with effect sizes in the 5–23 percentage point range on success and roughly 2× on peak contact force under delay.

7.2 The critical caveat: the comparison is against constant impedance, not against inferred impedance

Here the review must be careful, because the tele-impedance literature systematically compares against the wrong baseline.

There is a well-developed alternative that infers stiffness from demonstration variability, with no biosignals. Calinon et al. (2010) estimate variable stiffness from the inverse of the observed position covariance encapsulated in a GMM — formally KΣ1K \propto \Sigma^{-1}, so that high demonstration variability implies low inferred stiffness. This became the minimal-intervention control principle: be stiff only where the demonstrations agree [Zeestraten, Calinon et al., ICRA 2016]. Abu-Dakka & Saveriano (2020), Frontiers in Robotics and AI 7:590681 [link] [A] survey both branches and note that EMG methods "require a complex setup and a long calibration procedure" while covariance methods depend on demonstration quality.

They provide no quantitative head-to-head, and we could not find one anywhere. This is the single most consequential gap identified by this review. The entire case for EMG in robot learning rests on it supplying stiffness information, and nobody has tested whether that information is already recoverable, for free, from the variance of conventionally-collected demonstrations.

A hybrid exists — Li, Wu, Liu, Teng, Chen, Calinon, Caldwell, Chen [arXiv:2502.13707] [C] combine EMG-derived limb impedance with a six-direction perturbation calibration (0.02 m, 0.5 s) and a geometric endpoint-stiffness construction, reporting average Z-axis force reduced from 3.80 N to 2.74 N and 7.63 N to 3.29 N versus a constant-impedance baseline. But again the baseline is constant impedance.

Hypothesis H2 (the decisive experiment): on a matched task with matched demonstration counts, compare (i) EMG-derived stiffness, (ii) covariance-derived stiffness from the same demonstrations, and (iii) a learned stiffness policy from vision. Prediction: (ii) captures most of (i)'s benefit at zero hardware cost, and (iii) may exceed both.

Preliminary support for the (iii) branch already exists. Stiffness Copilot — Wang, Xu, Preechayasomboon, Abbatematteo, Memar, Colonnese & Chan (Meta Reality Labs / UW-Madison / Purdue), 2026 [arXiv:2603.14068] [C] — predicts a 3×3 direction-dependent stiffness matrix (eigenvalues 300–3000 N/m) from wrist-camera RGB, with the operator supplying pose only. Across 18 participants:

TaskMetricLow stiffnessCopilotHigh stiffness
Vase wipingmax force (N)60.58 ± 29.2359.28 ± 27.04117.46 ± 42.87
success0.76 ± 0.360.87 ± 0.170.24 ± 0.28
Peg-in-holemax force (N)49.00 ± 34.2533.41 ± 11.6780.60 ± 54.44
success0.67 ± 0.400.81 ± 0.310.70 ± 0.30

NASA-TLX 38.73 versus 49.33 (low) and 52.43 (high), all p < .05. This achieves tele-impedance's objective without any biosignal. It is the strongest single piece of evidence against the EMG-as-stiffness-source thesis, and it appeared in 2026 from a lab that also builds sEMG wristbands.

7.3 The reframing that works: EMG as annotation, not as interface

ForceBand's contribution is conceptual before it is empirical, and it is the most important idea in this review.

The pipeline: an 8-channel bipolar sEMG band (~$300, OpenBCI Cyton / ADS-1299) with anatomically guided placement — seven channels over finger-controlling forearm muscles, one over wrist flexors. A 15-minute per-user calibration collects paired sEMG and ground-truth fingertip forces. The fingertip force sensors are then removed. A pretrained EMG2Force model thereafter labels ordinary egocentric human video demonstrations with force traces. Those force-augmented human demonstrations train a flow-matching transformer policy whose action is

at=[pR3;  r6DR6;  gR;  fR]R11a_t = [\,p \in \mathbb{R}^3;\; r_{6\mathrm{D}} \in \mathbb{R}^6;\; g \in \mathbb{R};\; f \in \mathbb{R}\,] \in \mathbb{R}^{11}

— end-effector position, 6-D rotation, gripper aperture, and desired grip force, with a PD controller tracking the predicted force by modulating aperture.

Results: sEMG roughly halves hand-level force-regression error versus vision baselines; finger-level contact detection PR-AUC of 0.763 (ring) and 0.590 (pinky) versus 0.398 and 0.314 for a vision-based baseline, roughly 1.9× on each and over 6× above random. An electrode-placement ablation under a matched 30-minute protocol found muscle-aware 8-channel placement at MAE 0.77 N / RMSE 1.33 N versus evenly-spaced 8-channel at 0.94 N / 1.77 N — an 18% MAE reduction from anatomy-aware placement alone. On pick–squeeze–place across 9 objects (43–650 g, grasp widths 1–72 mm), overall success was 87%, with squeeze success 6–10/10 versus 0/10 for a binary-gripper baseline on every object and 0–4/10 for a continuous-gripper baseline. Predicted grip forces spanned 3.2 N to 19.3 N, object-specific.

Why this works when EMG-as-interface does not: the human demonstrates with their own hand at full human bandwidth. There is no decoder in the control loop, so latency (§4.2) is irrelevant — the labels are computed offline. Electrode shift and session drift degrade label quality rather than trajectory quality, and label noise is far more benign for a learned policy than state-correlated action noise. The 15-minute calibration is amortised over hours of subsequent collection. And the artifact that kills EMG teleoperation — closed-loop compensation hiding decoder error in the action log (§4.6) — does not arise, because there is no loop.

This reframing converts EMG from a bandwidth-limited actuator into a cheap sensor for an otherwise unobservable label. It is, as far as our search can establish, novel as of mid-2026 and has exactly one paper behind it.

7.4 Does force conditioning help downstream? Yes, on contact-rich tasks, with large effects

The relevant question then becomes whether a force channel earns its place in the action or observation space at all. The 2024–26 evidence is strong but must be disaggregated.

SystemForce signalEntry pointResult
DexForce (Chen, Yu, Choi, Cutkosky, Bohg; RA-L) [arXiv:2501.10356] [A]Fingertip forces during kinesthetic demosAction relabelling: xf=xo+kffx_f = x_o + k_f \cdot f, tracked by Cartesian impedance control76% avg across 6 tasks (57–90%) vs "near-zero" for the same policies on non-relabelled actions. 5–10 demos/task
ForceMimic / HybridIL (SJTU) [arXiv:2410.07554] [A]6-axis F/T on a handheld rigDiffusion policy predicts 20-step force–position trajectories; switches to hybrid control at ≥6 NPeel length >10 cm: 85% vs 55% (+54.5% rel.); mean interaction force ~9 N vs ~20 N; collection ~5 min vs >13 min for force-feedback teleop
FoAR (SJTU, RA-L) [arXiv:2411.15753] [A]Wrist F/T at 100 HzFuture-contact predictor gates multimodal fusionWiping 0.875 vs RISE 0.500, DP 0.400, RISE+force-token 0.575, RISE+force-concat 0.475. 50 demos/task. Force at 2 Hz or 10 Hz worse than 100 Hz
ManipForce / FMT (GIST) [arXiv:2509.19047] [C]Wrist F/T at 200+ Hz vs 30 Hz RGBFrequency + modality embeddings, bidirectional cross-attention83% avg vs 22% RGB-only across 6 tasks (box flipping 90 vs 5; open lid 100 vs 20). ~100 demos/task
Tactile-VLA (Tsinghua/UESTC/SJTU) [arXiv:2507.09160] [C]Tactile tokens + hybrid position–force control P=Ptarget+KΔFP = P_{\text{target}} + K\Delta FModel outputs force targetsUSB insertion 35% vs π₀-base 5%; charger 90% vs 40%; fragile-object OOD 90% vs 0%. Zero-shot language-to-force: "softly" 4.68 N vs "hard" 9.13 N (π₀-base showed no differentiation: 6.61 vs 5.69)
ForceVLA (NeurIPS 2025) [arXiv:2505.22159] [A]6-axis F/TForce-aware MoE fusing during action decoding, on π₀+23.2% over π₀ baselines; 80% on plug insertion
ForceVLA2 (Shanghai AI Lab et al.) [arXiv:2603.15169] [C]Force as prompt into the VLM expertCross-scale MoE, closed-loop hybrid force–position66.0% avg vs π₀ 18.0%, π₀.₅ 31.0%, ForceVLA 35.0% across 5 tasks
Adaptive Compliance Policy (Stanford + TRI) [arXiv:2410.09309] [C]F/T + kinesthetic teaching at low stiffnessPolicy predicts a stiffness value and a virtual target pose">50% improvement over SOTA visuomotor methods" on item flipping and vase wiping

The FoAR ablation is the most informative single row in this table. Naively concatenating force into the observation of a strong baseline yields 0.475–0.575 on wiping versus 0.500 for the same baseline without force — essentially zero benefit. Gating force on a learned contact predictor yields 0.875. Force in the observation is only useful if the policy knows when to trust it.

7.5 The null and deflationary results

A responsible review must weight these equally.

TaCo — Zorin et al. (2026) [arXiv:2605.21976] [C] benchmarked six tactile sensors across four modalities at two institutions on a Franka Panda, training ACT policies on identical data with and without tactile. Pick-and-place results: FSR 0.50 → 0.50 (no change); Daimon 0.95 → 0.80 — tactile actively hurt; FlexiTac 0.75 → 0.85; eGain 0.50 → 0.75; contact microphone 0.65 → 0.90; eFlesh 0.85 → 0.90. Plug insertion was uniformly positive but low (0.1 → 0.3, 0.2 → 0.7, 0.3 → 0.7). The benefit of tactile is sensor-dependent and task-dependent, and can be negative.

FELT — Li et al. (USC/Columbia), July 2026 [arXiv:2607.20683] [C] found that hallucinated tactile signals generated from RGB recover most of the benefit of real tactile: tube insertion 40% (vision-only) / 55% (real tactile) / 50% (generated); cup nesting 25 / 35 / 45; triangle peg 50 / 70 / 90. In two of four tasks the generated signal matched or beat the real sensor. If a vision model can predict the tactile signal, that signal was partly redundant with vision.

Gano, George & Barati Farimani (CMU) [arXiv:2406.15639] [C] performed contrastive visuo-tactile pretraining, then disabled the tactile sensor at inference, improving USB cable plugging "by up to 65%" with vision-only inference. The value was bankable into the visual encoder; the fragile sensor could be discarded.

The positive counterweight, under matched data: Funk, Chen, Schneider, Chalvatzaki, Calandra & Peters [arXiv:2504.13618] [C] used identical 20 expert demonstrations across conditions on robotic match lighting and found visuotactile policies improved "by over 40%" over vision-only. And Ablett et al. [arXiv:2311.01248] [C] decomposed the contributions across four door-opening tasks: force matching +62.5%, visuotactile mode switching +30.3%, visuotactile data as policy input +42.5%.

7.6 The three-way confound

That last decomposition exposes a systematic problem with how this literature is read. "Force helps" conflates three distinct interventions:

  1. Force as observation — concatenate or tokenise ftf_t into oto_t. FoAR's ablation shows this alone is worth roughly nothing.
  2. Force as action — the policy outputs a force or stiffness target executed by a hybrid or impedance controller. Tactile-VLA, ForceMimic, Adaptive Compliance Policy, ForceBand.
  3. Force as relabelling — the demonstration data itself is rewritten so that recorded actions encode the force the demonstrator applied. DexForce, Ablett's "force matching."

The evidence suggests (3) dominates and (1) is nearly worthless. DexForce moves from near-zero to 76%; Ablett's force matching (+62.5%) beats tactile-as-input (+42.5%) on the same tasks. This has a direct implication for our question. If the value of a force channel is realised through relabelling the action space, then what a prosthetic or EMG interface needs to supply is an accurate force estimate at demonstration time — which is exactly what ForceBand does — and not a control channel.

One further confound deserves flagging. Feng, Zheng, Wang et al. [arXiv:2602.23408] [C], with 500+ trained models and 13,000+ real rollouts, found that delta actions beat absolute actions 82.9% vs 71.9% averaged, and chunk-wise delta beats step-wise by ~10 points. Effects of that magnitude can swamp the force contribution. Any force-conditioning ablation that does not control for action parameterisation is confounded, and most do not report it.