Chapter 6
The experiment nobody has run
about 1 minutes
A decisive answer needs one model, one training recipe, one parameter budget, and a single binary toggle on the action→future-video attention edge, evaluated on both a saturated in-distribution benchmark and a distribution-shift benchmark, at two or three demonstration scales. GlanceWAM's mask design shows this is architecturally straightforward — its construction already guarantees that toggling the lookahead channel leaves backbone representations unchanged. UNIVERSE has run the toggle but only in driving, only in-distribution, and only on planning accuracy, where it came out null. Until someone runs it in manipulation across data scales and distribution shifts, the field's position on masking future frames from action tokens is best described as an unresolved and now sharply-drawn disagreement rather than a consensus.