Chapter 5
Caveats you should carry with these numbers
about 1 minutes
Reimplementation variance on this specific baseline is larger than many of the effects being claimed. Fast-WAM's LIBERO-Plus score is reported as 51.5 by the first Faster-WAM, 49.14 by the second, and 60.0 by LAWA's matched implementation — a spread of eleven points on the metric the entire OOD argument rests on. Its latency is reported as 190 ms by its own authors, 196.5 ms by LAWA, 211.7 ms and 320.97 ms by the two Faster-WAMs, 302 ms by ImageWAM, 493 ms by Enfold, 667 ms by ForeWAM, and 1418 ms by Beyond Task Success — that last figure is a 7.5× outlier with no configuration note, and its runtime column should not be relied on, though its behavioural and sparse-autoencoder results are independent of it. Speedup factors should only be compared within a single paper.
Several ablation tables that bear directly on the masking question could not be retrieved: Section 4.4 of the second Faster-WAM (which contains a subsection titled "Future Conditioning" and is the single most valuable missing piece), Section 4.3 of GlanceWAM, and Appendix D of Mask World Model. Where a figure came from a secondary rendering rather than the paper's own text — the second Faster-WAM's latency values and real-world aggregates in particular — I have said so.
Finally, most of these comparisons are matched on training data and schedule but not on parameter count or inference cost, and several of the strongest results come from a paper's own reimplementation of its baseline rather than the original authors' released checkpoint.