Robohouse ’26 Library

August 2026

World Action Models since Fast-WAM

Coverage: 17 March 2026 (Fast-WAM v1) to 28 August 2026.

Chapters
7
Words
4,331
Reading
20 min

Coverage: 17 March 2026 (Fast-WAM v1) to 28 August 2026. Semantic Scholar records 167 citations of Fast-WAM as of this date; roughly two dozen of those engage its argument substantively rather than citing it in passing, and those are the papers treated here.

The short version. Fast-WAM's first claim — that video co-training during training is what makes a World Action Model good — has survived five months of scrutiny completely intact, and no paper published since has argued for dropping it. Its second claim, that explicit future generation at test time can be discarded at negligible cost, has held only inside the regime it was measured in: large per-task demonstration counts, in-distribution evaluation, short-to-medium horizons, and terminal success rate as the sole metric. Outside that box, four independent groups now report deficits of eight to thirty points. The mechanistic explanation those groups converge on is not that future generation is necessary, but that Fast-WAM's attention mask — which forbids action tokens from attending to future video tokens at all — severs the only pathway by which predicted dynamics could reach the policy at inference. The distinction between "skip the rollout" and "sever the attention edge" is the single most important thing the field has learned since March, and it is a distinction Fast-WAM itself did not draw.