Chapter 5
Robot Foundation Models and "Brains"
17 sections · about 15 minutes
5.0 The category
These companies sell intelligence rather than bodies, or sell both but lead with intelligence. The category barely existed in 2023, absorbed several billion dollars in 2025–26, and has produced exactly one company with meaningful disclosed revenue (Dyna, and even that is contested). It is also the category with the highest rate of acquisition-before-product in modern technology: two of the highest-profile 2025–26 startups were absorbed by hyperscalers within twelve months of founding.
5.1 Physical Intelligence (π)
Founded 2024 in San Francisco by Sergey Levine, Karol Hausman, Chelsea Finn, Brian Ichter, Quan Vuong and Lachy Groom — the Google Brain / Stanford / Berkeley lineage. Funding: ~$400M seed (November 2024) at $2.4B, then $600M at a $5.6B post-money in November 2025 led by Alphabet's CapitalG. (A ~$1B round at $11B+ led by Founders Fund was reported in March 2026 but is not confirmed by any primary source. Treat $5.6B as the last verified valuation.)
Thesis: "ChatGPT for robots" — one generalist policy, hardware-agnostic, sold as an intelligence layer rather than a robot.
Model line, with dates all verified against pi.website:
- π0 (31 October 2024) — a VLA using flow matching on a PaliGemma backbone; open-sourced as openpi in February 2025.
- FAST (January 2025) — action tokenisation that made autoregressive VLAs practical.
- π0.5 (22 April 2025) — open-world generalisation to unseen homes, with ablations showing returns flattening around ~100 training homes and that "data from other robots... is important across all evaluation conditions."
- π*0.6 + RECAP (17 November 2025) — RL from experience and expert corrections via advantage-conditioned policies; 2× throughput and at least 2× fewer failures on espresso-making, laundry and box assembly.
- π0.7 (16 April 2026) — steerability via language, metadata and world-model-generated visual subgoals; claims "the first signs of compositional generalization," including folding laundry on a bimanual UR5e with no laundry data for that robot.
Also notable: online RL on real robots (March 2026) giving up to 3× speedups on contact-rich insertion from ~15 minutes of real robot data per phase, with "half of the trials from the final RL policy faster than any teleoperated demonstration"; and Multi-Scale Embodied Memory (March 2026), reaching tasks requiring up to 15 minutes of memory.
Data: pooled teleoperation across eight-plus platforms, human video, and autonomous on-robot experience. Weights are unusually open; the training corpus is not.
Skeptic: as of early 2026 the company reportedly had ~80 staff and no announced commercialisation timeline. Levine's own Sporks of AGI essay is the field's best argument against the surrogate-data strategies its competitors use — which is intellectually admirable and commercially expensive.
5.2 Skild AI
Founded 2023 in Pittsburgh by CMU roboticists Deepak Pathak (CEO) and Abhinav Gupta (President). Funding: $300M Series A (July 2024), then $1.4B at over $14B valuation in January 2026 led by SoftBank, with NVentures, Bezos Expeditions, Samsung, LG, Schneider Electric and Salesforce Ventures.
Thesis: the omni-bodied model — a single set of weights that controls quadrupeds, humanoids, arms and mobile manipulators without being told what body it inhabits, and that degrades gracefully under damage. The mechanism is morphological diversity as regulariser: train across ~100,000 simulated robot bodies and the model "cannot memorize the solution for one body, it must find a strategy that works across all of them." Skild reports zero-shot recovery from limb loss and jammed wheels.
Data: explicitly rejects teleoperation scaling. Its position paper is blunt: "Teleoperation happens in real-time. Even if we mobilized a global workforce to 'drive' robots 24/7, the time required to reach the trillions of tokens equivalent to an LLM is mathematically unfeasible" — and it is "trapped in sterile labs." The substitute is internet human video plus ~1,000 simulated years of physics.
Embodiment: aggressively hardware-agnostic. Skild sells brains; partners supply bodies for security, inspection, delivery, warehouses, data centres and construction.
Skeptic: ~$30M reported 2025 revenue (company-stated, unaudited) against a $14B valuation is roughly a 450× multiple. The omni-bodied claims rest on internal demos; no peer-reviewed cross-embodiment benchmark comparable to π or GR00T has been published.
5.3 Generalist AI
Founded 2024 by Pete Florence (ex-Google DeepMind; RT-2 and Dense Object Nets lineage) with co-founders from the same group. Reported at $400M at a ~$2B valuation in early 2026, with Fei-Fei Li among backers.
Thesis: scaling laws exist for embodied intelligence, and the corpus should be human, not robot.
- GEN-0 (4 November 2025) was released explicitly as a scaling-law demonstration, trained on over 270,000 hours of real manipulation data, with the claim that "harmonic reasoning" lets a model think and act simultaneously rather than alternating planner and controller.
- GEN-1 (2 April 2026) is the more radical claim: pretraining on over half a million hours of human wearable-device data containing no robot data at all, then adapting to new embodiments and tasks on first contact. Generalist reports a jump from 64% to 99% success, roughly 3× faster execution, with 10× less task-specific data.
Embodiment: maximally agnostic by construction — if the pretraining corpus has no robots in it, there is no body to be locked to.
Skeptic: "99% success" is self-reported on an undisclosed internal task suite. Wearable pretraining claims to sidestep the action-space and force-domain gap that has historically broken human-video transfer, and no third party has reproduced the scaling curve. If the claim holds, it is the most important result in the field; that is precisely why it needs independent replication.
5.4 Dyna Robotics
Founded 2024 by Lindon Gao (previously Caper AI, sold to Instacart), York Yang, and Jason Ma (ex-Google DeepMind, Eureka/DrEureka lineage). $120M Series A at over $600M valuation in September 2025, led by RoboStrategy with CRV, First Round, NVentures, the Amazon Industrial Innovation Fund, Samsung Next, LG Technology Ventures and Salesforce Ventures.
Thesis: deliberately anti-moonshot. Build a single-weight generalist model that is commercially useful today, monetise it as robots-as-a-service in laundries, hotels, restaurants and gyms, and let deployment revenue fund the path to physical AGI. DYNA-1 claimed 99% success over 24 hours of unattended operation on tasks like napkin folding.
Data: a deployment flywheel — fleets running 16+ hours a day generate the data that trains the next model.
Skeptic: reliability numbers come from narrow, repetitive, fixture-heavy tasks. This is closer to a very good task policy than a foundation model, and the "generalist" claim is the least tested in the cohort.
5.5 Google DeepMind Robotics
The deepest lineage in the field: RT-1 (2022) → RT-2 (2023) → Open X-Embodiment (2023) → ALOHA Unleashed and AutoRT (2024) → Gemini Robotics (March 2025).
Gemini Robotics 1.5 / ER 1.5 (September 2025) split the stack into an embodied-reasoning orchestrator and a VLA that "thinks before acting" and narrates its reasoning. The headline capability was motion transfer: skills trained only on ALOHA 2 worked zero-shot on Apptronik's Apollo and a bi-arm Franka.
Gemini Robotics 2 (late July 2026) added whole-body intelligence — "feet to fingertips," locomotion plus dexterous hands — in three variants: GR2 (VLA), GR ER 2 (reasoning and multi-robot teamwork), and GR On-Device 2, which adapts to a new morphology typically with fewer than 200 examples, even for bodies with "drastically different shapes, sensors, and degrees of freedom."
Thesis: robotics is a Gemini capability, not a separate stack.
Embodiment: hardware-agnostic by design and partnered rather than integrated — Apptronik, Boston Dynamics, Franka, Dexmate, SO-101.
Skeptic: Google has shipped VLAs since RT-1 without a commercial robot product. Access remains gated to trusted testers, and cross-embodiment transfer is demonstrated on a handful of curated platforms with broadly similar bi-arm kinematics. The "fewer than 200 examples" figure is itself the admission that transfer is a warm start, not a free lunch.
5.6 NVIDIA
Not a robot-brain startup but the substrate everyone builds on, and the company with the most obvious conflict of interest in this book — its incentive is to sell compute, not to win on policy quality.
The stack: Isaac GR00T humanoid VLAs (N1 in March 2025, billed as the first open humanoid foundation model; N1.5 in June 2025; N1.6 with Cosmos Reason at CES 2026; N1.7 entering early access in April 2026, pretrained on 20K hours of "EgoScale" human video). Cosmos world foundation models (Reason, Transfer, Predict) for synthetic data and policy evaluation. Newton, an open-source GPU physics engine co-developed with Google DeepMind and Disney Research and governed by the Linux Foundation, now selectable inside Isaac Lab 3.0 alongside PhysX, Warp and MuJoCo. Jetson Thor for on-robot inference.
Jensen Huang's framing is that "every industrial company will become a robotics company," and NVIDIA intends to be the Android of generalist robotics.
(Verification note: a "GR00T N2" was reported as previewed for end-2026; no N2 exists in the public repository as of August 2026.)
Skeptic: GR00T checkpoints are widely downloaded and rarely the model actually deployed by serious labs. "Open" GR00T ships alongside a commercial licence. And Newton's governance move is smart precisely because it makes NVIDIA's physics the default without NVIDIA owning it outright.
5.7 Covariant — the cautionary tale
Founded 2017 as Embodied Intelligence by Pieter Abbeel, Peter Chen, Rocky Duan and Tianhao Zhang — the Berkeley RL lineage that arguably invented the robot-foundation-model pitch. Raised ~$222M. Released RFM-1 in March 2024, a multimodal "any-to-any" model trained on years of warehouse picking data.
In August 2024 Amazon hired the three co-founders and roughly a quarter of the staff and took a non-exclusive licence to the models — a "reverse acquihire" structured to avoid antitrust review. The rump company still operates; no significant model has shipped since.
Why it belongs in a textbook: Covariant had the best real-deployment data flywheel of its era — millions of real grasps from live customer sites, not teleoperated demos — the earliest correct thesis, and top researchers. It still ended as a licensing deal, because bin-picking generality did not translate into pricing power against integrators. Every company in this chapter should be read against that outcome.
5.8 Amazon Robotics and Frontier AI & Robotics
Amazon operates the largest robot fleet on earth. In July 2025 it announced its millionth industrial mobile robot alongside DeepFleet, a generative foundation model coordinating fleet traffic that cut travel time roughly 10%. Vulcan (May 2025) is its first robot with a sense of touch — force-feedback end-of-arm tooling trained on real contact data, handling ~75% of stowed item types. Blue Jay (multi-arm sortation) and Project Eluna (agentic operations AI) followed in October 2025.
Amazon has also bought aggressively: Rivr (March 2026, wheeled-legged delivery), Fauna Robotics (March 2026, consumer humanoids, co-founded by Lerrel Pinto), and the 2024 Covariant team.
Skeptic: Blue Jay was cancelled in February 2026 after six months, and Amazon cut robotics jobs in March 2026. If the best-capitalised operator on earth, with the most deployment data on earth, struggles to convert foundation-model research into durable warehouse deployments, that is evidence about the category, not about Amazon.
5.9 Meta — the platform bet
Covered as a hardware entrant in §4.16; the model side deserves separate treatment because Meta's assets are unusual.
Meta owns the largest egocentric human-video apparatus in existence (Project Aria, Ego4D, Ego-Exo4D), the best open tactile stack (Digit 360, 8M+ taxels at 1 mN sensitivity; Sparsh, a general-purpose tactile encoder trained on 460k tactile images; Digit Plexus), a strong simulation platform (Habitat), a human-robot collaboration benchmark (PARTNR), and a world-model line (V-JEPA 2) whose headline result is directly relevant: 1M+ hours of video pretraining plus only 62 hours of robot data yields 65–80% zero-shot pick-and-place on unseen objects.
Thesis: be the software platform — an "Android for humanoids" — rather than the body maker, though Meta is hedging by doing both.
Skeptic: the egocentric-data advantage is worthless without action labels. Glasses see hands; they do not see torques. Meta has cycled through robotics strategies repeatedly since 2019 with nothing shipped.
5.10 World Labs
Founded 2024 by Fei-Fei Li with Justin Johnson, Christoph Lassner and Ben Mildenhall (the NeRF lineage). $230M at launch, then $1B reported raised in February 2026.
Thesis: set out in Li's From Words to Worlds manifesto (November 2025) — language is not enough; spatial intelligence, meaning persistent, 3D-consistent, generative world models, is AI's next frontier and the missing substrate for embodied agents.
Releases: RTFM real-time frame model (October 2025), Marble multimodal world model (November 2025), the World API (January 2026), 3D as code (March 2026), streaming 3D Gaussian splatting (April 2026), and — the explicit robotics bridge — "Building Worlds That Train Robots" (July 2026), a real-to-sim-to-real engine for policy training.
Embodiment: World Labs builds no robot and no policy. It sells the world.
Skeptic: visually gorgeous 3D worlds are not physically accurate ones. Real-to-sim-to-real still inherits the contact-dynamics gap that defeats photorealistic renderers, and most near-term revenue will likely come from games and media.
5.11 Sunday Robotics
Emerged from stealth 19 November 2025, founded by Tony Zhao (Stanford; ALOHA and ACT, later Physical Intelligence) and Cheng Chi (Diffusion Policy, UMI). Reported ~$35M from Benchmark and Conviction at launch, and a reported $165M raise in March 2026. Both figures unconfirmed by primary sources.
Thesis: the bottleneck is not architecture but data provenance. Teleoperation is a "scaling deadlock," so exploit "8 billion humans" instead. The mechanism is a Skill Capture glove whose geometry and sensing are matched to the robot's hand, plus Skill Transform software that strips embodiment-specific detail at roughly 90% fidelity.
Their headline model ACT-1 was trained on zero robot data yet completes a table-to-dishwasher sequence of 68 interactions across 21 objects, generalises zero-shot to unseen homes, folds socks and pulls espresso. Memo is the home robot; an ACT-2 preview appeared in July 2026.
Embodiment: deliberately vertically integrated — the glove only works because they control the hand.
Skeptic: hand-matched gloves make the model as embodiment-locked as any teleoperation dataset, which contradicts the field's cross-embodiment thesis. And "a robot in every home" is the graveyard slogan of consumer robotics.
5.12 Field AI
Founded 2023 by Ali Agha (ex-NASA JPL, led the DARPA SubT-winning CoSTAR team). Reported roughly $405M raised across 2025 rounds at a ~$2B valuation from Bezos Expeditions, Prosperity7, Temasek and Intel Capital. Figures unverified.
Thesis: the generalist-manipulation crowd is solving the wrong problem. Field AI targets unstructured, GPS-denied, safety-critical environments — construction, mines, substations, offshore — where failure is expensive and connectivity is absent. Its Field Foundation Models are "risk-aware," built around a Belief World Model that maintains explicit uncertainty and predicts what it does not know, running fully on-edge.
Embodiment: firmly hardware-agnostic; the EDGE "universal brain" deploys across quadrupeds, wheeled platforms and humanoids.
Skeptic: FFMs are largely navigation-and-inspection autonomy rebranded in foundation-model vocabulary. There is little published evidence of the dexterous manipulation or language-conditioned generalisation that defines the category, and almost no peer-reviewed benchmarking.
5.13 RAI Institute (Marc Raibert)
Founded 2022 as the Boston Dynamics AI Institute, renamed 2024, funded by Hyundai. Structurally the counterweight to the VLA consensus: a research institute with no obligation to ship.
Its stated pillars combine learning with "principled approaches" — model-based control and optimisation fused with learning rather than end-to-end imitation. Recent work: AthenaZero (April 2026), a bimanual robot that juggles barehanded from onboard vision; the ReLIC loco-manipulation framework (2025); whole-body manipulation on Spot combining RL with sampling-based optimisation (2025); an ultra-mobility wheeled platform; and a run of 2026 posts on scaling simulation and robot data collection.
Thesis: athletic, dynamic, contact-rich intelligence will not fall out of scaling teleoperation demonstrations.
Skeptic: Raibert's approach has produced the world's most agile robots and almost no generalisation. AthenaZero juggles but cannot be told to make coffee — which is precisely the bet the VLA labs are making.
5.14 Genesis AI
Founded December 2024 by CMU's Zhou Xian and ex-Mistral Théophile Gervet; launched with a $105M seed co-led by Eclipse and Khosla in July 2025.
Thesis: the purest sim-first bet in the field. A proprietary, very fast physics engine — spun out of an 18-university collaboration — makes synthetic data cheaper than teleoperation. The open-source Genesis engine unifies rigid, FEM, MPM, PBD/SPH and IPC solvers.
Skeptic: the original December 2024 speed claims drew community scepticism and the current documentation no longer foregrounds FPS numbers. Competing directly with NVIDIA's Isaac stack on physics is a difficult place to stand.
5.15 Open source: LeRobot, Open X-Embodiment, K-Scale
LeRobot (Hugging Face) is now the de facto open stack, with 26k+ GitHub stars. v0.5.0 (March 2026) and v0.6.0 (July 2026) added world-model policies, a model zoo spanning GR00T N1.7, MolmoAct2, EO-1 and EVO1, reward models, six benchmarks under a unified lerobot-eval CLI, FSDP multi-GPU training and DAgger-style human-in-the-loop rollouts. It unifies SO-100/SO-101, Koch, LeKiwi, ALOHA and Unitree G1 hardware with ACT, Diffusion Policy, π0, SmolVLA and GR00T. Hardware: Reachy Mini at $299/$449, a 3D-printed LeRobot Humanoid, and the Grabette open manipulation-data recorder. NVIDIA partnered to embed Isaac/GR00T directly.
Open X-Embodiment (October 2023) remains the canonical cross-embodiment corpus — 1M+ episodes, 22 embodiments, 527 skills, from 60 datasets across 34 labs — and has seen no comparable successor release from the Western academic community.
K-Scale Labs (YC W24) builds open-source humanoids. (A widely-circulated claim that K-Scale shut down in late 2025 could not be substantiated; its GitHub organisation shows repository activity into mid-2026. Treat reports of its closure as unverified.)
Skeptic: open weights without open data is hollow. No lab has released anything approaching π's or GEN-1's training corpus, so LeRobot democratises inference, not capability.
5.16 The absorption pattern
Worth stating as its own entry, because it is the category's defining structural fact.
- Covariant → Amazon reverse-acquihire, August 2024.
- Fauna Robotics (founded January 2026, Sprout humanoid) → acquired by Amazon 19 March 2026, roughly two months after emerging.
- Assured Robot Intelligence (Lerrel Pinto) → acquired by Meta, May 2026.
- Vayu Robotics → Serve Robotics, September 2025.
- Diligent Robotics → Serve Robotics, January 2026.
- Intrinsic → moved out of Alphabet's Other Bets into Google itself, February 2026.
- Kind Humanoid → 1X, December 2024.
The pattern says something uncomfortable: in a field where the binding constraint is data and compute, and where neither is available to a startup at hyperscaler scale, the realistic exit for a robot-brain company may be an acquihire rather than an independent business. Physical Intelligence, Skild and Generalist are all betting that they are the exceptions.