Who Builds Physical AI's Post-Training Factory?
If scaling laws hold for robot foundation models, the durable value isn't in pre-training. It's in the layer that fine-tunes them on data only the enterprise has.
Start with a correction, because it sets up everything after it. My earlier position was that the platforms for training embodied AI — the simulators, the world models, the synthetic-data pipelines — mostly didn't exist yet, and that absence was the reason to fund the category.
That's no longer true. NVIDIA's Isaac and Omniverse, Cosmos for world models, Jetson at the edge; DeepMind's Genie, World Labs' Marble, Tencent's HunyuanWorld. This stack went from research demo to shipping product faster than most people modeling it in 2024 expected, and it consolidated around one company harder than I would have guessed. At GTC 2026 NVIDIA pushed Cosmos 3, Isaac GR00T N1.7 and Alpamayo 1.5 as the default development environment for physical AI, and robotics companies across warehousing, humanoids and industrial inspection are building on that stack rather than against it.
Being wrong about that is what makes the more interesting question askable. If pre-training infrastructure is consolidating into a substrate, the question stops being who builds the simulator and becomes where does the marginal capability actually come from — and in the language-model stack we already know the answer to that one.
The premise: scaling laws appear to hold
This is the load-bearing assumption, so it deserves evidence rather than analogy.
As of early 2026 it looks real. EgoScale, published in February, gave the first strong empirical evidence that robot foundation models follow the same data-driven scaling laws as language models, with policy performance improving predictably as pre-training data grows. A separate result on vision-language-action models scaled real robot data from three thousand hours to twenty thousand and saw success rates climb monotonically with no saturation at the top end. That is the shape of curve that made language models work.
The caveat is the one that governs everything below. Language models scale on data that is essentially free to accumulate. Robot data is gated by physical constraints and annotation cost. The curve holds; the input is rationed.
Where the value went in language models, and why it should rhyme
Pre-training built the foundation, but post-training is now where usable capability comes from. Supervised fine-tuning, preference optimization and reinforcement learning against verifiable rewards account for the majority of what a deployed model can actually do. RLVR in particular — check the answer automatically, let the model grind against the checker — is the mechanism behind most of the recent step-changes in reasoning.
If that pattern transfers, the physical-AI equivalent is not the lab that pre-trains the biggest generalist policy. It is whoever can take an open, decent generalist and grind it against one specific enterprise's reality until it works there. Pre-training gets you generally competent. Post-training is what makes a policy work in this plant, on this hardware, with this product mix and this floor layout.
The precondition for a third party to do that work is already met. Physical Intelligence open-sourced pi-0 with LeRobot integration, NVIDIA shipped GR00T N1.5 as a production-grade open model, OpenVLA and SmolVLA are out, and Open X-Embodiment aggregates close to a million trajectories across twenty-two platforms. Open weights are exactly what created room for Fireworks, Together and Baseten to build businesses serving and fine-tuning other people's models — a layer that went from nothing to a seventeen-and-a-half-billion-dollar valuation on roughly eight hundred million of annualized revenue at Fireworks alone, most of that growth inside a single year. The same door is now open in robotics.
So who becomes that for physical enterprise?
The strongest leg: physical reality is a free verifier
Here is the part of the analogy that is better than it first looks.
RLVR works in language because some domains are cheaply checkable. Math has a ground-truth answer; code has a test suite. It works far less well in fuzzy domains, which is why there is an active research push to construct rubrics and graders for everything else. Verification is the scarce ingredient.
Physical tasks do not have that problem. Did the part seat. Did the weld hold. Did the cycle finish inside takt time. Did the robot drop it. Reality is the reward function, it is already instrumented in most industrial environments, and nobody has to write it. For a category that looks harder than language on every other axis, the verifier — the thing that made RL post-training take off — comes free.
That is a real reason to expect post-training to be where compounding happens in physical AI, and not merely an argument from resemblance.
Where the analogy breaks: you cannot rent physical rollouts
Now the part that should stop anyone from porting the Fireworks model wholesale.
A language rollout is nearly free at the margin and massively parallel. You can run millions, and a failure costs fractions of a cent. A physical rollout runs in wall-clock time on real hardware, and a failure can damage a machine, scrap a part or hurt somebody. The scarcity inverts: in language, data and environments are the constraint while compute is elastic. In physical AI the constraint is safe attempts on real hardware inside a real environment, and there is no spot market for that.
Two consequences follow, and together they make this a different business rather than the same business in a new vertical.
The factory has to be co-located. Fireworks works because weights are open, GPUs are fungible, and the customer's data comes to it over an API. Factory-floor telemetry is competitive intelligence and increasingly regulated; sovereignty has moved from a compliance checkbox to an architectural constraint. The robot is physically in the plant, and much of the data cannot leave it for legal or commercial reasons. The post-training factory for physical AI is on-premise and at the edge by necessity — which makes it look less like a usage-based cloud and more like an installed-base business, with the slower and stickier economics that implies.
The valuable data is the failure data. This is the detail that reframed the problem for me. Datasets collected by skilled teleoperators who rarely fail produce policies that do not recover when something goes wrong; deliberately injecting failure states substantially improves real-world resilience. So the highest-value training data is not the clean expert demonstration a lab can commission. It is the messy record of things going wrong and being recovered, and that accrues only to whoever runs the fleet in production, continuously, across sites they do not own. You can buy demonstrations. You cannot buy someone else's three years of edge cases.
My answer: the control plane grows down into training
Line up the requirements and the candidate list gets short. The winner has to sit on-premise where the data is, stay neutral across hardware and model vendors because no operator will accept lock-in on either axis, hold the safety and compliance surface already, and see every task attempt including the failures.
The foundation labs do not fit. They are structurally cloud-and-lab shaped, and they are the vendor the enterprise is trying not to get captured by. NVIDIA could reach up from the substrate and might, but neutrality across model providers is not in its interest. That leaves the least glamorous candidate: the fleet orchestration and RobOps layer, growing downward into training.
That layer is the only one that is simultaneously inside the building and indifferent to whose robot it is. It already sees every mission, every intervention, every failure and recovery, because that is what orchestration is. It already carries the audit trail and the approval workflow. The data flywheel is not a second product it has to go build — it is the exhaust of the job it already does.
That is also the strongest version of the argument behind General Robotics, a deal I built the pitch and IC deck for: a runtime that lets a robot call skills as an API across providers, so an operator is not betting a deployment on one vendor's roadmap. I underwrote it as an interoperability bet. The better reading is that the neutral runtime is an on-ramp — whoever holds that position ends up sitting on post-training data nobody else can assemble, and interoperability was the reason they were allowed in the door.
What would change my mind
Three things would break this.
If pre-training generalizes well enough that a stock policy works acceptably out of the box, post-training collapses into a thin configuration step and the factory is a feature rather than a company. The no-saturation scaling results cut both ways here.
If robot experience transfers across enterprises better than expected, then aggregating many customers' data beats owning one customer's deeply, the labs win on volume, and the proprietary-data moat is far shallower than I am assuming.
And the real bear case: hardware heterogeneity may be severe enough that nothing amortizes. If every deployment is bespoke integration work then there is no software margin in it, only a services business wearing a platform's clothes. That is precisely the trap that swallowed earlier attempts at a neutral robotics platform, and I do not think this generation has proven it escapes yet.
So, the falsifiable version. Within about three years, the best-performing physical AI policies in a given vertical should belong to whoever was running that vertical's fleets in production — not to whoever pre-trained the largest model. If the leaderboard in industrial manipulation ends up owned by whoever has the most GPUs rather than whoever has the most deployed sites, I have this wrong.