Robot Era's Chen Jianyu: world models, not VLA, are the next leap
Two founders and a researcher took the WRC 2026 stage in Beijing this week with the same message — embodied AI is hitting an inflection point, and the win condition is shifting from flashy hardware to brains, reliability, and recurring revenue.
Robot Era's Chen Jianyu says vision-language-action models are a stepping stone, not the destination. Speaking at the 2026 World Robot Conference, the founder argued that VLA — the now-mainstream approach that stitches vision, language, and motion into one imitation-trained model — tops out at tasks it has already seen. His bet is on world models: systems that learn the physics of the real world (a cup pours down, a cloth folds a certain way) so a robot can reason about situations it was never trained on. Robot Era claims to have shipped the approach early, demonstrating zero-shot task generalization and fine manipulation like unscrewing a cap or flipping a sock inside-out. It's a founder's pitch, not a benchmark result — but the underlying critique of imitation learning's ceiling is one the whole field is wrestling with.
ASU's Hua Wei warns LLM agents are repeating reinforcement learning's oldest mistake. In an IJCAI 2026 talk, the Arizona State professor drew a direct line between today's fragile foundation-model agents and the sim-to-real gap that broke RL systems a decade ago: train an agent in English and it crumbles in Chinese or a low-resource language, just as a traffic-light controller trained in a simulator fails in real weather. His fix borrows straight from robotics — "domain randomization," deliberately perturbing prompts, actions, and reward signals during training — which let a 3-billion-parameter model beat a 32-billion-parameter one on cross-environment tasks. The deeper point: agents need to know when they don't know, and escalate to a human.
Xinghai Tu's Gao Jiyang thinks the real money in robotics is selling "physical-world tokens." The co-founder laid out a three-stage business arc at WRC: sell hardware (40–60% margins today), then subscriptions (~20%), then a world where the robot body is nearly free and revenue comes from billing each successful action the "embodied brain" executes — much like OpenAI charges per text token. His lab backs the thesis with a unified autoregressive model handling zero-shot generalization, general grasping, and long-horizon tasks, plus real-world reinforcement learning that pushed success rates to 99.9%. (The token-for-actions framing was also picked up by Sina Finance.) It's a provocative analogy, but it captures where the margins are heading if manipulation actually scales.
What to watch: whether any of these "brain-first" theses translate into paying deployments beyond warehouses and logistics pilots.
If robots start billing by the action instead of the unit, who owns the token economy — the hardware makers or the model labs? Tell us in the comments.
Sources: Leiphone — Chen Jianyu on world models · Leiphone — Hua Wei on sim-to-real · Leiphone — Gao Jiyang on physical-world tokens · Sina Finance — Gao Jiyang token model