ShengShu lays out a five-level roadmap to general world models

Share
ShengShu lays out a five-level roadmap to general world models

ShengShu founder Zhu Jun used a World Robot Conference keynote to put a name and structure on where embodied AI is heading: a five-level roadmap for a general world model he argues must unify generating the world, interacting with it, and acting inside it — not a stack of separate models bolted together.

Zhu, the ACM/IEEE/AAAI fellow who built video model Vidu, described the general world model as a closed loop rather than a single generator or robot policy: a system that understands the world, predicts what comes next, and takes action — where each action changes the environment and feeds back into the next round of understanding. His ladder runs from L1 world generation (Vidu, 2024), through L2 interactive worlds (Vidu S1's live voice-driven editing, July 2026), to L3 acting in the world with the open-source motion model Motus and its follow-up Motubrain. The two top rungs remain unwon: L4, an autonomous world agent that sets goals and explores, and L5, a "world orchestrator" coordinating robots, digital agents, humans, and tools together.

The most concrete result is at L3. ShengShu says Motubrain runs about 10 times faster at inference than Motus, adapts to a new robot body with just 50 to 100 human demonstrations, and tops the RoboTwin 2.0 benchmark at 96.1 — already validated across nearly ten robot platforms including Galaxy General and Xingchen. The pitch is that most of the industry is still scattered across L1–L3, and the real frontier is the leap to autonomy and orchestration, which Zhu flags still needs breakthroughs in goal formation, exploration, continuous learning, and long-term memory. It's a framing argument as much as a technical one — a useful scoreboard for judging how far any embodied-AI team actually is.


Anthropic is asking job candidates blunt questions about their personal finances as it barrels toward an IPO. According to an Axios scoop, the AI lab — which publishers and investors increasingly value in the trillion-dollar range and is reportedly accelerating a public listing at a roughly $2 trillion target — has pressed some employment hopefuls about their money and spending habits during interviews.

The angle is telling: a company about to file a public S-1 is under intense scrutiny over insider conduct and spending, and Anthropic appears to be baking that diligence into its hiring pipeline rather than only its finance team. But it's also a pointed question about leverage — a candidate being asked to open their books years before any vesting schedule pays off. It's the clearest signal yet that Anthropic's fund-raising sprint has become an all-hands exercise in IPO-readiness.

Should pre-IPO AI labs grill candidates about their personal finances? Tell us in the comments.

Sources: 雷峰网 Leiphone · ShengShu · Axios · Dataonomy