WorldArena 2.0 puts world models to the test on real robots

The world-model field has argued for two years about video quality. This week the first full results landed for a benchmark that asks the harder question: can a robot actually use the prediction?
WorldArena 2.0's global challenge has closed its leaderboard, and for the first time world models are being graded on physical robots instead of plausible video. The benchmark — designed by a Tsinghua-led consortium with PKU, CMU, Stanford, Princeton and others — extends its 1.0 video scoring along three new axes: vision-plus-touch input, world models used as online reinforcement-learning environments, and execution on real hardware. The challenge ran as an IROS 2026 competition across three tracks, with final results released September 15 and an award ceremony at the IROS workshop: 81 models scored on video quality, 47 on serving as an RL environment, and a real-robot track run on AgileX dual-arm and Franka Panda manipulators. The total prize pool was 7,000 dollars.
The headline result is a route contest: JEPA-style latent prediction versus spatial-intelligence modeling, and the JEPA camp drew first blood. In the video-quality track, Visincept's WorldIncept took first at 72.33, with Tongji University's spatial-intelligence route BWM-Turbo second at 71.49 — a gap of 0.84 points. The split was telling: WorldIncept led on physics adherence and 3D accuracy, while BWM-Turbo topped background consistency and JEPA similarity. In the RL-environment track, Beijing Zhongguancun Academy's MW2 won with 72.93 and Twilight-line Tech's TTWM placed second at 72.40 — the only two of 47 entries above 72. Team affiliations come from Leiphone's reporting on the results; the scores themselves are on the public leaderboard.
On real robots, Xiaomi's MiRobot team won decisively — and the losing sub-score is the actual story. MiRobot's ViTacX topped the overall Track 3 standings with 76.64 against ShanghaiTech's OmniFlow at 67.37, driven by an 85.00 in vision-only manipulation where it notched perfect 100 scores on wiping, pouring and folding. But on visuo-tactile tasks — gripping a potato chip, peeling a cucumber, plugging in a two-pin plug — ViTacX fell to third at 66.67, and OmniFlow was the only team to place top-two in both sub-tracks. Seeing the scene clearly did not transfer to handling contact, which is exactly the gap the organizers say remains: in their own words, leading models still cannot accurately and continuously capture the state transitions an action causes, and small errors compound across a rollout.
What to watch: the organizers list longer closed-loop interaction, autonomous error recovery, cross-embodiment transfer, and tactile-geometry fusion as 3.0 candidate dimensions — plus whether any single model can eventually lead all three tracks instead of specialists trading wins.
For a neighboring scoreboard, we covered HiDream-O1-Embodied tops RoboColiseum's robustness chart earlier.
Should a world-model benchmark judge video plausibility, policy training, or real-robot success hardest? Tell us in the comments.




