HiDream-O1-World tops an interactive world-model benchmark
World models keep crossing out of the research demo into something you can actually steer — and the newest one to break the frontier comes from a Chinese lab better known for image generation. ZhiXiang WeiLai's HiDream-O1-World is an interactive, native all-modal world model, and its first trip onto a public benchmark just put it at the top.
HiDream-O1-World, the interactive world model from Chinese lab ZhiXiang WeiLai (the HiDream image-model maker), topped the Navi ranking of WBench — the new interactive-video world-model benchmark — with an average score of 80.9. The model generates a coherent, walkable world you can roam from first or third person, edit in real time, and build from text, an image, or direct interaction — a real-time sibling to what we saw from HelixWorld this morning (HelixWorld adds real-time sound to interactive AI worlds). What lifts it above the pack, per the lab and the benchmark numbers, is consistency: it scored 88.0 on spatial consistency and 73.3 on physics, the top mark on that axis. ZhiXiang WeiLai credits a native all-modal "UiT" backbone combined with memory of scene geometry and online test-time adaptation, so objects don't drift, vanish, or defy gravity as the camera swings — the classic failure mode that has made interactive worlds feel fake.
The headline number matters less than what it signals: WBench, built by Meituan's LongCat team with Fudan University, is the first real yardstick for interactive world models, spanning 289 multi-turn cases and 1,058 interaction rounds. HiDream-O1-World being the first to sit on top of it gives the field a reference point it didn't have. The underlying paper was also accepted to ECCV 2026. ZhiXiang WeiLai pitches the model at interactive film and games, embodied-AI simulation, and 3D scene production for creators — all natural next stops, but the honest read is that consistent, physics-respecting, long-horizon worlds are still the hard part, and topping one benchmark doesn't settle it.
ByteDance's Seedance 2.5 video model now generates native 1080P, and the API is open for it. The Volcengine-hosted model, pitched at "cinematic long-form" clips, had topped out at 480P and 720P; the new tier adds native 10-bit color output, sharper edges and fabric/hair detail, and more natural skin and lighting meant to cut the "AI look" from generated footage. It's a resolution and fidelity step rather than a new architecture, but it pushes long-form generation toward a place where output can actually pass through professional grading and post-production. The 1080P image-to-video API is live with a launch discount to roughly 72 percent of list price through September 17, which is ByteDance's way of seeding the higher tier into real workflows while the market for cinematic video models gets crowded.
What to watch: whether ZhiXiang WeiLai follows with open weights for HiDream-O1-World, and how quickly Meituan's WBench becomes the default leaderboard for interactive worlds.
Interactive world models can now hold a consistent, explorable scene — but do they need physics that's real before they're useful, or is believability enough? Tell us in the comments.
Sources: Leiphone — HiDream-O1-World · Leiphone — Seedance 2.5 1080P · DreamWorld paper page