HiDream-O1-Embodied tops RoboColiseum's robustness chart
A Chinese world-model lab's first move into embodied robotics lands on top of the toughest sub-score, while Alibaba ships a "multi-user workbench" that turns a sentence into a working group app.
HiDream-O1-Embodied debuted at number one on the RoboColiseum robustness leaderboard, scoring 0.692 in the benchmark's most adversarial sub-track on its first submission. HiDream.ai (智象未来) framed the release as the "real-world" leg of a world-model strategy that already covers image (HiDream-O1-Image) and interactive video (HiDream-O1-World, which topped the WBench Navi sub-leaderboard a month ago). The new model closes the loop: the interactive world-model handles "understand and predict," the embodied one handles "operate and execute." RoboColiseum's robustness track is the hard one — it changes background, lighting, materials, camera position, image quality and rewrites the instructions, then checks whether the policy still works in a "greenhouse" the lab never saw. CTO Yao Ting tied the score to three engineering bets: instruction understanding that survives paraphrasing, multi-camera visual fusion that holds up when one view is occluded, and a training pipeline that feeds the model noisy, partially corrupted inputs on purpose. The data side is the moat. HiDream uses mocap-grade human motion from partner Noitom as the "real base," then has the model itself generate the long tail of background, lighting and object variations — making the model both "examinee" and "examiner." For a Chinese lab that has been quietly racking up image and world-model wins since last summer, an embodied number-one on a tough benchmark is the strongest signal yet that "native omnimodal" — one stack from pixels to actions — is a real alternative to bolted-on visuomotor policies.
Alibaba's Qianwen Office shipped what it calls the industry's first "multi-user workbench": type a sentence, get a hosted web app that up to a hundred people can run at once with role permissions, a backend and cloud storage. Most AI workspace tools stop at single-user generators — a resume page, a personal check-in tracker, a portfolio. Qianwen's version adds four enterprise-grade pieces: roles (admin vs. member vs. guest), a shared database, a management console, and one-click publish to a shareable link. Use cases the team demoed read like a roll-call of friction every small org already knows: a market organizer running vendor sign-ups, deposits and booth assignments for 100 stalls; a teacher building a homework and grading tool that students and parents each see only their own slice; a brand managing influencer and supplier workflows across cities. The pitch — explicit in Alibaba's own copy — is that you can replace a ¥10,000-a-year vertical SaaS contract (or a six-figure custom dev cycle) with a natural-language prompt, then keep editing the workflow in plain language as the business changes. Qianwen Office says it crossed 30 million users in its first month, with enterprise accounts over half of that base. The interesting bet is the positioning: Alibaba is not selling "AI for the individual knowledge worker," it is selling "AI that builds your internal tools." If the workbench holds up, every other Chinese agent suite now has a "build a working app from a sentence" reference point to match or beat.
What to watch: HiDream has topped image and interactive-world leaderboards before, but a single embodied win is one data point — the real question is whether robustness holds when real robots run the policy outside simulation. For Qianwen Office, the test is whether 30 million sign-ups translate into paying teams, or whether the workbench is a productivity demo that does not survive the second business process.
Which of these two — embodied world models or agent-built internal tools — do you think will reshape Chinese enterprise AI first? Tell us in the comments.
Sources: QbitAI (量子位) · Zhidx (智东西) · Leiphone (雷峰网)