NIO's MM-Future plans driving by imagining 64 futures at once
A NIO paper turns consequence-aware planning into a benchmark score, Waymo stacks 271 million driverless miles behind its safety claim, and the AlphaFold Database quietly became pandemic infrastructure.
Ren Shaoqing's name is on an autonomous-driving paper for the first time in a decade, and the method is the one the world-model field has been circling: let the car imagine many futures and pick the one it likes best. MM-Future, posted to arXiv on September 17 by a NIO-led team with Ren — a co-author of Faster R-CNN and ResNet, now a chair professor at USTC and head of NIO's intelligent driving business — as corresponding author, generates dozens of paired scene-and-action hypotheses at once. Each candidate trajectory carries its own predicted future scene; the two evolve together inside the model, and a final scorer that can read only a candidate's own paired future picks the winner. That is the gap it targets. Cascade models predict the future and then plan behind it, one direction only; joint world-action models let scene and action influence each other but typically produce a single result.
The scores are what will get quoted. On NAVSIM's navtest split MM-Future reaches 94.0 PDMS against DriveFuture's 90.7, takes 91.5 EPDMS on NAVSIM-v2 against UniTeD's 90.1, and hits 32.3 average HD-Score across 436 HUGSIM scenarios without any fine-tuning. The ablations explain where the gain comes from: a single action mode scores 84.1, thirty-two modes reach 92.3, joint scene-action generation adds 0.6, and letting the scorer read each candidate's paired future adds another 0.4 — and multi-mode training converges in about 3,800 steps to 0.80 validation PDM where a single mode needs 17,500. The costs are real and the paper states them: sampling 64 hypotheses takes roughly 233 ms end to end on one H800 at batch size 1, and on HUGSIM's extreme difficulty it loses badly to Latent-WAM (8.6 versus 18.1 HD-Score) while finishing slightly fewer routes. NIO's version of this bet is already on the road — by the company's count the world-model-plus-reinforcement-learning stack is in more than 950,000 vehicles. For the vocabulary behind it, we have AI 101 — What is a world model?.
Waymo published a safety report covering 271 million driverless miles and claims 82 percent fewer injury-causing crashes than human drivers, with 95 percent fewer crashes causing serious injury or worse. The data runs through the end of June across Atlanta, Austin, Phoenix, Los Angeles and the Bay Area — 50 million miles past its previous report — and the rest of the list holds the pattern: 93 percent fewer pedestrian injury crashes, 86 percent fewer cyclist, 82 percent fewer motorcycle. Two things stop this from being a closed case. First, it is Waymo's own data. The Insurance Institute for Highway Safety found a smaller 68 percent reduction in its independent study, and its president David Harkey warns the current data system "isn't good enough to allow continuous monitoring of a large-scale expansion"; under the federal standing general order Waymo still reported 797 crashes from March through September, including one fatality where a pedestrian struck by a human driver was thrown into a stopped Waymo. Second, a sensor-laden robotaxi compared against the average human driver — older cars, distracted drivers — mostly demonstrates that automatic emergency braking and lidar work.
The AlphaFold Database now holds predicted 3D structures for the protein complexes of more than 2,800 viruses, released by EMBL-EBI, Google DeepMind and NVIDIA with Seoul National University, the University of Glasgow, the Swiss Institute of Bioinformatics and CEPI. The release leans on NVIDIA's BioNeMo Inference Runtime to scale AlphaFold2 across thousands of viral proteomes, and NVIDIA says roughly 30 percent of the protein interactions added had never been documented in the Protein Data Bank — the company also open-sourced the pipeline that produced them. The caveat comes from a collaborator, not a critic: Joe Grove, a molecular virologist at Glasgow, points out that a structure alone says nothing about what a mutation does and does not make it easier to engineer a more dangerous pathogen. The clock is the argument. The Center for Global Development puts close to a 50 percent chance on a pandemic as severe as COVID-19 by 2050, and the 100-day target for vaccines and treatments only works if the structural groundwork happened long before the outbreak.
What to watch: whether MM-Future survives peer review and reaches NIO's next production model, and whether the cities still holding Waymo out read a self-reported dashboard as evidence.
Would you ride in a car whose plan came from a simulated future — or should "the model imagined it" need a regulator's signature first? Tell us in the comments.
Sources: arXiv — MM-Future: Multi-Mode Joint World–Action Modeling for Autonomous Driving · QbitAI — 时隔十年,AI大牛署名新论文 · 36Kr (智能车参考) — Ren Shaoqing's new paper · Waymo — Safety Impact · The Verge — Waymo's driverless cars continue to crash less often than people · IIHS — Waymo's driverless cars crash less often than people · EMBL-EBI — AlphaFold Database adds viral protein complexes to support pandemic preparedness · NVIDIA — How Open Science Can Help Researchers Prepare for the Next Pandemic · AlphaFold Protein Structure Database