Unisound's U2-Flash builds a working agent in 18 minutes
Two stories from China's model race this morning, one about the model and one about the machine it runs on — both arguing the same thing: the frontier metric has moved from raw size to cost per finished task.
Unisound released U2-Flash, a "flash"-class model it says handles mainline agent, coding and office work at roughly 10 billion active parameters. The company claims under 4% of total parameters fire on any given call, so inference cost anchors to the activated slice rather than the whole checkpoint. Its post-training stack is the interesting part: multi-teacher online distillation pulls maths, code, agent and instruction-following behaviour into one model, asynchronous agent RL harvests execution traces from long-chain tasks in parallel and pinpoints the action that decided success or failure, and the model inspects its own training environment for faults — what Unisound calls a first step toward recursive self-improvement, explicitly sandboxed and validation-gated. Zhidx's hands-on found the pitch holds up in the mundane direction: an agent scaffold (including OpenCode subscription wiring, memory, model selector and preview page) built in 18 minutes, a nine-paper DeepSeek technical lineage summarised into a report plus a deck in about 40 minutes, and a 10-metre-scale Blender slide — spiral run, supports, railings, real material thickness — rendered and self-corrected without human edits. Unisound's own forecast is that Flash-tier models cover more than 90% of enterprise tasks with a flagship kept as fallback, and the launch ships with a limited-time 40% discount plus tuning for domestic Chinese accelerators, where it claims throughput and latency are closing on Nvidia GPU deployments. The catch, from the same hands-on: hard maths, frontier research and spatial-reasoning tasks still need a bigger ceiling.
SemiAnalysis benchmarked Nvidia's Vera Rubin NVL72 at 67 times Blackwell's performance per dollar on agentic inference. The run used the firm's open-source InferenceX harness against a DeepSeek V4 Pro 1.6T workload, and reported roughly seven times Blackwell's tokens per megawatt, alongside a claim that a Rubin fleet earns twice the annual profit per gigawatt. The headline number is an operating point, not a constant — the same piece credits Nvidia's own software team for Rubin bring-up, which is a polite way of saying the result assumes the best current stack — but the direction matches Nvidia's marketing of "up to 30x more work per watt" for agent serving. For anyone buying inference at scale, the useful signal is that competitive numbers for Rubin hardware are already circulating two months before delivery.
What to watch: whether U2-Flash's domestic-accelerator throughput claims survive an independent run on the same workload.
Is a small model that finishes the task better than a big model that reasons about it — where do you draw the line? Tell us in the comments.
Sources: Zhidx (智东西) hands-on · Unisound U2-Flash announcement · Sina China reprint of the Unisound release · SemiAnalysis: Rubin NVL72 Agentic Inference · AI Weekly summary of the InferenceX run · Nvidia blog on Vera Rubin NVL72 efficiency