Sugon validates GPU-direct RDMA across a 100,000-card AI cluster

Share
Sugon validates GPU-direct RDMA across a 100,000-card AI cluster

The most important AI infrastructure news out of China this week is not a model — it is the plumbing underneath one.

Sugon says IBGDA, a networking technique that lets GPUs drive the network card directly instead of routing every transfer through the CPU, has completed full-scale validation on a cluster of 100,000 accelerator cards. The company's token-acceleration stack, built on its scaleFabric RDMA fabric, marks the first at-scale deployment of InfiniBand GPUDirect Async inside China — filling what the company calls a blank in the country's large-scale AI networking.

The why is arithmetic. GPUs waste more time waiting than computing: Sugon's own figures put data-center model FLOP utilization at 38–43%, with network communication eating 30–50% of training time on large distributed runs. Standard RDMA already let network cards read GPU memory directly, but the CPU still had to initiate every transfer — and in MoE training's All-to-All exchange, thousands of tiny messages mean CPU overhead dwarfs the data movement itself. IBGDA hands the job to the GPU kernels, which stage data in video memory and ring the card's doorbell themselves; Sugon says small-message efficiency improved by two orders of magnitude.

On the 100,000-card cluster, the company reports 260-nanosecond single-hop switching latency and under one microsecond end-to-end across the usual three-hop switch path, with 4KB-plus messages running roughly 63% below mainstream RoCE latency — a figure Sugon puts on par with Nvidia's NDR generation. scaleFabric is also the first Chinese network product adapted to DeepEP, the open-source MoE communication library, and the stack adds storage shortcuts too: NVMe over RDMA cut storage access latency by more than 80%, while NFS over RDMA reportedly lifted GPU utilization in one deployment from about 55% to over 95%.

Timing explains the push. China's intelligent compute capacity grew 177% year over year to 2,185 exaFLOPS by the end of June, daily token consumption passed 140 trillion in March — up 40% from the end of 2025 — and the 15th Five-Year plan explicitly orders 10,000-card-plus clusters on domestic silicon. At that scale, utilization is won or lost in the network, not the accelerator. The parity claim deserves independent benchmarks before anyone celebrates, but the harder proof is already in: this ran in production, not on a slide. The interconnect layer — the same one Ayar Labs just raised $150 million against as copper hits its limits (Ayar Labs adds $150M more as copper interconnects hit the wall) — is becoming the real battleground.

What to watch: independent numbers on the NDR-parity claim, and whether DeepSeek-class labs adopt the stack for their next training runs.

Is the interconnect, not the chip, the real frontier of the AI compute race? Tell us in the comments.

Sources: Huanqiu.com (via NetEase) · Xinhua Finance · 21jingji · Sohu news