DeepSeek ports its kernels to Huawei's Ascend 950

Share
DeepSeek ports its kernels to Huawei's Ascend 950

OpenAI's DevDay ate the wire overnight — 25 announcements, a new model, a new agent. The more consequential release landed 4,000 miles away, with no English-language coverage, and it is the piece of software that decides what Chinese models can run on.

DeepSeek has open-sourced the kernel and communication layer its models actually execute on — rewritten for Huawei's Ascend chips instead of Nvidia's. The lab published two new repositories on Wednesday, DeepGEMM-Ascend and DeepEP-Ascend, and added Ascend backends in place to three existing ones: FlashMLA, TileKernels and DeepSelect. These are not model weights. They are the pieces that make GPUs fast: matrix-multiply kernels, attention kernels, top-K sampling kernels and the all-to-all communication library MoE models use to shuffle tokens between experts across a cluster. DeepSeek's READMEs say the Ascend versions are API-compatible with the GPU originals and ship under the same package names, so a user installs the Ascend package and keeps the same code and workflow. The validated stack is Huawei's CANN 9.20 toolkit on the Ascend 950 series.

The numbers are the point. On Ascend 950DT, DeepGEMM-Ascend's dense BF16 matrix multiply at a 4096×7168×16384 shape runs at 431 of a 432 TFLOPS hardware limit — 99.8% — with FP8 at 861 of 865 TFLOPS, or 99.5%. DeepSeek's sparse-attention kernels, the ones that carry its Sparse Attention design, run at up to 410 TFlops for prefill (95% of peak) and 360 TFlops for decoding (83%) on the same chip. DeepEP-Ascend's expert-parallel communication hits 373–375 GB/s for dispatch and 345–347 GB/s for combine at EP8, which the repo says is roughly 90–95% of the physical payload bandwidth limit for EP sizes up to 32. Those are the specific operations that have kept non-Nvidia accelerators roughly a generation behind on large-model inference. We covered the scale problem in DeepSeek plans a 160,000-chip Huawei cluster in Inner Mongolia and the systems-workaround version in China Mobile Cloud runs attention on GPUs and FFN on brain chips.

Read the fine print before believing the bandwidth. DeepEP-Ascend's own README states the measurements came from "a PoC HDK supplied to DeepSeek, with additional manual configuration" that "is not a publicly distributed release," and warns that users on earlier PoC hardware may see lower bandwidth. The recommended public baseline is Huawei's commercial Q3 release for the Atlas 850E, planned for around October 15, which is still unreleased — and the repo says these are not results from it. Combine and EP sizes above 32 remain under optimization. The kernel benchmark on the other hand is a standard matrix shape with a stated hardware ceiling, and it reproduces easily.

Why this is the story rather than another model port: DeepSeek has published open weights before, but weights are portable — anyone can compile a model for a new accelerator. The kernel stack is where the lock-in lives, because it encodes years of hand-tuned memory paths and scheduling for one vendor's hardware. A frontier lab rewriting that layer for domestic silicon, in parallel with the codebases it maintains for Nvidia, means its next models are being prepared to run first-class on chips it can buy without an export licence. The port usually trails a model by months; here it landed days after Huawei's own toolchain update, and Huawei's engineers are named as collaborators in the READMEs. Chinese coverage describes it as a milestone for the domestic software ecosystem. From outside China, it is a measurable dent in a moat rather than a broken one.

What to watch: whether Huawei's commercial HDK ships the full-bandwidth configuration on schedule around October 15 — and whether DeepSeek's next model release leads with Ascend numbers.

If the kernel layer is where the lock-in lives, does a port to Ascend change the calculation on export controls? Tell us in the comments.

Sources: DeepGEMM-Ascend (GitHub) · DeepEP-Ascend (GitHub) · QbitAI