AI 101 — What is CUDA?
CUDA is the software layer that lets ordinary programs run on NVIDIA's graphics chips — and it is the foundation nearly all modern AI is built on. Short for Compute Unified Device Architecture, it first shipped on February 16, 2007, and it turned graphics cards — machines designed to draw video games — into general-purpose computers that any developer could program.
Before CUDA, using a graphics chip for serious math meant writing code in the chip's own graphics language, in fragments, with little support. NVIDIA's bet was different: give developers a compiler, a programming model, and a toolbox of pre-built math routines, and the same hardware that renders games could also simulate weather, train models, and serve answers to chatbots. Almost every AI framework since — PyTorch, TensorFlow, and the training stacks at every major lab — calls CUDA under the hood when it touches an NVIDIA GPU.
You have probably seen the name without looking for it. Graphics cards advertise "CUDA cores" in their spec sheets. Install guides for AI tools ask which version of CUDA you have. And when news stories describe one company's dominance in AI chips, CUDA is usually the reason being described — the moat is not just the silicon.

Why it matters right now
For two decades, the practical answer to "where do I run AI?" has been NVIDIA hardware, and the reason is rarely the chip alone — it is the software stack sitting on top of it. A GPU is useless without the libraries that turn matrix multiplications into speed, and CUDA ships them: cuBLAS for linear algebra, cuDNN for the operations neural networks lean on, TensorRT for squeezing inference onto one card, NCCL for coordinating thousands of GPUs in a training run. Each one is a problem that took NVIDIA years to solve, and every one of them assumes you are running on NVIDIA's own platform.
That assumption is now being poked from every direction. AMD's open-source ROCm stack — at version 10 since August 2026 — is the most established alternative, and it runs the major frameworks. In September, a Windows project built on the open-source ZLUDA translation layer completed a real PyTorch reinforcement-learning run on an AMD Radeon RX 9060 XT, which we covered in CUDA training just ran on an AMD Radeon in Windows — reproducibly. Even NVIDIA's library code is no longer sacred: as What is a learned kernel? explains, AI programs now write parts of this math code themselves — and beat NVIDIA's hand-tuned cuBLAS routines at it.
The one-paragraph mental model
Think of three layers. At the bottom, the hardware: a CPU is a handful of extremely fast, extremely flexible cores; a GPU is thousands of simpler cores that do the same small calculation over and over. In the middle, CUDA: the translator that takes a job written for the CPU, splits it into thousands of identical pieces, hands them to the GPU, and gathers the results — plus the pre-written libraries for the calculations AI needs most. At the top, the framework you actually touch, such as PyTorch, which asks CUDA for GPU acceleration without you ever seeing it. When someone says a model "runs on CUDA," they mean the middle layer is NVIDIA's — which is why competitors have to rebuild not just a chip, but the whole translator.
The kitchen analogy
Picture a professional kitchen. The CPU is the head chef: can do anything, improvises beautifully, but works one dish at a time. The GPU is a hundred line cooks: each one can only do simple, repetitive tasks — chop, stir, plate — but do them simultaneously. Before CUDA, the head chef handed out orders in a language the line cooks did not speak, so most kitchens never used them. CUDA is the ticketing system and the training manual: every order gets written in a form any cook can execute, dispatched to whoever is free, and the plates come back assembled. The library recipes are the thousand dishes the restaurant serves every night, already written down — which is why switching to a rival kitchen means re-writing the entire recipe book, not just hiring new cooks.
Common misconceptions
"You need to learn CUDA to use AI." Almost nobody writes CUDA directly anymore, in the same way almost nobody writes machine code. You call a framework, the framework calls CUDA. Learning it helps when you are optimizing how a model runs — deciding what the GPU does with every millisecond — not when you are using one.
"CUDA cores tell you how fast a GPU is." The count is NVIDIA's own marketing unit — the arithmetic units inside its chips — and there is no clean way to compare it against AMD's or Apple's equivalents, which count differently. What predicts AI performance better is the whole package: chip architecture, memory bandwidth (our explainer on What is HBM? covers why memory is usually the binding constraint), and how well the libraries exploit them.
"CUDA is NVIDIA's AI product." It predates the current AI boom by more than a decade — it is infrastructure, not a model. NVIDIA sells chips; CUDA is why people buy the chips. The company's AI story runs entirely on top of a graphics-computing tool it launched in 2007.
"The alternatives are ready and CUDA is finished." Real movement exists — ROCm is genuinely open-source, ZLUDA proved the translation path works — but gaps remain: the ZLUDA team's own notes show multi-GPU communication and some inference tooling still assume NVIDIA. Being viable in one workload is not being a drop-in replacement, which is exactly why the escape attempts make headlines.
Where to learn more
NVIDIA's CUDA documentation is the definitive reference — and the fastest way to see how much sits on top of the platform. AMD's ROCm documentation is the best view of what building the alternative actually requires. The ZLUDA repository, where a mostly-solo effort documents each CUDA feature it manages to translate, is a good measure of how deep the compatibility problem runs.
Related reading: What is a TPU? covers Google's answer to buying NVIDIA's stack — building its own chip and its own software, on its own terms.
Would you buy an AMD card today if the software stack were identical? Tell us in the comments.
Sources: NVIDIA — CUDA documentation · Wikipedia — CUDA · AMD — ROCm documentation · ZLUDA (GitHub)