AI 101 — What is AI inference?

Every answer an AI gives you is inference: running a trained model on new input to produce an output. Training is how a model learns; inference is how it works. If training is teaching, inference is doing.
Why it matters right now. Training makes the headlines — a new model, a bigger cluster, a record run — but training happens once per model, while inference happens every time anyone asks a question. That asymmetry is why the money has shifted: Google Cloud describes inference as the phase "where AI delivers business value," and this week's news — Amazon pledging $1 billion to host data-center towns, an AI unit reinventing itself as an inference cloud — is infrastructure being built for that steady, always-on demand, not the one-off act of training. When people talk about "inference economics," this is what they mean: the cost of answers, multiplied by billions of answers.
The mental model. A trained model is a finished instrument. Its weights — the billions of numbers learned during training — are frozen at release; nothing new is learned when you type a message. When your request arrives, the model reads your prompt in one pass, then writes its reply one token at a time, each new token predicted from everything before it. Two phases, both called inference: reading the question (engineers call it prefill) and writing the answer (decode). Your whole chat window is a sequence of these cycles, repeated as fast as the hardware allows.

The analogy. Think of a driver who spent years in driving school. That was training — expensive, slow, done once. Now she drives a taxi route eight hours a day. Every trip is inference: same knowledge, new street, instant judgment. Each individual trip costs a tiny fraction of what the school cost, but the taxi company's entire budget is trips, because she does thousands of them. The school made her a driver; the trips are the business.
Common misconceptions. First: "the model is learning from our conversation." It isn't. Inference runs the trained weights as-is — which is why a chatbot can be confidently wrong and why your corrections don't stick (when a vendor does let a model learn from feedback, that's a separate step, usually fine-tuning). Second: "inference is the cheap part." Per request, yes — Google Cloud notes each prediction is far less computationally demanding than a training run. But training a frontier model happens a handful of times a year, while inference runs continuously at global scale, so it now dominates how AI compute is actually spent. Third: "inference" doesn't mean the model is reasoning or concluding anything deep — the word is borrowed from statistics, where it simply means drawing an output from a model given data.
Why inference has its own industry. Because it runs constantly and users watch the clock, engineers optimize it differently from training. Requests from many users are batched onto the same hardware; a model can be quantized — its numbers stored with less precision — to fit more answers per second; and repeat portions of your prompt can be cached so the model doesn't re-read your whole document every message (as we covered in What is prompt caching?). Even the economics of thinking changed: reasoning models spend extra inference — tokens generated before the visible reply — to work through hard problems, which is why turning "thinking" up makes the same question cost more. Every token you see is produced one at a time by this machinery, one predicted from what came before.
Related reading: What is a large language model? · What is a token in AI? · What is model quantization?
When a chatbot gets something basic wrong, do you blame the model or the prompt? Tell us in the comments.




