AI 101 — What is AI inference?

Share
AI 101 — What is AI inference?

Every answer an AI gives you is inference: running a trained model on new input to produce an output. Training is how a model learns; inference is how it works. If training is teaching, inference is doing.

Why it matters right now. Training makes the headlines — a new model, a bigger cluster, a record run — but training happens once per model, while inference happens every time anyone asks a question. That asymmetry is why the money has shifted: Google Cloud describes inference as the phase "where AI delivers business value," and this week's news — Amazon pledging $1 billion to host data-center towns, an AI unit reinventing itself as an inference cloud — is infrastructure being built for that steady, always-on demand, not the one-off act of training. When people talk about "inference economics," this is what they mean: the cost of answers, multiplied by billions of answers.

The mental model. A trained model is a finished instrument. Its weights — the billions of numbers learned during training — are frozen at release; nothing new is learned when you type a message. When your request arrives, the model reads your prompt in one pass, then writes its reply one token at a time, each new token predicted from everything before it. Two phases, both called inference: reading the question (engineers call it prefill) and writing the answer (decode). Your whole chat window is a sequence of these cycles, repeated as fast as the hardware allows.

Contemporary computer with black screen placed on stand near row of server steel racks in data center

The analogy. Think of a driver who spent years in driving school. That was training — expensive, slow, done once. Now she drives a taxi route eight hours a day. Every trip is inference: same knowledge, new street, instant judgment. Each individual trip costs a tiny fraction of what the school cost, but the taxi company's entire budget is trips, because she does thousands of them. The school made her a driver; the trips are the business.

Common misconceptions. First: "the model is learning from our conversation." It isn't. Inference runs the trained weights as-is — which is why a chatbot can be confidently wrong and why your corrections don't stick (when a vendor does let a model learn from feedback, that's a separate step, usually fine-tuning). Second: "inference is the cheap part." Per request, yes — Google Cloud notes each prediction is far less computationally demanding than a training run. But training a frontier model happens a handful of times a year, while inference runs continuously at global scale, so it now dominates how AI compute is actually spent. Third: "inference" doesn't mean the model is reasoning or concluding anything deep — the word is borrowed from statistics, where it simply means drawing an output from a model given data.

Why inference has its own industry. Because it runs constantly and users watch the clock, engineers optimize it differently from training. Requests from many users are batched onto the same hardware; a model can be quantized — its numbers stored with less precision — to fit more answers per second; and repeat portions of your prompt can be cached so the model doesn't re-read your whole document every message (as we covered in What is prompt caching?). Even the economics of thinking changed: reasoning models spend extra inference — tokens generated before the visible reply — to work through hard problems, which is why turning "thinking" up makes the same question cost more. Every token you see is produced one at a time by this machinery, one predicted from what came before.

Related reading: What is a large language model? · What is a token in AI? · What is model quantization?

When a chatbot gets something basic wrong, do you blame the model or the prompt? Tell us in the comments.

Read more

Deep Dive — The intelligence explosion, in the labs' own numbers

Deep Dive — The intelligence explosion, in the labs' own numbers

On Monday, the Cambridge Programme on AI Science & Policy published a 22-author report arguing that automating AI research could compress years of progress into months or less, and that the mechanics for it are closer than the field's usual hedging admits. The author list runs from Geoffrey Hinton and Yoshua Bengio to OpenAI chief scientist Jakub Pachocki, Anthropic co-founder Jack Clark, Microsoft's Eric Horvitz, UMass Amherst's Andrew Barto, Berkeley's Dawn Song and UBC's Jeff Clune — every si

Hinton, Bengio among 22 authors warning of an intelligence explosion

Hinton, Bengio among 22 authors warning of an intelligence explosion

The people building frontier AI are now publishing the warnings about it — plus Musk recruits a second foundry for Terafab, and Apple moves to wall AI agents off the Mac's most powerful permission. Twenty-two AI researchers — including Geoffrey Hinton, Yoshua Bengio, OpenAI chief scientist Jakub Pachocki and Anthropic co-founder Jack Clark — argue that automating AI research could trigger an intelligence explosion, and that the mechanics for one are closer than the field's usual hedging admits.

Kaiming He's harness gives Claude a perfect ARC-AGI-3 score

Kaiming He's harness gives Claude a perfect ARC-AGI-3 score

A benchmark saturated twice in a month, a security program buried by machine-generated reports, and Anthropic spending real money on human training — three stories, one theme: the stack around the model is now where the action is. A visual harness from Kaiming He's MIT lab pushed Claude Opus 5.0 RHAE to a perfect 100.00 on ARC-AGI-3, all 25 public games — using 57.4% fewer actions than first-time human players. The paper (VISTA, posted October 1) is careful about what it did and didn't do: it d

Airbnb's Chesky says AI agents need their own OS

Airbnb's Chesky says AI agents need their own OS

Airbnb CEO Brian Chesky says AI agents need their own operating system — and he's told Sam Altman so. In an interview following Airbnb's fall update, which shipped AI-powered search, Chesky argued that the chatbot is the wrong interface for e-commerce entirely — "what you're seeing today is also not the endgame for e-commerce or for travel or shopping" — and sketched where the company is actually heading: an Airbnb agent in the explore tab, another in customer service, and eventually "a macro Ai