AI 101 — What is a large language model?

A large language model (LLM) is a very big neural network trained on enormous amounts of text to do one job well: predict the next piece of a text, given everything that came before it — then tuned so that when you talk to it, it actually helps instead of rambling. ChatGPT, Claude, Gemini, Llama, DeepSeek: every chatbot in the news is an LLM, and most AI features in the software you already use have one quietly bolted inside.
That prediction job sounds too simple to matter, and it is the whole trick. The training has two stages. First comes pretraining: the model reads trillions of words of websites, books, and code, and plays an endless game of fill-in-the-blank — guess the next token, check the answer, adjust its internal numbers, repeat billions of times. Along the way it quietly picks up grammar, facts, style, and patterns of reasoning, because those are what make it good at guessing. Then comes post-training, where the model is taught to follow instructions and hold a conversation rather than merely continue a text — the human-feedback step we unpack in What is RLHF?.

Why it matters right now
Practically every AI story of the last two years is an LLM story in disguise. The data centers breaking ground this year exist mainly to train and run them; the software feature launches are LLMs with a product wrapper; the "agents" that companies are racing to ship are LLMs given tools and allowed to take turns — the subject of What is an AI agent?. Understanding what an LLM actually does turns a lot of breathless coverage into something predictable: a system that generates plausible text has predictable strengths (fluency, breadth, speed) and predictable weaknesses (confident nonsense, no built-in lookup), and most of the debates in AI right now are debates about those two facts.
The one-paragraph mental model
When you send a message, the model chops it into tokens and runs one loop: given everything so far, score every possible next token by likelihood, pick one, append it, repeat. Your prompt is read as the opening of a sentence the model is in the middle of writing, and its answer is that sentence being written, word by word, in one pass after another. The window of text it can see at once is its context window — everything outside it is, for that moment, invisible. Nothing else is hiding under the hood: no database query, no lookup table, just that guessing loop running very fast over a very large network — the architecture being the transformer, as our What is a transformer? explainer covers.
The exam analogy
Picture a student who has read more than any person alive — every textbook, every forum thread, every cookbook — and who sits an exam by always writing the most plausible next word rather than stopping to think about what's true. When the exam asks "The capital of France is," a million prior texts make "Paris" the overwhelming favorite, and the student aces the question. When it asks something rare or half-asked before, the student still produces a fluent, confident sentence — because plausible was always the instruction, not accurate. That single habit explains both the magic ("it can finish anything") and the failure mode we call hallucination. "Large," incidentally, originally meant the size of the network: GPT-3, released in 2020, was called large because it had 175 billion parameters — the network's tunable numbers. The word has been a moving target ever since.
Common misconceptions
"It's a searchable database." It isn't, and nothing is looked up during generation — the text comes entirely from the model's parameters. That is why a model can state a false fact with total confidence: there is no fact-checking step, only next-token prediction. It's also why products bolt on separate machinery — retrieval, search, tools — when accuracy matters.
"Bigger always means smarter." Size was the first lever, not the last word. DeepMind's 2022 Chinchilla work showed that models of the time were starved of training data — for a given compute budget, doubling the network size should mean doubling the data too — and a 70-billion-parameter model trained properly beat far larger rivals, GPT-3's 175 billion included. Data quality and training method matter as much as raw size.
"The LLM is the whole product." Usually it's the engine. A useful AI product layers retrieval, memory, guardrails, and sometimes a second model judging the first on top of the LLM. When a product "gets dumber," it's often this scaffolding changing — or the context window filling up — not the model itself.
Where to learn more
The original "Attention Is All You Need" paper by Vaswani and colleagues at Google introduced the transformer architecture behind nearly every LLM — dense, but the diagrams are worth a look. Andrej Karpathy's "Let's build GPT from scratch" builds a working language model line by line with no hidden steps. For a gentler start, Hugging Face's free NLP course walks through what language models are and how they're trained.
Related reading: What is a token in AI? breaks text into the pieces an LLM actually sees · What is a context window? covers how much it can hold at once · What is an AI hallucination? explains the price of guessing instead of knowing.
If an LLM is only predicting the next token, is the reasoning you see in its answers real thinking — or are we reading thought into a very good guesser? Tell us in the comments.




