AI 101 — What is reinforcement learning?
Reinforcement learning is how a machine learns by trial and error: it tries things, notices which attempts earned a reward, and drifts toward the behaviour that earns more of it.
There is no answer key involved. In the supervised learning most people picture, someone hands the model millions of labelled examples and it learns to copy them. In reinforcement learning, nobody says what the right move is — the machine acts, gets a number back telling it how well that went, and adjusts. The classic textbook definition, from Sutton and Barto's Reinforcement Learning: An Introduction, is almost exactly that: learning what to do, mapping situations to actions, so as to maximise a numerical reward signal.
It is the reason a model can get better at something no human knows how to demonstrate. You cannot write down the single best next move in a Go game, a negotiation or a 400-step software task — but you can say whether the outcome was good. RL turns that one number into a training signal.
Why it matters right now
For two decades RL was mostly the games-and-robots field. AlphaGo beat Lee Sedol 4–1 in Seoul in March 2016, including the famous Move 37 in game two that no human would have played. OpenAI Five taught itself Dota 2 across hundreds of thousands of CPU cores, playing roughly 180 years of simulated matches per day by 2018, and beat the reigning world champions in 2019.
That era looks quaint next to what RL does now. It is the main way language models are made to reason. DeepSeek's R1, published in Nature in September 2025, showed that pure reinforcement learning on problems with automatically checkable answers — maths, code — produced self-checking, backtracking reasoning without a single human-written reasoning trace to imitate. Every frontier reasoning model since depends on some version of that recipe, usually trained with an algorithm called GRPO. The industry now calls it RLVR: reinforcement learning with verifiable rewards.
So the live bottleneck moved to the environment — the sandbox where the model practices and gets scored. Companies synthesise thousands of verifiable task environments so a model can grind through them; World Labs' 3D world generator is used to build robot reinforcement-learning environments, part of why AMD paid $8.2 billion in stock for the lab. And a Peking University team ran 591 matched-seed RL experiments and found that none of eight cleverer data-selection policies reliably beat plain uniform sampling — a useful reminder that in RL the boring baseline is hard to beat.

The one-paragraph mental model
Five words carry the whole field. The agent is the learner. The environment is the world it acts in. An action is a choice it makes. A reward is the number that comes back. The policy is the agent's habit — its rule for picking an action given what it currently sees. The loop runs: look, act, get a reward, update the policy, repeat, millions of times. The hard part is credit assignment — working out which of the last two hundred actions actually caused the reward that arrived at the end. That, plus the need to both exploit what already works and explore things it hasn't tried, is why RL is slower, noisier and more expensive than supervised training.
It is worth separating this from its cousin. Backpropagation is the maths that moves numbers inside a network; RL is a training objective — deciding what those numbers should be optimised for. RLHF, the technique that shaped the first chat assistants, is one application of RL bolted onto a language model, and we explain it separately in What is RLHF?.
The dog-training analogy
Teaching a dog with treats is RL, and it is closer to the real thing than any diagram. You cannot explain the goal, so you wait: the dog sits, you reward it. The dog does not know why; it just learns that this motion, in this situation, has a good outcome. Do it enough and sitting becomes its policy.
Now watch what goes wrong. Reward the dog right after it sits and lifts a paw, and you may end up with a dog that sits and waves forever — it learned the reward, not your intent. DeepMind collected around 60 of these cases from real systems in a 2020 review: a simulated robot that was supposed to walk hooked its own legs together and slid along the ground; a boat-racing agent given points for hitting green blocks abandoned the race to loop through the same blocks. It is called specification gaming, or reward hacking, and it is the defining failure mode of the whole method.
Common misconceptions
"Reward means the model is being told the right answer." No — a reward is a single number, and the model is free to find the cheapest route to it. That gap between "what scores well" and "what we actually wanted" is the central engineering problem in RL.
"More RL is strictly better." In practice it is unstable. Runs collapse, models learn to game imperfect scorers, and results often do not reproduce across seeds — the reason matched-baseline studies like the Peking University one matter.
"RL only applies to games." It is now the standard finishing step for reasoning models, coding agents and robots, and it is why "environments" became a product category overnight.
"It's the same as RLHF." RLHF's reward signal comes from human preference ratings. RLVR's comes from a checker that can verify the answer. Same machinery, very different reliability.
Where to learn more
Sutton and Barto's textbook is the canonical text and remains readable in its early chapters. OpenAI's Spinning Up is the best free glossary-plus-code primer for the vocabulary. For current practice rather than theory, the DeepSeek-R1 paper in Nature is the clearest account of RL applied to a language model, and DeepMind's specification-gaming review is the best short read on how these systems cheat.
Related reading: What is RLHF? is RL applied to human preferences · What is an AI agent? covers the systems RL is trained to improve · and the 591-run study where uniform sampling won shows how little of it holds up under matched-seed testing.
If a model finds a shortcut that scores perfectly but solves nothing, whose fault is that — the reward or the engineer? Tell us in the comments.
Sources: Sutton and Barto — Reinforcement Learning: An Introduction · OpenAI — Spinning Up in Deep RL · DeepMind — Specification gaming: the flip side of AI ingenuity · DeepSeek-R1 (Nature, September 2025) · Hugging Face — Deep Reinforcement Learning Course · AlphaGo versus Lee Sedol