AI 101 — What are scaling laws?

Scaling laws are the empirical rules describing how an AI model's performance improves in a smooth, predictable curve as three things grow: the size of the model, the amount of data it trains on, and the computing power spent training it. They are not physics — nobody proved them in a lab the way gravity was proven — they are patterns that kept holding every time someone measured them, which is why the entire AI industry now plans around them.
The pattern got its canonical statement in a January 2020 paper from OpenAI, "Scaling Laws for Neural Language Models." Train a few hundred small models, plot each one's error rate against its size, and something remarkable falls out: on the right kind of graph, the points land on a straight line. That linearity — a power law, where multiplying the input by ten always buys the same percentage drop in error — held across more than seven orders of magnitude, from tiny models to the largest that could be built at the time. Details like making the network wider or deeper barely mattered next to raw scale.
Why it matters right now
Scaling laws are the reason the AI economy looks the way it does. If quality were unpredictable, nobody could justify the data-center loans, chip orders, and gigawatt power contracts that fill our news pages — you do not borrow billions against a lottery ticket. You borrow against a curve. When a lab commits to a training run, the scaling-law line is the forecast on the slide deck: spend this much compute, get roughly that much capability, with error bars small enough to bet a company on.
The curve also explains the industry's shifting strategy. For years, the game was simply making models bigger. Then, in March 2022, DeepMind's "Training Compute-Optimal Large Language Models" — the paper that produced the Chinchilla model — showed the industry had been doing it wastefully: for compute-optimal training, model size and training data should double together, and most existing models were undertrained. A 70 billion-parameter model trained on four times as much data beat its 280 billion-parameter sibling running on the same compute budget. And after raw pretraining scale started delivering visibly smaller gains, labs found a new axis entirely: spending compute at answer time, letting a model think longer before replying. A 2024 study from Snell and colleagues showed that, done well, this test-time scaling can beat simply buying a bigger model — and it is the logic behind every reasoning model shipped since.
The one-paragraph mental model
Picture a straight line drawn on graph paper — except both axes grow by multiples, not steps. That is what researchers plot: model size times ten, error down by a fixed fraction, again and again, as far as anyone has measured. The line is the product of countless small experiments, and its power is prediction: fill in where you want to land on the error axis, and the line tells you what size model, how much data, and how many chips it should take to get there. Every training run announced by a major lab is, at bottom, an attempt to place a point exactly where a line says it will go.
The analogy
Think of a person learning a foreign language. The first thousand words transform you — you go from helpless to ordering dinner. The next thousand help less in absolute terms; the thousand after that less still. But the gains never stop, and the curve is so regular that a good teacher can predict roughly where you will be after the next batch of vocabulary. Each doubling of study buys a smaller slice of fluency than the last, yet a consistent slice — consistent enough to plan a curriculum around. That is a power law in everyday clothing, and a scaling law says the same thing about models: predictable, diminishing, but never running dry — at least, not yet.
Common misconceptions
"Bigger is always better." Not since Chinchilla. Scaling three things unequally wastes compute: a giant model starved of data underperforms a smaller one fed properly. Post-2022 labs scale parameters and tokens together — which is also why training data became a scarce, fought-over commodity.
"It's a law of nature." It is a pattern, not a guarantee. The line was measured on the data, hardware, and methods available at the time; it says nothing about what happens when quality training text runs short, or when a new technique shifts the curve. Researchers keep re-measuring precisely because the line could bend.
"Scaling laws died when pretraining slowed." They changed axes. When making the base model ever larger stopped paying off at the same rate, the frontier moved to reasoning — allocating more compute to thinking before answering, the territory we cover in What are reasoning tokens?. That is still scaling; the budget just moved from training time to answer time. And none of it changes what a model fundamentally is under the hood — see What is a large language model?.
"The line measures intelligence." It measures loss — error at predicting text. Turning that into "the model is smart" happens through tests like those in What is an AI benchmark?, and lower loss does not automatically mean better judgment, safety, or usefulness on the task you care about.
Where to learn more
Start with the original: OpenAI's "Scaling Laws for Neural Language Models" is short, surprisingly readable, and its graphs make the whole idea obvious in one glance. For the correction that reshaped training budgets, read the Chinchilla paper's abstract — the finding fits in three sentences. And for the modern twist, Snell and colleagues' test-time compute paper explains why "thinking longer" became the industry's favorite new lever.
Related reading: What is a large language model? · What are reasoning tokens? · What is an AI benchmark?
Is the whole AI buildout just a bet that the line keeps holding? Tell us in the comments.




