AI 101 — What is a neural network?

A neural network is a computing system made of layers of simple math functions whose internal numbers are tuned by examples until the whole thing gets the right answers — a design loosely inspired by the brain, and the foundation of every AI system in the news. ChatGPT, Midjourney, AlphaFold: all neural networks, all variations on the same idea. The concept is older than computers.

Share
AI 101 — What is a neural network?

A neural network is a computing system made of layers of simple math functions whose internal numbers are tuned by examples until the whole thing gets the right answers — a design loosely inspired by the brain, and the foundation of every AI system in the news. ChatGPT, Midjourney, AlphaFold: all neural networks, all variations on the same idea.

The concept is older than computers. In 1943, Warren McCulloch and Walter Pitts described an artificial neuron — a thing that takes inputs and fires when they add up past a threshold. In 1958, psychologist Frank Rosenblatt built one at Cornell, the Mark I Perceptron, and wired it to a 400-pixel camera; it learned to tell left from right, and the press treated it as the first step toward an electronic brain. Nothing about the core idea has changed since. What changed is that we finally got the data, the chips, and the training trick — What is backpropagation? — to make it work at scale.

Why it matters right now

Every model announcement is a neural-network announcement. When a lab says a new model has billions of "parameters," that count is the network's internal numbers — the settings it learned. GPT-3, released in 2020, had 175 billion of them; today's frontier models run far beyond that, and the race to fund data centers is essentially a race to build bigger networks and run them longer.

It also explains the vocabulary of the whole field. Fine-tuning adjusts a network's numbers on new examples. Model quantization shrinks them to save memory. Knowledge distillation copies a big network's behavior into a small one. Distillation, MoE, RLHF — every technique you read about is operating on the same object: a stack of layers full of tunable numbers.

The one-paragraph mental model

A network is organized in layers. The first layer receives the input — the pixels of a photo, or a sentence broken into tokens. The last layer produces the output — a label, a score, the next word. In between sit the hidden layers, and every connection between them carries a number called a weight. Each neuron does one small job: take what arrives from the layer before, multiply it by its weights, add the results, pass the sum onward if it clears a threshold. A single neuron is almost useless; stacked in the millions and billions, with weights tuned by training, the collective produces fluent text and recognisable faces. The knowledge lives entirely in those weights — the code that runs them is generic and identical across a million models.

The mixing-desk analogy

Picture a professional recording studio's mixing desk: thousands of knobs and faders, one for every channel. A raw track runs in one side and comes out the other. When the desk is brand new, every knob is in a random position and the output sounds like noise. Training is an engineer twisting knobs, listening, and adjusting — helped by an assistant who calculates exactly which knobs to turn and in which direction after every pass. When the song finally sounds right, the knob positions themselves are the neural network; the song's quality isn't stored anywhere else. A new song through the same desk starts rough again, and small tweaks get it there fast — which is what fine-tuning a model feels like from the inside. "Deep" simply means the desk has many stages of knobs in series, not just one row.

Common misconceptions

"It works like a brain." Only in loose metaphor. Artificial neurons are arithmetic — multiply, add, threshold. Nobody knows whether real biology does anything of the sort; the brain inspiration is the arrangement (layers of connected units), not the mechanism.

"More layers always means smarter." Depth helps only when the network is trained well on enough data. A bigger network trained badly can memorise its examples and fall apart on new ones — the failure mode our What is overfitting? explainer covers. Scale, data quality, and training method all matter; scale alone doesn't win.

"The model stores knowledge like a filing cabinet." It doesn't, and no single neuron "holds" the concept of a cat. OpenAI's neuron-explanation research found that GPT-4's behaviour is spread across many neurons acting together — knowledge is distributed across weights, which is why networks can be compressed, and why extracting what one "knows" is genuinely hard.

"Neural networks are a new invention." The idea is 80-plus years old; the field survived two "AI winters" when funding dried up because networks couldn't scale. The recent revolution is engineering — more data, GPUs like the ones behind What is CUDA?, and better training recipes — not a new concept.

Where to learn more

3Blue1Brown's "But what is a Neural Network?" is the best visual walkthrough — it builds the layers on screen in fifteen minutes. Google's Machine Learning Crash Course starts with the same ideas in plain terms and free exercises. For the field's own summary, the 2015 Nature review "Deep learning" by LeCun, Bengio, and Hinton remains the standard reference.

Related reading: What is backpropagation? covers how the knobs actually get tuned · What is a transformer? is the network design behind today's chatbots · What is an embedding? explains what networks do with words before thinking begins.

If knowledge lives in millions of tuned numbers rather than in code, does a trained model "know" anything? Tell us in the comments.

Read more

US charges California man over $300M in smuggled Nvidia chips

US charges California man over $300M in smuggled Nvidia chips

Export enforcement leads: federal prosecutors have put a price tag on the China chip-smuggling trade, while Suno pushes its music model into spoken word. US authorities arrested Greg Lui, a 38-year-old California tech executive, on charges of smuggling more than $300 million worth of export-controlled computer servers containing Nvidia GPUs to China. Prosecutors say Lui, who runs Earthmade Computer Inc. in the City of Industry, shipped servers through Malaysia and Singapore between 2023 and 202

Deep Dive — Meta wants a cut of what Muse buys, not your eyeballs

Deep Dive — Meta wants a cut of what Muse buys, not your eyeballs

"We believe that Muse will make you money," Mark Zuckerberg told the room at Meta Connect on September 24, "and we are standing behind this by making Muse free for a huge number of tokens with the expectation that over time we will profit by taking a small fee from transactions." Read it again next to what the rest of the industry is doing and it lands as something close to heresy: the world's largest advertising company, announcing that its hottest new product will not carry a single ad.