How Neural Networks Learn: Neurons, Weights and Backpropagation Explained Simply

Every time a photo app recognizes your friend's face or a chatbot finishes your sentence, a neural network is doing the work. The name sounds intimidating, but the core idea is surprisingly simple: a neural network is a big collection of adjustable numbers, and learning means nudging those numbers until the network's mistakes get small.

In the AI Primer we saw that machine learning finds rules from examples. This post opens the box one level deeper. We'll build up from a single artificial neuron to a full network, then walk through exactly how it learns, ending with about 40 lines of Python you can run yourself.

Start with one neuron

An artificial neuron is a tiny calculator. It takes a few numbers in, gives one number out, and it does that in three steps.

Imagine a neuron deciding whether you should go for a run. It looks at three inputs: is it sunny (1 = yes), do you have free time (1 = yes), and are you tired (0 = no). Each input has a weight that says how much it matters. Sunshine and free time push toward yes, so they get positive weights. Tiredness pushes toward no, so it gets a negative weight.

1Sunnyw = 0.81Free timew = 0.60Tiredw = −0.9Σ + bsumReLUactivation0.4output

One neuron: multiply each input by its weight, add them up with a bias, then pass the total through an activation function.

Here is the arithmetic. Multiply each input by its weight and add them up: (1 × 0.8) + (1 × 0.6) + (0 × −0.9) = 1.4. Then add a bias, a number that shifts how easily the neuron says yes. With a bias of −1.0 the total is 0.4. Finally, the total passes through an activation function, which decides what the neuron actually outputs.

That's it. Weights, a bias, and an activation function. The magic is not in any single neuron but in having thousands or billions of them, with their weights set just right.

The key insight: nobody sets those weights by hand. The network starts with random weights and learns good values from examples. The rest of this post is about how.

Why the activation function matters

Without an activation function, a neuron is just a weighted sum, and stacking weighted sums on top of each other only ever gives you another weighted sum. In other words, the whole network could only draw straight lines, and the real world is not made of straight lines.

An activation function adds a bend. Two common choices:

ReLUkeep positives, zero out negativesinputSigmoidsquash anything into 0 to 1input

Two popular activation functions. The bend is what lets networks learn curved, complicated patterns.

  • ReLU (rectified linear unit) outputs zero for any negative input and passes positive inputs through unchanged. It is simple, fast, and the default choice inside most modern networks.
  • Sigmoid squashes any number into the range 0 to 1, which makes it handy for outputs that should look like probabilities, such as "how likely is this email spam?"

Stack neurons into layers

A neural network arranges neurons in layers. The input layer receives the raw data, such as the pixels of an image. One or more hidden layers transform it. The output layer produces the answer, for example the probability that the image shows a cat or a dog. When a network has many hidden layers, we call it deep learning.

cat 0.9dog 0.1InputHidden layer 1Hidden layer 2OutputForward pass: make a predictionBackpropagation: send the error back and adjust weights

Data flows forward to make a prediction. The error flows backward so every weight learns how to change.

Each layer learns to detect something a little more abstract than the layer before it. In an image network, early layers respond to edges and color blobs, middle layers combine those into textures and shapes, and later layers recognize things like ears, wheels or faces. Nobody programs these features. They emerge during training because they help reduce the error.

The forward pass: making a prediction

Making a prediction is called the forward pass. Data enters on the left, every neuron does its multiply-add-activate routine, and the results flow layer by layer to the output. At the start of training, with random weights, the prediction is basically a guess. Show the network a cat photo and it might answer "cat 0.52, dog 0.48".

Measuring mistakes with a loss function

To improve, the network needs a single number that says how wrong it was. That number is the loss. If the photo really is a cat and the network said "cat 0.52", the loss is fairly high. If it said "cat 0.97", the loss is close to zero.

Different tasks use different loss functions. Predicting a number, such as a house price, often uses mean squared error: the average of the squared differences between predictions and true values. Classification tasks usually use cross-entropy, which punishes confident wrong answers especially hard. The details vary, but the goal is always the same: make the loss as small as possible.

Gradient descent: walking downhill

Picture the loss as a landscape of hills and valleys, where your position is set by the network's weights and your altitude is the loss. Training means finding a low valley. The catch is that you're standing in thick fog and can only feel the slope under your feet.

So you do the sensible thing: feel which way is downhill, take a small step that way, and repeat. That's gradient descent. The gradient is simply the slope: for every weight, it says whether increasing that weight would raise or lower the loss, and by how much.

error (loss)weight valuestartminimum: best weighteach step = learning rate × slope

Gradient descent: measure the slope, step downhill, repeat. Steps shrink as the slope flattens near the bottom.

The size of each step is controlled by the learning rate. Too small, and training crawls. Too large, and you overshoot the valley and bounce around, or even climb out of it. Choosing a good learning rate is one of the most important practical skills in deep learning.

Backpropagation: sharing out the blame

Gradient descent needs the slope for every single weight, and a modern network can have billions of them. Computing each one separately would take forever. Backpropagation is the clever shortcut that makes it practical.

Think of a restaurant kitchen. A customer sends back a dish that's too salty. The complaint goes to the head chef, who works out which station added the salt, and that station figures out which step went wrong. Blame flows backward from the result to the causes, and each station learns how much it contributed.

Backpropagation does the same with the error. It starts at the output, where the mistake is known, and works backward layer by layer, using the chain rule from calculus to calculate how much each weight contributed. One backward sweep gives the gradient for every weight at once, at roughly the same cost as the forward pass. That efficiency, popularized by Rumelhart, Hinton and Williams in 1986, is why training deep networks is possible at all.

The complete training loop

Put it all together and training is a loop that repeats thousands or millions of times:

  1. Forward pass: feed a batch of examples through the network to get predictions.
  2. Compute the loss: compare predictions with the correct answers.
  3. Backward pass: use backpropagation to get the gradient for every weight.
  4. Update: nudge every weight a small step downhill, scaled by the learning rate.
  5. Repeat with the next batch. One full pass through the training data is called an epoch.

See it in code: a network that learns XOR

Here is a complete neural network written with nothing but NumPy. It learns XOR, a classic puzzle where the answer is 1 only when exactly one of two inputs is 1. A single neuron can't solve XOR, because no straight line separates the answers, but a network with one hidden layer can. Each numbered comment matches a step of the training loop above.

import numpy as np

# XOR: the output is 1 only when exactly one input is 1
X = np.array([[0, 0], [0, 1], [1, 0], [1, 1]])
y = np.array([[0], [1], [1], [0]])

rng = np.random.default_rng(42)
W1 = rng.normal(size=(2, 4))   # input -> hidden weights
b1 = np.zeros((1, 4))
W2 = rng.normal(size=(4, 1))   # hidden -> output weights
b2 = np.zeros((1, 1))

def sigmoid(z):
    return 1 / (1 + np.exp(-z))

learning_rate = 1.0
for epoch in range(5000):
    # 1. Forward pass: make predictions
    h = sigmoid(X @ W1 + b1)
    pred = sigmoid(h @ W2 + b2)

    # 2. Loss: how wrong are we? (mean squared error)
    loss = np.mean((pred - y) ** 2)

    # 3. Backward pass: how much did each weight contribute to the error?
    d_pred = (pred - y) * pred * (1 - pred)
    d_W2 = h.T @ d_pred
    d_h = (d_pred @ W2.T) * h * (1 - h)
    d_W1 = X.T @ d_h

    # 4. Update: nudge every weight a little bit downhill
    W2 -= learning_rate * d_W2
    b2 -= learning_rate * d_pred.sum(axis=0, keepdims=True)
    W1 -= learning_rate * d_W1
    b1 -= learning_rate * d_h.sum(axis=0, keepdims=True)

    if epoch % 1000 == 0:
        print(f"epoch {epoch:4d}  loss {loss:.4f}")

print(pred.round(2).ravel())

When you run it, the loss falls from about 0.28 to about 0.0003, and the final predictions come out close to the correct answers of 0, 1, 1, 0:

epoch    0  loss 0.2775
epoch 1000  loss 0.0025
epoch 2000  loss 0.0008
epoch 3000  loss 0.0004
epoch 4000  loss 0.0003
[0.01 0.98 0.99 0.02]

Try changing learning_rate to 0.01 or 10 and watch what happens to the loss. It's the fastest way to build intuition for everything above. Real projects use libraries such as PyTorch or TensorFlow, which compute the backward pass for you automatically, but under the hood they do exactly this.

What can go wrong

  • Overfitting: the network memorizes the training examples instead of learning general patterns, so it fails on new data. More data, simpler models and techniques such as dropout help.
  • A bad learning rate: too high and the loss jumps around or explodes; too low and training takes forever.
  • Vanishing gradients: in very deep networks the error signal can shrink to almost nothing before reaching the early layers, so they stop learning. ReLU activations and careful architecture design were key fixes.
  • Poor data: a network can only learn what its examples show. Biased or mislabeled data produces a confidently wrong model.

Key terms, decoded

Weight
A number that sets how strongly one neuron's input affects its output. Learning means adjusting weights.
Bias
An extra number added to a neuron's sum that shifts how easily it activates.
Activation function
The bend applied to a neuron's total, such as ReLU or sigmoid, that lets networks learn non-linear patterns.
Loss
A single number measuring how wrong the network's predictions are.
Gradient
The slope of the loss with respect to each weight: which way and how much to change it.
Learning rate
How big a step each update takes.
Backpropagation
The algorithm that sends the error backward through the network to compute every gradient efficiently.
Epoch
One complete pass through the training data.

A neural network is just neurons doing multiply-add-activate, arranged in layers, with weights tuned by walking downhill on the loss. Scale that idea up to billions of weights and trillions of words of text, and you get the large language models behind today's chatbots. That's exactly where we'll go next: how large language models turn your prompt into an answer.

Questions about any step? Leave a comment and I'll answer it, or suggest what you'd like explained next.

Comments