Neural networks are often presented as complicated systems, but the entry-level view can be simple: a neural network is a trainable composition of functions. Each layer transforms its input, and multiple layers together can represent more complex relationships.
This article starts with a single neuron and explains weights, bias, activation functions, forward propagation, and the intuition behind backpropagation. The goal is not to derive every formula, but to make neural network training code easier to read.
While reading, keep one main loop in mind: the network predicts with current parameters, measures loss, then updates parameters in the direction that reduces loss.
1. Start With One Neuron
A simple neuron can be written as:
z = w1 * x1 + w2 * x2 + ... + b
output = activation(z)
The parts are:
x: input featuresw: weightsb: biasactivation: an activation function
Without activation functions, multiple linear layers can still be collapsed into one linear transformation. Activation functions give the network nonlinear expressive power.
2. What a Perceptron Can Do
A perceptron can be viewed as an early simple neural network. It computes a weighted sum of inputs, then applies a threshold to produce a class label.
if w1 * x1 + w2 * x2 + b > 0:
predict 1
else:
predict 0
This can solve linearly separable problems, where classes can be separated by a line, plane, or higher-dimensional hyperplane.
Real data often contains nonlinear relationships, so we need multi-layer networks and nonlinear activation functions.
3. What Is a Layer?
A layer sends a group of inputs through multiple neurons and returns a group of outputs. Common layer roles include:
- Input layer: receives raw features
- Hidden layer: performs intermediate transformations
- Output layer: returns class probabilities or numeric predictions
A small multi-layer network can be represented as:
input features -> hidden layer 1 -> hidden layer 2 -> output layer
Parameterised layers such as a dense layer have weights and biases. The input layer is a data entry point, and an activation such as ReLU has no trainable parameters. Training updates the enabled parameters that participate in the loss.
4. Forward Propagation
Forward propagation means computing from input to output, layer by layer.
x -> layer1 -> activation -> layer2 -> activation -> output
In code, this usually corresponds to a model’s forward function. It answers:
Given the current parameters and a batch of input, what does the model predict?
Both training and inference use forward propagation. During training, the prediction is also used to compute loss and update parameters.
5. The Intuition Behind Backpropagation
Backpropagation calculates how each parameter affects the loss. Intuitively, it asks:
If this weight became slightly larger or smaller, how would the final loss change?
With that information, an optimizer can update parameters in a direction that reduces loss.
prediction -> compute loss -> backpropagate gradients -> update parameters
You do not need to hand-write backpropagation at the beginning. Frameworks such as PyTorch and TensorFlow compute gradients automatically. But you should understand why training code contains steps such as loss.backward() and optimizer.step().
6. A Typical Training Loop
In pseudocode, neural network training often looks like this:
for epoch in range(num_epochs):
for X_batch, y_batch in train_loader:
y_pred = model(X_batch)
loss = loss_fn(y_pred, y_batch)
optimizer.zero_grad()
loss.backward()
optimizer.step()
The loop can be read as five steps:
- Take a batch of training data
- Run forward propagation to get predictions
- Compute loss against the true labels
- Backpropagate gradients
- Let the optimizer update parameters
7. Why Deep Learning Needs More Data and Compute
Neural networks can express complex patterns, but the cost is real:
- They have many parameters and can overfit
- They usually need more data
- They have a larger tuning space
- Training speed depends more heavily on hardware
This is why it is useful to learn the traditional machine learning workflow first. Once data, features, training, and evaluation are clear, neural networks become easier to reason about.
8. Neural Networks and Large Models
Large language models, image generation systems, and speech recognition systems are deep learning systems. They use more complex architectures, larger datasets, and longer training processes.
Even when the model is large, the foundation questions remain similar:
- How is input represented as numbers?
- How does the model transform input into output?
- How does the loss function measure prediction error?
- How does training update parameters?
- Does the evaluation method reflect real use?
Learning neural network basics is not only about training a network immediately. It gives you the shared language behind modern AI systems.
9. Common Beginner Misunderstandings
When first learning neural networks, these misunderstandings are common:
- Assuming more layers are always better while ignoring data size, overfitting, and training cost
- Treating activation functions as minor details instead of understanding their nonlinear role
- Focusing only on architecture while ignoring the loss function and evaluation metrics
- Assuming the model is reliable just because training loss goes down
Neural networks are powerful because of their expressive capacity, but reliability still depends on data splits, evaluation, and error analysis.
10. Neural Network Training Evidence Checklist
A beginner neural network experiment should leave behind enough evidence for someone else to reproduce the result and identify failure modes. The checklist below connects the concepts in this article to practical training records.
| Evidence item | What to record | Why it matters | Failure signal |
|---|---|---|---|
| Input shape | Batch size, feature count, tensor layout, and normalization range | Most silent neural network bugs are shape or scale mistakes | Loss changes when only the batch dimension or image channel order changes |
| Loss curve | Training loss, validation loss, and learning rate per epoch | The curve shows underfitting, overfitting, or optimizer instability | Training loss falls while validation loss rises for many epochs |
| Gradient health | Gradient norm, exploding or vanishing activations, and optimizer step size | Backpropagation can fail even when the code has no syntax error | Weights become NaN, gradients collapse to zero, or updates oscillate wildly |
| Error analysis | Confusion matrix, hard examples, and examples outside the training distribution | Aggregate accuracy hides systematic mistakes | The model is strong on common classes but unreliable on rare or shifted inputs |
Work one forward and backward pass by hand
No amount of description substitutes for computing the smallest example once. The network below has two inputs, one hidden neuron and one output, with squared error as the loss. Every number can be checked mentally.
input x = [1.0, 2.0] target y = 1.0
hidden w = [0.3, -0.1], b = 0.2 ReLU activation
output v = 0.5, c = 0.1 no activation
Forward:
z = 0.3×1.0 + (-0.1)×2.0 + 0.2 = 0.3
a = ReLU(0.3) = 0.3
ŷ = 0.5×0.3 + 0.1 = 0.25
L = (ŷ - y)² = (0.25 - 1.0)² = 0.5625
Backward: start at the loss and multiply derivatives back layer by layer.
∂L/∂ŷ = 2(ŷ - y) = 2 × (-0.75) = -1.5
∂L/∂v = ∂L/∂ŷ × a = -1.5 × 0.3 = -0.45
∂L/∂c = ∂L/∂ŷ × 1 = -1.5
∂L/∂a = ∂L/∂ŷ × v = -1.5 × 0.5 = -0.75
∂L/∂z = ∂L/∂a × ReLU'(0.3) = -0.75 × 1 = -0.75 ← z>0, so the derivative is 1
∂L/∂w₁ = ∂L/∂z × x₁ = -0.75 × 1.0 = -0.75
∂L/∂w₂ = ∂L/∂z × x₂ = -0.75 × 2.0 = -1.5
∂L/∂b = ∂L/∂z × 1 = -0.75
At the current parameter point, these negative derivatives indicate local descent for small increases along each coordinate. They do not guarantee improvement for any step size. A simultaneous update at learning rate 0.1 gives [w1,w2,b,v,c]=[0.375,0.05,0.275,0.545,0.25], followed by prediction=0.65875 and loss=0.1164515625.
All gradients use the same pre-update parameter snapshot. Updating the output weight and then reusing its new value in this backward pass would no longer compute the gradient of the original forward pass. Changing only w1 to 0.375 gives prediction 0.2875. The simultaneous update improves more in this example, not as a general rule.
For this one sample, dL/dw_j=(dL/dz) x_j: both weights share the upstream derivative, so doubling the input doubles that component. Batch gradients also sum or average over samples and can cancel; a larger feature does not guarantee a larger gradient in every network. Feature scaling can improve the scale of the optimisation problem.
Setting hidden bias b=-0.2 gives z=-0.1 and a zero ReLU output. The w1, w2 and b gradients are blocked, and the output weight v has zero gradient because its input is zero. But dL/dc=-1.8 still updates the output bias. Inactivity on one sample is not permanent neuron death; check other samples and later steps.
What actually goes wrong with zero initialisation
Zero is not a number that cannot learn. The problem concerns symmetry between hidden units and activation derivatives. In this dense MLP, two units with identical incoming weights, biases and outgoing connections receive identical activations and gradients under identical deterministic updates. They remain unable to specialise. Random initialisation is a common way to break that symmetry, not the only possible way.
The package checks three identical hidden units over 10 updates and confirms they remain identical. A separate all-zero ReLU case explicitly chooses ReLU'(0)=0: hidden gradients and output-weight gradients are zero, but four targets [0,0,0,1] give output-bias gradient [-0.25,0.25]. One update reduces mean cross-entropy from 0.6931471806 to 0.6809596480. It adjusts class priors without learning an input feature.
A linear softmax classifier without hidden layers can start at zero weights and still have nonzero data-dependent weight gradients. Biases can usually start at zero once weights break hidden-unit symmetry. Xavier and He initialisation also account for layer width and activation-related scale; neither guarantees successful training for every network.
Vanishing gradients, floating point and residual connections
A scalar chain multiplies local derivatives; branched graphs also add contributions from different paths, and vector networks use Jacobian products. Multiplying by 0.5 fifty times gives 2^-50 = 8.881784197e-16. This is a positive representable float32 value, not an underflow to zero.
Representing a gradient and changing a parameter are different questions. In the float32 experiment, 1.0 - 0.1 * 2^-50 still rounds to 1.0: the update is far smaller than the spacing near that parameter. The smallest normal positive float32 is about 1.175e-38, while the gap immediately above 1.0 is about 1.192e-7. Dynamic range and relative precision are not interchangeable.
- Activations. Sigmoid’s derivative is at most 0.25, but a layer derivative also contains its weight matrix. This does not prove fourfold attenuation for the entire layer. ReLU’s positive-side derivative of 1 does not eliminate weight-product effects or its inactive region.
- Normalisation. It can change numerical scale and optimisation conditions, without ensuring gradient norms near 1. BatchNorm and LayerNorm also differ in their normalisation axes and training/inference behaviour.
- Residuals. The scalar derivative of
y=x+f(x)is1+f'(x); the vector form isI+J_f. The identity path can help, but f(x)=-x makes the total derivative zero. A residual is not a guarantee of undiminished gradients at arbitrary depth.
Inspect activations, gradients and actual parameter updates alongside loss and validation results. Nonzero gradients do not prove a correct implementation, and decreasing loss does not establish generalisation.
Run the original example: learning rate and local gradients
These are executed one-step updates of the five parameters above. Each row restarts at the same initial point; they are not three consecutive steps. This checks an update rule, not training accuracy.
| Learning rate | Prediction after | Squared error after |
|---|---|---|
| 0.01 | 0.2890525 | 0.5054463478 |
| 0.1 | 0.65875 | 0.1164515625 |
| 1.0 | 6.16 | 26.6256 |
The initial loss is 0.5625. Learning rate 1.0 follows the negative gradient but overshoots the local descent region. Direction alone is not enough; step size matters.
Show the runnable one-step Python example
"""The article's five-parameter ReLU example; gradients use one snapshot."""
def forward_and_grad(params):
w1, w2, b, v, c = params
z = w1 + 2 * w2 + b
a = max(0.0, z)
prediction = v * a + c
loss = (prediction - 1.0) ** 2
d_prediction = 2 * (prediction - 1.0)
dz = d_prediction * v * (z > 0)
grad = [dz, 2 * dz, dz, d_prediction * a, d_prediction]
return prediction, loss, grad
def one_step(learning_rate):
before = [0.3, -0.1, 0.2, 0.5, 0.1]
prediction, loss, grad = forward_and_grad(before)
after = [p - learning_rate * g for p, g in zip(before, grad)]
next_prediction, next_loss, _ = forward_and_grad(after)
return {"learning_rate": learning_rate, "before": before, "grad": grad,
"after": after, "prediction_before": prediction, "loss_before": loss,
"prediction_after": next_prediction, "loss_after": next_loss}
if __name__ == "__main__":
for rate in (0.01, 0.1, 1.0):
row = one_step(rate)
print(f"lr={rate:g}: prediction={row['prediction_after']:.8f}, loss={row['loss_after']:.10f}")
Download the neural gradient lab. The scalar example above needs only Python. The extended audit uses NumPy 2.3.5 and includes all inputs, per-coordinate difference CSVs for 1029 parameters, initialisation counterexamples and reference JSON. Run python audit_gradients.py --output run-results and expect NEURAL_GRADIENT_AUDIT_PASSED. There is no training-data download, GPU test or generalisation benchmark.
Continue with the batched backpropagation experiment. See NumPy finfo for precision/range definitions and PyTorch autograd notes for derivative choices at nondifferentiable points.
11. What to Read Next
The previous article is Model Training and Evaluation. To connect the whole series in one runnable exercise, continue with Python AI Mini Practice.