This is a CIFAR-10 Tiny CNN tutorial in C. By the end, you can build and train a small convolutional neural network for CIFAR-10 image classification: load the binary dataset, run 3 x 3 convolution, ReLU, 2 x 2 max pooling, fully connected logits, softmax, and one small training run.
This article is based on the local cifar10_tiny_cnn.c project. The goal is not to reach state-of-the-art CIFAR-10 accuracy. The goal is to make each part of a CNN visible as code you can inspect, compile, and modify.
If you are still learning the foundations, start with Neural Network Basics. If you have already read the handwritten digit softmax classifier in C, this article is the next step from linear classification into convolutional image models.
Correction, 2026-09-08: this revision separates the historical log from current downloads and corrects the training-accuracy protocol. The aim is to bind reported numbers to identifiable files and update times, not merely run a CNN.
1. What The CIFAR-10 Input Looks Like
Each CIFAR-10 image is a 32 x 32 color image with red, green, and blue channels. The input shape in the program is therefore 3 x 32 x 32. Each sample has one label from 10 classes.
The source fixes that shape with a few constants:
#define IMG_C 3
#define IMG_H 32
#define IMG_W 32
#define NCLASS 10
Each binary record contains one label byte and 3,072 pixel bytes, with separate channel blocks rather than interleaved RGB. The loader maps 0..255 values to approximately -1..1; it does not use per-channel mean standardization.
This is different from the handwritten digit softmax project. The digit project flattens a grayscale image into one vector. The CNN keeps the spatial layout first, so convolution filters can scan local regions and detect color, edge, and texture patterns.
2. Tiny CNN Architecture And Parameter Count
The current macro is NF=16. The introductory source comment still says eight filters, but the macro controls computation and weight layout. The implementation uses 16 filters of size 3 x 3. Because the convolution is valid convolution, a 32 x 32 input becomes 30 x 30 after convolution. Then 2 x 2 max pooling reduces it to 15 x 15.
#define NF 16
#define K 3
#define CONV_H 30
#define CONV_W 30
#define POOL_H 15
#define POOL_W 15
#define FEAT (NF * POOL_H * POOL_W)
The full path is:
- Input: 3 x 32 x 32
- Convolution: 16 filters, 3 x 3, output 16 x 30 x 30
- ReLU: clips negative activations to 0
- Max pooling: 2 x 2, output 16 x 15 x 15
- Fully connected layer: maps pooled features to 10 logits
- Softmax: converts logits into class probabilities
The model is intentionally small. The convolution layer has 16 * 3 * 3 * 3 + 16 = 448 parameters. The fully connected layer has 10 * 3600 + 10 = 36010 parameters. The total is 36458 parameters, which is useful for learning and debugging rather than competing with modern CIFAR-10 models.
3. Convolution Written Directly In C
The convolution layer is nested loops. Each filter scans every output location, then multiplies and accumulates over the 3 input channels and the 3 x 3 local window.
float sum = net->conv_b[f];
for (int c = 0; c < IMG_C; c++) {
for (int r = 0; r < K; r++) {
for (int t = 0; t < K; t++) {
sum += net->conv_w[f][c][r][t] *
get_pixel(s->x, c, i + r, j + t);
}
}
}
conv[f][i][j] = sum > 0.0f ? sum : 0.0f;
The kernel is not flipped: this is cross-correlation, commonly called convolution in neural networks. The last line also applies ReLU. This plain implementation is useful because it exposes the relationship between filters, channels, local windows, and output positions.
4. Max Pooling Keeps The Strongest Local Response
The pooling layer compresses every 2 x 2 window into one value by keeping the maximum activation. The code also records where the maximum came from, because backpropagation only sends the gradient back to that winning position.
float best = conv[f][base_i][base_j];
int best_idx = 0;
for (int di = 0; di < 2; di++) {
for (int dj = 0; dj < 2; dj++) {
float v = conv[f][base_i + di][base_j + dj];
if (v > best) {
best = v;
best_idx = di * 2 + dj;
}
}
}
Equal maxima retain the first scanned position because the comparison is strictly greater; backpropagation follows that selected position. Pooling reduces spatial size but does not guarantee identical outputs under arbitrary input translations.
5. Fully Connected Layer And Softmax
The pooled features are flattened into one vector and passed into a fully connected layer. Each class has its own weights, producing 10 logits.
for (int k = 0; k < NCLASS; k++) {
float z = net->fc_b[k];
for (int p = 0; p < FEAT; p++) {
z += net->fc_w[k][p] * feat[p];
}
logits[k] = z;
}
Softmax normalizes these raw scores into probabilities. Training uses cross-entropy loss, and prediction selects the class with the highest probability.
6. What Backpropagation Updates
This program does not hide backpropagation inside a framework. It explicitly does four things:
- subtracts the one-hot target vector from softmax probabilities: subtract one at the true class and leave the other class probabilities unchanged
- accumulates dfeat using pre-update fully connected weights, then updates those weights and biases
- sends gradients back through max pooling and ReLU
- updates every convolution filter from the matching input window
net->fc_w[k][p] -= LR * dz[k] * feat[p];
net->fc_b[k] -= LR * dz[k];
...
net->conv_w[f][c][r][t] -= LR * dw;
net->conv_b[f] -= LR * db;
This is the most valuable part of the project. CNN training is not magic: each layer computes its gradients and updates parameters along the negative gradient; the chosen learning rate does not guarantee that every update reduces loss.
7. Compile And Run Locally
The site publishes the source file, sample weights, sample predictions, and explanation notes. It does not publish the full CIFAR-10 dataset. Download the binary version from the official CIFAR page (published MD5 c32a1d4ab5d03f1284b67883e8d87530), extract cifar-10-batches-bin, and run the program against that folder.
gcc -O2 -std=c11 cifar10_tiny_cnn.c -lm -o cifar10_tiny_cnn
./cifar10_tiny_cnn ./cifar-10-batches-bin 1 2000 1000
The four command-line arguments are:
- data directory
- number of epochs
- training sample limit
- test sample limit
Start with a small sample limit first. After the full flow works, increase the training sample count and epoch count gradually.
8. Check Whether the Log and Downloads Belong to the Same Run
The earlier article showed a log for 2,000 training images, 1,000 test images, and one epoch, reporting train_acc=0.835 and test_acc=0.284. As checked on 2026-09-08, however, the actual public prediction CSV contains 5,000 rows, with 2,351 correct predictions: 47.02%. That file must not be presented as the output of the quoted 1,000-test-image run.
| Evidence | What can be established |
|---|---|
| Historical log text | The old article recorded a 2,000/1,000, one-epoch configuration. A complete run manifest bound to that text was not found, so this audit does not claim to reproduce its 28.4%. |
| Current CSV | IDs run consecutively from 0 through 4,999; predictions and labels are in 0 through 9; 2,351/5,000=47.02%. This is a deterministic recount of the public file. |
| Current weights | 145,832 bytes interpreted as 36,458 finite float32 values under the matched ABI. The size matches the parameter count but stores no training sample count, epoch count, or checkpoint-selection history. |
| Artifact identity | Source, weights, and CSV fetched from their ordinary public URLs match the lab’s legacy copies byte for byte. Full SHA-256 hashes are recorded in audit.json. |
Inspect the historical log text retained from the old article
Loaded train=2000, test=1000
epoch 1 step 1000/2000 loss=2.0507 train_acc_recent_est=0.809
epoch 1 step 2000/2000 loss=1.9440 train_acc_recent_est=0.835
epoch 1 done: avg_loss=1.9440 train_acc=0.835 test_acc=0.284
Model saved to model_weights.bin
Predictions saved to test_predictions.csv
This text is retained to explain the discrepancy, not to label an unbound historical record as a new result. Even successfully matching a checkpoint to predictions cannot recover its entire training history.
9. Why 83.5% Is Not Final-Checkpoint Training Accuracy
The decisive detail is the order of operations in the original loop:
sum_loss += train_one(&net, s);
int pred = predict(&net, s);
if (pred == s->label) correct++;
The forward pass inside train_one records loss before updating the parameters. The subsequent predict uses the updated parameters on the very same sample. Loss and accuracy therefore refer to different times, and successive accuracy observations use different model states. Despite its name, train_acc_recent_est is also cumulative from the start of the epoch, not a recent sliding window.
| Protocol | Interpretation |
|---|---|
| Before update | Predict before learning from the current sample. This describes a sequence of changing models, still not the training accuracy of the final checkpoint. |
| After update | The original train_acc. The current label has just influenced the model, so the score can look optimistic. Comparing it directly with final-checkpoint test_acc does not establish a generalization gap. |
| Final model | Finish all updates, freeze one parameter state, then evaluate the training and test sets in separate passes. Model state is now comparable; data selection and tuning history still need documentation. |
A controlled counterexample makes this concrete. Set all weights and the input to zero and use label 7. Initially every logit is equal, so the original first-maximum rule predicts 0. After one update, the class-7 bias is approximately 0.009 and the other biases are -0.001, making that same sample correct immediately. Initial loss is log(10)=2.302585. A one-sample post-update score of 100% does not show that the model has learned image classification.
This is a synthetic metric-timing check, not a CIFAR-10 score. The problem requires a corrected experimental interpretation, not a more attractive label for 83.5%.
10. What Aggregate Accuracy Hides

For example, automobile recall is 309/505=61.19%, while cat recall is 138/497=27.77%. Of the 497 cat records, 107 are predicted as dog. Those rows are a useful starting point for input inspection, but the counts alone do not identify background, texture similarity, or a particular filter as the cause.
Deer illustrates a different denominator: 254 correct out of 507 true deer gives recall 50.10%, whereas 722 total deer predictions give precision only 35.18%. Precision and recall answer different questions. One aggregate accuracy cannot replace class-level inspection.
Per-class CSV records support, predicted count, correct count, recall, and precision. The confusion CSV uses true labels as rows and predicted labels as columns. Every count can be recomputed from the original prediction file without training a model.
11. Check the Numerical Meaning of the Logged Loss
The source subtracts the largest logit before softmax, which is a stabilizing step. It then reports -logf(prob[y] + 1e-8f). When the true-class probability is extremely small, this addition limits the displayed loss to approximately 18.42 instead of reflecting increasingly wrong predictions.
A separate constructed boundary test sets all parameters to zero except a class-0 bias of 100 and uses target 1. The original returned loss is approximately 18.420681; cross-entropy computed directly from the logits is 100. This exposes clipping in the report, not evidence that the public model encountered the same condition on real inputs.
m = max(logits)
cross_entropy = m + log(sum(exp(logits - m))) - logits[y]
A future training-code revision should align logged loss with the gradient objective and check finite values. This audit preserves the original source rather than silently replacing it; the independent test wrapper computes the stable reference value.
12. Reproduce the Audit and Respect Its Limits
Download the audit package. It contains hash-locked copies of the three original artifacts, a C test wrapper, Python standard-library checks, class statistics, and plotting code. It does not contain the full CIFAR-10 dataset or precompiled executables.
The Run notes above still refer to the original cifar10-cnn project. Run the following commands inside the new cifar-audit-lab package instead of mixing the two working directories:
python3 audit_cifar.py --output reproduced
# Optional: official binary archive with the published checksum
python3 audit_cifar.py --output reproduced-full \
--archive /path/to/cifar-10-binary.tar.gz
The first command requires Python 3.11+ and a C11 compiler. It checks the CSV, weight payload, and two synthetic counterexamples, and rejects eight malformed CSV cases and five invalid sample counts. Its success marker is CIFAR_ARTIFACT_AUDIT_OK, which does not imply official-data replay was performed.
The second command validates the complete official archive’s MD5 before extracting the required batches. It compares the archived labels with the first 5,000 official test records and generates predictions from the public weights. Only complete row agreement produces CIFAR_DATA_REPLAY_OK; differences are reported explicitly. Consult audit.json for actual results and environment. Matching file sizes cannot substitute for this check.
Full mode also runs a separate 128-training/128-test, one-epoch experiment, recording pre-update, post-update, and frozen-final-checkpoint scores, then repeats it on the same host. It does not update the public weights or reconstruct the historical 2,000/1,000 log. srand(1) does not guarantee identical random sequences across different C standard libraries.
Measured Replay Results, 2026-09-08
These results were measured on arm64 macOS with Apple clang 21, Python 3.13.9, and NumPy 2.3.5. Compilation disables fast-math and floating-point contraction. The package records the full environment, file hashes, and unrounded results.
| Check | Observed result and scope |
|---|---|
| Public checkpoint | All 5,000 labels match the official test prefix. Replaying the public weights reproduces all 5,000 archived predictions: 2,351/5,000=47.02%. Mean logit cross-entropy is 2.035592. This is frozen-checkpoint inference, not a new training run. |
| NumPy forward | NumPy float64 independently computes 160 logits for the first 16 images. Maximum absolute difference from C float32 is 9.07e-06; all 16 predicted classes agree. Declared tolerances are atol=1e-3 and rtol=2e-4. This checks forward calculation, not all training gradients. |
| Layout controls | Deliberately interpreting planar channels as interleaved RGB gives maximum logit error 32.393047. Deliberately reshaping pooling windows incorrectly gives error 13.720984. Both exceed the predefined detection threshold of 0.01. |
| New run | One epoch on the first 128 records of the first training batch gives online pre-update score 19/128 and same-sample post-update score 92/128. The frozen final model scores 88/128 on those training records and 17/128 on the first 128 test records. Counts and step traces repeat exactly on the same host. These are not the published checkpoint’s training scores or the old log’s results. |
The independent forward check needs NumPy. From the extracted audit package, run:
python3 -m venv .venv
.venv/bin/python -m pip install -r requirements-numpy.txt
.venv/bin/python audit_cifar.py --output reproduced-numpy \
--archive /path/to/cifar-10-binary.tar.gz --numpy-reference
The step trace CSV retains source IDs, pre-update predictions, post-update predictions, and pre-update loss. The probe JSON records the final evaluation counts and denominators. Agreement establishes correspondence among these public artifacts; it does not establish an untouched test set or recover missing training history.
The weight file is a raw C struct with no header, architecture version, or byte-order declaration. The audit matches a little-endian, four-byte-float layout; it is not a portable model format. The original dataset reader also lacks comprehensive malformed-file validation and should not be exposed as an arbitrary-upload service.
13. Improve the Experiment Record Before the Model
- Record source, dataset, and checkpoint hashes together with compiler flags, sample ranges, seed, and hyperparameters. A .bin file alone cannot reproduce a training procedure.
- Evaluate both sets using the same frozen final checkpoint. Name online pre-update and post-update metrics separately.
- Create a validation split within the training data before comparing rates, epochs, or augmentation. Change one main variable at a time; repeatedly inspected public test output is not an untouched validation source.
- Report class-level statistics and failure cases with explicit dataset coverage. These 5,000 rows do not automatically represent all 10,000 test images.
- Before adding mini-batches or deeper layers, perform small gradient checks and input validation. The artifact, counting, and counterexample checks here do not establish full backpropagation correctness or hardware speed.
More data, additional layers, and a framework port are hypotheses to test, not guaranteed accuracy improvements. The public checkpoint’s complete recipe and selection history are unavailable, so its 47.02% alone cannot establish why particular predictions fail.
14. Original Resources and Further Reading
The original C source, model weights, predictions, and historical explanation PDF remain available. Actual parameter counts here are tied to macros, compiled layout, and weight payload rather than the introductory comment or filename.
Continue with the reproducible convolution and receptive-field experiment for local operations and the two-layer MLP backpropagation article for gradient checking. This project’s value is an inspectable link between inputs, update timing, artifacts, and interpretation, not a high-performance label for an educational implementation.