A handwritten-digit project does not have to start with a training loop. Here a 28 by 28 grayscale image is stored as one CSV row: 784 pixels go in, and a digit from 0 to 9 comes out. A misplaced field or a mixture of project versions can make otherwise plausible loss and accuracy logs answer the wrong question. This walkthrough checks the actual downloadable files before connecting them to the C classifier and browser demo.
Artifact check: 2026-09-08. This work ran complete row checks, checksum comparisons and small numerical probes. It did not retrain a model or obtain a new test accuracy. The original source, data and model downloads remain unchanged.
1. Establish the file and column contract
train.csv: label,pixel0,pixel1,...,pixel783 785 columns
test.csv: pixel0,pixel1,...,pixel783 784 columns
Image coordinate (r, c) maps to pixel[28*r + c], using zero-based indices
Normalized input x[j] = pixel[j] / 255.0
Column zero is a label in training data but already a pixel in the test file. A supervised classification loss needs a true label; this unlabeled test file can produce predictions, but cannot directly provide accuracy. For tuning, first freeze a training/validation split of the labeled data and retain its row IDs. Predictions in a submission are not ground-truth labels.
sample_submission.csv is an output-format example; submission.csv is an existing prediction artifact. Both should have the header ImageId,Label and consecutive IDs from 1 through the test count. A valid format is not evidence of correct predictions, and a filename does not identify the training run that produced it.
2. Results from the actual files
The checker examined the complete header, exact field count, label range and integer range of every pixel. These are measured counts from the current materials, not expected numbers copied from a competition description. Row counts exclude the header.
| Check | Train | Test |
|---|---|---|
| Data rows | 42,000 | 28,000 |
| Fields per row | 785 | 784 |
| Raw pixel range | 0 to 255 | 0 to 255 |
| Total pixels | 32,928,000 | 21,952,000 |
| Zero pixels | 26,621,312 | 17,753,313 |
| Nonzero fraction | 19.152964% | 19.126672% |
The counts for labels 0 through 9 are [4132, 4684, 4177, 4351, 4072, 3795, 4137, 4401, 4063, 4188]. Every class is represented, but the classes are not perfectly balanced. Always predicting the most frequent class, 1, would already give about 11.15% on these training rows. Merely exceeding 10% is therefore a weak success criterion.

Both existing submission files pass the 28,000-row, sequential-ID and label-range checks. Structural validity does not establish absence of duplicates, overlap between training and validation, writer overlap or annotation errors. Those independence and labeling checks were not performed here.
3. Parsing a CSV is not validating it
The preserved C loader uses atoi for labels and atof for pixels before dividing by 255. This excerpt shows the mapping; it is not a production input validator.
y_train[sample_count] = atoi(tokens[0]);
for (int j = 0; j < FEATURES; j++) {
X_train[sample_count][j] = atof(tokens[j + 1]) / 255.0;
}
Its strtok helper collapses consecutive delimiters. Calling the original helper on 1,,2 returns two tokens, losing the empty middle field. With a capacity of three, 1,2,3,4 returns three tokens: the count does not prove the original row had only three fields. Some malformed rows are skipped by the loader, but label bounds and numeric validity are not fully checked. Do not feed untrusted CSV files directly into this training program.
The companion checker uses standard-library csv.reader to preserve field structure, then validates the schema explicitly. Eleven fixed fixtures are rejected: empty file, header only, wrong header, missing field, extra field, label 10, empty pixel, nan, pixel 256, negative pixel and fractional pixel. Accepted integer text follows the canonical spelling used by this dataset, such as 0 through 255; it does not silently repair 0.0. CSV parsing alone does not impose these domain constraints, as the Python csv documentation explains.
4. Reconstruct a row and inspect the indexing
Run the following program inside the companion package with python3 preview_first.py ../digit-assets/train.csv.zip. It imports the package’s validate_row, reads the first data row explicitly and never relies on a variable left by an earlier loop. It does not unpack or modify the dataset.
Complete first-row text preview
"""Print one validated training row without unpacking or changing the data."""
import csv
import io
import sys
import zipfile
from audit_digits import PIXEL_HEADER, validate_row
with zipfile.ZipFile(sys.argv[1]) as archive:
with archive.open("train.csv") as raw:
with io.TextIOWrapper(raw, encoding="utf-8", newline="") as stream:
reader = csv.reader(stream, strict=True)
if next(reader, None) != ["label"] + PIXEL_HEADER:
raise ValueError("unexpected training header")
row = next(reader, None)
if row is None:
raise ValueError("no training rows")
label, pixels = validate_row(row, True, 2)
print("label:", label)
for r in range(28):
print("".join(".:*#"[pixels[28 * r + c] * 3 // 255]
for c in range(28)))
The inspected file starts with label 1; the program prints 28 lines of 28 characters. A coarse text picture cannot replace field checks. In particular, treating row[1] in the test file as pixel0 is not necessarily a clean one-pixel image translation. It can lose an element or cross a row boundary in the flattened representation; whether it raises an error depends on the reader. Verify the header, length and 28*r+c correspondence directly.
5. Scaling pixels is not one learning-rate adjustment
Changing the input also changes logits, probabilities and gradients. In a fixed two-class example, take z=[0.01*x,0] with true class 0. Then dW0=(p0-1)*x and db0=p0-1. At x=1, the measured values are p0=0.502500 and dW0=-0.497500. At x=255, they become p0=0.927574 and dW0=-18.468754. The weight gradient is not simply 255 times larger, and the bias gradient changes differently.
Keep training and inference scaling consistent, then inspect optimization behavior instead of claiming that unnormalized input just multiplies the learning rate by 255. Division by 255 is this project’s input contract, not the only preprocessing suitable for every image model. Any additional mean or standard deviation must be fitted on the training portion and then applied to validation and test inputs.
6. Separate the two same-named project versions
The standalone digit_softmax_classifier.c is now a 784 to 128 ReLU to 10 Softmax MLP. The same-named source inside the preserved handwritten-digits-materials.zip is a 784 to 10 linear Softmax model. Their training/test ZIPs and existing submission files are byte-identical, but their source and model JSON are not. The old archive remains useful for studying the linear baseline; mixing some standalone files into it does not produce a coherent historical experiment.
The first 12 SHA-256 characters are eddd081ee163 for the standalone source and 88b8953bac9f for the archived source. Full hashes, archive-member comparisons and raw counts are in the audit JSON. A checksum identifies bytes, not training lineage or data authorization.
7. Reproduce and continue
Download the artifact and numerical audit package and follow its bilingual README to obtain the existing inputs. The data checker needs only the Python 3.10+ standard library; this run used Python 3.13.9. Plotting dependencies are optional. The package contains no replacement model weights or dataset copies.
python3 audit_digits.py --assets ../digit-assets --out ../digit-results
python3 preview_first.py ../digit-assets/train.csv.zip
Continue with the two C versions, gradients and losses, or the browser model and probability interpretation. Original materials remain on the downloads page; the interactive module is in the playground.