This article explains the C handwritten-digit project, but there are two same-named versions to distinguish first. The preserved archive contains linear Softmax regression; the standalone download is now an MLP with 128 ReLU hidden units. Both use Softmax for ten-class output, so the existing article title and URL remain. Their internal structure, learning rate and reproducibility limits must nevertheless be described separately.
Checked on 2026-09-08: the public download bytes were compared, both source files were read, and the unchanged standalone source was compiled for bounded helper-function probes. Full training on 42,000 samples was not rerun. No new accuracy is reported, and no old source, model or submission was overwritten.
1. Identify the C program you downloaded
The standalone source and the file inside the old project archive share the name digit_softmax_classifier.c, not the contents. Hashes below show the first 12 SHA-256 characters; full values are in the artifact record.
| Version | Layers | SGD |
|---|---|---|
| Archive | 784 to 10 | 20 epochs Rate 0.01 |
| Direct | 784 to 128 to 10 |
20 epochs Rate 0.005 |
The archived source has hash prefix 88b8953bac9f and 7,850 parameters; the standalone source has prefix eddd081ee163 and 101,770 parameters. Both use Softmax outputs, while the standalone source also has a ReLU hidden layer.
Both programs use single-example SGD in fixed data order, without a separate validation set, per-epoch shuffling or mini-batches. The standalone MLP defines W1[128][784], b1[128], W2[10][128] and b2[10]. Its parameter count is 128*784+128+10*128+10=101770. Explaining this whole file as one W[10][784] matrix is no longer correct.
2. Match the input, hidden layer and output layer
The archived linear version directly computes a score for each class. This fragment corresponds to that path:
z[k] = b[k];
for (int j = 0; j < FEATURES; j++) {
z[k] += W[k][j] * x[j];
}
The standalone version computes a hidden representation before its output scores. With column vectors, its forward pass is:
x: 784 normalized pixels
z1 = W1*x + b1 shape 128
a1 = max(0, z1) ReLU, elementwise
z2 = W2*a1 + b2 shape 10
p = softmax(z2) shape 10
Both paths scale pixels by /255.0. Before Softmax, z2 contains logits; afterward, p contains scores normalized across ten classes. Prediction chooses the largest probability. There is no “not a digit” class and no automatic rejection of a blank board. A sum near one validates normalization, not classification correctness.
3. Why backpropagation needs the old weights
For the linear version, the single-example cross-entropy signal is p[k]-onehot(y)[k]. The source updates weights directly from pixels:
double error = p[k] - (k == y_train[i] ? 1.0 : 0.0);
for (int j = 0; j < FEATURES; j++) {
W[k][j] -= LEARNING_RATE * error * X_train[i][j];
}
b[k] -= LEARNING_RATE * error;
As an explicitly hypothetical example, take true label 7, p[3]=0.62, p[7]=0.21 and one pixel x[j]=0.4. The corresponding gradients are 0.248 and -0.316. At the linear version’s learning rate of 0.01, the weight changes are -0.00248 and +0.00316. These are teaching inputs, not an observed training record.
The MLP output layer uses hidden activations a1 instead of raw pixels. Its backpropagation is:
dz2 = p - onehot(y)
dW2 = dz2 * a1^T
db2 = dz2
dz1 = (W2^T * dz2) * (z1 > 0)
dW1 = dz1 * x^T
db1 = dz1
Computing dz1 requires W2 from the same forward pass. The standalone code calculates da1 and dz1 before updating W2 and W1, which preserves that ordering. Moving the W2 update earlier would propagate a gradient through different parameters. At zero, the source’s z1 > 0 convention selects a zero ReLU derivative.
4. Stable probabilities do not guarantee faithful loss logs
Both sources subtract the maximum logit before exponentiation. For finite inputs, this avoids positive exponential overflow; tiny terms can still underflow to zero. It does not automatically repair NaN or infinite inputs. A suddenly invalid loss also does not rule out an excessive learning rate, malformed data, indexing errors or already-corrupted parameters.
The probe harness called the original standalone functions without replacing them. It compiled without warnings using Apple Clang 21.0.0 and -std=c11 -Wall -Wextra -O2. Results:
| Input or check | Result | Meaning |
|---|---|---|
[1000,-1000] |
Probabilities [1,0] |
The smaller exponential underflows |
| True class 1; source log formula | Loss 34.538776 |
-log(p[y]+1e-15) hides a larger loss |
| Same input; logsumexp form | Loss 2000 |
Preserves the wrong-class logit gap |
[+inf,0] |
Output is nonfinite | Validate finite inputs first |
The log uses an epsilon-added loss, while updates use p-onehot. That signal is the standard cross-entropy gradient with respect to logits, not the exact derivative of the epsilon-added expression across all numerical ranges. Under extreme inputs, the printed curve is therefore not a faithful standard cross-entropy record.
The following separate teaching program returns 2000 on the same fixed input; it does not patch the original C source. It follows the numerical idea behind logsumexp, but uses only the Python standard library, not SciPy.
Complete stable cross-entropy example
"""A finite-logit cross-entropy example; no model training or dependencies."""
import math
def cross_entropy(logits, label):
if not logits or type(label) is not int or not 0 <= label < len(logits):
raise ValueError("nonempty logits and a valid integer label are required")
if any(not math.isfinite(z) for z in logits):
raise ValueError("logits must be finite")
maximum = max(logits)
shifted_sum = sum(math.exp(z - maximum) for z in logits)
return (maximum - logits[label]) + math.log(shifted_sum)
if __name__ == "__main__":
loss = cross_entropy([1000.0, -1000.0], 1)
assert loss == 2000.0
print("logsumexp cross-entropy:", loss)
Rejecting nonfinite logits is not a promise of a finite loss for every extreme floating-point range. If the difference between the maximum and the true-class score itself exceeds the representable range, the result can still be infinite. This is a specific counterexample check, not comprehensive numerical certification.
5. Training logs are not validation results
Epoch logs accumulate predictions made before each sample update while parameters keep changing. They are not metrics of one frozen end-of-epoch model on the whole dataset. The final evaluate_model does revisit training data with fixed final weights to print training accuracy and a confusion matrix. It still does not measure validation or unlabeled-test accuracy.
A confusion matrix reveals label confusions, but cannot by itself establish that model capacity has been exhausted. Check feature mapping, labels, distribution differences and optimization before comparing architectures. A generalization claim needs frozen validation row IDs, training settings, preprocessing, random state and a model artifact, with per-example predictions retained. Missing historical records are not filled in here as completed experiments.
6. Use a separate run directory
After preparing digit-assets as described in the package README, run the read-only audit and bounded helper checks:
python3 audit_digits.py --assets ../digit-assets --out ../digit-results
python3 run_c_probe.py --source ../digit-assets/digit_softmax_classifier.c --out ../digit-results
python3 stable_ce.py
run_c_probe.py verifies the complete source hash and compiles an alternative entry point. It tests Softmax and CSV tokenization, never calling the training main. The commands below are the full-training path in a new directory; that training was not run in this audit:
set -e
mkdir ../digit-mlp-run
cd ../digit-mlp-run
unzip ../digit-assets/train.csv.zip
unzip ../digit-assets/test.csv.zip
cc -std=c11 -Wall -Wextra -O2 ../digit-assets/digit_softmax_classifier.c -lm -o digit_classifier
./digit_classifier
The program looks for data in the working directory and writes or overwrites submission.csv there. Its static training/test pixel arrays alone occupy about 419 MiB as doubles. Full training and the small helper probes have very different resource needs. A fixed srand(42) also does not guarantee identical random sequences across C runtime implementations.
7. Matching architecture is not matching provenance
The standalone browser JSON has the same MLP shapes as the standalone source, but this C file has no JSON-weight export code. The old archive’s browser model records different training settings again. The existing submission.csv is byte-identical in the archive and standalone downloads; it cannot prove that current MLP training has been reproduced. That connection requires an export script, checkpoint, split and run logs.
See the C probe record, runnable audit package and README. Continue with strict dataset checks and browser versions and probability limits, or open the playground and original downloads.