The handwritten-digit browser module performs inference: it loads weights and turns 784 grayscale values into ten class probabilities rather than retraining the full dataset on each visit. That distinction is necessary but incomplete. We also need to identify the file being loaded. The preserved project archive and the standalone model JSON have different structures, metadata and embedded samples.
Checked on 2026-09-08. This article reports an artifact comparison, offline inference on three fixed inputs and the live blank-board behavior. It adds no training score, transfers no historical accuracy to the current model, and changes neither online weights nor playground interactions.
1. The current model is not the archived model
The page configuration loads the standalone digit-playground-model.json. Its keys are W1, b1, W2, b2 and meta, forming a 784 to 128 ReLU to 10 Softmax MLP. All 101,770 parameters pass shape and finite-number checks. Its first 12 SHA-256 characters are 51073f30e1b5.
| File | Model | Samples |
|---|---|---|
| Direct | MLP 784 / 128 / 10 |
0 |
| Archive | Linear 784 / 10 |
40 |
The current JSON metadata contains only a model name, 20 epochs and learning rate 0.005, with no validation accuracy. The archived JSON has 7,850 parameters and hash prefix 6b256e9455cc; its metadata records 10,000 training samples, 2,000 validation samples, 18 epochs, learning rate 0.35 and validation accuracy 90.65%. The table’s sample count refers to embedded records available for selection, not training-set size.
90.65% is a metadata entry in the old file; its split and training process were not reproduced here. It is neither the current MLP’s validation accuracy nor an estimate for hand-drawn inputs. The archived C source uses 20 epochs and learning rate 0.01, different again from that JSON record. These files cannot simply be described as inputs and outputs of one training run.
The standalone C source has the current JSON’s MLP shapes, but no JSON-weight export code. Available checkpoints and logs do not establish a binding between them. Matching filenames or layer counts do not prove lineage. Full hashes and archive comparisons are in the audit record.
2. Why the sample list and validation score are empty
The browser script supports both weight structures. Sample options come from the JSON samples array, and the score comes from meta.validation_accuracy. Neither field exists in the current standalone file. The live check therefore found zero sample options and a dash for validation accuracy. That dash is not a hidden 90.65% result.
Drawing and probability inference still work, but this artifact does not provide the embedded sample browsing described by the older article. The nearby sample-preview image is a separate display asset, not evidence that selectable records exist inside the model JSON. This article records that limitation without switching back to old weights or inventing samples.
The brush generates grayscale values clamped to 0 through 1. The current inference path does not automatically crop, resize to a particular bounding box or center by mass. A hand-drawn approximation is not identical to the binary arrays used below, so it should not be expected to produce identical probability digits.
3. Measured blank and translated-input outputs
The offline checker reads the current JSON and evaluates its two matrix operations, ReLU and maximum-shifted Softmax using Python floats. It defines three exact inputs: all zeros; a vertical line of ones at column 14, rows 5 through 22 inclusive; and the same line at column 7. Coordinates are zero-based, all other pixels are zero, and there is no interpolation or automatic centering.
| Fixed input | Top-1 | Top-2 |
|---|---|---|
| Blank | 5; 52.87% | 6; 10.23% |
| Line, column 14 | 1; 99.96% | 9; 0.0159% |
| Line, column 7 | 4; 40.59% | 0; 21.93% |

On the live Chinese interface, clicking Clear and then Predict displayed digit 5 at 52.9%, matching the offline blank result at the UI’s precision. The two line results are offline array probes, not screenshots of exact inputs reproduced with the browser brush. The package records all ten probabilities, their sums and input hashes.
A blank still leaves something for the network to compute: a1=ReLU(b1), followed by W2*a1+b2. Translating the line activates different pixel weights, changing the preferred class. This demonstrates sensitivity of these weights to one specified perturbation; it does not estimate an overall handwritten-digit error rate.
4. Probability, margin and accuracy answer different questions
Top-1, Top-2 and their difference describe one output, not proof that an input belongs to the training distribution. The output must distribute its mass across ten digits even when the input is blank. A score near one cannot automatically be interpreted as nearly 100% correctness for similar inputs. A near-uniform distribution also does not, by itself, establish that an input was never seen.
Accuracy requires a collection with true labels. Probability calibration compares score intervals with observed correctness on representative samples. Out-of-distribution detection needs a separately designed set of non-digits, blanks and abnormal inputs. A historical 90.65% field or a single probability margin cannot substitute for these different evaluations.
Keeping failure inputs teaches more than showing only successful digits, but a favorable set selected after observing predictions is not a defensible test set. The three arrays here explain mechanisms; they have no ground-truth labels and contribute to no accuracy denominator.
5. A centroid diagnostic is not an implemented fix
Stroke position, size, grayscale and thickness may differ from training inputs. The model itself, parsing or weight provenance can also cause mistakes; we cannot start by declaring the model faultless. The following function only calculates the grayscale centroid offset from the grid center. It does not translate an image and is not connected to the online inference path.
Centroid diagnostic with a blank-input guard
"""Diagnostic offset only: this does not resample or improve a classifier."""
import math
def centroid_offset(pixels):
if len(pixels) != 784 or any(not math.isfinite(v) or not 0 <= v <= 1 for v in pixels):
raise ValueError("expected 784 finite grayscale values in 0..1")
mass = sum(pixels)
if mass == 0:
return None
cx = sum((i % 28) * v for i, v in enumerate(pixels)) / mass
cy = sum((i // 28) * v for i, v in enumerate(pixels)) / mass
return 13.5 - cx, 13.5 - cy
if __name__ == "__main__":
from audit_digits import probe_inputs
for name, pixels in probe_inputs().items():
print(name, centroid_offset(pixels))
For the fixed inputs, the outputs are None, (-0.5,0.0) and (6.5,0.0). A blank has no defined grayscale centroid, so dividing by zero is invalid. With pixel-center coordinates 0 through 27, the geometric center is 13.5; a half-pixel translation needs an explicit interpolation convention.
Before adding cropping, proportional scaling or centering, define preprocessing consistent with training. Then compare raw and transformed inputs on a frozen hand-drawn validation set not used for tuning, retaining per-example outcomes. Cropping can remove strokes and interpolation changes grayscale. No experiment here establishes that centering must substantially improve accuracy; a proposed comparison is not a completed result.
6. Reproduce and retain the minimum record
Download the artifact audit package, prepare standalone resources using its README, and run:
python3 audit_digits.py --assets ../digit-assets --out ../digit-results
python3 centroid_offset.py
Data and inference checks require only the Python standard library; this run used Python 3.13.9. Retain the full model hash, input pixels and their hash, normalization convention, inference-code version, ten probabilities, and whether a ground-truth label exists. Recompute after replacing weights rather than treating this page’s numbers as universal answers.
Continue with dataset checks and spatial indexing and the two C versions and loss counterexample. Original materials remain in downloads, and the current module is in the playground. Separating measured behavior, historical metadata and proposed experiments makes individual claims testable rather than asking readers to trust a single score.