How I Fooled Myself Validating Quantisation: IoU 0.98 That Was Really 0.30
How I Fooled Myself Validating Quantisation: IoU 0.98 That Was Really 0.30
Search
Ask the AI

How I Fooled Myself Validating Quantisation: IoU 0.98 That Was Really 0.30

Agreement with FP32 is not the same as agreement with the intended object. This article retains the leads from a historical browser-matting investigation, but separates three questions: what IoU is compared against, where thresholding occurs, and how much a local check can actually rule out.

Revision, 2026-09-08: the 0.98 and 0.30 in the title come from earlier notes, not a new measurement. The existing FP16 file can be checked today; the original failure image, FP32/INT8 files, complete outputs and timing logs are still needed to reproduce that comparison. This revision adds runnable metric counterexamples, not replacement evidence for the old model results.

1. What the historical record says

The earlier notes describe ONNX Runtime Web matting followed by thresholding, colour layers and contour extraction. A dynamically quantized 8-bit version agreed closely with FP32 on some anime illustrations, then retained much less of a subject whose colours resembled its background. That is a lead worth investigating, not a population accuracy estimate or proof of a unique cause without the sample list and tensors.

Record Value and limit
Earlier examples INT8 versus FP32 IoU 0.976–0.984. Complete sample count, sampling method and per-image results were not recorded.
Failure input INT8 versus FP32: [email protected] 0.296, maximum absolute difference 0.388, mean absolute difference 0.0547. These are not comparisons with human labels.
FP16 comparison IoU versus FP32 recorded as 0.9996, maximum difference about 0.001. These cannot be confirmed in this revision without the output arrays.
Speed and size The earlier 35% improvement, 79/337 ms timings and 167.9/84.0/42.1 MB sizes lack complete timing boundaries, environment or unit checks. They are not used as a ranking.
Show the historical coverage table, pending reproduction
threshold    int8     fp16     fp32
      48    32.4%    33.3%    33.3%
      80    25.9%    32.0%    32.0%
     112    11.8%    29.2%    29.2%
     127     7.6%    25.9%    25.9%

These percentages are coverage, not IoU. Equal coverage does not locate foreground at the same coordinates. Nor can an integer cut of 127 be labelled a raw floating-point cut of 0.5 without specifying rounding and the comparison operator. A visual estimate of how much of the frame the subject occupies is not a pixel-labelled target.

2. Four checks are not four universal exclusions

Check Limit
Shape Matching names, types, dimensions and 1,048,576 elements supports an interface contract. It does not rule out a transpose, offset, wrong tensor or crop.
Input A good result on one image validates that case. Check the same failing input’s decoding, orientation, colour handling, interpolation, padding, channel order and actual tensor hash.
Contour Matching bounding boxes does not establish correct area, holes or connectivity. Rasterize the contour back to the same grid and compare pixels, with a declared boundary convention.
Backend Matching max=0.9727 and mean=0.1264 establishes only two rounded summaries. Use the same model and input, compare complete outputs, nonfinite values, maximum error and error locations.

Even full-array agreement constrains a specified model, input, version and option set, not all backend behaviour. Failure to support four candidate explanations does not prove a remaining explanation uniquely correct. Freeze model, input, preprocessing, runtime and postprocessing, then change one factor at a time.

Two models failing on one input does not prove that model choice can never help. Larger coverage need not be more accurate. Without checking architectures, training data and representations, a feature-space explanation is a hypothesis, not an observation. A soft-mask score is also not automatically a calibrated uncertainty estimate or proof that the model “saw” a particular part.

3. Sixteen pixels make the metric distinction explicit

Here foreground means score s >= t. IoU divides foreground intersection by foreground union. Prediction-to-prediction IoU measures agreement; prediction-to-independent-label IoU evaluates quality for that labelled task. For two empty foreground masks this lab reports null and counts the case separately, rather than silently awarding a perfect score.

Construct 4×4 arrays with every row of A equal to [0.8,0.8,0.2,0.2], and B reversed. Their complete histograms, means, maxima and foreground coverage match. At threshold 0.5 each has eight foreground pixels, but there is no overlap: IoU is 0. If A is the constructed target, two identical copies of B have mutual IoU 1 while both score 0 against that target.

Four synthetic mask controls: a spatial permutation, a shared error, and perturbations far from or near a threshold; red marks changed decisions
Rendered directly from the fixture files at threshold 0.5. Teal is foreground; red marks differences between predictions. Target is defined by the construction rule, not a human-labelled photograph.
Show the complete spatial_counterexample.py
"""Equal histograms and coverage do not imply spatial or label agreement."""
def iou(a, b):
    if len(a) != len(b) or not a:
        raise ValueError('Expected equally sized, nonempty masks')
    intersection = sum(x and y for x, y in zip(a, b, strict=True))
    union = sum(x or y for x, y in zip(a, b, strict=True))
    return intersection / union if union else None


a = [0.8, 0.8, 0.2, 0.2] * 4
b = [0.2, 0.2, 0.8, 0.8] * 4
mask_a = [x >= 0.5 for x in a]
mask_b = [x >= 0.5 for x in b]
assert sorted(a) == sorted(b)
assert sum(a) == sum(b) and max(a) == max(b)
assert sum(mask_a) == sum(mask_b) == 8
assert iou(mask_a, mask_b) == 0.0
assert iou(mask_b, mask_b) == 1.0
assert iou([False]*16, [False]*16) is None
print('same histogram; foreground 8/16 in each; A versus B IoU = 0')
print('if truth is A, two copies of B agree perfectly but both have truth IoU = 0')
print('SPATIAL_COUNTEREXAMPLE_OK')

Aggregation matters too. If one image has intersection/union 1/1 and another 0/99, the macro average is 0.5 while pooled, micro IoU is 0.01. State denominators and empty-case handling. At 0.5, this lab’s five synthetic cases give macro agreement 0.5 after excluding one empty union, and micro agreement 12/32=0.375. These are artificial examples, not a model’s validation scores.

4. Threshold margin is more testable than “confidence”

Let the reference score be s, the changed score s+δ, and the threshold t. If each pixel satisfies |δ| < |s-t|, its binary decision cannot flip. This is a sufficient condition, not a guarantee that quantization error is small. Real errors may be structured and depend on operators and inputs; independent random jitter is not a universal quantization model.

The lab defines the central 2×2 square as its target. Subtracting 0.02 preserves foreground/background decisions for scores 0.95/0.05 at threshold 0.5. Starting from 0.51/0.49 instead removes all four foreground pixels and changes prediction agreement from 1 to 0. No quantizer is executed in this score-perturbation example.

Lowering the second case’s threshold to 48/255 restores agreement to 1, but now the entire image is foreground and candidate-target IoU is only 0.25. A lower cut can recover foreground while admitting background. Greater coverage alone is not evidence of a fix.

The site’s current implementation adds a concrete complication: cropped scores are written into a Uint8ClampedArray, resized through Canvas, and then tested with full[k*4] > cut. This lab measures raw floating-point scores before those steps. Isolating rounding and comparison, scaled values 80, 80.25, 80.5 and 80.75 become 80, 80, 80 and 81. Byte >80 keeps only the last; raw score >=80/255 keeps all four.

Show the complete uint8_threshold.cjs
'use strict';
const assert = require('node:assert/strict');
const scaled = [80, 80.25, 80.5, 80.75];
const score = scaled.map(x => x / 255);
const byte = Array.from(new Uint8ClampedArray(score.map(x => x * 255)));
const byteMask = byte.map(x => x > 80);
const rawMask = score.map(x => x >= 80 / 255);
assert.deepEqual(byte, [80, 80, 80, 81]);
assert.deepEqual(byteMask, [false, false, false, true]);
assert.deepEqual(rawMask, [true, true, true, true]);
console.log(JSON.stringify({ scaled, byte, byte_gt_80: byteMask, raw_ge_80_over_255: rawMask }));
console.log('UINT8_THRESHOLD_OK');

Export raw output, the resized 8-bit image and the final mask separately, then compare like stages. The interactive threshold lab helps illustrate perturbations, but its artificial scores and noise are not the original model and cannot replace a quantization evaluation.

5. Preserving the FP16 interface is not an accuracy guarantee

keep_io_types=True retains float32 inputs and outputs during conversion. It does not imply every internal operation becomes FP16: blocked operators can remain float32 with inserted casts. Compatibility, speed and acceptable error must be tested for the actual graph and runtime. See the ONNX Runtime FP16 conversion guide.

The old explanation equating a clamp warning at ±1e-7 with FP16’s smallest subnormal was inaccurate. IEEE binary16’s smallest positive subnormal is 2^-24, about 5.96e-8; the converter’s min_positive_val=1e-7 is a separate clipping parameter. Values are also rounded to FP16. A warning threshold is not the representation’s physical lower limit; see the converter implementation. A warning alone proves neither harmful output changes nor their absence.

The current FP16 file is 88,070,593 bytes. The ONNX deployment article publishes its hash, two synthetic inputs and complete WASM outputs. Those checks establish assets and interfaces, not this article’s historical 0.9996 or an FP16/INT8 speed ranking.

6. Design the next validation before inspecting its scores

Define the intended use and sampling rules first. Use an independent validation set representative of expected inputs to estimate routine performance, and separate stress strata for low contrast, fine structures, occlusion and atypical styles. Report both, without treating an intentionally inflated hard-case proportion as the production distribution.

Record input and model hashes, preprocessing version, runtime, threshold and export stage per sample. Report raw-score errors, prediction agreement, independent-label quality, foreground counts and failure cases, not just an average. Freeze graph-optimization options in backend comparisons. Time download, session creation, warm-up and repeated inference separately.

Choose a threshold on tuning data, freeze it, then evaluate untouched holdout data. A repeatedly debugged failure image remains useful as a regression test, but is no longer independent confirmation. Finding no failure proves neither robustness nor sampling bias by itself; inspect coverage, sample size and missing input types.

7. Reproduce the metrics or evaluate your own exports

Download the mask metrics lab. Five synthetic 4×4 cases at five thresholds produce 25 rows, with raw files, CSV, JSON, plotting source and validation tests. Measurement uses Python’s standard library; the rounding example needs Node. No model download or GPU is required.

python3 spatial_counterexample.py
node uint8_threshold.cjs
python3 make_fixtures.py --output reproduced-fixtures
python3 evaluate_masks.py --manifest reproduced-fixtures/pairs.json --output reproduced
python3 -B test_metrics.py

Output directories must not already exist. Successful measurement prints MASK_METRICS_OK 5 samples 25 rows. For your own data, use the README to declare H×W, little-endian raw float32 scores, optional binary targets and thresholds. The evaluator rejects wrong lengths, nonfinite values, out-of-range scores and invalid labels. It cannot establish that source coordinates are aligned or data are independently sampled.

This revision verifies metric definitions, counterexample results and the current code’s thresholding stages. Reproducing historical model quality still requires the original materials. Neither “quantization was the only cause” nor “FP16 is generally lossless on this input class” follows from the evidence available here.

Leave a Reply

Scroll down