Replacing a manually approved form with an automated anti-abuse gate turns out not to hinge on the algorithm. It hinges on one sentence: the score sets a price, it never renders a verdict. No reputation score causes a rejection; it only decides how long you spend hashing. Local signals will misjudge people, and the cost of a misjudgement should be a few extra seconds — not being unable to do the thing.
1. Work out what that manual step was actually deciding
The original flow: a reader submits a mailbox request, an administrator glances at it and clicks approve. That looks like one action. In practice the administrator was making two decisions at once:
- Whether to create the mailbox — defending against automated bulk registration;
- Whether to grant outbound sending (
send_allowed, default0).
These carry risks of entirely different magnitudes. The bad case for the first is some junk accounts in a database. The bad case for the second is damaged domain reputation — and domain reputation is shared with every existing mailbox under it. One abusive sender takes the whole mail system down with them.
So the automation covers only the first. Passing the proof of work sets the request to approved with a default quota, while send_allowed stays at 0 and is still granted separately from the back end.
That is deliberate, not unfinished. When automating a manual step that bundles several decisions, split it first and ask “what is the bad case for this one” about each. Automating away the part that should have stayed manual is the classic accident in this kind of migration.
2. A hand-written SHA-256 does not error when it is wrong — it just never verifies
Proof of work means hashing a lot in the browser. The instinct is crypto.subtle.digest, but it returns a Promise per call — one microtask hop per hash, measured at a few tens of thousands per second. That cannot carry a puzzle needing millions of attempts. So the worker has to contain a synchronous SHA-256.
And a hand-written SHA-256 has a particularly nasty failure mode: if it has a bug, the client and server compute different digests, verification fails forever, and nothing errors. The user sees “it just never accepts my submission.” The server log shows an ordinary verification failure. Nothing points at the hash implementation.
So this implementation cannot be accepted on the grounds that it runs. It has to be checked against vectors. The worker source is extracted from the JS and run standalone under node:
- 21 digest vectors generated with Python’s
hashlib, chosen to cover block boundaries: 55, 56, 63, 64 and 65 bytes. SHA-256 processes 64-byte blocks and the length field occupies 8 of them, so the 55/56 pair straddles exactly the line between “the padding still fits” and “we need another block.” That is where nearly every hand-written implementation’s bug lives. - 400 leading-zero-bit cases to check the difficulty predicate. Counting leading zero bits is unusually easy to get wrong across byte boundaries.
- End to end: node mines a solution using the worker’s logic, then node’s own crypto verifies it using the server’s PHP logic. Two independent implementations that have to agree.
All three layers are necessary. The first two establish that the hash is correct; the third establishes that both ends agree on what is being hashed — concatenation order, separators, encoding. Any mismatch there and a perfectly correct hash still never verifies.
3. Difficulty has to be measured; a guessed number is a punishment
PoW difficulty is expressed in leading zero bits, and the expected number of attempts is 2bits. Converting that into “how long does the user wait” needs a measured hash rate:
node, single thread ~720,000 hashes/sec
browser worker 40-60% of that (~350,000/sec)
Which gives:
| Difficulty | Expected attempts | In-browser time | How it feels |
|---|---|---|---|
| 17 bits | 131k | ~0.4 s | imperceptible |
| 21 bits | 2.1M | ~6 s | clearly waiting, but acceptable |
| 8.4M | ~24 s | my original ceiling; far too long |
23 bits was the number I guessed first. Twenty-four seconds is punishing for a legitimate reader the signals happened to misjudge — someone who did nothing wrong and merely arrived from a particular address range. The ceiling ended up at 21 bits.
The point being: difficulty is not a security parameter, it is a user-experience parameter. Its ceiling should be set by “how long does an innocent user wait in the worst case,” not by “how expensive should this be for an attacker” — the latter always produces a larger number.
4. IP scoring: only signals the site can observe itself
Difficulty comes from a local reputation score built only from what this site can see on its own:
| Signal | Points |
|---|---|
| Reverse DNS matches hosting-provider keywords | +30 |
| No PTR record | +5 |
| Recent failures | +10 each (cap 30) |
| Repeated solves in a short window | +8 each (cap 25) |
| New / young account | +15 / +8 |
| Missing or obviously automated user agent | +20 |
The three things deliberately left out say more about the design than the list above:
- No external reputation API. That means sending every visitor’s IP to a third party — manufacturing a larger privacy problem in order to solve an abuse one.
- No third-party blocklists. Their false positives are neither visible to you nor fixable by you, and your readers pay for them.
- No country weighting. This one deserves spelling out: a country’s correlation with “will this person abuse the service” is far weaker than its correlation with “where does this person live.” Weighting by country penalises geography rather than identifying behaviour.
IPs are not stored in the clear; counter keys are HMAC(IP, salt). The system needs the fact “this address failed recently,” not the address.
5. The score sets a price, never a verdict
This is the one part of the design that does not bend: no score causes a rejection. A high score only means a harder puzzle.
The reason is simple. Every signal above will misjudge people. A shared corporate egress, carrier-grade NAT, a freshly registered account, an unusual browser’s user agent — all of these describe perfectly ordinary readers, and all of them add points.
In a system that rejects, the cost of a false positive is “this person cannot do the thing, and does not know why.” In a system that only prices, the cost is “this person waited a few extra seconds.” The first has to drive its false-positive rate near zero before you dare ship it; the second can operate with a known false-positive rate.
This generalises well beyond proof of work. Any heuristic-driven abuse control built on local, incomplete signals should consider changing its output from pass/reject to a cost.
6. Two kinds of puzzle must not be interchangeable
Two entry points on the site use this gate: the mailbox request, and a public IP reputation lookup tool. Their puzzles are not interchangeable — the challenge payload carries a purpose field compared verbatim at verification.
Without that, an attacker farms puzzles at the cheapest entry point (the public tool, low difficulty) and spends them at the most sensitive one (the mailbox request). The cheapest entry point ends up pricing the most sensitive one.
The challenge itself is a stateless HMAC-signed string, deliberately not persisted — the issuing endpoint is public, and persisting would hand out a free write interface. Single-use is enforced by writing one short-lived record at the moment verification succeeds.
7. The public diagnostic tool has to solve a puzzle too
The site’s IP reputation tool tells you the score assigned to your current address and the weight of every signal that contributed. That disclosure is intentional — the gate’s security does not rest on keeping the signals secret.
But the tool itself still requires a proof of work. Otherwise it is a free IP reputation lookup API: point a proxy pool at it, enumerate which address ranges this site considers clean, and then use exactly those.
This is easy to miss: a diagnostic tool that exposes your abuse-control criteria is itself part of your abuse control.
8. What was measured, and what was not
Measured:
curl user agent score 20
browser user agent score 0
fetch 17-bit puzzle -> solve 0.05s -> submit returns 200
replay the same puzzle 403 already_used
What was not measured belongs in the write-up too: the logged-in mailbox request has not been exercised end to end, because there was no test account to hand. The code path is unambiguous (verification failure redirects with an error code), but “an unambiguous code path” and “something that has been run once” are different claims. Writing it down is how it gets covered next time there is an account, rather than quietly becoming “probably fine.”
References
- RFC 6234: SHA-256 specification and test vectors
- Back (2002), Hashcash — A Denial of Service Counter-Measure
- MDN: SubtleCrypto.digest (and why it is asynchronous)
Related: the IP reputation scorer (the tool described above), threat modelling for AI systems, and an auth gate covers paths, not data (another case of a control that looked like it was blocking something and was not).
