The loopback interface cannot physically drop a packet. No cable, no switch, no queue to overflow. And yet ss reported 62 retransmissions on lo. That one impossible observation was the only entry point into the whole failure chain — and following it led to a firewall rule that had been faithfully protecting the wrong address.
1. The symptom: dies under load, fine after a reboot, nothing left to inspect
A home router running a transparent proxy. The symptom was the kind that is hard to even describe: during evening peak, new connections failed en masse, clients retried furiously, the router’s CPU pegged, and a reboot fixed everything.
The worst part is that the reboot destroyed the evidence. On this machine:
logread 128KB in-memory ring, empty after reboot
/var/log tmpfs
collectd/RRD not installed
pstore/ramoops absent
wtmp empty
Worse still, kernel.panic_on_oops=1 together with kernel.panic=3 means any kernel oops silently reboots after three seconds, leaving nothing at all. Before investigating a machine’s history of failures, check whether it keeps any logs that survive a reboot. If it does not, the first job is not diagnosis — it is getting logs onto disk and waiting for a recurrence.
2. The first clue: a number that should not exist
Start with CPU. Only two of four cores were busy, and load average sat around 2 — which reads as “normal load.” That is a trap: this machine’s NIC is single-queue with interrupts pinned to specific cores, so network processing can only use half the cores. By the time load average looks alarming, the network died long ago. Do not use load average to judge network health on a box like this.
Then connection state, which produced the number that should not exist:
ss -tin | grep -c retrans # retransmissions on loopback connections
62.
What makes this number meaningful is that it should be impossible. Loopback traverses no physical medium: no link-layer loss, no queue overflow, no middlebox. Once a packet enters loopback, it cannot need retransmitting unless something deliberately dropped it.
So the question stopped being “why is the network slow” — an unanswerable description — and became something far more precise: who is dropping packets the machine sends to itself?
3. Which packets: reading backwards from the state distribution
Those 62 retransmissions were not spread evenly. Broken down by TCP state:
LAST_ACK 53
CLOSE_WAIT 9
everything else 0
That distribution nearly writes the answer out for you. Both states belong to connection teardown:
CLOSE_WAIT: the peer’s FIN arrived, the local side has not closed yetLAST_ACK: the local side sent its FIN and is waiting for the final ACK
Retransmissions piling up exclusively in teardown means the teardown signals are what is being dropped. In this data path those signals are mostly RSTs — the proxy processes tear connections down with a reset rather than completing a full four-way close.
4. Root cause: the rule protected the wrong address
The router carried an nftables rule that drops RST packets originating from a set of upstream addresses (a common technique when a middlebox injects forged resets):
ip saddr @protected_node_v4 tcp flags & rst == rst drop
That @protected_node_v4 set is synced every two minutes by a cron script, which awks the server: field out of the upper proxy’s configuration.
Here is the problem: the chain has two layers. The upper proxy hands traffic to a local process on the same box, so all six of its configured nodes carry server: 127.0.0.1. The real upstream addresses live in the address field of a different config file entirely.
So the synced set was:
protected_node_v4 = { 127.0.0.1 }
From then on the rule did exactly what it was told: drop every RST originating from 127.0.0.1. Measured over 21 minutes, it dropped 1,074 of them — roughly 73,500 a day. Those RSTs were the teardown signals between two local processes.
5. The full failure chain
RSTs dropped
-> sockets on both ends stick in LAST_ACK / CLOSE_WAIT, never released
-> file descriptors and socket buffers accumulate
-> the processes have an nofile limit of only 4096
-> peak traffic approaches the ceiling, new connections fail to be created
-> clients retry furiously
-> the one or two cores that can handle networking peg
-> presents as "the network is dead"
The fd magnitude is worth recording: 252 connections consumed 1,983 file descriptors, a ratio of about 7.8. That ratio is unremarkable in a layered proxy — each logical connection opens several sockets along the chain — but against a 4096 ceiling it means roughly 520 concurrent connections is the hard limit. For a household router that number is absurdly low.
6. The fork in the road that sends you after the wrong bug
Staring at a climbing fd count, it is very easy to conclude “the upstream process leaks descriptors” and disappear into its issue tracker.
But the fd count came back down: 1968 → 980.
A leak is monotonic. A number that falls again means those descriptors were eventually reclaimed — just too slowly to survive peak load. This is “release is blocked,” not “release never happens.” The two demand completely different fixes: the first requires finding what is blocking, and only the second justifies chasing an upstream bug.
The test is simple: watch long enough to see whether the curve is monotonic. A single instantaneous reading makes the two look identical.
7. The other half of the same bug: an experiment that had never run
That same wrong set was referenced by a second rule, which adds a fixed delay to traffic bound for the upstream nodes as a controlled experiment:
oifname "eth0" ip daddr @protected_node_v4 ...
The set contained only 127.0.0.1, and packets bound for loopback never leave eth0. So this rule’s counter sat at zero forever — the delay experiment that had been running for a long time had never taken effect on a single packet, while raising no error and simply, quietly, doing nothing.
This half caused no outage, but its lesson generalises better: a rule that never matches and a rule that works perfectly look identical without a counter. Glancing at the counters in nft list ruleset -a when you add a rule is far cheaper than reasoning about it later.
8. Fix and verification
The fix itself is small: point the sync script at the correct config file, and add a public-address allowlist that rejects loopback, RFC1918 space, link-local, multicast, and the fake-IP ranges proxy software commonly uses.
What matters is verifying against specific numbers rather than “it feels better now”:
| Metric | Before | After |
|---|---|---|
| Loopback retransmissions | 62 | 0 |
| Loopback connection states | 11 CLOSE_WAIT + 7 LAST_ACK + 6 FIN_WAIT2 | 110 ESTABLISHED + 49 TIME_WAIT |
| Set contents | { 127.0.0.1 } |
4 real upstream addresses |
The appearance of TIME_WAIT is the key signal. It means connections are completing teardown — the side that closed actively is waiting out any duplicate FIN. Previously there was not a single TIME_WAIT, everything stuck in LAST_ACK, which by itself said teardown had never once succeeded.
The processes’ nofile limit also went from 4096 to 65535. That is not a fix; it is moving an absurd ceiling out of the way.
9. The second trap: nftables chain priority decides what state you can read
With that fixed, the next step was narrowing the RST-dropping rule — forged resets are typically injected around the handshake, while legitimate teardown resets appear at the end of a connection. So the ideal condition is “only drop RSTs early in a connection”:
ct original packets <= 20
The original rule hangs off prerouting priority raw, which is -300. Adding the condition in place is the obvious, minimal edit.
And it silently does nothing at all.
Conntrack does not run until priority -200. In a raw chain at -300 the connection tracking state does not exist yet, ct original packets has no value to read, and the condition quietly fails to match.
That conclusion was not reasoned out, it was measured: build two identical rules, one in the raw chain and one in mangle, each with a counter, and let them run:
raw chain counter = 0
mangle chain counter = 108
So the new chain hangs at priority mangle - 10 instead. The narrowed rule has measurable results too: connections stuck in LAST_ACK went from 40–90 down to 0, retransmissions to upstream went to zero, and the early-RST drop counter kept climbing — the protection still works.
This one is worth memorising on its own: whether nftables can read a given piece of state depends on the priority your chain hangs at, and when it cannot, nothing errors — the condition just quietly never matches. Before writing any rule with a ct condition, confirm it runs after conntrack.
10. What this cost and what it taught
- Find a physically impossible observation and start there. Loopback retransmission is exactly that — it compresses an uninvestigable complaint (“the network is slow”) into an investigable question (“who is dropping the machine’s own packets”).
- The state distribution carries more information than the total. 62 retransmissions says nothing by itself; 62 retransmissions all sitting in LAST_ACK and CLOSE_WAIT points straight at teardown signals.
- Check whether the curve is monotonic before chasing an upstream bug. A rising fd count is not a leak.
- A rule that never matches will not tell you it never matched. Counters are the only way to tell, and the time to look is at deployment, not after an outage.
- When you suspect a rule fails silently, run a twin. Two rules differing by one variable, running side by side, beats reading the documentation and is more reliable than reasoning.
- Without logs that survive a reboot there is no debugging. The first thing added to this machine was not the fix — it was persistent logging plus a one-line-per-minute resource trend, because that is the only thing the next crash will leave behind.
References
- nft(8): chain types, hooks and priorities
- RFC 9293: TCP specification (state machine and teardown)
- ss(8): socket statistics and the
-iretransmit counters
Related: TCP reliability and the congestion window (what a retransmission means at the protocol level), forward and reverse proxy trust boundaries (how layered proxies are structured), and SOCKS5 proxies and DNS boundaries.
