Your disconnect handler deletes the connection that replaced it
The client did everything right: it noticed the network change, came back in under a second, and announced itself again. Forty seconds later the server forgot it existed, because the connection it had replaced finally got round to dying.
On this page
TL;DR: A cleanup that removes its entry by key rather than by connection deletes its own replacement. The client reconnects in a second; the server declares the abandoned socket dead forty-five seconds later, and only then runs the handler that erases the registration the new connection made. The client is connected, the server has no record of it, and nothing errors on either side. We measured it: identical code is 0% on a graceful close, 0% on a reset and 100% on a silent disappearance — so a stale disconnect handler is harmless in every test you are likely to write.
Start with what this is not, because the neighbouring failure is better known and the two look identical from the outside.
The famous one is a client that comes back wrong: the socket reconnects, the transport recovers, and the application-layer subscription does not — nobody re-issues the subscribe, so the client sits attached to a feed that will never send it anything. That is a client-side bug with a client-side fix, and we have written it up separately.
This is the opposite failure. Here the client does everything right. It notices the network change, comes back in under a second, announces itself, and the server accepts the registration and stores it. The client is correct, the server is correct, and the record is gone anyway — because the connection this one replaced had not finished dying yet, and when it did, its cleanup took the new one with it.
Two sightings at opposite ends of the same pipe
The first was a directory of connected devices, keyed by device id. Register on connect, delete on disconnect — the shape everyone writes:
// on connect
devices[deviceId] = socket
// on disconnect
delete devices[deviceId]The second was three services away and did not look related: an ingest worker holding its upstream feed socket in a module-level binding, whose error handler closed the socket the binding currently pointed at. Reconnect logic elsewhere had already replaced that binding.
The two have nothing in common except the thing that matters. Both name their target by a location that something else is allowed to overwrite, and both assume that the thing found there when the cleanup runs is the thing the cleanup was created for. That assumption holds for exactly as long as the cleanup runs promptly, and the whole point of what follows is that it does not.
- Device connectsstored under its device id
- Network changesit vanishes without closing
- Server notices nothingwrites still succeed
- Device returns~1s, re-registers, same id
- Everything worksthe directory is right
- Heartbeat gives upon the connection from step 2
- Its cleanup deletes the idand the entry is step 4's
Why doesn't the old handler know it has been replaced?
Because nothing tells it, and nothing can. The handler closes over a device id, not over a connection: at the moment it was registered, the id and the connection were the same thing, and the code that reads naturally is the code that names the id. By the time it fires, the id has been re-pointed at a socket it has never seen.
The second half is the timing, and it is not an accident either. A server cannot notice that a peer
has vanished without closing — there is no event for "the other end stopped existing". It finds out
by not hearing from it, which means a timer. Socket.IO's shipped defaults are
pingInterval 25000 and
pingTimeout 20000, and its documentation
states the sum plainly: a peer is declared gone after pingInterval + pingTimeout. Forty-five
seconds.
Both halves of the ratio come from the same vendor's defaults, which is worth being pedantic
about, because the argument is the ratio and half of it asserted would be half an argument. On the
client side, reconnectionDelay is
1000 ms and
randomizationFactor is 0.5, and the
documentation works the example through: "1st reconnection attempt happens between 500 and 1500
ms."
So a device that noticed its network change is back in about a second, against a window of forty-five. The replacement arrives, registers, and is served for the rest of that window before the handler that erases it is even scheduled to run.
The obvious objection is that a real reconnect is slower than a documented backoff, and it is — particularly when a whole fleet returns at once and contends for the connect path. We have measured that separately: a ten-thousand-client fleet reconnecting on those same defaults waited a median of 10.5 seconds to get back in. That is ten times the single-client figure and still under a quarter of the window — 10.5 against 45 — so the ordering does not change. It is the objection that strengthens the claim rather than weakening it.
The teardown mode decides whether it fires at all
This is the part that explains why the defect survives review. We built a harness that severs one connection three different ways in front of a server with a heartbeat and a directory, and ran the same unguarded cleanup against each. Six trials per arm, the client returning 150 ms after the cut:
| How the connection ended | Live registration erased | Client connected, server has no record |
|---|---|---|
| Graceful close | 0% | 0% |
| Reset | 0% | 0% |
| Vanished silently | 100% | 100% |
| Vanished silently, cleanup guarded | 0% | 0% |
An announced death is reported to the server immediately — the read fails, the handler runs, and it runs before the client can possibly come back. There is nothing for it to delete but its own entry, which is correct. Only the silent disappearance leaves the cleanup pending long enough to land on a successor.
So every teardown a test suite performs is the safe one. Closing a socket in a test closes it. Stopping a container sends a signal. Restarting a service in staging is a graceful shutdown. The one teardown that fires this is a peer that stops existing without saying so, which is what a laptop closing its lid, a phone changing towers or a NAT dropping a mapping actually looks like — and which is unusually hard to stage on purpose.
Is it a race? Only inside one ping interval
"Race condition" is the natural label and it is worth being precise about, because it implies a coin toss and this is not one. Sweeping how fast the client returns, against the same silent teardown:
| Client returns after | Live registration erased |
|---|---|
| 50 ms | 100% |
| 200 ms | 100% |
| 450 ms | 67% |
| 1 s | 0% |
| 2 s | 0% |
The transition band is exactly one ping interval wide, and inside it the outcome really is decided by chance — specifically by where in the heartbeat cycle the connection happened to die. Outside it, there is no chance involved in either direction.
The rig's timers are scaled down about fortyfold so the sweep runs in minutes; what matters is the ratio, which is why the sweep exists rather than a single arm. Put the real numbers back and a one-second reconnect sits an order of magnitude inside the certain region. In production this is not a race that sometimes bites. It is a scheduled deletion that happens every time.
The fix is one comparison, and it belongs at the delete
Delete by identity, not by key:
// on disconnect
if (devices[deviceId] === socket) {
delete devices[deviceId]
}That is the whole change, and it is the one arm in the table that reads 0% under a silent teardown.
None of which is new, and pretending otherwise would be the wrong claim to make. Guarding a disconnect handler against a stale socket is a fix people have already written, in pull requests, in review comments, in code you may already have. What has not been written down is when it fires — that the same unguarded code is harmless on every disconnect a test performs and certain on the one production performs, and that "race condition" is the wrong label for something with a coin-toss band one ping interval wide and certainty on both sides of it.
The general form is worth stating because the directory is only where we happened to find it: a cleanup may not name its target by a location another writer can update. Capture the thing, compare before you remove it, and if it does not match, the cleanup has already been superseded and its correct action is nothing.
Two habits fall out of that. Prefer a compare-and-delete over a bare delete anywhere the key is
shared and the handler is deferred — every disconnect path, every timeout callback, every
finally that tidies a resource identified by name. And be suspicious of module-level bindings read
inside error handlers, which is the same bug with the map spelled differently: the binding is the
shared mutable location, and the reconnect logic is the other writer.
Every teardown a test suite performs is the safe one
FAQ
Isn't this just the reconnect bug where the client fails to re-subscribe? No, and they are worth keeping apart because the fixes live on opposite sides. In that one the client comes back without re-establishing its application state, and the fix is on the client. Here the client comes back and re-establishes everything correctly; the server then deletes it. You can have either, or both, and fixing one does nothing for the other.
Would a shorter heartbeat interval fix it? It shrinks the window, and it cannot close it. The failure needs the client to return before the server gives up, and clients return in about a second — you would need a liveness timeout under that to make the ordering flip, which trades a rare correctness bug for constant false disconnections of slow or busy peers. The comparison at the delete costs nothing and closes it entirely.
Doesn't connection state recovery handle this?
No, and it is worth knowing why rather than assuming either way. Socket.IO's
connection state recovery restores, in its own
words, "the id, the rooms and the data attribute of the socket" — server-side socket state. The
entry this failure deletes is not any of those. It is a key in an application's own map, put there by
application code, and nothing in the feature knows it exists. The documentation says as much in the
general case: "you will still need to handle the case where the states of the client and the server
must be synchronized." Turn it on if you want your rooms back; it will not put your directory entry
back.
How would this show up in monitoring? As nothing, mostly. Both connections were healthy, one closed on schedule, and the delete succeeded. What a user reports is that a device shows offline while it is plainly online, or that messages addressed to it are silently dropped — and the gap between the reconnect and the deletion means the symptom appears up to a heartbeat cycle after the event that caused it, which is long enough to break the association for whoever is looking.
Does this need two processes to happen? No. Everything above is one process and one server. Multiple instances make it worse in a different way, because the directory may then be shared and the two connections may land on different nodes, but nothing here requires that.
The three things worth taking away
The teardown mode decides whether the bug exists. Identical code is fine on every disconnect a test suite performs and broken on the one production performs. If you have never seen this fire, that is weak evidence — it may only mean nothing has vanished silently while you were watching.
Delete by identity, not by key. A cleanup that names a shared mutable location is correct only while it runs promptly, and a heartbeat guarantees it does not. One comparison converts it from correct-by-timing to correct.
Look for the same shape where the map is spelled differently. A module-level binding read inside an error handler is the same defect: something else is allowed to write it, and the handler assumes what it finds is what it was made for.
On a market feed the same scheduled deletion would drop a subscriber rather than a device; what else goes quiet there is on market feeds that reconnect without their symbols.
If a device in your fleet has ever gone missing from a directory while the device itself insisted it was connected, that is the kind of thing we come in for.