The previous post named a gap as future work: horizontal scale-out for the signaling layer worked, but a crash on any instance left the surviving peer stranded, with no way to know.

The session state kept pointing at a peer that no longer existed.

This post closes that gap, then goes further: it builds two different ways to detect a failed instance, measures both, and looks at what each actually costs.

Dangling session

The signaling cluster runs N backend instances behind nginx, sharing session state through Redis. A 1:1 call session tracks two peers by socket ID. When an instance goes down (kill -9), its sockets drop from the cluster with no announcement.

The surviving peer’s session is still CONNECTED. A reload doesn’t help: the rejoin handler checks the session state before allowing a rejoin, and CONNECTED is the same state whether the peer is alive or not, so the rejoin gets rejected. As the previous post’s live test found, the only thing that eventually cleans up this dangling session is the Redis key’s 6-hour TTL. For six hours, the call looks alive to the system but is unusable for both users.

A live 2-instance cluster confirmed this: a kill -9 mid-call left the other peer connected to a call that no longer existed.

Two designs

The reaper

The reaper already existed before this feature, built to scan every live session on a tick and clean up ones stuck too long in AWAITING_REJOIN. A single instance in the cluster holds a lock and runs it on an interval; “single-owner” means only one runs at a time, even with N backends, so there’s no race doing the same cleanup twice.

It’s the natural place to add liveness detection: the tick already runs on an interval and already scans every session, so adding a second check there is cheaper than building a new mechanism.

Both designs: sweep vs heartbeat lease

Cost scales proportionally both ways: double the count, double the cost.

Sweep: reuse the adapter’s own presence primitive. Socket.IO’s Redis adapter exposes fetchSockets(): every instance publishes a request, live nodes reply with their local sockets, the caller merges the replies (or times out after a few seconds). A failed node simply doesn’t answer, so its sockets are absent from the merged set. No new state, no new keys. Cost is proportional to however many sockets currently exist, since every reply carries a full description (socket ID, room, connection metadata) per connection.

Heartbeat: a per-instance lease. Each instance writes a short-TTL key to Redis (sig:instance:<id>) and refreshes it on an interval. Every session, when a peer’s socket is assigned, records which one owns it (ownerInstanceId). The reaper’s tick becomes: SCAN for those live keys, then for each session check whether its peers’ owners are still in that set. Cost is proportional to the total count, not sockets, plus one Set lookup per session (does the owner’s key still exist?).

Both designs share the same failure mode: a node that’s alive but stalled, paused by a garbage-collection cycle (the runtime briefly stopping to free memory) or a CPU spike (the machine too busy to run the process), can look gone for a moment even though it will recover on its own. Both handle it the same way: don’t act on the first gap. Confirm on the next tick before declaring a peer gone. Two-tick confirmation, one reaper interval apart.

I didn’t reach for a heavier failure-detection scheme on purpose: at three or four instances, the reaper can already see every session in a single tick.

Measurement

I built a harness, a tool that spins up its own N-instance fleet, introduces failures, and records results: it measures three metrics for each design:

  • Detection latency. Establish live calls, crash an instance (kill -9), record how long until the surviving peer is notified. Repeated 20 times, rotating which instance goes down, decomposed into kill_to_first_gap_ms and first_gap_to_confirm_ms using structured logs the backend emits at each step.

    const notified = new Map(); // cid -> t_received
    for (const a of affected) {
    	a.survivor.socket.once("call-peer-disconnected", () =>
    		notified.set(a.cid, Date.now()),
    	);
    }
    
    const tKill = Date.now();
    fleet[target].kill9();
    
  • False positives under a slow node. Freeze an instance with SIGSTOP, a POSIX signal that pauses a process instead of terminating it, for a multiple of the reaper interval, then resume it with SIGCONT. Any notification sent while the instance was merely paused, not failed, counts as a misread.

    fleet[target].signal("SIGSTOP");
    await settle(pauseMs);
    fleet[target].signal("SIGCONT");
    
    // let the reaper tick a few times, then check recovery
    await settle(REAPER_MS * 3);
    
  • Cost under load. No fault at all, just the reaper’s steady-state tick duration and how much it has to look at, measured across 50 to 1000 concurrent calls.

    // per-tick cost = delta between two cumulative Prometheus scrapes
    const durationMs = histTotals(txt, "signaling_liveness_sweep_duration");
    const deltaMs = histDelta(prev.get(`d${i}`), durationMs);
    

Run on a 8-core cloud box, three backend instances, enough to show one instance going down while the other two keep serving unrelated calls without needing a bigger box, 5-second reaper interval, 6-second heartbeat TTL.

Three instances is a lower bound, not a guarantee for bigger clusters: the same box-size limitation that the scale-out article ran into at N=4. Heartbeat’s Redis SCAN cost would need re-measuring past a few more instances.

Results

Detection latency

Detection latency, sweep vs heartbeat

p50p95max
sweep8.7s8.8s9.4s
heartbeat13.4s13.4s14.4s

Detection latency splits into two segments: how long until the first gap opens, and how long the confirmation tick takes after that. Both designs share the same second segment, about 5 seconds to confirm, one reaper interval, identical mechanism in both. The difference is in the first segment.

Sweep’s first segment averages 3.6 seconds. fetchSockets() has no separate timer: an instance that’s actually down just doesn’t answer on the next tick, so the gap is bounded by the reaper interval itself.

Heartbeat’s first segment averages 8.3 seconds, over twice as long. It’s gated by the lease TTL, not the tick interval: its heartbeat key doesn’t expire in Redis until the TTL runs out, so the earliest the gap can open is roughly the TTL plus however far into the last refresh cycle the instance went down.

The heartbeat design’s detection speed is tunable by shortening the TTL, but in this configuration it’s the slower of the two.

Shortening the TTL enough would make heartbeat faster than sweep, but the same short TTL is what causes the false positives in the next section.

False positives under a paused node

False positives vs pause duration

Both designs wait for two consecutive reaper ticks to agree an instance is gone before acting on it, 10 seconds at this cadence, so a shorter pause shouldn’t cause a false positive at all. The tested durations sit at or under those 10 seconds, plus one past it, to check whether that holds and what breaks once a pause runs longer.

pause duration× reaper intervalsweep falseheartbeat false
2.5s0.5x0 / 1000 / 100
5s1x0 / 1000 / 100
7.5s1.5x0 / 1000 / 100
10s2x0 / 1000 / 100
15s3x0 / 100100 / 100

This is the starker finding. Sweep just asks “who’s here now” every tick, so a paused instance simply answers late once it wakes up. Heartbeat relies on a Redis key with a fixed TTL: once that expires mid-pause, only the instance itself can renew it, and it can’t while frozen.

That’s why a long pause reliably causes a slip for heartbeat but not for sweep.

The shared two-tick confirmation doesn’t save heartbeat here, because it only filters out a one-off blip, not a sustained absence. Sweep asks it directly on every tick, so it gets an answer the moment it’s back. Heartbeat needs it to refresh its own key first, an extra step on its own schedule, so if it’s still catching up when the confirmation tick runs, the key still reads as expired.

Cost under load

Reaper-tick cost vs concurrent calls

concurrent callssweep p95 (ms)heartbeat p95 (ms)
503.31.2
2007.23.9
100029.015.8

Heartbeat is consistently cheaper, roughly half the tick duration at every level tested. Both curves still climb with call count, since both still walk every live session each tick; what stays flat for heartbeat is the Redis-side scan itself; the number of keys it reads never grows past the instance count, regardless of how many calls are running. Sweep’s cost is tied to the number of connected sockets, since fetchSockets()’s replies carry a full socket description per connection, and that count grows directly with call volume.

No collateral damage

Both designs were also checked for safety, via unit and integration tests, not the harness above. Three things had to hold: calls on other instances keep running undisturbed, only one instance acts on a terminated session, and each of those gets cleaned up once detection finishes.

The single-owner property, no race between instances doing the same cleanup twice, comes from the reaper’s lock:

it("acquireReaperTick is an NX lock: single owner per interval", async () => {
	const other = new RedisSessionStore(redis);
	expect(await store.acquireReaperTick(15_000)).toBe(true);
	expect(await other.acquireReaperTick(15_000)).toBe(false);
	other.stop();
});

An integration test that terminates a backend process mid-call on a live 2-instance cluster covers the remaining two:

const droppedAtB = ctx.answerer.waitFor("call-peer-disconnected", 25_000);
killPort(KILL_PORT); // terminates the process, not disconnect()

const drop = await droppedAtB;
assert.equal(drop.role, "offer");
// surviving peer notified, reaper marked the call AWAITING_REJOIN

const rejoinAtB = ctx.answerer.waitFor("call-rejoin-offer", 5000);
reOfferer.emit("call-rejoin", {
	conversationId: ctx.conversationId,
	offer: OFFER,
});
const rejoin = await rejoinAtB;
assert.equal(rejoin.conversationId, ctx.conversationId);
// dropped peer rejoined against the surviving instance

When to use each

Sweep detects faster and tolerates pauses, but costs more per tick as the cluster grows. Heartbeat is cheaper to run continuously and its detection speed can be tuned, at the cost of a stricter limit on how slow a node can be before it’s treated as failed.

I’ll use sweep: no TTL to tune, and the cost gap is a few milliseconds at this size. Heartbeat wins once cost becomes the constraint.