The earlier posts (scale and liveness) measured 1:1 call signaling on Node.

Each call runs a state machine (FSM): IDLE → OFFERING → CONNECTED → RENEGOTIATING, plus AWAITING_REJOIN and ENDED. I wrote a minimal Go version of that state machine (gosignal) to benchmark it next to Node, and completed the port afterwards.

The prototype benchmark looked lopsided. At 400 simultaneous calls the load driver measured a median offer-to-answer time of 446ms on Node and 37ms on gosignal, about 12x lower. Then I replayed Node’s lifecycle tests on the port, and three gaps fell out that the benchmark had no way to see: a rejoin that reversed the caller and callee roles, a second offer that overwrote a live call, and renegotiation accepted from any state.

The 12x also shrank once both sides were timed the same way. gosignal needs a per-call ticket: a signed pass that Node issues after a Mongo membership query, which Node also runs on every offer. The driver fetched the ticket before starting the clock, so the port’s number left that query out and Node’s included it. Timed the same way on both sides, the gap at 400 setups is 427ms vs 264ms at p50, 1.6x, and nearly all of the 264ms is the ticket request to Node. The signaling step itself takes about 0.5ms. The case here is isolation, plus a signaling step that stayed under 1ms across the ramp, not a latency win a user would notice.

The post covers the prototype benchmark, the port and the gaps Node’s tests found, and the rerun that cut 12x to 1.6x.

Benchmark setup

The prototype gosignal was the smallest slice a benchmark needed: the state machine, an in-memory session store and the wire events, over raw WebSocket (gorilla/websocket) instead of Socket.IO. Node looks up the caller’s membership in Mongo on every call-offer.

The prototype had no database, so a bench-only route on Node ran the same query and issued a ticket.

The driver fetched that ticket before starting the stopwatch, and gosignal verified it on connect. Node’s timing included the membership query; the port’s did not. I knew that going in: the benchmark was there to time the state machine itself.

I reused the driver from the scale post: backend/test/load/, a Node client that plays both sides of N simulated calls, extended with a raw-WebSocket mode for the port. The load ran on one box (8 dedicated vCPU). A level is N established calls: all sessions connect, N offers go out at once, the answers come back, then ICE candidates flow. Levels run from 25 to 400, 3 trials each, with a fresh server process per run. The prototype run had three targets: node-http (the production baseline), node-uws (the same Socket.IO handlers on uWebSockets.js, a transport-only change tried because the earlier posts traced Node’s limit to its event loop) and gosignal over raw WebSocket.

Prototype results

The prototype run, zero errors across 81 runs, median of 3 trials (driver-side p50 / p95, ms):

levelnode-httpnode-uwsgosignal
2553 / 6562 / 636 / 7
100161 / 172159 / 16111 / 12
400446 / 504439 / 46137 / 40

gosignal came out 8-15x lower at every level, and node-uws stayed within about 5% of node-http from 100 calls up. That is the 12x from the opening: a minimal prototype with its ticket fetch off the clock, next to Node with the membership query on it.

A 12x lead was enough to finish the port.

A faster WebSocket layer barely moved Node’s latency

node-uws was the smallest change on the list: uWebSockets.js is a faster WebSocket implementation, and Node’s Socket.IO stack attaches to it with one call (io.attachApp(app) on a uWS.App), with the signaling handlers called unchanged. If the WebSocket layer was costing time, it would show here. It didn’t: from 100 concurrent calls up, node-uws landed within about 5% of node-http on driver-side p50 (439ms vs 446ms at 400). I had expected more. The transport isn’t where the time goes, so I dropped it and the rewrite became the next candidate.

Two caveats on the prototype numbers

The driver’s headline number comes from a stopwatch: start a timer before sending call-offer, stop it when call-answer arrives. That works when latencies run in the hundreds of milliseconds, like Node’s.

It breaks for gosignal, because the driver is one Node.js process playing every simulated caller and callee inside a single event loop. When that loop is busy, the driver reads the clock late, so a latency it reports can include time spent waiting for its own turn. The harness measures this with Node’s monitorEventLoopDelay (a histogram of how late a 20ms timer runs) and reports the 99th percentile per level as drv_el_p99_ms: 1% of the time, a callback waited at least that long.

targetdriver lag (p99)measured latency (p50, 25 to 400 setups)
node-http50-60ms at every level53-446ms
gosignal80-133ms from 150 setups up6-38ms

For Node the lag is small next to the latency. For gosignal it is larger than the latency the stopwatch was timing, so its driver-side numbers are upper bounds.

The server also times each offer from storing it to relaying the answer, where the driver’s lag cannot reach. Node already recorded that span in an OTel histogram, signaling.offer_answer.duration; the port now records the same histogram.

The histogram never resets between levels, which flattered the level-400 numbers until I started restarting the server for every run.

Porting the state machine

Node’s state machine is the spec. gosignal/internal/fsm ports backend/signaling/fsm.js: the six states (IDLE to ENDED) and the transitions between them. Node’s lifecycle tests in backend/test/signaling/ serve as the acceptance test for the port; I replay them on gosignal below.

With the port in place, the browser keeps its Socket.IO connection to Node for chat and group calls, and opens another WebSocket connection to gosignal for 1:1 call events. Node stays the front door for auth: it hands the browser a token, and gosignal verifies it without querying Mongo.

The browser keeps Socket.IO to Node for chat and group calls, and opens another WebSocket to gosignal for 1:1 call events; Node hands over a signed token that gosignal verifies offline

Replaying Node’s tests

The prototype benchmark exercised one path: offer, answer, then a burst of made-up ICE candidates once the call was up. It left out authorization and every edge case, such as a peer dropping, ICE candidates arriving early, or a second offer. So the benchmark showed how fast gosignal is on the happy path (every message valid and in order, every peer staying connected), not whether it handles the edge cases the way Node does.

The session-existence guard was the exception: while smoke-testing the rig I had noted that handleCallOffer lacked it, and the benchmark never sent a second offer to expose it.

Next I replayed Node’s tests. Its lifecycle and characterization suites gave 15 scenarios, ported into gosignal/internal/server/parity_test.go. Three gaps fell out, each closed with a test.

A rejoin that reversed the roles

When a peer drops, the call parks in AWAITING_REJOIN. The dropped peer reconnects, sends a rejoin offer, and the survivor answers it. CompleteRejoin recorded whoever answered as the answerer. That holds when the offerer drops, since the original answerer survives and answers. When the answerer dropped, the surviving offerer answered the rejoin, got recorded as the answerer, and the roles reversed. Node keeps the roles and refreshes only the answering socket, so the change does the same:

switch answeringUserID {
case f.session.OffererUserID:
	f.session.OffererSocketID = answeringSocketID
case f.session.AnswererUserID:
	f.session.AnswererSocketID = answeringSocketID
}

Both end-to-end rejoin tests now assert that the roles are unchanged after recovery.

A second offer overwrote a live session

handleCallOffer rejected offers from outside the pair and nothing else, so a participant’s second call-offer while the call was live or recovering replaced the session. Node rejects both cases (“already in progress”, “recovering, rejoin”); the change rejects them too, as soon as a session exists:

if existing := s.store.Get(req.ConversationID); existing != nil {
	if !isParticipant(existing, c.userID) {
		return errNotParticipant
	}
	if existing.State == fsm.AwaitingRejoin {
		return errCallRecovering
	}
	return errCallInProgress
}

Renegotiation from any state

Node’s handler allows renegotiation only from CONNECTED. The port’s handler verified participation, then handed the session to the FSM without reading its state, so a renegotiation offer was accepted from any state, including before the call had connected. The FSM’s tiebreak for simultaneous offers (glare: both peers renegotiating at once) exists for CONNECTED and was never reached through this path. The patch adds Node’s guard:

if sess.State != fsm.Connected {
    return fmt.Errorf("cannot renegotiate from state %s", sess.State)
}

Latency never showed any of them; they change behavior, not speed, and no load test could have found them. Before I read another benchmark number from gosignal, it has to pass Node’s tests.

What the benchmark port skipped

Replaying Node’s tests covered behavior. Four other parts of the prototype were shortcuts a live call can’t use: a token scoped to one conversation, ICE sent only after the call was up, no acks (replies confirming a message arrived), and a load driver in place of the real client.

The token was scoped to a conversation, not a user

In the benchmark, Node minted a token scoped to one conversation and gosignal verified it on call-offer.

That was enough for the benchmark, not for a live call. The frontend delivers the incoming-call toast to an idle user over one global socket, which a per-conversation token can’t serve.

So auth split in two. On connect, a short-lived user token carries only the user id, so an idle callee can receive a call. Each call-offer carries a per-call ticket: Node queries membership once and signs the conversation and its participants, and gosignal verifies the signature offline, with no query. The participants come from the signed ticket, not from the request. A test has Node’s jsonwebtoken sign a token and gosignal verify it.

I didn’t reuse the login JWT. A browser can’t send an Authorization header on a WebSocket, so the token rides in the URL, and nginx’s access logs would record a one-hour token that works on the whole API.

A user-scoped connection can send events for any conversation: gosignal compares the sender of each call event with the two participants in the ticket.

At offer time the session records who may answer. Outsiders cannot answer, end the call or send it ICE candidates. Each guard has a test.

ICE candidates, lost twice

The offerer starts gathering ICE candidates (network routes for the media) as soon as it calls setLocalDescription, before the callee answers. Node buffers them on the session and returns them in the call-answer ack. The prototype gosignal aimed them at the still-empty answerer socket and dropped them, and the benchmark never noticed because it never sent ICE candidates from a browser. The change buffers on the session, holds the offerer’s candidates until the callee answers, clears them on rejoin as Node does, and routes by caller identity instead of the client’s didIOffer flag. Requests carrying a reqId get an ack, and call-answer’s ack is the buffered list, or [] plus an error event on failure, so a waiting client never hangs.

The other loss was in the browser. The caller’s code ran setLocalDescription, which starts candidate gathering, then fetched the ticket over HTTP, then sent call-offer, which is what creates the session. Candidates were already flowing during the ticket fetch; the early ones reached gosignal before the session existed and were silently dropped, and the callee received fewer routes than the caller found. Node drops them too, but it has no ticket fetch to open that window. A run with two Chromium contexts showed it. The change fetches the ticket before creating the offer, so the window is no longer than Node’s.

A client that looks like Socket.IO

To keep the frontend diff small, SignalingClient exposes the same methods as Socket.IO (on, off, removeAllListeners, emit, emitWithAck) over the plain WebSocket. The 1:1 call code switches from socket to signaling, and chat and group calls stay on Socket.IO.

No heartbeat, no retry

gorilla/websocket has no heartbeat. A peer that dropped without sending a close frame, such as a lost network or a sleeping laptop, left its read goroutine (a lightweight thread) blocked forever. conn.pingLoop adds a read deadline plus a ping ticker (60s deadline, 54s ping). The frontend client had no retry at all, so SignalingClient got exponential backoff (1s growing to 5s, 5 attempts) on the same instance, which keeps its listeners alive across a drop.

Scale-out: built, not enabled

gosignal ships as a single instance. The Redis path is built and tested across two instances, but I haven’t enabled it in production or measured it, and a piece is missing before it can run on several instances (the last paragraph of this section).

The release runs as a single instance, which is also what the benchmarks measured. A crashed instance can strand its calls on several, with nothing yet to detect it.

RedisStore ports Node’s key layout for sessions. Node delivers events across instances through Socket.IO’s Redis adapter, which is part of Socket.IO: it plugs into Socket.IO’s rooms and sends Socket.IO packets over Redis. gosignal speaks raw WebSocket with a JSON envelope, so there was nothing to plug in, and I wrote delivery and presence on Redis directly.

Every push goes through deliverToUser. It tries the local connection registry, which costs no network hop, and publishes to Redis only when the user isn’t connected to this instance. Each instance also runs a subscriber, which delivers a message when the user it is addressed to is connected to that instance:

func (s *Server) deliverToUser(userID, event string, data any) error {
	if peer, ok := s.conns.byUserID(userID); ok {
		return peer.send(event, data)
	}
	if err := s.broker.Publish(context.Background(), userID, event, data); err != nil {
		s.log.Warn("broker publish failed", "event", event, "err", err)
	}
	return nil
}

func (s *Server) StartBroker(ctx context.Context) {
	s.broker.Subscribe(ctx, func(userID, event string, data json.RawMessage) {
		if peer, ok := s.conns.byUserID(userID); ok {
			_ = peer.send(event, data)
		}
	})
}

All instances share a Redis channel, sig:deliver. The published message carries the userId, the event and its data. Every instance receives it, the instance holding that user’s socket delivers it, and the others ignore it. A failed publish is logged and not returned, and the subscriber resubscribes after any Redis error, so a dropped connection doesn’t stop delivery for the life of the process.

Node’s adapter solves the same two problems, delivery and presence. Delivery works the same way in both. The adapter publishes each broadcast to a channel named after its room, but each instance subscribes with a single pattern that covers all of them, so it receives every message too:

// @socket.io/redis-adapter 8.3.0, dist/index.js (trimmed), on broadcast
let channel = this.channel; // "socket.io#/#"
if (opts.rooms && opts.rooms.size === 1) {
  channel += opts.rooms.keys().next().value + "#";
}
this.pubClient.publish(channel, msg);

// on startup, a single pattern covers every room's channel
this.subClient.pSubscribe(this.channel + "*", ...);

Presence works differently. disconnectSeat decides between recovering and ending a call based on whether the peer is still connected, and it originally consulted only the local registry. A peer on another instance looked gone, so the call ended instead of waiting in AWAITING_REJOIN. Node’s isConnected() answers this question with the adapter’s fetchSockets(), which queries every instance and waits for the replies:

const sockets = await this.socketServer.in(socketId).fetchSockets();
return sockets.length > 0;

gosignal uses a Redis key instead. Each connection sets sig:online:<userId> with a 2-minute TTL, and the ping loop that already runs renews it on every ping. Pings go out every 54 seconds, 90% of the 60-second read deadline, so a single missed renewal still leaves the key alive:

func (b *Redis) MarkOnline(ctx context.Context, userID string) error {
	return b.rdb.Set(ctx, onlineKey(userID), "1", onlineTTL).Err() // 2 minutes
}

func (s *Server) isUserReachable(userID string) bool {
	if _, ok := s.conns.byUserID(userID); ok {
		return true
	}
	return s.broker.IsOnline(context.Background(), userID)
}

fetchSockets() comes from Socket.IO’s adapter, and plain WebSocket has no cluster-wide socket query. A sweep would need a new request and reply protocol over the broker, so I used the key: the ping loop already runs and can renew it. It was less work to build, and it can be replaced later: disconnectSeat only asks the broker IsOnline(userID), so a query across instances can answer the same call.

Two-instance tests cover offer, answer, end and disconnect recovery across processes.

The key covers a peer disconnecting, not an instance crashing. The reaper from the liveness post, which finds calls stranded when an instance crashes, was not ported, so gosignal cannot detect that case today.

A crashed instance’s keys also keep its users marked online for up to 2 minutes, until the TTL expires. That is why scale-out stays off until gosignal can detect a failed instance.

Rerun: timing the ticket too

The prototype run had an asymmetry I knew about. Node queries the caller’s membership inside the handler, so the query was inside its timed span. gosignal has no Mongo: Node mints a token (later, a per-call ticket) after the same query, and the driver fetched it before the timer started. Same query, same once-per-call frequency, opposite sides of the clock.

The rerun moves the ticket request inside the window. The driver also stores every duration, which gives exact percentiles:

const t0 = performance.now();
offerPayload.ticket = await s.off.fetchTicket(cid);
const t1 = performance.now(); // ticket round trip = t1 - t0
s.off.emit("call-offer", offerPayload);
// ...answer comes back at t2
metrics.offerAnswer.push(t2 - t0); // the comparable number
metrics.signaling.push(t2 - t1); // gosignal's part

Two targets (node-http and gosignal with the ticket gate on), nine levels from 25 to 400 setups, median of 3 trials per level (driver-side p50 / p95, ms; four levels shown). At 25 the node-http median has 2 trials, since a trial timed out and I dropped it:

levelnode-httpgosignalratio p50 / p95
2565 / 7355 / 651.2x / 1.1x
100158 / 16389 / 1161.8x / 1.4x
200260 / 275138 / 2191.9x / 1.3x
400427 / 485264 / 4121.6x / 1.2x

Where the time goes at 400 setups (p50 / p95, ms). gosignal makes no database call in its handler: the membership query runs in a separate ticket request, which the driver has to time.

piececlockp50 / p95
node-http Mongo lookupserver309 / 374
node-http setupserver366 / 402
node-http dispatchserver50 / 64
gosignal ticket round tripdriver263 / 412
gosignal signalingdriver0.5 / 4.0
  • Mongo lookup: the findOne plus populate in handleCallOffer, as the handler sees it, so event-loop wait is included.
  • Setup: handleCallOffer entry to the answer relayed to the offerer; it contains the query.
  • Dispatch: offer stored to answer relayed. It includes the answerer’s turnaround through the shared driver, so it is not pure server work.
  • Ticket round trip: the driver’s GET for the call ticket: HTTP hop, auth middleware, findOne and JWT signing, all on Node.
  • Signaling: the driver’s call-offer emit to call-answer received.

Across all nine levels the 8-15x gap is gone: gosignal is 1.6-1.9x faster at p50 from 50 setups up and 1.1-1.5x at p95. Its time is the ticket round trip, 99%+ of its total at every level. Node’s time is the Mongo lookup, 84% of its setup at 400 (309 of 366ms). Both paths run a membership query per call; the ticket path runs it in an HTTP request to the same Node process.

Was the port worth it?

At 25 concurrent setups the driver sees 65ms on Node vs 55ms on gosignal. At 400 it is 427 vs 264ms. A user wouldn’t notice either next to ICE gathering (network route discovery for the media), which takes seconds. On latency alone the rewrite doesn’t pay.

Where it pays is a signaling step that stays flat under load, because it runs in a separate process: from 25 to 400 setups gosignal’s signaling step stayed under 1ms at p50, while Node’s dispatch climbed from 14 to 50ms, consistent with a single event loop queuing work, the mechanism the scale post measured.

That fits the Go runtime: each WebSocket connection runs in a separate goroutine, scheduled across every core by the runtime.

I can’t tell from this run whether the Go runtime keeps that 0.5ms from growing with load: the port saw a lighter load than Node, and the driver’s timer can’t resolve a step that small (details under Limitations).

Limitations

The 0.5ms was measured under a lighter load than Node’s. The driver fetched a ticket before each offer to gosignal, so those offers arrived at different times; Node’s all arrived at once. The driver’s event-loop lag reaches 32-95ms (p99), too much to time a 0.5ms step. The comparable figure is the ticket-plus-signaling total (264ms at 400 setups vs Node’s 427ms), since both include the membership query. The lag adds time to both totals, so the true ratios are probably larger than shown (untested).

gosignal’s session store takes the same sync.RWMutex for all calls: a write to any call blocks the rest.

Sharding it into fixed buckets keyed by a hash of the conversation ID would stop that. At level 400 this run shows no sign that it is needed: the signaling step stays under 1ms at p50 across the ramp, and its p95 reaches 4-6ms, which a driver with 32-95ms of event-loop delay can’t resolve.

Cross-instance delivery through Redis is correctness-tested, not perf-tested. scaleout_test.go runs two Server processes that share only Redis, connects the offerer to one and the callee to the other, and confirms the answer arrives. That proves the hand-built broker delivers, the same guarantee Node gets from @socket.io/redis-adapter. It says nothing about latency or throughput: every number in this post is single-instance. Before scale-out is enabled, gosignal needs a way to detect a failed instance and a benchmark of broker latency under a burst (how long a message takes to cross Redis between two instances when many offers arrive at once).

Where the rollout stands

Node’s FSM went to production first, as a protocol change on the same runtime. The port followed as a runtime change behind the same protocol, on a single instance. Deleting Node’s 1:1 handlers comes next.