The earlier posts (scale and liveness) measured 1:1 call signaling on Node.
Each call runs a state machine (FSM):
IDLE → OFFERING → CONNECTED → RENEGOTIATING, plus AWAITING_REJOIN and
ENDED. I wrote a minimal Go version of that state machine (gosignal) to benchmark it next to Node, and completed the port afterwards.
The prototype benchmark looked lopsided. At 400 simultaneous calls the load driver
measured a median offer-to-answer time of 446ms on Node and 37ms on gosignal,
about 12x lower. Then I replayed Node’s lifecycle tests on the port,
and three gaps fell out that the benchmark had no way to see: a rejoin that reversed the caller and callee roles, a second offer that overwrote a live call,
and renegotiation accepted from any state.
The 12x also shrank once both sides were timed the same way. gosignal needs a per-call ticket: a signed pass that Node issues after a Mongo membership query, which Node also runs on every offer. The driver fetched the ticket before
starting the clock, so the port’s number left that query out and Node’s included it.
Timed the same way on both sides, the gap at 400 setups is 427ms vs 264ms at
p50, 1.6x, and nearly all of the 264ms is the ticket request to Node. The signaling step itself takes about 0.5ms. The case here is isolation, plus a signaling step that stayed under 1ms across the ramp,
not a latency win a user would notice.
The post covers the prototype benchmark, the port and the gaps Node’s tests found, and the rerun that cut 12x to 1.6x.
Benchmark setup
The prototype gosignal was the smallest slice a benchmark needed: the
state machine, an in-memory session store and the wire events, over raw
WebSocket (gorilla/websocket) instead of Socket.IO. Node looks up the
caller’s membership in Mongo on every call-offer.
The prototype had no database, so a bench-only route on Node ran the same query and issued a ticket.
The driver fetched that ticket before starting the stopwatch, and gosignal verified it on
connect. Node’s timing included the membership query; the port’s did not. I knew
that going in: the benchmark was there to time the state machine itself.
I reused the driver from the scale post:
backend/test/load/, a Node client that plays both sides of N simulated calls,
extended with a raw-WebSocket mode for the port. The load ran on one box (8
dedicated vCPU). A level is N established calls: all sessions connect, N offers
go out at once, the answers come back, then ICE candidates flow. Levels run from
25 to 400, 3 trials each, with a fresh server process per run. The prototype run had three targets:
node-http (the production baseline), node-uws (the same Socket.IO handlers on uWebSockets.js, a transport-only change tried because the earlier posts traced Node’s limit to its event loop) and gosignal over raw WebSocket.
Prototype results
The prototype run, zero errors across 81 runs, median of 3 trials (driver-side p50 / p95, ms):
| level | node-http | node-uws | gosignal |
|---|---|---|---|
| 25 | 53 / 65 | 62 / 63 | 6 / 7 |
| 100 | 161 / 172 | 159 / 161 | 11 / 12 |
| 400 | 446 / 504 | 439 / 461 | 37 / 40 |
gosignal came out 8-15x lower at every level, and node-uws stayed within
about 5% of node-http from 100 calls up. That is the 12x from the opening: a minimal prototype with its ticket fetch off the clock, next to Node with the membership query on it.
A 12x lead was enough to finish the port.
A faster WebSocket layer barely moved Node’s latency
node-uws was the smallest change on the list: uWebSockets.js is a faster WebSocket implementation, and Node’s Socket.IO stack attaches to it with one call (io.attachApp(app) on a uWS.App), with the signaling handlers called unchanged. If the WebSocket layer was costing time, it would show
here. It didn’t: from 100 concurrent calls up, node-uws landed within about
5% of node-http on driver-side p50 (439ms vs 446ms at 400). I had
expected more. The transport isn’t where the time goes, so I dropped it and the
rewrite became the next candidate.
Two caveats on the prototype numbers
The driver’s headline number comes from a stopwatch: start a timer before
sending call-offer, stop it when call-answer arrives. That works when
latencies run in the hundreds of milliseconds, like Node’s.
It breaks for gosignal, because the driver is one Node.js process playing
every simulated caller and callee inside a single event loop. When that loop is
busy, the driver reads the clock late, so a latency it reports can include time
spent waiting for its own turn. The harness measures this with Node’s monitorEventLoopDelay (a histogram of
how late a 20ms timer runs) and reports the 99th percentile per
level as drv_el_p99_ms: 1% of the time, a callback waited at least that long.
| target | driver lag (p99) | measured latency (p50, 25 to 400 setups) |
|---|---|---|
node-http | 50-60ms at every level | 53-446ms |
gosignal | 80-133ms from 150 setups up | 6-38ms |
For Node the lag is small next to the latency. For gosignal it is larger than
the latency the stopwatch was timing, so its driver-side numbers are upper
bounds.
The server also times each offer from storing it to relaying the answer, where
the driver’s lag cannot reach. Node already recorded that span in an OTel
histogram, signaling.offer_answer.duration; the port now records the same histogram.
The histogram never resets between levels, which flattered the level-400 numbers until I started restarting the server for every run.
Porting the state machine
Node’s state machine is the spec. gosignal/internal/fsm ports
backend/signaling/fsm.js: the six states (IDLE to ENDED) and the
transitions between them. Node’s lifecycle tests in backend/test/signaling/ serve as the acceptance
test for the port; I replay them on gosignal below.
With the port in place, the browser keeps its Socket.IO connection to Node for
chat and group calls, and opens another WebSocket connection to gosignal
for 1:1 call events. Node stays the front door for auth: it hands the browser a
token, and gosignal verifies it without querying Mongo.
Replaying Node’s tests
The prototype benchmark exercised one path: offer, answer, then a burst of made-up
ICE candidates once the call was up. It left out authorization and every edge
case, such as a peer dropping, ICE candidates arriving early, or a second offer.
So the benchmark showed how fast gosignal is on the happy path (every
message valid and in order, every peer staying connected), not whether it
handles the edge cases the way Node does.
The session-existence guard was the exception: while smoke-testing the rig I had
noted that handleCallOffer lacked it, and the benchmark never sent a second
offer to expose it.
Next I replayed Node’s tests. Its lifecycle and characterization suites
gave 15 scenarios, ported into gosignal/internal/server/parity_test.go.
Three gaps fell out, each closed with a test.
A rejoin that reversed the roles
When a peer drops, the call parks in AWAITING_REJOIN. The dropped peer
reconnects, sends a rejoin offer, and the survivor answers it. CompleteRejoin
recorded whoever answered as the answerer. That holds when the offerer drops,
since the original answerer survives and answers. When the answerer dropped,
the surviving offerer answered the rejoin, got recorded as the answerer, and
the roles reversed. Node keeps the roles and refreshes only the answering
socket, so the change does the same:
switch answeringUserID {
case f.session.OffererUserID:
f.session.OffererSocketID = answeringSocketID
case f.session.AnswererUserID:
f.session.AnswererSocketID = answeringSocketID
}
Both end-to-end rejoin tests now assert that the roles are unchanged after recovery.
A second offer overwrote a live session
handleCallOffer rejected offers from outside the pair and nothing else, so a
participant’s second call-offer while the call was live or recovering
replaced the session. Node rejects both cases (“already in progress”,
“recovering, rejoin”); the change rejects them too, as soon as a session
exists:
if existing := s.store.Get(req.ConversationID); existing != nil {
if !isParticipant(existing, c.userID) {
return errNotParticipant
}
if existing.State == fsm.AwaitingRejoin {
return errCallRecovering
}
return errCallInProgress
}
Renegotiation from any state
Node’s handler allows renegotiation only from CONNECTED. The port’s handler
verified participation, then handed the session to the FSM without reading
its state, so a renegotiation offer was accepted from any state, including
before the call had connected. The FSM’s tiebreak for simultaneous offers
(glare: both peers renegotiating at once) exists for CONNECTED and was never
reached through this path. The patch adds Node’s guard:
if sess.State != fsm.Connected {
return fmt.Errorf("cannot renegotiate from state %s", sess.State)
}
Latency never showed any of them; they change behavior, not speed, and no load
test could have found them. Before I read another benchmark number from
gosignal, it has to pass Node’s tests.
What the benchmark port skipped
Replaying Node’s tests covered behavior. Four other parts of the prototype were shortcuts a live call can’t use: a token scoped to one conversation, ICE sent only after the call was up, no acks (replies confirming a message arrived), and a load driver in place of the real client.
The token was scoped to a conversation, not a user
In the benchmark, Node minted a token scoped to one conversation and
gosignal verified it on call-offer.
That was enough for the benchmark, not for a live call. The frontend delivers the incoming-call toast to an idle user over one global socket, which a per-conversation token can’t serve.
So auth split in two. On connect, a short-lived user token carries only the user
id, so an idle callee can receive a call. Each call-offer carries a per-call
ticket: Node queries membership once and signs the conversation and its
participants, and gosignal verifies the signature offline, with no query.
The participants come from the signed ticket, not from the request. A test has
Node’s jsonwebtoken sign a token and gosignal verify it.
I didn’t reuse the login JWT. A browser can’t send an Authorization header on
a WebSocket, so the token rides in the URL, and nginx’s access logs would record
a one-hour token that works on the whole API.
A user-scoped connection can send events for any conversation: gosignal
compares the sender of each call event with the two participants in the ticket.
At offer time the session records who may answer. Outsiders cannot answer, end the call or send it ICE candidates. Each guard has a test.
ICE candidates, lost twice
The offerer starts gathering ICE candidates (network routes for the media) as
soon as it calls setLocalDescription, before the callee answers. Node buffers them
on the session and returns them in the call-answer ack. The prototype gosignal
aimed them at the still-empty answerer socket and dropped them, and the
benchmark never noticed because it never sent ICE candidates from a browser. The
change buffers on the session, holds the offerer’s candidates until the callee
answers, clears them on rejoin as Node does, and routes by caller identity
instead of the client’s didIOffer flag. Requests carrying a reqId get an
ack, and call-answer’s ack is the buffered list, or [] plus an error
event on failure, so a waiting client never hangs.
The other loss was in the browser. The caller’s code ran setLocalDescription,
which starts candidate gathering, then fetched the ticket over HTTP, then sent
call-offer, which is what creates the session. Candidates were already flowing
during the ticket fetch; the early ones reached gosignal before the session
existed and were silently dropped, and the callee received fewer routes than the
caller found. Node drops them too, but it has no ticket fetch to open that
window. A run with two Chromium contexts showed it. The change
fetches the ticket before creating the offer, so the window is no longer than
Node’s.
A client that looks like Socket.IO
To keep the frontend diff small, SignalingClient exposes the same methods as
Socket.IO (on, off, removeAllListeners, emit, emitWithAck) over the
plain WebSocket. The 1:1 call code switches from socket
to signaling, and chat and group calls stay on Socket.IO.
No heartbeat, no retry
gorilla/websocket has no heartbeat. A peer that dropped without sending a
close frame, such as a lost network or a sleeping laptop, left its read
goroutine (a lightweight thread) blocked forever. conn.pingLoop adds a read deadline plus a ping
ticker (60s deadline, 54s ping). The frontend client had no retry at all,
so SignalingClient got exponential backoff (1s growing to 5s, 5 attempts) on
the same instance, which keeps its listeners alive across a drop.
Scale-out: built, not enabled
gosignal ships as a single instance. The Redis path is built and tested across
two instances, but I haven’t enabled it in production or measured it, and a
piece is missing before it can run on several instances (the last
paragraph of this section).
The release runs as a single instance, which is also what the benchmarks measured. A crashed instance can strand its calls on several, with nothing yet to detect it.
RedisStore ports Node’s key layout for sessions. Node delivers events across
instances through Socket.IO’s Redis adapter, which is part of Socket.IO: it
plugs into Socket.IO’s rooms and sends Socket.IO packets over Redis. gosignal
speaks raw WebSocket with a JSON envelope, so there was nothing to plug in, and
I wrote delivery and presence on Redis directly.
Every push goes through deliverToUser. It tries the local connection registry,
which costs no network hop, and publishes to Redis only when the user
isn’t connected to this instance. Each instance also runs a subscriber, which
delivers a message when the user it is addressed to is connected to that
instance:
func (s *Server) deliverToUser(userID, event string, data any) error {
if peer, ok := s.conns.byUserID(userID); ok {
return peer.send(event, data)
}
if err := s.broker.Publish(context.Background(), userID, event, data); err != nil {
s.log.Warn("broker publish failed", "event", event, "err", err)
}
return nil
}
func (s *Server) StartBroker(ctx context.Context) {
s.broker.Subscribe(ctx, func(userID, event string, data json.RawMessage) {
if peer, ok := s.conns.byUserID(userID); ok {
_ = peer.send(event, data)
}
})
}
All instances share a Redis channel, sig:deliver. The published message
carries the userId, the event and its data. Every instance receives it,
the instance holding that user’s socket delivers it, and the others ignore it.
A failed publish is logged and not returned, and the subscriber resubscribes
after any Redis error, so a dropped connection doesn’t stop delivery for the
life of the process.
Node’s adapter solves the same two problems, delivery and presence. Delivery works the same way in both. The adapter publishes each broadcast to a channel named after its room, but each instance subscribes with a single pattern that covers all of them, so it receives every message too:
// @socket.io/redis-adapter 8.3.0, dist/index.js (trimmed), on broadcast
let channel = this.channel; // "socket.io#/#"
if (opts.rooms && opts.rooms.size === 1) {
channel += opts.rooms.keys().next().value + "#";
}
this.pubClient.publish(channel, msg);
// on startup, a single pattern covers every room's channel
this.subClient.pSubscribe(this.channel + "*", ...);
Presence works differently. disconnectSeat decides between
recovering and ending a call based on whether the peer is still connected, and
it originally consulted only the local registry. A peer on another instance looked
gone, so the call ended instead of waiting in AWAITING_REJOIN. Node’s
isConnected() answers this question with the adapter’s fetchSockets(), which
queries every instance and waits for the replies:
const sockets = await this.socketServer.in(socketId).fetchSockets();
return sockets.length > 0;
gosignal uses a Redis key instead. Each connection sets sig:online:<userId>
with a 2-minute TTL, and the ping loop that already runs renews it on every ping.
Pings go out every 54 seconds, 90% of the 60-second read deadline, so a single missed
renewal still leaves the key alive:
func (b *Redis) MarkOnline(ctx context.Context, userID string) error {
return b.rdb.Set(ctx, onlineKey(userID), "1", onlineTTL).Err() // 2 minutes
}
func (s *Server) isUserReachable(userID string) bool {
if _, ok := s.conns.byUserID(userID); ok {
return true
}
return s.broker.IsOnline(context.Background(), userID)
}
fetchSockets() comes from Socket.IO’s adapter, and plain WebSocket has no
cluster-wide socket query. A sweep would need a new request and reply protocol
over the broker, so I used the key: the ping loop already runs and can renew it.
It was less work to build, and it can be replaced later: disconnectSeat only
asks the broker IsOnline(userID), so a query across instances can answer the
same call.
Two-instance tests cover offer, answer, end and disconnect recovery across processes.
The key covers a peer disconnecting, not an instance crashing. The reaper from the
liveness post, which finds calls stranded
when an instance crashes, was not ported, so gosignal cannot detect that case today.
A crashed instance’s keys also keep its users marked online for up to 2 minutes,
until the TTL expires. That is why scale-out stays off until gosignal can
detect a failed instance.
Rerun: timing the ticket too
The prototype run had an asymmetry I knew about. Node queries the caller’s membership inside the handler, so the query was inside its timed span. gosignal has no Mongo: Node mints a token (later, a per-call ticket) after the same query, and the driver fetched it before the timer started. Same query, same once-per-call frequency, opposite sides of the clock.
The rerun moves the ticket request inside the window. The driver also stores every duration, which gives exact percentiles:
const t0 = performance.now();
offerPayload.ticket = await s.off.fetchTicket(cid);
const t1 = performance.now(); // ticket round trip = t1 - t0
s.off.emit("call-offer", offerPayload);
// ...answer comes back at t2
metrics.offerAnswer.push(t2 - t0); // the comparable number
metrics.signaling.push(t2 - t1); // gosignal's part
Two targets (node-http and gosignal with the ticket gate on), nine levels
from 25 to 400 setups, median of 3 trials per level (driver-side p50 / p95, ms;
four levels shown). At 25 the node-http median has 2 trials, since a trial timed
out and I dropped it:
| level | node-http | gosignal | ratio p50 / p95 |
|---|---|---|---|
| 25 | 65 / 73 | 55 / 65 | 1.2x / 1.1x |
| 100 | 158 / 163 | 89 / 116 | 1.8x / 1.4x |
| 200 | 260 / 275 | 138 / 219 | 1.9x / 1.3x |
| 400 | 427 / 485 | 264 / 412 | 1.6x / 1.2x |
Where the time goes at 400 setups (p50 / p95, ms). gosignal makes no database call
in its handler: the membership query runs in a separate ticket request, which
the driver has to time.
| piece | clock | p50 / p95 |
|---|---|---|
node-http Mongo lookup | server | 309 / 374 |
node-http setup | server | 366 / 402 |
node-http dispatch | server | 50 / 64 |
gosignal ticket round trip | driver | 263 / 412 |
gosignal signaling | driver | 0.5 / 4.0 |
- Mongo lookup: the
findOnepluspopulateinhandleCallOffer, as the handler sees it, so event-loop wait is included. - Setup:
handleCallOfferentry to the answer relayed to the offerer; it contains the query. - Dispatch: offer stored to answer relayed. It includes the answerer’s turnaround through the shared driver, so it is not pure server work.
- Ticket round trip: the driver’s
GETfor the call ticket: HTTP hop, auth middleware,findOneand JWT signing, all on Node. - Signaling: the driver’s
call-offeremit tocall-answerreceived.
Across all nine levels the 8-15x gap is gone: gosignal is 1.6-1.9x faster at
p50 from 50 setups up and 1.1-1.5x at p95. Its time is the ticket round trip,
99%+ of its total at every level. Node’s time is the Mongo lookup, 84% of its
setup at 400 (309 of 366ms). Both paths run a membership query per call; the
ticket path runs it in an HTTP request to the same Node process.
Was the port worth it?
At 25 concurrent setups the driver sees 65ms on Node vs 55ms on gosignal.
At 400 it is 427 vs 264ms. A user wouldn’t notice either next to ICE
gathering (network route discovery for the media), which takes seconds. On
latency alone the rewrite doesn’t pay.
Where it pays is a signaling step that stays flat under load, because it runs
in a separate process: from 25 to 400 setups gosignal’s signaling step stayed
under 1ms at p50, while Node’s dispatch climbed from 14 to 50ms, consistent
with a single event loop queuing work, the mechanism
the scale post
measured.
That fits the Go runtime: each WebSocket connection runs in a separate goroutine, scheduled across every core by the runtime.
I can’t tell from this run whether the Go runtime keeps that 0.5ms from growing with load: the port saw a lighter load than Node, and the driver’s timer can’t resolve a step that small (details under Limitations).
Limitations
The 0.5ms was measured under a lighter load than Node’s. The driver
fetched a ticket before each offer to gosignal, so those offers arrived at
different times; Node’s all arrived at once. The driver’s event-loop lag
reaches 32-95ms (p99), too much to time a 0.5ms step. The comparable figure is
the ticket-plus-signaling total (264ms at 400 setups vs Node’s 427ms), since
both include the membership query. The lag adds time to both totals, so the true
ratios are probably larger than shown (untested).
gosignal’s session store takes the same sync.RWMutex for all calls: a write
to any call blocks the rest.
Sharding it into fixed buckets keyed by a hash of the conversation ID would stop that. At level 400 this run shows no sign that it is needed: the signaling step stays under 1ms at p50 across the ramp, and its p95 reaches 4-6ms, which a driver with 32-95ms of event-loop delay can’t resolve.
Cross-instance delivery through Redis is correctness-tested, not perf-tested.
scaleout_test.go runs two Server processes that share only Redis, connects
the offerer to one and the callee to the other, and confirms the answer arrives.
That proves the hand-built broker delivers, the same guarantee Node gets from
@socket.io/redis-adapter. It says nothing about latency or throughput: every
number in this post is single-instance. Before scale-out is enabled, gosignal
needs a way to detect a failed instance and a benchmark of broker latency under
a burst (how long a message takes to cross Redis between two instances when many
offers arrive at once).
Where the rollout stands
Node’s FSM went to production first, as a protocol change on the same runtime. The port followed as a runtime change behind the same protocol, on a single instance. Deleting Node’s 1:1 handlers comes next.