TL;DR: I built a Go call-transcriber for Jitsi group calls: audio in from JVB, transcription via Deepgram, storage in MongoDB, events out on NATS.
Five protocol-level bugs shaped how it works.
Studying opus-transcriber-proxy
opus-transcriber-proxy is Jitsi’s official transcription proxy: multi-provider (OpenAI, Deepgram, Gemini, xAI), built-in failover, translation. I studied it as a reference for how JVB’s transcription protocol works (a WebSocket JVB opens to stream conference audio out to a listener), and how a production relay handles multi-provider routing. Then I built a smaller version.
Its code revealed the protocol. More importantly, it showed me what this call flow actually needs: one provider, plus the app integration on top of it (correlating a transcript to the right conversation, since Jicofo’s MEETING_ID is ephemeral and changes every reconnect; attributing each segment to a speaker as participants join and leave; storing all of it).
Reading opus-transcriber-proxy’s code settled a design question before writing any Go: which STT provider needs which audio format.
Deepgram accepts raw Opus passthrough; OpenAI, Gemini, and xAI all need decoded PCM. The reference implementation needs a native C/C++ Opus addon for that path. So I went with Deepgram: no decoding needed.
Two other things carried over: the multiplexing pattern (one connection, tagged streams, described below), and how sessionId and tag are used.
sessionId identifies which conference the connection belongs to; tag identifies which participant within it. The reference implementation treats both as opaque query params (it doesn’t interpret their contents, just passes them through). It doesn’t specify how sessionId is assigned, just that the caller provides one. Jicofo fills it in, with the conference’s MEETING_ID.
Storing transcripts tied to a conversation needs a stable identity across reconnects. MEETING_ID changes on every reconnect. That’s the problem a later section covers.
Architecture
JVB dials in as the client, and the transcriber runs an HTTP server waiting for it: ws://agent:8090/transcribe?sessionId={{MEETING_ID}}&sendBack=true.
JVB also retries a dropped connection indefinitely, with exponential backoff (0/1/2/4/8/16s, capped at 30s), for as long as the room’s metadata has transcription enabled. The transcriber has no way to tell Prosody/Jicofo to turn that off.
To bound STT costs (Deepgram bills by the minute), the transcriber caps STT usage per session, configurable per deployment. To avoid reconnect storms, it doesn’t reject JVB’s connections. Instead, it tracks cumulative transcribing time per sessionId across reconnects and stops dialing Deepgram once the budget is exhausted:
func (s *Server) sessionElapsed(sessionID string) time.Duration {
start, ok := s.sessionStartAt[sessionID]
if !ok {
start = time.Now()
s.sessionStartAt[sessionID] = start
}
return time.Since(start)
}
The transcriber turns unreliable signals from JVB and Deepgram into stored, queryable state.
Internally, one WebSocket connection carries every participant’s audio for a conference. JVB multiplexes each speaker’s stream onto it, distinguished by a tag field, rather than opening one connection per participant. The session’s read loop dispatches on message type:
switch env.Event {
case "media":
sess.handleMedia(raw, logger) // base64 Opus payload, tagged per participant
case "start":
// a new participant stream began
}
Each media event carries base64-encoded Opus for a tag. The session keeps a map[string]Provider keyed per participant (not per conference), so handleMedia decodes the payload and forwards it to that tag’s stream.
Five bugs
Each of these was silent. The system looked healthy until I measured.
Deepgram’s KeepAlive timeout (30–64 seconds of speech)
Transcription would cut off mid-call after ~30–64 seconds of speech, even though the connection looked healthy. Deepgram closes idle WebSocket connections after ~30 seconds of no activity, and natural speech has pauses. Without a KeepAlive message, the connection died during that kind of gap.
A goroutine now sends Deepgram’s documented KeepAlive message on a fixed ticker, whether or not audio is flowing (I found this requirement by reading the real-time streaming guide, not the REST API docs):
func (p *deepgramProvider) keepAliveLoop() {
ticker := time.NewTicker(keepAliveInterval)
defer ticker.Stop()
for {
select {
case <-ticker.C:
p.writeMu.Lock()
err := p.conn.WriteJSON(map[string]string{"type": "KeepAlive"})
p.writeMu.Unlock()
if err != nil {
return
}
case <-p.stopKeepAlive:
return
case <-p.done:
return
}
}
}
The keepAlive loop needs its own goroutine so it sends KeepAlive messages on a fixed schedule, independent of audio flow. If it were serialized with SendAudio, a pause in incoming frames could delay the KeepAlive and trigger Deepgram’s idle timeout.
This loop and SendAudio both write to the same WebSocket connection from different goroutines, and gorilla/websocket only allows one writer at a time, so writeMu guards every write.
Two cancellation channels (stopKeepAlive, done) let the loop exit cleanly instead of leaking a goroutine after the session ends.
Deepgram’s 30-second idle timeout isn’t a bug. It’s a constraint to design around.
JVB’s WebSocket ping/pong ID echo
The WebSocket connection to JVB was reconnecting every ~13 seconds, even though data flowed fine and there were no error messages. JVB sends periodic WebSocket ping frames, and the protocol expects the client to echo back the exact id from the ping message. Replies were going out correctly, but with a different id, so JVB never saw the pong. After 13 seconds of unanswered pongs, it closed the connection.
Extracting the id field from JVB’s ping message and echoing it back in the pong fixed it:
if msg.Type == "ping" {
pong := map[string]interface{}{
"type": "pong",
"id": msg.ID, // ← This was missing
}
ws.WriteJSON(pong)
}
The WebSocket RFC requires pong responses to carry identical payload to the ping, but the transcriber was sending a pong with a mismatched id.
Not in the spec, there in the reference implementation’s code.
Missing audio parameters (0-duration transcriptions)
Deepgram connections would “succeed” (the handshake completed, no errors logged), but transcription would silently produce 0 seconds of audio processed, with empty transcripts. Deepgram’s WebSocket endpoint for raw Opus accepts ?encoding=opus&channels=2&sample_rate=48000, and the connection request sent encoding=opus but omitted channels and sample_rate. Deepgram accepted the connection with no error, but couldn’t decode the frames without knowing the channel count and sample rate, so it silently processed 0 seconds.
Including all required parameters solved it:
q.Set("channels", "2")
q.Set("sample_rate", "48000")
q.Set("encoding", "opus")
This is documented: “Raw Opus doesn’t self-describe these values.” I caught it by observing the 0-duration metadata response. Zero-duration transcription is harder to debug than a connection error, since the system looks fine while producing no output.
Unhandled WebSocket errors
The transcriber would dial Deepgram, start receiving frames, then suddenly stop transcribing. No error in the logs, just silence, and no reconnect logic triggered. Deepgram sends error messages over the WebSocket as JSON {"type": "error", "message": "..."}, but only two message types were being read: {"type": "transcript_started", ...} and {"type": "results", ...}. When Deepgram sent an error (e.g., invalid API key format, auth failure), it went unhandled and the connection just hung.
Reading all message types and surfacing errors stopped it:
for {
var msg map[string]interface{}
ws.ReadJSON(&msg)
msgType, ok := msg["type"].(string)
if !ok {
continue
}
switch msgType {
case "error":
return fmt.Errorf("deepgram error: %v", msg["message"])
case "results":
handleTranscript(msg)
}
}
Like the missing audio parameters, this took longer to catch. The system just quietly did less than it should.
Reconnect storm (hundreds of attempts per second)
When Deepgram closed the connection for any reason, the transcriber would immediately try to reconnect, fail, retry, fail, retry, hundreds of attempts per second, saturating the connection pool until it couldn’t recover. The reconnect logic had no backoff, so every failed dial attempt immediately triggered the next attempt.
Exponential backoff with jitter solved it. Exponential backoff: each failed reconnect doubles the wait time (100ms, 200ms, 400ms…), giving the service time to recover. Jitter adds randomness (±25% variance) so multiple clients don’t all retry at the exact same moment and hammer the service again:
backoff := time.Duration(math.Pow(2, float64(attempts))) * 100 * time.Millisecond
backoff = backoff + time.Duration(rand.Int63n(int64(backoff/4)))
time.Sleep(backoff)
Without exponential backoff, the transcriber would hammer Deepgram hundreds of times per second instead of backing off and letting it recover.
Session correlation: the ephemeral MEETING_ID problem
Jicofo only knows its MEETING_ID, not one-chat’s conversationId, so passing the conversation ID isn’t an option. Jicofo generates that MEETING_ID for each conference, but it’s ephemeral. If a participant reconnects, Jicofo assigns a different MEETING_ID. This breaks correlation: I couldn’t query “all transcripts for conversation X” because each reconnect created a new “conversation” to the transcriber.
Querying JVB’s /debug endpoint resolves the stable Room ID, cached since the mapping is static for the life of a conference:
func (r *ConversationIDResolver) Resolve(meetingID string) (string, bool) {
resp, err := r.client.Get(r.baseURL + "/debug")
// ...
for _, conf := range body.Conferences {
if conf.MeetingID != meetingID {
continue
}
// name is a room JID like "<conversationId>@muc.meet.example.com"
conversationID, _, ok := strings.Cut(conf.Name, "@")
if !ok || conversationID == "" {
return "", false
}
r.cache[meetingID] = conversationID
return conversationID, true
}
return "", false
}
JVB’s /debug response only exposes the conversation ID inside a room JID (<conversationId>@muc.meet.example.com), so the identifier is the local part before the @.
The MEETING_ID is resolved once per conference; everything stored after that is keyed on the conversation ID.
Speaker attribution
Once transcripts are stored, the next challenge is speaker attribution. Jitsi’s participant metadata (usernames, IDs) is in-memory in JVB. Deepgram’s endpointId (which participant is speaking) is ephemeral.
A new MongoDB collection, CallParticipant, tracks participant lifecycle:
{
conversationId: ObjectId,
callId: string,
userId: ObjectId,
participantId: string, // JVB's participant ID
joinedAt: Date,
leftAt: Date,
}
When a transcript segment arrives, look up the participantId in this collection to find the userId. Store it with every segment.
Then publish to NATS:
{
"type": "transcription.segment",
"conversationId": "...",
"userId": "...",
"text": "...",
"timestamp": 1695412567890,
"isFinal": true
}
Downstream consumers (frontends, analytics, summarization workflows) subscribe to this event and build their views.
Publishing events decouples the source of truth (the transcriber) from its consumers (UI, analytics, workflows). Each one reacts on its schedule, without the transcriber needing to know who’s listening.
Takeaways
What I learned building this:
Protocol layers are opinionated. Deepgram, JVB, and NATS each have hidden assumptions about timing, message format, and error handling. Integrating with each one, then seeing where it broke, surfaced them.
Silent failures are difficult to catch. All five bugs above were silent: nothing broke or logged an error. I noticed them by timing and counting.
Distributed systems are about state reconciliation. The transcriber’s job is not “call Deepgram,” it’s “maintain consistent state about who’s speaking, who’s listening, what was said, and where we are in the process.” Deepgram/Mongo/NATS are just the tools for that. The MEETING_ID problem above is a version of this: JVB’s ephemeral session identity has to line up with the conversation’s stable identity.
Timeouts are constraints, not bugs. KeepAlive, reconnect backoff, session TTLs: these aren’t edge cases, they’re the normal operating mode. They’re built into the operating mode; the transcriber works within them, not around them.
What’s next
It’s deployed and running.
Segments already publish to NATS as they’re finalized, but nothing subscribes yet. The frontend gets its live updates from a MongoDB Change Stream instead. Next is wiring NATS consumers so they can react to those events.