How do I set up Jitsi bridge-based transcription?

Short answer

Jitsi bridge-based transcription sends participant audio from JVB to opus-transcriber-proxy, which calls a speech-to-text backend and returns captions. The stable-11248 components support this path, but Docker still needs a custom Jicofo configuration and a Prosody module. Set the server-controlled asyncTranscription flag and let a moderator start transcription through the meeting UI. A cloud backend receives meeting audio even when the proxy runs on your server. [S1][S4][S5][S6][S15]

Who this is for

The handbook deprecates Jigasi transcription and recommends the bridge path for new captions deployments. For the older offline architecture, keep the Jigasi and Vosk guide. [S1][S18]

How it works

Jicofo reads room metadata and gives one selected JVB a WebSocket URL through Colibri2. JVB exports each participant’s Opus audio with source identity over a conference WebSocket. The proxy opens backend connections for the separate audio streams and returns transcription results to the conference. There is no Jigasi participant dialing into the room. [S4][S6][S15]

Two flags must be true: asyncTranscription and recording.isTranscribingEnabled. The first is server-controlled; browsers cannot set it. The module below makes rooms eligible for this path, while the moderator UI controls the second flag. It does not automatically transcribe every room. [S4][S6]

Before you start

Latest Docker release checked 2026-10-06: stable-11248. Use its matching stack. This source baseline does not establish the earliest compatible release. [S3]

Component Covered baseline
Jitsi Meet 2.0.11248-1 [S9]
Jicofo 1.0-1205-1 [S9]
JVB 2.3-318-gbf271b11f-1 [S9]
Meet web and Prosody plugins 1.0.9442-1 [S9]
Prosody server Docker build selects 13.*; package runtime is distribution-dependent, with a dependency accepting prosody >= 0.12.0 or listed alternatives. This is separate from the plugin version. [S3][S9]

Have working meetings, moderator access and a backend account. The example selects Deepgram Nova-3 for English. JVB must reach proxy TCP port 8080 internally; the proxy needs outbound access to its selected provider. Remote bridges need authenticated WSS and a reachable TLS endpoint. [S1][S2][S15]

Budget by summed billable participant audio: ten participants each supplying one billable hour produce ten audio hours. Attendance alone does not establish usage. Prices are USD, checked 2026-10-06; Deepgram rates are PAYG streaming. Hosting, egress and translation are excluded. [S10][S11][S12][S13][S15]

Backend One audio hour Ten audio hours
OpenAI gpt-4o-mini-transcribe about $0.18 about $1.80 [S10]
OpenAI gpt-4o-transcribe about $0.36 about $3.60 [S10]
Deepgram Nova-3 monolingual $0.288 promotional, $0.462 regular $2.88 promotional, $4.62 regular [S11]
xAI streaming STT $0.20 $2.00 [S13]
Gemini 3.5 Transcribe Live about $0.54, reference estimate about $5.40, compatibility with this proxy unconfirmed [S12]

OpenAI’s per-minute figures are estimates for token billing. Deepgram’s multilingual promotion is $0.0058/minute, versus $0.0048 for monolingual; these describe language mode, not audio channels. Gemini’s current model is not a verified substitution for the proxy’s old default. [S10][S11][S12][S15]

Steps

  1. Docker stable-11248: add the proxy. Merge this new service into compose.override.yaml beside docker-compose.yml. The inspected official image has amd64 and arm64 builds. It listens internally, with no published host port. Keep any other Compose overlays in your usual command list. [S2][S5][S16]

    YAML
    services:
      opus-transcriber-proxy:
        image: jitsi/opus-transcriber-proxy:6747f187c0e7f4fa17f7744e0d76a5ce6d7d4b24
        restart: unless-stopped
        environment:
          - PROVIDERS_PRIORITY
          - DEEPGRAM_API_KEY
          - DEEPGRAM_MODEL
          - DEEPGRAM_LANGUAGE
          - DEEPGRAM_MIP_OPT_OUT
        networks:
          meet.jitsi:

    Edit existing .env assignments. Replace the API-key placeholder locally. Append the module to existing XMPP_MUC_MODULES. Ordinary Jicofo and JVB passwords remain required; no Jigasi transcriber password is needed. [S2][S5]

    dotenv
    ENABLE_TRANSCRIPTIONS=1
    PROVIDERS_PRIORITY=deepgram
    DEEPGRAM_API_KEY=REPLACE_WITH_YOUR_API_KEY
    DEEPGRAM_MODEL=nova-3
    DEEPGRAM_LANGUAGE=en
    DEEPGRAM_MIP_OPT_OUT=true
    XMPP_MUC_MODULES=force_async_transcription

    Keep .env out of Git and private: [S8][S17]

    Terminal
    chmod 600 .env
  2. Docker: configure Jicofo and room metadata. Set shell CONFIG to match .env, using its documented default below. Merge the HOCON into ${CONFIG}/jicofo/custom-jicofo.conf. Startup includes it after generated configuration. JVB initiates the connection. [S4][S5][S8]

    Terminal
    CONFIG="$HOME/.jitsi-meet-cfg"
    Text
    jicofo.transcription {
      url-template = "ws://opus-transcriber-proxy:8080/transcribe?sessionId={{MEETING_ID}}&sendBack=true"
      http-headers {}
      ping {
        enabled = true
        interval = 10 seconds
        timeout = 3 seconds
      }
    }

    Create ${CONFIG}/prosody/prosody-plugins-custom/mod_force_async_transcription.lua. This adaptation of the handbook module skips health-check rooms and runs after metadata initialization. [S1][S5][S6]

    Lua
    local is_healthcheck_room = module:require("util").is_healthcheck_room;
    module:hook("muc-room-created", function(event)
        local room = event.room;
        if is_healthcheck_room(room.jid) then return; end
        room.jitsiMetadata = room.jitsiMetadata or {};
        room.jitsiMetadata.asyncTranscription = true;
    end, -2);

    Rootless Jitsi services use UID 1000. Make files readable by that user and keep secret-bearing HOCON private: [S8][S17]

    Terminal
    sudo chown 1000:1000 "${CONFIG}/jicofo/custom-jicofo.conf"
    sudo chmod 600 "${CONFIG}/jicofo/custom-jicofo.conf"
    sudo chmod 644 "${CONFIG}/prosody/prosody-plugins-custom/mod_force_async_transcription.lua"
  3. Docker: apply the configuration. Validate and recreate changed services. ENABLE_TRANSCRIPTIONS=1 enables transcription and captions. Use a new meeting, because the module runs on room creation. [S5][S6][S16]

    Terminal
    docker compose -f docker-compose.yml -f compose.override.yaml config --quiet
    docker compose -f docker-compose.yml -f compose.override.yaml up -d --force-recreate opus-transcriber-proxy web prosody jicofo
  4. Debian/Ubuntu matching packages: run the same proxy locally. Put the five backend assignments from step 1 in a private .env dedicated to the proxy. If JVB shares this host, bind the published port to loopback: [S1][S2][S16]

    Terminal
    docker run -d --name transcriber --restart unless-stopped \
      -p 127.0.0.1:9090:8080 --env-file .env \
      jitsi/opus-transcriber-proxy:6747f187c0e7f4fa17f7744e0d76a5ce6d7d4b24

    Merge the Jicofo block into /etc/jitsi/jicofo/jicofo.conf, using this URL instead: [S4][S9]

    Text
    url-template = "ws://127.0.0.1:9090/transcribe?sessionId={{MEETING_ID}}&sendBack=true"

    Save the Lua module in /usr/share/jitsi-meet/prosody-plugins/mod_force_async_transcription.lua. Add "force_async_transcription"; to the existing conference MUC modules_enabled list in /etc/prosody/conf.avail/meet.example.com.cfg.lua. Preserve the metadata component and other modules. In /etc/jitsi/meet/meet.example.com-config.js, set the existing transcription.enabled to true. [S1][S6][S9]

    Terminal
    sudo systemctl restart prosody
    sudo systemctl restart jicofo
  5. Either install: select alternatives and protect remote access. OpenAI uses OPENAI_API_KEY and OPENAI_MODEL; xAI uses XAI_API_KEY with PROVIDERS_PRIORITY=xai. Gemini uses GEMINI_API_KEY and GEMINI_MODEL, but its current compatibility needs testing. Audio goes to OpenAI as 24 kHz PCM, Deepgram normally as Opus, and Gemini/xAI as 16 kHz PCM. Cloud retention and training terms are provider-specific. [S2][S15][S19]

    Docker: forward the selected variables in the proxy’s environment list, change PROVIDERS_PRIORITY to openai, gemini or xai, and recreate it. For packages, update the dedicated environment file and recreate the container. Forward the custom-provider enable variable too when using that option. [S2][S16]

    A custom provider needs ENABLE_OPENAI_CUSTOM_PROVIDER=true, URL parameters provider=openai_custom and URL-encoded openaiCustomUrl, plus the X-Custom-Openai-Api-Key HTTP header. It must implement Realtime transcription over WebSocket; a Whisper HTTP upload endpoint is insufficient. [S2][S15]

    For remote bridges, use WSS behind a gateway that validates Jicofo’s http-headers. The standalone proxy does not authenticate incoming connections. Adding an authorization header alone supplies no enforcement. [S1][S2]

Configuration reference

Defaults: stable-11248 or the pinned proxy. [S2][S4][S5][S6]

Setting Location Default Purpose
CONFIG Docker .env ~/.jitsi-meet-cfg Host configuration directory. [S8]
ENABLE_TRANSCRIPTIONS Docker .env 0 Enables client controls. [S5]
XMPP_MUC_MODULES Docker .env Empty Extra conference modules. [S5]
asyncTranscription Server room metadata Absent Selects bridge path. [S6]
recording.isTranscribingEnabled Room metadata Absent Start/stop gate. [S4][S6]
transcription.enabled Browser config false Enables transcription UI. [S9]
jicofo.transcription.url-template HOCON Unset JVB destination; substitutes {{MEETING_ID}} and optional {{REGION}}. [S4]
jicofo.transcription.http-headers HOCON {} Gateway/custom-backend headers. [S4]
jicofo.transcription.ping.enabled HOCON true WebSocket keepalive. [S4]
jicofo.transcription.ping.interval HOCON 10 seconds Ping interval. [S4]
jicofo.transcription.ping.timeout HOCON 3 seconds Pong timeout. [S4]
sendBack URL query false Returns results to JVB. [S2]
sessionId, provider, openaiCustomUrl URL query Unset Meeting identity, provider override, custom endpoint. [S2]
X-Custom-Openai-Api-Key HTTP header Unset Custom backend secret. [S2]
PROVIDERS_PRIORITY Proxy environment openai,deepgram,gemini Default provider selection. [S2]
DEEPGRAM_API_KEY, OPENAI_API_KEY, GEMINI_API_KEY, XAI_API_KEY Proxy environment Unset Backend secrets. [S2]
DEEPGRAM_MODEL, DEEPGRAM_LANGUAGE, DEEPGRAM_MIP_OPT_OUT Proxy environment nova-2, multi, false Model, language, data-program opt-out. [S2]
OPENAI_MODEL, GEMINI_MODEL Proxy environment gpt-4o-mini-transcribe, gemini-2.0-flash-exp Model selection, Gemini default needs compatibility review. [S2]
ENABLE_OPENAI_CUSTOM_PROVIDER Proxy environment false Enables custom Realtime backend. [S2]
PORT, HOST, OPUS_BACKEND Proxy environment 8080, 0.0.0.0, wasm Listener and decoder. [S2]

Common mistakes

  • Invented Docker variables: JICOFO_TRANSCRIPTION_URL_TEMPLATE belongs to open PR #2326, not stock stable-11248. Use the custom file. #2293 remains open without a maintainer solution. [S7]
  • Editing generated files: startup regenerates Docker Jicofo configuration. Keep additions in custom-jicofo.conf. [S5]
  • Only setting async: the second metadata gate must also be enabled. Ordinary moderators receive transcription permission; custom permissions or JWT features can deny it. [S6]
  • Treating the OpenAI crash as fixed: #17742 remains open and reports RangeError: Maximum call stack size exceeded. No released fix was verified. [S14]

Verify

Docker: test reachability from JVB using the same Compose files: [S2][S5][S16][S17]

Terminal
docker compose -f docker-compose.yml -f compose.override.yaml exec -T jvb curl --fail --silent --show-error http://opus-transcriber-proxy:8080/health

Packages with the local proxy: [S2][S17]

Terminal
curl --fail --silent --show-error http://127.0.0.1:9090/health

Both return HTTP 200 with exact body OK. Proxy startup should report Default provider: deepgram. These prove reachability and selection, not captions. [S2]

Join a new room as moderator, start transcription, and have two people speak distinct sentences. Confirm captions contain both sentences with correct speaker attribution, then stop and confirm transcription stops. JVB’s Websocket connected: true is useful connection evidence, but does not prove returned speech results. [S4][S6][S15]

If it still fails

Read Docker proxy, Jicofo, Prosody and JVB logs with your usual Compose file list: [S16]

Terminal
docker compose -f docker-compose.yml -f compose.override.yaml logs --tail=100 opus-transcriber-proxy jicofo prosody jvb

On packages, default logs include /var/log/jitsi/jicofo.log and /var/log/jitsi/jvb.log; check Prosody’s configured log destination, commonly /var/log/prosody/prosody.log. Also read docker logs transcriber. Jicofo’s exact Transcription enabled, but no URL is configured. identifies missing URL configuration. [S4][S9][S16]

Check API authentication, backend availability, both room flags and proxy connectivity from each bridge. Avoid debug logging and media dumps with real meeting data: the proxy’s debug messages can expose headers and provider secrets. For older Jigasi failures, use transcription troubleshooting. [S2][S15][S18]

FAQ

Do I need Jigasi or a transcriber password?

No Jigasi process is required for this bridge route. Stock Docker still generates some legacy configuration, but a missing Jigasi transcriber password is not fatal here. [S5]

Can I keep all audio on my server?

Not with a cloud backend. A local custom backend must implement the required Realtime WebSocket protocol; generic OpenAI-compatible REST support does not prove compatibility. [S15]

Does self-hosting make captions free?

The proxy runs on your infrastructure, but provider charges follow billed audio usage. For implementation help, see the transcription setup service. [S10][S11][S12][S13][S18]

Sources

All checked 2026-10-06.

[S1] Bridge transcription handbook, updated 2026-10-05, official doc.

[S2] Proxy snapshot 2026-10-05: configuration, server, publishing workflow, backend guide, source code and official doc.

[S3] Latest Docker release, stable-11248, 2026-09-14; Prosody build, release note and source code.

[S4] 11248 Jicofo defaults, room gate, bridge selection, conference logic, 2026-09-14, source code.

[S5] stable-11248 Compose, Jicofo initialization, Jicofo template, Prosody template, account registration, web template, 2026-09-14, source code.

[S6] 11248 room metadata, permissions, caption controls, 2026-09-14, source code.

[S7] Docker PR #2326, updated 2026-09-22, proposed source code; #2293, updated 2026-10-05, community report.

[S8] Docker handbook, volume checks, gitignore, stable-11248, official doc and source code.

[S9] Package index, quickstart, Jicofo launcher, browser defaults, web installer, Prosody installer, plugin mapping, Jicofo logging, JVB service, Prosody logging, checked 2026-10-06, official repository, doc and source code.

[S10] OpenAI pricing, checked 2026-10-06, official doc.

[S11] Deepgram pricing, MIP pricing change, checked 2026-10-06, official docs.

[S12] Gemini pricing, checked 2026-10-06, official doc.

[S13] xAI streaming STT, updated 2026-10-03, official doc.

[S14] Meet #17742, updated 2026-09-18, community report.

[S15] Proxy backend implementations, per-source connections, JVB exporter, serializer, 2026-09-14 to 2026-10-05, source code.

[S16] Docker Compose merging, Compose commands, container run, checked 2026-10-06, official docs.

[S17] curl, chmod, chown, checked 2026-10-06, official docs.

[S18] jitsi.help Vosk guide, troubleshooting, setup service, checked 2026-10-06, independent guides.

[S19] OpenAI data controls, Deepgram data handling, Gemini terms, xAI security, checked 2026-10-06, official docs.

Open questions

The earliest fully compatible release, a working current Gemini model substitution, and the complete captions flow on a real server remain unverified. This guide does not claim a live deployment or meeting test. [S1][S12][S14]

Recently updated