How do I run self-hosted Whisper transcription in Jitsi?

Short answer

Jigasi can send meeting audio to self-hosted Skynet streaming Whisper at /streaming-whisper/ws, with Jigasi appending the connection ID. Use the published CPU image or a GPU build with matching CUDA libraries, and configure its Whisper service explicitly. Signing keys authenticate Jigasi to Skynet and are separate from the transcriber XMPP password. The bridge replacement needs OpenAI Realtime WebSocket compatibility; a working local Whisper example was not verified. [S1][S2][S3][S5][S11]

Who this is for

Use this guide for local Whisper recognition behind an existing Jigasi transcriber. It complements the Vosk setup; it does not repeat XMPP installation. Jigasi transcription is deprecated without an announced removal date, so keep a migration plan. [S2][S11][S13]

How it works

Jigasi sends Skynet one WebSocket connection per conference. Binary messages contain a 60-byte speaker/language header and raw mono 16 kHz audio. Skynet runs Faster Whisper locally and sends interim/final results back for captions. Its model is shared, with separate participant state. [S2][S3]

Configure Jigasi’s base URL as ws://skynet:8000/streaming-whisper/ws. It appends its own UUID. The full server route is /streaming-whisper/ws/{meeting_id}; do not append a meeting ID yourself. README issue #576 was corrected by PR #628, merged March 10, 2026, but the Java fallback still lacks the module prefix. [S2][S3][S6]

The bridge path uses a different protocol. An OpenAI-compatible HTTP upload endpoint cannot replace Skynet’s binary protocol or satisfy the proxy’s Realtime contract. [S3][S11]

Before you start

As checked October 6, 2026, Docker’s latest release is stable-11248. Use matching Jitsi images and Compose files. Package instructions cover Jigasi 1.1-415-g2750d54-1. This recipe pins the published 2026.4.1-cpu Skynet image; a newer GitHub release does not establish that its CPU image exists. [S1][S2][S3]

Start with tiny.en or base.en for an English pilot. Whisper’s original PyTorch guidance lists approximate VRAM of 1 GB for tiny/base, 2 GB for small, 5 GB for medium and 10 GB for large. These are model references, not Skynet memory per concurrent speaker. .en models are English-only; test multilingual models for other languages. [S8]

Faster Whisper’s published benchmark processes 13 minutes of audio with the small model on an i7-12700K, eight threads, beam size 5 and int8 in 1 minute 42 seconds, using 1477 MB RAM. Large-v2 GPU int8 uses 2926 MB in its RTX 3070 Ti benchmark. These are offline file benchmarks, not live conference capacity. No dependable VRAM-per-speaker rule was found. [S9]

Model downloads initially require network access. Recognition can remain local, but verify a restart without internet after caching. Maintain storage readable/writable by Skynet’s UID/GID 1001. Apply Jitsi service changes with no active meetings. [S3][S12]

Steps

  1. Docker stable-11248: add a private CPU backend. From the Compose directory, set CONFIG to match .env, using the default below. This guide introduces the host directory skynet-models. [S1][S3][S12]

    Terminal
    CONFIG="$HOME/.jitsi-meet-cfg"
    mkdir -p "${CONFIG}/skynet-models"
    sudo chown 1001:1001 "${CONFIG}/skynet-models"

    Save as whisper.yml. No host port is published; authentication bypass is for this private-network pilot. Streaming-only inference does not require Redis. [S3][S4][S12]

    YAML
    services:
      skynet:
        image: jitsi/skynet:2026.4.1-cpu@sha256:8b02c6ed4bd3cc60c659f78ac304be5618ce375ea9cbf2011d34a1bfc312db13
        restart: unless-stopped
        environment:
          ENABLED_MODULES: streaming_whisper
          WHISPER_MODEL_NAME: base.en
          WHISPER_MODEL_PATH: /models
          WHISPER_DEVICE: cpu
          WHISPER_COMPUTE_TYPE: int8
          BEAM_SIZE: "1"
          BYPASS_AUTHORIZATION: "true"
          LOG_LEVEL: INFO
        volumes:
          - ${CONFIG}/skynet-models:/models
        networks:
          meet.jitsi:
  2. Docker stable-11248: select Jigasi’s Whisper service. Edit existing .env assignments. Retain your existing JIGASI_TRANSCRIBER_PASSWORD, other authentication and optional overlays. The Whisper signing variables may remain unset for the private bypass pilot. [S1][S2]

    dotenv
    ENABLE_TRANSCRIPTIONS=1
    JIGASI_TRANSCRIBER_CUSTOM_SERVICE=org.jitsi.jigasi.transcription.WhisperTranscriptionService
    JIGASI_TRANSCRIBER_WHISPER_URL=ws://skynet:8000/streaming-whisper/ws

    Use stock transcriber.yml, which defines service transcriber. Add existing overlays to this helper, then validate and recreate during maintenance. [S1][S12]

    Terminal
    dc() {
        docker compose -f docker-compose.yml -f transcriber.yml -f whisper.yml "$@"
    }
    dc config --quiet
    dc up -d --force-recreate skynet transcriber prosody web
  3. Either install: create optional Skynet signing keys. These are RS256 credentials, not the XMPP password. Run in a private directory outside your repository. Example filenames are chosen here. [S2][S4][S12]

    Terminal
    umask 077
    mkdir -p "$HOME/skynet-keys"
    cd "$HOME/skynet-keys"
    openssl genpkey -algorithm RSA -pkeyopt rsa_keygen_bits:2048 -out skynet-private.pem
    openssl pkcs8 -topk8 -nocrypt -in skynet-private.pem -outform DER |
        base64 -w 0 > skynet-private.base64
    KEY_HASH=$(printf '%s' 'jitsi-whisper' | sha256sum | cut -d ' ' -f 1)
    openssl pkey -in skynet-private.pem -pubout -out "${KEY_HASH}.pem"
    cd -

    Publish only the public ${KEY_HASH}.pem at https://meet.example.com/skynet-keys/${KEY_HASH}.pem. In Skynet’s environment, set BYPASS_AUTHORIZATION=false, ASAP_PUB_KEYS_REPO_URL=https://meet.example.com, ASAP_PUB_KEYS_FOLDER=skynet-keys, ASAP_PUB_KEYS_AUDS=jitsi. For Docker, edit whisper.yml under skynet.environment; these are not automatically forwarded from Jitsi’s .env. Set these Jigasi Docker assignments locally: [S2][S4][S12]

    dotenv
    JIGASI_TRANSCRIBER_WHISPER_PRIVATE_KEY_NAME=jitsi-whisper
    JIGASI_TRANSCRIBER_WHISPER_PRIVATE_KEY=REPLACE_WITH_BASE64_PKCS8_DER

    Replace the placeholder with the Base64 file’s contents, not its filename or PEM wrapper. The key name is JWT kid; Skynet fetches its hashed public filename. Jigasi sends a five-minute token through Authorization: Bearer, which Skynet accepts. Protect .env with chmod 600 .env, then recreate Skynet/transcriber. For remote access, use authenticated WSS. [S2][S4][S12]

  4. Docker GPU option: build Skynet source 7099ed2. Install NVIDIA’s driver and Container Toolkit using its official instructions. Build locally, preserving the Dockerfile’s CUDA bases and omitting vLLM: [S5][S12]

    Terminal
    git clone https://github.com/jitsi/skynet.git
    cd skynet
    git checkout 7099ed2acc596d74a5fcae231e7e23bf36459446
    docker build --build-arg BUILD_WITH_VLLM=0 -t skynet-whisper:7099ed2-gpu .
    cd -

    In whisper.yml, replace the image with skynet-whisper:7099ed2-gpu, set WHISPER_DEVICE=cuda, WHISPER_COMPUTE_TYPE=int8_float16, and add this reservation under skynet. Validate and recreate with dc again. [S3][S5][S9][S12]

    YAML
    deploy:
      resources:
        reservations:
          devices:
            - driver: nvidia
              count: 1
              capabilities: [gpu]
  5. Debian/Ubuntu Jigasi baseline: use the same backend. Create a private skynet.env containing the eight environment assignments from the CPU overlay in dotenv syntax. On the same host as Jigasi, publish only loopback ports: [S2][S3][S12]

    Terminal
    mkdir -p "$HOME/skynet-models"
    sudo chown 1001:1001 "$HOME/skynet-models"
    docker run -d --name skynet --restart unless-stopped \
        -p 127.0.0.1:8000:8000 -p 127.0.0.1:8001:8001 \
        --env-file skynet.env -v "$HOME/skynet-models:/models" \
        jitsi/skynet:2026.4.1-cpu@sha256:8b02c6ed4bd3cc60c659f78ac304be5618ce375ea9cbf2011d34a1bfc312db13

    Edit existing assignments in /etc/jitsi/jigasi/sip-communicator.properties: [S2]

    Properties
    org.jitsi.jigasi.ENABLE_TRANSCRIPTION=true
    org.jitsi.jigasi.transcription.customService=org.jitsi.jigasi.transcription.WhisperTranscriptionService
    org.jitsi.jigasi.transcription.whisper.websocket_url=ws://127.0.0.1:8000/streaming-whisper/ws

    Authenticated packages use org.jitsi.jigasi.transcription.whisper.private_key and org.jitsi.jigasi.transcription.whisper.private_key_name with the same values from step 3; org.jitsi.jigasi.transcription.whisper.jwt_audience defaults to jitsi. Update Skynet’s settings in skynet.env and recreate its container. Preserve SIP and existing XMPP settings. Run sudo systemctl restart jigasi during maintenance. [S2][S12]

  6. Bridge-based Whisper: verify compatibility before wiring. The current proxy offers ENABLE_OPENAI_CUSTOM_PROVIDER=true, provider=openai_custom, URL-encoded openaiCustomUrl and X-Custom-Openai-Api-Key. A local server must accept the Realtime WebSocket negotiation and nested session.update, 24 kHz PCM through input_audio_buffer.append/input_audio_buffer.commit, and compatible transcript delta/completed events. [S11]

    No working local Whisper deployment was verified against this contract. Skynet’s binary endpoint is incompatible. Speaches implements a Realtime route, but the inspected handshake/session schema differs. Use the bridge guide after proving compatibility; do not treat an illustrative endpoint as a working example. [S11][S13]

Configuration reference

Defaults: published CPU source, GPU snapshot and stable-11248 mappings. [S1][S2][S3][S5]

Setting Where Default Purpose
CONFIG Docker .env ~/.jitsi-meet-cfg Host files. [S1]
ENABLE_TRANSCRIPTIONS Docker .env 0 Caption controls. [S1]
JIGASI_TRANSCRIBER_PASSWORD Docker .env Unset XMPP login. [S1]
JIGASI_TRANSCRIBER_CUSTOM_SERVICE Docker .env Unset Select Whisper class. [S1]
JIGASI_TRANSCRIBER_WHISPER_URL Docker .env Unset; Java fallback ws://localhost:8000/ws Correct Skynet base URL. [S1][S2]
JIGASI_TRANSCRIBER_WHISPER_PRIVATE_KEY, JIGASI_TRANSCRIBER_WHISPER_PRIVATE_KEY_NAME Docker .env Unset DER Base64 and JWT kid. [S1][S2]
org.jitsi.jigasi.ENABLE_TRANSCRIPTION Package properties false Start transcriber gateway. [S2]
org.jitsi.jigasi.transcription.customService Package properties Unset Select backend. [S2]
org.jitsi.jigasi.transcription.whisper.websocket_url, org.jitsi.jigasi.transcription.whisper.private_key, org.jitsi.jigasi.transcription.whisper.private_key_name, org.jitsi.jigasi.transcription.whisper.jwt_audience Package properties Fallback URL above, empty keys, jitsi Package equivalents. [S2]
ENABLED_MODULES Skynet environment summaries:dispatcher,summaries:executor,assistant,customer_configs Set streaming_whisper. [S3]
WHISPER_MODEL_NAME, WHISPER_MODEL_PATH Skynet environment Unset, /app/models/streaming_whisper in CPU image Model name or converted local model/cache. [S3]
WHISPER_DEVICE, WHISPER_COMPUTE_TYPE, BEAM_SIZE Skynet environment auto, int8, 5 Device, precision, decoding search. [S3]
BYPASS_AUTHORIZATION, LOG_LEVEL Skynet environment false, DEBUG Private pilot bypass; reduce sensitive debug output. [S3][S4]
ASAP_PUB_KEYS_REPO_URL, ASAP_PUB_KEYS_FOLDER, ASAP_PUB_KEYS_AUDS Skynet environment Unset, unset, empty Public-key lookup and audiences. [S4]
BUILD_WITH_VLLM Image build argument 1 Set 0 for this Whisper-only build. [S5]
ENABLE_OPENAI_CUSTOM_PROVIDER, provider, openaiCustomUrl, X-Custom-Openai-Api-Key Bridge proxy environment/query/header false, unset, unset, unset Conditional custom Realtime backend. [S11]

Common mistakes

Skynet #214 reports: [S7]

Text
Could not load library libcudnn_ops_infer.so.8. Error: libcudnn_graph.so.9: cannot open shared object file: No such file or directory

The maintainer explained that make local_build uses Ubuntu without CUDA libraries. The reporter confirmed that using the CUDA build worked. Our GPU command preserves those bases without requiring the Makefile’s registry push. The issue remains open; no fixing release/PR was identified. [S5][S7]

The pinned GPU source uses CTranslate2 4.4.0 with CUDA 12.2.2/cuDNN 8. Latest Faster Whisper dependency guidance differs, so do not upgrade one library blindly. Wrong /ws URLs, missing public-key files, incorrect audience, and using PEM text as DER Base64 also prevent connection. [S2][S4][S5][S9]

Verify

Docker: [S3][S12]

Terminal
dc exec -T jvb curl --fail --silent --show-error http://skynet:8001/healthz
dc logs --tail=100 skynet transcriber

Packages: curl --fail --silent --show-error http://127.0.0.1:8001/healthz. Expect HTTP 200 and exact {"status":"ok"}. Model logs should show Using cpu and Model: base.en, or Using cuda for GPU. Jigasi reports Successfully connected to followed by its complete URL. These checks do not prove captions. [S2][S3]

Start captions in a new room, speak distinct sentences from two participants, and confirm both names and words. Pause, then stop captions and check the final sentence. Measure delay and errors with your expected number of simultaneous speakers. [S2][S3]

If it still fails

Read Skynet/transcriber logs, verify writable model storage and the actual selected device, then check the full URL and public-key lookup. Package defaults use /var/log/jitsi/jigasi.log; also read docker logs skynet. Keep real speech and signing material out of debug logs. [S2][S3][S4][S12]

For XMPP failures, follow transcription troubleshooting. Jigasi issue #527 remains open; its community HTTP/gRPC suggestions do not describe the released Skynet adapter. Bridge issue #17742 also remains open, with no released fix verified for RangeError: Maximum call stack size exceeded. [S6][S11][S13]

FAQ

Is Whisper more accurate than Vosk?

No controlled Jitsi comparison was verified. Vosk supports continuous streaming; Skynet buffers and repeatedly decodes audio. Compare both on your languages, microphones and accents rather than promising universal accuracy or latency. [S3][S8][S9][S10]

Can I run it without a GPU?

Yes, the published CPU image and int8 configuration support that route. Batch benchmarks show CPU can be practical, but they do not establish your live meeting capacity. [S3][S9]

Does local recognition require a cloud API key?

No cloud recognition key is needed for Skynet. Its optional RSA key authenticates your own service. For help implementing the setup, see our transcription service. [S2][S4][S13]

Sources

All checked 2026-10-06.

[S1] Latest Docker release, transcriber overlay, property template, Docker handbook, 2026-09-14 release, release note/source code/official doc.

[S2] Released Jigasi adapter, JWT signing, gateway switches, service, README, 2026-09-11 snapshot, source code/official doc.

[S3] Skynet published CPU tag, matching environment, model loading, Dockerfile, protocol, metrics, connections, 2026.4.1 snapshot, official repository/source code.

S3 also includes Skynet’s participant buffering and decoding, same released snapshot, source code.

[S4] Skynet 2026.4.1 authorization, verification, header extraction, official doc/source code.

[S5] GPU snapshot Dockerfile, Makefile, dependencies, 2026-07-02 snapshot, source code.

[S6] Jigasi #576, community report; PR #628, merged 2026-03-10, source change; #527, community report, checked 2026-10-06.

[S7] Skynet #214, 2025-06-19 to 2025-06-20, community report and maintainer comments.

[S8] Whisper model guidance, checked 2026-10-06, official repository documentation.

[S9] Faster Whisper benchmarks and CUDA compatibility, checked 2026-10-06; benchmark version 1.1.0, official repository documentation.

[S10] Vosk, checked 2026-10-06, official doc.

[S11] Bridge handbook, proxy backend, custom validation, Speaches route, session handler, Meet #17742, checked 2026-10-06, official doc/source code/community report.

[S12] Docker CLI, Compose GPU access, NVIDIA Toolkit, OpenSSL key generation, PKCS8, public-key export, base64, sha256sum, cut, chmod, chown, Git, systemctl, curl, checked 2026-10-06, official docs.

[S13] jitsi.help Vosk, migration, bridge setup, troubleshooting, setup service, checked 2026-10-06, independent guides.

Open questions

Real-server CPU/GPU performance, concurrent-speaker memory, offline restart behavior, and a working local Realtime-compatible Whisper deployment remain unverified. No controlled Jitsi Whisper/Vosk accuracy or latency comparison, live meeting test or GPU build test is claimed. [S3][S8][S9][S10][S11]

Recently updated