Who this is for
Use this guide for local Whisper recognition behind an existing Jigasi transcriber. It complements the Vosk setup; it does not repeat XMPP installation. Jigasi transcription is deprecated without an announced removal date, so keep a migration plan. [S2][S11][S13]
How it works
Jigasi sends Skynet one WebSocket connection per conference. Binary messages contain a 60-byte speaker/language header and raw mono 16 kHz audio. Skynet runs Faster Whisper locally and sends interim/final results back for captions. Its model is shared, with separate participant state. [S2][S3]
Configure Jigasi’s base URL as ws://skynet:8000/streaming-whisper/ws. It appends its own UUID. The full server route is /streaming-whisper/ws/{meeting_id}; do not append a meeting ID yourself. README issue #576 was corrected by PR #628, merged March 10, 2026, but the Java fallback still lacks the module prefix. [S2][S3][S6]
The bridge path uses a different protocol. An OpenAI-compatible HTTP upload endpoint cannot replace Skynet’s binary protocol or satisfy the proxy’s Realtime contract. [S3][S11]
Before you start
As checked October 6, 2026, Docker’s latest release is stable-11248. Use matching Jitsi images and Compose files. Package instructions cover Jigasi 1.1-415-g2750d54-1. This recipe pins the published 2026.4.1-cpu Skynet image; a newer GitHub release does not establish that its CPU image exists. [S1][S2][S3]
Start with tiny.en or base.en for an English pilot. Whisper’s original PyTorch guidance lists approximate VRAM of 1 GB for tiny/base, 2 GB for small, 5 GB for medium and 10 GB for large. These are model references, not Skynet memory per concurrent speaker. .en models are English-only; test multilingual models for other languages. [S8]
Faster Whisper’s published benchmark processes 13 minutes of audio with the small model on an i7-12700K, eight threads, beam size 5 and int8 in 1 minute 42 seconds, using 1477 MB RAM. Large-v2 GPU int8 uses 2926 MB in its RTX 3070 Ti benchmark. These are offline file benchmarks, not live conference capacity. No dependable VRAM-per-speaker rule was found. [S9]
Model downloads initially require network access. Recognition can remain local, but verify a restart without internet after caching. Maintain storage readable/writable by Skynet’s UID/GID 1001. Apply Jitsi service changes with no active meetings. [S3][S12]
Steps
-
Docker stable-11248: add a private CPU backend. From the Compose directory, set
CONFIGto match.env, using the default below. This guide introduces the host directoryskynet-models. [S1][S3][S12]TerminalCONFIG="$HOME/.jitsi-meet-cfg" mkdir -p "${CONFIG}/skynet-models" sudo chown 1001:1001 "${CONFIG}/skynet-models"Save as
whisper.yml. No host port is published; authentication bypass is for this private-network pilot. Streaming-only inference does not require Redis. [S3][S4][S12]YAMLservices: skynet: image: jitsi/skynet:2026.4.1-cpu@sha256:8b02c6ed4bd3cc60c659f78ac304be5618ce375ea9cbf2011d34a1bfc312db13 restart: unless-stopped environment: ENABLED_MODULES: streaming_whisper WHISPER_MODEL_NAME: base.en WHISPER_MODEL_PATH: /models WHISPER_DEVICE: cpu WHISPER_COMPUTE_TYPE: int8 BEAM_SIZE: "1" BYPASS_AUTHORIZATION: "true" LOG_LEVEL: INFO volumes: - ${CONFIG}/skynet-models:/models networks: meet.jitsi: -
Docker stable-11248: select Jigasi’s Whisper service. Edit existing
.envassignments. Retain your existingJIGASI_TRANSCRIBER_PASSWORD, other authentication and optional overlays. The Whisper signing variables may remain unset for the private bypass pilot. [S1][S2]dotenvENABLE_TRANSCRIPTIONS=1 JIGASI_TRANSCRIBER_CUSTOM_SERVICE=org.jitsi.jigasi.transcription.WhisperTranscriptionService JIGASI_TRANSCRIBER_WHISPER_URL=ws://skynet:8000/streaming-whisper/wsUse stock
transcriber.yml, which defines servicetranscriber. Add existing overlays to this helper, then validate and recreate during maintenance. [S1][S12]Terminaldc() { docker compose -f docker-compose.yml -f transcriber.yml -f whisper.yml "$@" } dc config --quiet dc up -d --force-recreate skynet transcriber prosody web -
Either install: create optional Skynet signing keys. These are RS256 credentials, not the XMPP password. Run in a private directory outside your repository. Example filenames are chosen here. [S2][S4][S12]
Terminalumask 077 mkdir -p "$HOME/skynet-keys" cd "$HOME/skynet-keys" openssl genpkey -algorithm RSA -pkeyopt rsa_keygen_bits:2048 -out skynet-private.pem openssl pkcs8 -topk8 -nocrypt -in skynet-private.pem -outform DER | base64 -w 0 > skynet-private.base64 KEY_HASH=$(printf '%s' 'jitsi-whisper' | sha256sum | cut -d ' ' -f 1) openssl pkey -in skynet-private.pem -pubout -out "${KEY_HASH}.pem" cd -Publish only the public
${KEY_HASH}.pemathttps://meet.example.com/skynet-keys/${KEY_HASH}.pem. In Skynet’s environment, setBYPASS_AUTHORIZATION=false,ASAP_PUB_KEYS_REPO_URL=https://meet.example.com,ASAP_PUB_KEYS_FOLDER=skynet-keys,ASAP_PUB_KEYS_AUDS=jitsi. For Docker, editwhisper.ymlunderskynet.environment; these are not automatically forwarded from Jitsi’s.env. Set these Jigasi Docker assignments locally: [S2][S4][S12]dotenvJIGASI_TRANSCRIBER_WHISPER_PRIVATE_KEY_NAME=jitsi-whisper JIGASI_TRANSCRIBER_WHISPER_PRIVATE_KEY=REPLACE_WITH_BASE64_PKCS8_DERReplace the placeholder with the Base64 file’s contents, not its filename or PEM wrapper. The key name is JWT
kid; Skynet fetches its hashed public filename. Jigasi sends a five-minute token throughAuthorization: Bearer, which Skynet accepts. Protect.envwithchmod 600 .env, then recreate Skynet/transcriber. For remote access, use authenticated WSS. [S2][S4][S12] -
Docker GPU option: build Skynet source
7099ed2. Install NVIDIA’s driver and Container Toolkit using its official instructions. Build locally, preserving the Dockerfile’s CUDA bases and omitting vLLM: [S5][S12]Terminalgit clone https://github.com/jitsi/skynet.git cd skynet git checkout 7099ed2acc596d74a5fcae231e7e23bf36459446 docker build --build-arg BUILD_WITH_VLLM=0 -t skynet-whisper:7099ed2-gpu . cd -In
whisper.yml, replace the image withskynet-whisper:7099ed2-gpu, setWHISPER_DEVICE=cuda,WHISPER_COMPUTE_TYPE=int8_float16, and add this reservation underskynet. Validate and recreate withdcagain. [S3][S5][S9][S12]YAMLdeploy: resources: reservations: devices: - driver: nvidia count: 1 capabilities: [gpu] -
Debian/Ubuntu Jigasi baseline: use the same backend. Create a private
skynet.envcontaining the eight environment assignments from the CPU overlay in dotenv syntax. On the same host as Jigasi, publish only loopback ports: [S2][S3][S12]Terminalmkdir -p "$HOME/skynet-models" sudo chown 1001:1001 "$HOME/skynet-models" docker run -d --name skynet --restart unless-stopped \ -p 127.0.0.1:8000:8000 -p 127.0.0.1:8001:8001 \ --env-file skynet.env -v "$HOME/skynet-models:/models" \ jitsi/skynet:2026.4.1-cpu@sha256:8b02c6ed4bd3cc60c659f78ac304be5618ce375ea9cbf2011d34a1bfc312db13Edit existing assignments in
/etc/jitsi/jigasi/sip-communicator.properties: [S2]Propertiesorg.jitsi.jigasi.ENABLE_TRANSCRIPTION=true org.jitsi.jigasi.transcription.customService=org.jitsi.jigasi.transcription.WhisperTranscriptionService org.jitsi.jigasi.transcription.whisper.websocket_url=ws://127.0.0.1:8000/streaming-whisper/wsAuthenticated packages use
org.jitsi.jigasi.transcription.whisper.private_keyandorg.jitsi.jigasi.transcription.whisper.private_key_namewith the same values from step 3;org.jitsi.jigasi.transcription.whisper.jwt_audiencedefaults tojitsi. Update Skynet’s settings inskynet.envand recreate its container. Preserve SIP and existing XMPP settings. Runsudo systemctl restart jigasiduring maintenance. [S2][S12] -
Bridge-based Whisper: verify compatibility before wiring. The current proxy offers
ENABLE_OPENAI_CUSTOM_PROVIDER=true,provider=openai_custom, URL-encodedopenaiCustomUrlandX-Custom-Openai-Api-Key. A local server must accept the Realtime WebSocket negotiation and nestedsession.update, 24 kHz PCM throughinput_audio_buffer.append/input_audio_buffer.commit, and compatible transcript delta/completed events. [S11]No working local Whisper deployment was verified against this contract. Skynet’s binary endpoint is incompatible. Speaches implements a Realtime route, but the inspected handshake/session schema differs. Use the bridge guide after proving compatibility; do not treat an illustrative endpoint as a working example. [S11][S13]
Configuration reference
Defaults: published CPU source, GPU snapshot and stable-11248 mappings. [S1][S2][S3][S5]
| Setting | Where | Default | Purpose |
|---|---|---|---|
CONFIG |
Docker .env |
~/.jitsi-meet-cfg |
Host files. [S1] |
ENABLE_TRANSCRIPTIONS |
Docker .env |
0 |
Caption controls. [S1] |
JIGASI_TRANSCRIBER_PASSWORD |
Docker .env |
Unset | XMPP login. [S1] |
JIGASI_TRANSCRIBER_CUSTOM_SERVICE |
Docker .env |
Unset | Select Whisper class. [S1] |
JIGASI_TRANSCRIBER_WHISPER_URL |
Docker .env |
Unset; Java fallback ws://localhost:8000/ws |
Correct Skynet base URL. [S1][S2] |
JIGASI_TRANSCRIBER_WHISPER_PRIVATE_KEY, JIGASI_TRANSCRIBER_WHISPER_PRIVATE_KEY_NAME |
Docker .env |
Unset | DER Base64 and JWT kid. [S1][S2] |
org.jitsi.jigasi.ENABLE_TRANSCRIPTION |
Package properties | false |
Start transcriber gateway. [S2] |
org.jitsi.jigasi.transcription.customService |
Package properties | Unset | Select backend. [S2] |
org.jitsi.jigasi.transcription.whisper.websocket_url, org.jitsi.jigasi.transcription.whisper.private_key, org.jitsi.jigasi.transcription.whisper.private_key_name, org.jitsi.jigasi.transcription.whisper.jwt_audience |
Package properties | Fallback URL above, empty keys, jitsi |
Package equivalents. [S2] |
ENABLED_MODULES |
Skynet environment | summaries:dispatcher,summaries:executor,assistant,customer_configs |
Set streaming_whisper. [S3] |
WHISPER_MODEL_NAME, WHISPER_MODEL_PATH |
Skynet environment | Unset, /app/models/streaming_whisper in CPU image |
Model name or converted local model/cache. [S3] |
WHISPER_DEVICE, WHISPER_COMPUTE_TYPE, BEAM_SIZE |
Skynet environment | auto, int8, 5 |
Device, precision, decoding search. [S3] |
BYPASS_AUTHORIZATION, LOG_LEVEL |
Skynet environment | false, DEBUG |
Private pilot bypass; reduce sensitive debug output. [S3][S4] |
ASAP_PUB_KEYS_REPO_URL, ASAP_PUB_KEYS_FOLDER, ASAP_PUB_KEYS_AUDS |
Skynet environment | Unset, unset, empty | Public-key lookup and audiences. [S4] |
BUILD_WITH_VLLM |
Image build argument | 1 |
Set 0 for this Whisper-only build. [S5] |
ENABLE_OPENAI_CUSTOM_PROVIDER, provider, openaiCustomUrl, X-Custom-Openai-Api-Key |
Bridge proxy environment/query/header | false, unset, unset, unset |
Conditional custom Realtime backend. [S11] |
Common mistakes
Skynet #214 reports: [S7]
Could not load library libcudnn_ops_infer.so.8. Error: libcudnn_graph.so.9: cannot open shared object file: No such file or directoryThe maintainer explained that make local_build uses Ubuntu without CUDA libraries. The reporter confirmed that using the CUDA build worked. Our GPU command preserves those bases without requiring the Makefile’s registry push. The issue remains open; no fixing release/PR was identified. [S5][S7]
The pinned GPU source uses CTranslate2 4.4.0 with CUDA 12.2.2/cuDNN 8. Latest Faster Whisper dependency guidance differs, so do not upgrade one library blindly. Wrong /ws URLs, missing public-key files, incorrect audience, and using PEM text as DER Base64 also prevent connection. [S2][S4][S5][S9]
Verify
Docker: [S3][S12]
dc exec -T jvb curl --fail --silent --show-error http://skynet:8001/healthz
dc logs --tail=100 skynet transcriberPackages: curl --fail --silent --show-error http://127.0.0.1:8001/healthz. Expect HTTP 200 and exact {"status":"ok"}. Model logs should show Using cpu and Model: base.en, or Using cuda for GPU. Jigasi reports Successfully connected to followed by its complete URL. These checks do not prove captions. [S2][S3]
Start captions in a new room, speak distinct sentences from two participants, and confirm both names and words. Pause, then stop captions and check the final sentence. Measure delay and errors with your expected number of simultaneous speakers. [S2][S3]
If it still fails
Read Skynet/transcriber logs, verify writable model storage and the actual selected device, then check the full URL and public-key lookup. Package defaults use /var/log/jitsi/jigasi.log; also read docker logs skynet. Keep real speech and signing material out of debug logs. [S2][S3][S4][S12]
For XMPP failures, follow transcription troubleshooting. Jigasi issue #527 remains open; its community HTTP/gRPC suggestions do not describe the released Skynet adapter. Bridge issue #17742 also remains open, with no released fix verified for RangeError: Maximum call stack size exceeded. [S6][S11][S13]
FAQ
Is Whisper more accurate than Vosk?
No controlled Jitsi comparison was verified. Vosk supports continuous streaming; Skynet buffers and repeatedly decodes audio. Compare both on your languages, microphones and accents rather than promising universal accuracy or latency. [S3][S8][S9][S10]
Can I run it without a GPU?
Yes, the published CPU image and int8 configuration support that route. Batch benchmarks show CPU can be practical, but they do not establish your live meeting capacity. [S3][S9]
Does local recognition require a cloud API key?
No cloud recognition key is needed for Skynet. Its optional RSA key authenticates your own service. For help implementing the setup, see our transcription service. [S2][S4][S13]
Sources
All checked 2026-10-06.
[S1] Latest Docker release, transcriber overlay, property template, Docker handbook, 2026-09-14 release, release note/source code/official doc.
[S2] Released Jigasi adapter, JWT signing, gateway switches, service, README, 2026-09-11 snapshot, source code/official doc.
[S3] Skynet published CPU tag, matching environment, model loading, Dockerfile, protocol, metrics, connections, 2026.4.1 snapshot, official repository/source code.
S3 also includes Skynet’s participant buffering and decoding, same released snapshot, source code.
[S4] Skynet 2026.4.1 authorization, verification, header extraction, official doc/source code.
[S5] GPU snapshot Dockerfile, Makefile, dependencies, 2026-07-02 snapshot, source code.
[S6] Jigasi #576, community report; PR #628, merged 2026-03-10, source change; #527, community report, checked 2026-10-06.
[S7] Skynet #214, 2025-06-19 to 2025-06-20, community report and maintainer comments.
[S8] Whisper model guidance, checked 2026-10-06, official repository documentation.
[S9] Faster Whisper benchmarks and CUDA compatibility, checked 2026-10-06; benchmark version 1.1.0, official repository documentation.
[S10] Vosk, checked 2026-10-06, official doc.
[S11] Bridge handbook, proxy backend, custom validation, Speaches route, session handler, Meet #17742, checked 2026-10-06, official doc/source code/community report.
[S12] Docker CLI, Compose GPU access, NVIDIA Toolkit, OpenSSL key generation, PKCS8, public-key export, base64, sha256sum, cut, chmod, chown, Git, systemctl, curl, checked 2026-10-06, official docs.
[S13] jitsi.help Vosk, migration, bridge setup, troubleshooting, setup service, checked 2026-10-06, independent guides.
Open questions
Real-server CPU/GPU performance, concurrent-speaker memory, offline restart behavior, and a working local Realtime-compatible Whisper deployment remain unverified. No controlled Jitsi Whisper/Vosk accuracy or latency comparison, live meeting test or GPU build test is claimed. [S3][S8][S9][S10][S11]