Whisper on premise means running OpenAI's open-weight speech recognition model on your own hardware, usually through faster-whisper or whisper.cpp, so audio never reaches a hosted transcription API. OpenAI releases Whisper's code and weights under the MIT License. The turbo model needs about 6 GB of VRAM, and a voice activity detector in front of it reduces hallucinated text during silence.
Whisper model sizes, memory and licence
OpenAI’s repository lists six sizes. The VRAM figures and relative speeds below are OpenAI’s own, with speed measured against the large model (verified October 2026):
| Size | Parameters | English-only variant | VRAM | Relative speed |
|---|---|---|---|---|
| tiny | 39M | tiny.en | ~1 GB | ~10x |
| base | 74M | base.en | ~1 GB | ~7x |
| small | 244M | small.en | ~2 GB | ~4x |
| medium | 769M | medium.en | ~5 GB | ~2x |
| large | 1,550M | none | ~10 GB | 1x |
| turbo | 809M | none | ~6 GB | ~8x |
Three details shape the choice:
- large-v3 is the accuracy reference. The original Whisper paper trained on 680,000 hours of multilingual, multitask audio. The large-v3 card adds 1 million hours of weakly labelled and 4 million hours of pseudo-labelled audio, moves to 128 Mel frequency bins from 80, and reports 10% to 20% fewer errors than large-v2.
- turbo trades a little quality for a lot of speed. It is a fine-tuned, pruned large-v3 with 4 decoder layers instead of 32. OpenAI notes that it was not trained for translation and returns the original language even when asked to translate.
- Audio is processed in 30-second windows. The reference transcribe() method slides a 30-second window across the file, so long recordings are handled in pieces and live audio needs extra machinery.
On licensing, OpenAI states that Whisper’s code and model weights are released under the MIT License. On Hugging Face the large-v3 card carries an Apache 2.0 tag and the turbo card shows MIT, and both licences permit commercial use.
Runtimes for self-hosted Whisper compared
Production deployments rarely run the reference Python package. These are the usual alternatives (verified October 2026):
| Runtime | What it is | Best fit | Licence | Watch for |
|---|---|---|---|---|
| faster-whisper | Whisper reimplemented on the CTranslate2 inference engine | Python services on NVIDIA GPUs, or CPUs with 8-bit models | MIT | Current CTranslate2 releases support only CUDA 12 with cuDNN 9 |
| whisper.cpp | C/C++ inference on the ggml library | CPU-only hosts, Apple Silicon, Vulkan, AMD ROCm, OpenVINO, desktop and embedded apps | MIT | Uses converted ggml models; integer quantization lowers memory |
| WhisperX | faster-whisper plus wav2vec2 alignment and pyannote diarization | Word-level timestamps and speaker labels | BSD-2-Clause | Diarization needs a Hugging Face token and accepted model conditions |
| openai-whisper | The reference implementation | Experiments and reproducing OpenAI’s published numbers | MIT | Requires ffmpeg; the baseline other runtimes compare against |
faster-whisper claims up to four times the speed of openai/whisper at the same accuracy while using less memory, and its BatchedInferencePipeline drops in for the standard transcribe call. whisper.cpp treats Apple Silicon as a first-class target through Metal and Core ML, and also supports NVIDIA GPUs, Vulkan, AMD ROCm, OpenVINO, several NPUs and plain CPUs.
WhisperX reports 70x real-time transcription with large-v2 through batched inference, and says its faster-whisper backend needs under 8 GB of GPU memory for large-v2 at beam size 5. The project is candid about its limits: overlapping speech is handled poorly, tokens such as “£13.60” that contain characters outside the alignment dictionary get no timestamp, diarization is far from perfect, and every language needs its own alignment model.
NVIDIA alternatives. Parakeet TDT 0.6B v2 (English) and v3 (25 European languages), along with Canary-1B-v2 (the same 25 languages plus speech translation), are CC-BY-4.0 models that run on NVIDIA’s Apache-licensed NeMo toolkit. Parakeet v2 transcribes up to 24 minutes in one pass with punctuation, capitalisation and word timestamps. If your languages fall inside that set, test them alongside Whisper.
Throughput on GPU and CPU, from the projects’ benchmarks
faster-whisper publishes the most useful numbers because it names both the hardware and the audio: 13 minutes of speech, run on an NVIDIA RTX 3070 Ti with 8 GB for the GPU tests and an Intel Core i7-12700K on 8 threads for the CPU tests.
| Setup in the faster-whisper benchmark | Time for 13 minutes of audio | Memory |
|---|---|---|
| large-v2, GPU, fp16, beam size 5 | 1 min 03 s | 4,525 MB VRAM |
| large-v2, GPU, int8, beam size 5 | 59 s | 2,926 MB VRAM |
| large-v2, GPU, fp16, batch size 8 | 17 s | 6,090 MB VRAM |
| small, CPU, fp32, beam size 5 | 2 min 37 s | 2,257 MB RAM |
| small, CPU, int8, beam size 5 | 1 min 42 s | 1,477 MB RAM |
| small, CPU, fp32, batch size 8 | 1 min 06 s | 4,230 MB RAM |
Converted into real-time factors, batched large-v2 on that consumer GPU got through the 780 seconds of audio about 46 times faster than real time, and the small model on a desktop CPU ran about 12 times faster than real time with batching. Those are the project’s measurements on its own test file, not a promise for your recordings, yet they suggest one GPU goes a long way for batch transcription and that CPU-only sites can work with smaller models.
whisper.cpp lists memory instead of speed: about 273 MB for tiny, 388 MB for base, 852 MB for small, 2.1 GB for medium and 3.9 GB for large, before quantization trims it further. When a language model will also summarise the transcripts, size the combined host with our server sizing guide.
Azure speech to text on premise, and other vendor routes
Microsoft ships Azure Speech in Foundry Tools as Docker containers you can run on your own hosts. The current list (verified October 2026):
| Container | Latest version | Disconnected mode |
|---|---|---|
| Speech to text | 5.1.0 | Yes, after approval |
| Custom speech to text | 5.1.0 | Yes, after approval |
| Neural text to speech | 4.1.0 | Yes, after approval |
| Fast transcription | 1.0.0, public preview | Not listed |
| Speech language identification | 1.18.0, public preview | Not listed |
The conditions deserve a careful read:
- Connected mode still depends on Azure. Containers report usage about every 10 to 15 minutes. A container that loses its billing connection keeps running but stops answering queries until the connection returns, and Microsoft documents ten reconnection attempts before it stops serving requests.
- Disconnected mode is granted by application. You submit a request form tied to an Azure subscription, Microsoft replies within 10 business days, and access is limited to organisations identified as strategic customers or partners with offline, remote or strictly regulated use cases.
- Commitment plans are annual and paid up front. Disconnected containers use commitment-tier pricing with a calendar-year term, charged in full at purchase. The downloaded license file carries an expiry date, so renewal belongs in the operations calendar.
- The hardware is CPU-based. Hosts must be x64 with AVX2. Speech to text needs at least 4 cores and 4 GB of memory, with 8 cores and 8 GB recommended, plus 4 to 8 GB for the speech model, and each core must run at 2.6 GHz or faster.
- Images are per locale. The latest tag pulls en-US, and other locales have their own tags.
- Live diarization needs Redis. From version 5.1.0, speaker diarization in the container requires a customer-operated Redis-compatible cache, which keeps four hours of diarization data by default.
Google Distributed Cloud air-gapped includes Speech-to-Text among its three Vertex AI pre-trained APIs. It offers synchronous recognition for up to a minute of audio, asynchronous recognition for up to 480 minutes, and streaming with interim results, and Google says the recognition models are under 1 GB. It comes as part of that air-gapped platform.
NVIDIA Riva packages recognition, synthesis and translation as GPU-accelerated containers for on-premises, cloud, edge and embedded deployment. Continued production use is sold as part of NVIDIA AI Enterprise, with a 90-day trial licence available.
Where Whisper transcription goes wrong
Hallucinated text. The large-v3 card warns that, because Whisper was trained on large-scale noisy data, its output may contain text that was never spoken, and that the architecture is prone to repetition. The “Careless Whisper” paper at ACM FAccT 2024 found entire invented phrases or sentences in roughly 1% of transcriptions made through OpenAI’s Whisper API in 2023. It judged 38% of those hallucinations to carry explicit harms, such as invented violence or false authority, and found them more often for speakers with longer non-vocal stretches, a common symptom of aphasia.
The practical defence is to strip silence before decoding. faster-whisper integrates the Silero VAD model to drop audio without speech, whisper.cpp passes only detected speech segments to the model when run with —vad, and WhisperX credits its VAD preprocessing with reducing hallucination and enabling batched inference with no loss in word error rate. Route transcripts of clinical, legal or HR recordings to human review before anyone acts on them.
Uneven language accuracy. OpenAI notes that performance varies widely by language and publishes per-language error rates for large-v3 and large-v2 on Common Voice 15 and FLEURS. Check your languages in that breakdown first, then test on your own audio. For the 25 European languages NVIDIA covers, Parakeet v3 and Canary-1B-v2 are worth a side-by-side run.
Batch versus live audio. Whisper was designed around 30-second windows. whisper_streaming adds a local-agreement policy and reports 3.3 seconds of latency on long-form speech, and its authors now point users to a successor, SimulStreaming. whisper.cpp’s stream example samples the microphone every half second and transcribes continuously. Those approaches suit live captions; conversational assistants need the tighter budget described in our guide to on-premise voice AI.
Testing word error rate on your own recordings
Leaderboards rank models on public datasets. The Open ASR Leaderboard, for instance, reports word error rate and an RTFx speed figure for each model. Your recordings have their own microphones, accents, noise and vocabulary, so a short evaluation on them tells you more than any public ranking.
Word error rate counts substitutions (S), deletions (D) and insertions (I) against the number of words in a human reference transcript (N): WER = (S + D + I) / N. Lower is better, and because insertions are unbounded, a model that invents whole sentences can score above 100%.
- Sample real traffic. Pull recordings from every channel you will transcribe, such as headsets, meeting rooms, phone lines and field recorders, including the accents and jargon your people actually use.
- Write reference transcripts by hand. Have people transcribe the sample against one style guide for numbers, names and filler words.
- Normalise both sides. Whisper ships an EnglishTextNormalizer and a BasicTextNormalizer; apply the same normalisation to references and model output so casing and punctuation do not count as errors.
- Score with a standard tool. jiwer, an Apache 2.0 Python package, computes WER, CER and related measures.
- Break the results down. Report WER per channel, speaker group and language rather than one average, and count insertions on silent or music-only clips separately.
- Time it on target hardware. Record the real-time factor and memory use at the batch size you plan to run.
- Re-run after every change. New model versions, quantization levels and VAD settings can move accuracy in either direction.
An air-gapped Whisper deployment, step by step
The pattern follows any air-gapped AI deployment: build and verify outside, transfer once, then run with no outbound path.
- Pin the stack on a connected staging host. Fix the versions of the runtime, CUDA and cuDNN (CUDA 12 with cuDNN 9 for current faster-whisper) and build a container image.
- Download only the weights you need. faster-whisper loads converted CTranslate2 models from the Hugging Face Hub or from a local directory, whisper.cpp ships a download script for its ggml models, and the pyannote diarization pipeline is cloned with git-lfs once its conditions are accepted on Hugging Face.
- Record checksums and licences. Hash every file and keep each model’s card and licence beside it.
- Transfer through your approved media process. Re-verify the hashes inside the enclave before loading anything.
- Force offline mode. Set HF_HUB_OFFLINE=1 so the Hugging Face libraries make no HTTP calls and read only cached files, set HF_HUB_DISABLE_TELEMETRY=1, and load models from local paths.
- Block egress at the network layer. Library settings are a second line of defence; transcription hosts should have no outbound route at all.
- Decide what happens to the audio. Set retention periods for source recordings and transcripts, and decide who can access each.
- Keep the evaluation set inside. Re-run your word error rate test after every model promotion, against the same references.
How VDF AI fits
VDF AI gives agents three transcription tools. The speech-to-text tool handles short clips with automatic language detection, the audio transcription tool handles long recordings with timestamps and speaker separation, and the video transcription tool transcribes the audio track of MP4, WebM or MOV files into timestamped segments. Transcription can run on a local model inside your perimeter, so recordings are not uploaded to a hosted transcription service.
The transcripts come back as text an agent can summarise, search or cite, and each video transcription run is logged with the recording, the agent and the model used. The tools work on recordings you supply; VDF AI is not a telephony or meeting-bot product.
Sources
- OpenAI Whisper repository
- Radford et al., Robust Speech Recognition via Large-Scale Weak Supervision (arXiv)
- Whisper large-v3 model card
- Whisper large-v3-turbo model card
- faster-whisper
- whisper.cpp
- WhisperX
- pyannote speaker-diarization-community-1
- Silero VAD
- NVIDIA Parakeet TDT 0.6B v2 model card
- NVIDIA Parakeet TDT 0.6B v3 model card
- NVIDIA Canary-1B-v2 model card
- NVIDIA NeMo
- Koenecke et al., Careless Whisper (arXiv, ACM FAccT 2024)
- whisper_streaming
- Open ASR Leaderboard
- WER definition, Hugging Face evaluate
- jiwer
- Hugging Face Hub environment variables
- Microsoft Learn, Speech containers overview
- Microsoft Learn, install and run Speech containers
- Microsoft Learn, speech to text containers
- Microsoft Learn, containers in disconnected environments
- Google, Speech-to-Text on GDC air-gapped
- Google, Vertex AI on GDC air-gapped
- NVIDIA Riva, getting started