AI Infrastructure

Whisper On-Premise: How to Self-Host Speech to Text Without a Cloud API

A practical guide to on-premise speech to text with OpenAI Whisper in 2026: model sizes and VRAM, faster-whisper, whisper.cpp and WhisperX, NVIDIA and Azure alternatives, hallucination and accuracy caveats, word error rate testing and an air-gapped setup.

Whisper on premise means running OpenAI's open-weight speech recognition model on your own hardware, usually through faster-whisper or whisper.cpp, so audio never reaches a hosted transcription API. OpenAI releases Whisper's code and weights under the MIT License. The turbo model needs about 6 GB of VRAM, and a voice activity detector in front of it reduces hallucinated text during silence.

Whisper model sizes, memory and licence

OpenAI’s repository lists six sizes. The VRAM figures and relative speeds below are OpenAI’s own, with speed measured against the large model (verified October 2026):

SizeParametersEnglish-only variantVRAMRelative speed
tiny39Mtiny.en~1 GB~10x
base74Mbase.en~1 GB~7x
small244Msmall.en~2 GB~4x
medium769Mmedium.en~5 GB~2x
large1,550Mnone~10 GB1x
turbo809Mnone~6 GB~8x

Three details shape the choice:

  • large-v3 is the accuracy reference. The original Whisper paper trained on 680,000 hours of multilingual, multitask audio. The large-v3 card adds 1 million hours of weakly labelled and 4 million hours of pseudo-labelled audio, moves to 128 Mel frequency bins from 80, and reports 10% to 20% fewer errors than large-v2.
  • turbo trades a little quality for a lot of speed. It is a fine-tuned, pruned large-v3 with 4 decoder layers instead of 32. OpenAI notes that it was not trained for translation and returns the original language even when asked to translate.
  • Audio is processed in 30-second windows. The reference transcribe() method slides a 30-second window across the file, so long recordings are handled in pieces and live audio needs extra machinery.

On licensing, OpenAI states that Whisper’s code and model weights are released under the MIT License. On Hugging Face the large-v3 card carries an Apache 2.0 tag and the turbo card shows MIT, and both licences permit commercial use.

Runtimes for self-hosted Whisper compared

Production deployments rarely run the reference Python package. These are the usual alternatives (verified October 2026):

RuntimeWhat it isBest fitLicenceWatch for
faster-whisperWhisper reimplemented on the CTranslate2 inference enginePython services on NVIDIA GPUs, or CPUs with 8-bit modelsMITCurrent CTranslate2 releases support only CUDA 12 with cuDNN 9
whisper.cppC/C++ inference on the ggml libraryCPU-only hosts, Apple Silicon, Vulkan, AMD ROCm, OpenVINO, desktop and embedded appsMITUses converted ggml models; integer quantization lowers memory
WhisperXfaster-whisper plus wav2vec2 alignment and pyannote diarizationWord-level timestamps and speaker labelsBSD-2-ClauseDiarization needs a Hugging Face token and accepted model conditions
openai-whisperThe reference implementationExperiments and reproducing OpenAI’s published numbersMITRequires ffmpeg; the baseline other runtimes compare against

faster-whisper claims up to four times the speed of openai/whisper at the same accuracy while using less memory, and its BatchedInferencePipeline drops in for the standard transcribe call. whisper.cpp treats Apple Silicon as a first-class target through Metal and Core ML, and also supports NVIDIA GPUs, Vulkan, AMD ROCm, OpenVINO, several NPUs and plain CPUs.

WhisperX reports 70x real-time transcription with large-v2 through batched inference, and says its faster-whisper backend needs under 8 GB of GPU memory for large-v2 at beam size 5. The project is candid about its limits: overlapping speech is handled poorly, tokens such as “£13.60” that contain characters outside the alignment dictionary get no timestamp, diarization is far from perfect, and every language needs its own alignment model.

NVIDIA alternatives. Parakeet TDT 0.6B v2 (English) and v3 (25 European languages), along with Canary-1B-v2 (the same 25 languages plus speech translation), are CC-BY-4.0 models that run on NVIDIA’s Apache-licensed NeMo toolkit. Parakeet v2 transcribes up to 24 minutes in one pass with punctuation, capitalisation and word timestamps. If your languages fall inside that set, test them alongside Whisper.

Throughput on GPU and CPU, from the projects’ benchmarks

faster-whisper publishes the most useful numbers because it names both the hardware and the audio: 13 minutes of speech, run on an NVIDIA RTX 3070 Ti with 8 GB for the GPU tests and an Intel Core i7-12700K on 8 threads for the CPU tests.

Setup in the faster-whisper benchmarkTime for 13 minutes of audioMemory
large-v2, GPU, fp16, beam size 51 min 03 s4,525 MB VRAM
large-v2, GPU, int8, beam size 559 s2,926 MB VRAM
large-v2, GPU, fp16, batch size 817 s6,090 MB VRAM
small, CPU, fp32, beam size 52 min 37 s2,257 MB RAM
small, CPU, int8, beam size 51 min 42 s1,477 MB RAM
small, CPU, fp32, batch size 81 min 06 s4,230 MB RAM

Converted into real-time factors, batched large-v2 on that consumer GPU got through the 780 seconds of audio about 46 times faster than real time, and the small model on a desktop CPU ran about 12 times faster than real time with batching. Those are the project’s measurements on its own test file, not a promise for your recordings, yet they suggest one GPU goes a long way for batch transcription and that CPU-only sites can work with smaller models.

whisper.cpp lists memory instead of speed: about 273 MB for tiny, 388 MB for base, 852 MB for small, 2.1 GB for medium and 3.9 GB for large, before quantization trims it further. When a language model will also summarise the transcripts, size the combined host with our server sizing guide.

Azure speech to text on premise, and other vendor routes

Microsoft ships Azure Speech in Foundry Tools as Docker containers you can run on your own hosts. The current list (verified October 2026):

ContainerLatest versionDisconnected mode
Speech to text5.1.0Yes, after approval
Custom speech to text5.1.0Yes, after approval
Neural text to speech4.1.0Yes, after approval
Fast transcription1.0.0, public previewNot listed
Speech language identification1.18.0, public previewNot listed

The conditions deserve a careful read:

  • Connected mode still depends on Azure. Containers report usage about every 10 to 15 minutes. A container that loses its billing connection keeps running but stops answering queries until the connection returns, and Microsoft documents ten reconnection attempts before it stops serving requests.
  • Disconnected mode is granted by application. You submit a request form tied to an Azure subscription, Microsoft replies within 10 business days, and access is limited to organisations identified as strategic customers or partners with offline, remote or strictly regulated use cases.
  • Commitment plans are annual and paid up front. Disconnected containers use commitment-tier pricing with a calendar-year term, charged in full at purchase. The downloaded license file carries an expiry date, so renewal belongs in the operations calendar.
  • The hardware is CPU-based. Hosts must be x64 with AVX2. Speech to text needs at least 4 cores and 4 GB of memory, with 8 cores and 8 GB recommended, plus 4 to 8 GB for the speech model, and each core must run at 2.6 GHz or faster.
  • Images are per locale. The latest tag pulls en-US, and other locales have their own tags.
  • Live diarization needs Redis. From version 5.1.0, speaker diarization in the container requires a customer-operated Redis-compatible cache, which keeps four hours of diarization data by default.

Google Distributed Cloud air-gapped includes Speech-to-Text among its three Vertex AI pre-trained APIs. It offers synchronous recognition for up to a minute of audio, asynchronous recognition for up to 480 minutes, and streaming with interim results, and Google says the recognition models are under 1 GB. It comes as part of that air-gapped platform.

NVIDIA Riva packages recognition, synthesis and translation as GPU-accelerated containers for on-premises, cloud, edge and embedded deployment. Continued production use is sold as part of NVIDIA AI Enterprise, with a 90-day trial licence available.

Where Whisper transcription goes wrong

Hallucinated text. The large-v3 card warns that, because Whisper was trained on large-scale noisy data, its output may contain text that was never spoken, and that the architecture is prone to repetition. The “Careless Whisper” paper at ACM FAccT 2024 found entire invented phrases or sentences in roughly 1% of transcriptions made through OpenAI’s Whisper API in 2023. It judged 38% of those hallucinations to carry explicit harms, such as invented violence or false authority, and found them more often for speakers with longer non-vocal stretches, a common symptom of aphasia.

The practical defence is to strip silence before decoding. faster-whisper integrates the Silero VAD model to drop audio without speech, whisper.cpp passes only detected speech segments to the model when run with —vad, and WhisperX credits its VAD preprocessing with reducing hallucination and enabling batched inference with no loss in word error rate. Route transcripts of clinical, legal or HR recordings to human review before anyone acts on them.

Uneven language accuracy. OpenAI notes that performance varies widely by language and publishes per-language error rates for large-v3 and large-v2 on Common Voice 15 and FLEURS. Check your languages in that breakdown first, then test on your own audio. For the 25 European languages NVIDIA covers, Parakeet v3 and Canary-1B-v2 are worth a side-by-side run.

Batch versus live audio. Whisper was designed around 30-second windows. whisper_streaming adds a local-agreement policy and reports 3.3 seconds of latency on long-form speech, and its authors now point users to a successor, SimulStreaming. whisper.cpp’s stream example samples the microphone every half second and transcribes continuously. Those approaches suit live captions; conversational assistants need the tighter budget described in our guide to on-premise voice AI.

Testing word error rate on your own recordings

Leaderboards rank models on public datasets. The Open ASR Leaderboard, for instance, reports word error rate and an RTFx speed figure for each model. Your recordings have their own microphones, accents, noise and vocabulary, so a short evaluation on them tells you more than any public ranking.

Word error rate counts substitutions (S), deletions (D) and insertions (I) against the number of words in a human reference transcript (N): WER = (S + D + I) / N. Lower is better, and because insertions are unbounded, a model that invents whole sentences can score above 100%.

  1. Sample real traffic. Pull recordings from every channel you will transcribe, such as headsets, meeting rooms, phone lines and field recorders, including the accents and jargon your people actually use.
  2. Write reference transcripts by hand. Have people transcribe the sample against one style guide for numbers, names and filler words.
  3. Normalise both sides. Whisper ships an EnglishTextNormalizer and a BasicTextNormalizer; apply the same normalisation to references and model output so casing and punctuation do not count as errors.
  4. Score with a standard tool. jiwer, an Apache 2.0 Python package, computes WER, CER and related measures.
  5. Break the results down. Report WER per channel, speaker group and language rather than one average, and count insertions on silent or music-only clips separately.
  6. Time it on target hardware. Record the real-time factor and memory use at the batch size you plan to run.
  7. Re-run after every change. New model versions, quantization levels and VAD settings can move accuracy in either direction.

An air-gapped Whisper deployment, step by step

The pattern follows any air-gapped AI deployment: build and verify outside, transfer once, then run with no outbound path.

  1. Pin the stack on a connected staging host. Fix the versions of the runtime, CUDA and cuDNN (CUDA 12 with cuDNN 9 for current faster-whisper) and build a container image.
  2. Download only the weights you need. faster-whisper loads converted CTranslate2 models from the Hugging Face Hub or from a local directory, whisper.cpp ships a download script for its ggml models, and the pyannote diarization pipeline is cloned with git-lfs once its conditions are accepted on Hugging Face.
  3. Record checksums and licences. Hash every file and keep each model’s card and licence beside it.
  4. Transfer through your approved media process. Re-verify the hashes inside the enclave before loading anything.
  5. Force offline mode. Set HF_HUB_OFFLINE=1 so the Hugging Face libraries make no HTTP calls and read only cached files, set HF_HUB_DISABLE_TELEMETRY=1, and load models from local paths.
  6. Block egress at the network layer. Library settings are a second line of defence; transcription hosts should have no outbound route at all.
  7. Decide what happens to the audio. Set retention periods for source recordings and transcripts, and decide who can access each.
  8. Keep the evaluation set inside. Re-run your word error rate test after every model promotion, against the same references.

How VDF AI fits

VDF AI gives agents three transcription tools. The speech-to-text tool handles short clips with automatic language detection, the audio transcription tool handles long recordings with timestamps and speaker separation, and the video transcription tool transcribes the audio track of MP4, WebM or MOV files into timestamped segments. Transcription can run on a local model inside your perimeter, so recordings are not uploaded to a hosted transcription service.

The transcripts come back as text an agent can summarise, search or cite, and each video transcription run is logged with the recording, the agent and the model used. The tools work on recordings you supply; VDF AI is not a telephony or meeting-bot product.

Sources

Frequently asked questions

Can I run OpenAI Whisper on premise?

Yes. OpenAI publishes Whisper's code and model weights under the MIT License, so you can download them once and transcribe on your own servers with no connection to OpenAI. Most production deployments swap the reference package for a faster runtime: faster-whisper on NVIDIA GPUs, or whisper.cpp for CPUs, Apple Silicon and other accelerators. For an air-gapped site, carry the converted model files across with checksums and switch the Hugging Face libraries to offline mode so nothing tries to reach the internet when models load.

How much VRAM does Whisper need?

OpenAI's own table gives roughly 1 GB for tiny and base, 2 GB for small, 5 GB for medium, 10 GB for large and 6 GB for turbo. Runtimes change those figures. In faster-whisper's benchmark, large-v2 transcribed 13 minutes of audio using 4,525 MB of VRAM at 16-bit precision and 2,926 MB with 8-bit quantization on an RTX 3070 Ti. whisper.cpp lists about 3.9 GB of memory for the large model. Batching raises memory use in exchange for higher throughput.

Is Whisper free for commercial use?

OpenAI's repository states that Whisper's code and weights are released under the MIT License, which allows commercial use, and the large-v3-turbo model card also shows MIT. The large-v3 card on Hugging Face carries an Apache 2.0 tag, which is permissive as well. Check what sits around the model too: faster-whisper and whisper.cpp are MIT, WhisperX is BSD-2-Clause, and the pyannote diarization pipeline that WhisperX uses is CC-BY-4.0 and gated behind accepted conditions on Hugging Face.

Why does Whisper make up text when nobody is speaking?

Whisper learned from large amounts of weakly labelled audio, and its model card warns that it can output text that was never spoken and can repeat itself. A paper presented at ACM FAccT 2024 found entire invented phrases or sentences in roughly 1% of transcriptions made through OpenAI's Whisper API in 2023, more often for speakers with long non-vocal stretches. Remove silence before decoding: faster-whisper integrates Silero VAD, whisper.cpp has a --vad option and WhisperX applies VAD preprocessing by design.

Can Azure speech to text run on premise?

Yes, as a Docker container on your own x64 hosts with AVX2 support. In standard mode the container still needs an Azure resource and reports usage about every 10 to 15 minutes, and it stops answering queries when it cannot reach the billing endpoint. A fully disconnected mode exists for speech to text, custom speech to text and neural text to speech, but it requires an approved application and a commitment-tier plan bought for a calendar year, with a license file that expires.

Should I use faster-whisper or whisper.cpp?

Choose faster-whisper for Python services on NVIDIA GPUs: it reimplements Whisper on CTranslate2, claims up to four times the speed of the reference code for the same accuracy and supports 8-bit quantization and batched inference. Choose whisper.cpp for C or C++ integration, CPU-only servers, Apple Silicon, Vulkan, AMD ROCm or a small footprint. Add WhisperX on top of faster-whisper when you need word-level timestamps and speaker labels. Benchmark the shortlist on your own recordings before committing.

Filed under
on-premises AIair-gapped AIopen-weight modelslocal AI infrastructuredata sovereigntymodel serving
On-Prem AI

Plan your on-prem AI deployment

Book an architecture call and we will scope a private, on-prem AI deployment for your environment — integrations, hardware, and governance included.

Keep reading