On-premise voice AI runs speech recognition, a language model and text to speech on servers you control, so recordings, transcripts and generated speech never pass through a cloud speech API. A workable 2026 stack pairs Whisper or NVIDIA Parakeet for recognition with a local LLM and Kokoro or Piper for synthesis, with voice activity detection judging when a speaker has finished.
What an on-premise voice AI stack contains
Voice projects come in a few shapes, and each one loads the same components differently:
- Dictation. A person speaks, the system writes, the person edits. Recognition accuracy decides success, and a second of delay barely matters.
- Turn-based voice assistants. The user asks aloud and hears the answer aloud. Every stage sits on the path of every reply, so latency becomes the main design constraint.
- Full-duplex voice agents. The system keeps listening while it talks and copes with interruptions. That needs streaming at every stage, and some teams replace the chain with a single speech-to-speech model.
Transcribing recorded calls, interviews and meetings is a batch job with different trade-offs, covered in our self-hosted transcription guide.
| Stage | Job in the pipeline | Open-source options (verified October 2026) |
|---|---|---|
| Voice activity detection | Finds speech in the stream and decides when a turn has ended | Silero VAD (MIT) |
| Speech recognition (ASR) | Converts speech to text, ideally with partial results while the user is still talking | Whisper and its runtimes, NVIDIA Parakeet and Canary |
| Language model | Interprets the request, calls tools and retrieval, drafts the reply | An open-weight LLM served on your own GPUs |
| Text to speech (TTS) | Speaks the reply, ideally one sentence at a time | Kokoro, Piper, Coqui XTTS-v2 |
| Speaker diarization | Labels who spoke when in multi-speaker audio | pyannote.audio |
| Orchestration | Moves audio between stages and handles interruptions and transport | Pipecat, LiveKit Agents |
Why voice data stays inside the perimeter
A recording holds more than the words in it. The European Data Protection Board’s guidelines on virtual voice assistants point out that voice data carries personal data both in what is said and in meta-information such as the speaker’s sex or age, and they recall that voice data is inherently biometric personal data. Where it is processed to uniquely identify someone, the controller needs a derogation under Article 9 of the GDPR as well as an Article 6 legal basis. The same guidelines note that assistants can capture people who never meant to use them.
Clinical dictation brings in HIPAA. Its definition of health information covers information “whether oral or recorded in any form or medium”, so a physician’s dictated note about a patient is protected health information from the moment it is recorded, whichever system transcribes it.
Recording rules vary by jurisdiction. US federal law permits a recording when one party to the conversation consents, under 18 U.S.C. § 2511(2)(d), while California Penal Code § 632 requires the consent of all parties to a confidential communication. Running capture and transcription in-house leaves those consent duties in place, but it keeps the recordings on systems governed by your own retention and access rules.
Customer-facing voices also fall under the EU AI Act. Since 2 August 2026, Article 50(1) has required providers to design systems that interact directly with people so that those people are told they are dealing with an AI system, unless that is obvious, and Article 50(2) requires synthetic audio to be marked in a machine-readable, detectable way. Our Article 50 explainer covers the marking grace period and who counts as the provider.
Speech recognition and text-to-speech models compared
For recognition, three families cover most on-premise projects. Sizes and licences come from each project’s own repository or model card (verified October 2026):
| Model | Parameters | Languages | Licence | Worth knowing |
|---|---|---|---|---|
| Whisper large-v3-turbo | 809M | Multilingual | MIT | Large-v3 with the decoder cut from 32 layers to 4; about 6 GB of VRAM; not trained for translation |
| NVIDIA Parakeet TDT 0.6B v3 | 600M | 25 European languages | CC-BY-4.0 | Punctuation and word timestamps; up to 24 minutes of audio per pass with full attention |
| NVIDIA Canary-1B-v2 | 978M | 25 European languages | CC-BY-4.0 | Also translates into and out of English; runs on NVIDIA’s NeMo toolkit |
Runtimes, hallucination fixes and word error rate testing for Whisper are covered in the transcription guide linked above.
For text to speech, licence terms separate the options more sharply than voice quality does:
| Option | Languages and voices | Licence | Watch for |
|---|---|---|---|
| Kokoro-82M | 8 languages, 54 voices, 82M parameters | Apache 2.0 | Its card lists the training audio, which includes synthetic speech generated by closed TTS models from large providers |
| Piper | Voices published per language, each with its own model card | Engine GPL-3.0 (the archived original repository was MIT); each voice links its dataset’s licence | Embeds espeak-ng; the Open Home Foundation is looking for maintainers |
| Coqui XTTS-v2 | 17 languages; clones a voice from a 6-second clip | Coqui Public Model License, non-commercial purposes only | Ruled out for commercial products under that licence |
| Azure Neural TTS container | Voices and locales listed in Microsoft’s container registry | Commercial, metered through an Azure resource | 6 cores and 12 GB of memory at minimum |
| NVIDIA Riva | GPU-accelerated speech synthesis alongside recognition and translation | Production use through NVIDIA AI Enterprise | Ships as containers for Docker or Kubernetes |
Read Piper’s voice cards before a commercial launch. The US English lessac voice, for example, links the licence of the Blizzard 2013 dataset it was trained on, while the British alba voice links CC BY 4.0. Voice cloning raises a separate question: record whose voice was cloned and on what permission before anyone generates speech with it.
Serving and orchestration choices
Models are half the build. Something has to stream audio in, hand partial transcripts to the language model, start speaking before the model has finished, and stop speaking when the user interrupts.
Open frameworks. Pipecat, licensed BSD-2-Clause, is an open-source framework for voice agents and multimodal apps, and it lists local services such as Whisper, Piper and Ollama among its integrations. LiveKit Agents is Apache 2.0 and can run with LiveKit’s open-source WebRTC media server on your own machines. Its semantic turn detection uses a transformer model to judge when a user has finished speaking, which LiveKit says helps reduce interruptions.
Packaged speech services. NVIDIA Riva is a GPU-accelerated SDK for speech recognition, synthesis and translation that deploys with Docker or Helm charts, on premises or in the cloud. NVIDIA sells ongoing production use as part of NVIDIA AI Enterprise and offers a 90-day trial licence. Microsoft publishes Docker containers for Azure Speech in Foundry Tools covering speech to text, custom speech to text and neural text to speech, plus two containers in preview. They report usage to Azure in standard mode, and only approved customers on a commitment plan can run them disconnected; the transcription guide walks through those conditions.
Speech-to-speech models. Kyutai’s Moshi is a full-duplex alternative to the chain. Its repository reports a theoretical latency of 160 ms and a practical latency as low as 200 ms on an L4 GPU, with model weights under CC-BY 4.0. The trade-off is architectural: the text hand-off between stages is also where most teams attach retrieval, tool calls, policy checks and audit logging.
A latency budget for one spoken turn
People take turns quickly. A 2009 PNAS study of conversations in ten languages found that the most common gap between one speaker finishing and the next starting fell between 0 and 200 milliseconds in every language, and the mean across the whole dataset was 208 milliseconds. A cascaded assistant has to notice that the turn is over, transcribe, generate and synthesise before any sound comes back, so the work is to shrink each stage and overlap them.
| Stage | What sets the delay | How to shorten it |
|---|---|---|
| End-of-turn detection | How long the system waits in silence before treating the turn as finished | Tune the silence threshold on real recordings, or add a semantic turn detector |
| Final transcript | Model size, decoding settings and whether partial results stream | Streaming recognition, smaller or turbo models, models kept loaded on the GPU |
| First token from the LLM | Prompt length, model size and queueing on a busy GPU | Short prompts, smaller models for routine turns, capacity reserved for voice traffic |
| First audio from TTS | Synthesis speed and how much text the engine waits for | Synthesise sentence by sentence while the LLM is still writing |
| Playback | Audio buffers and network hops | Keep media servers close to users and avoid extra relays |
Whisper itself reads audio in 30-second windows, which is why streaming wrappers exist. One of them, whisper_streaming, reports 3.3 seconds of latency on unsegmented long-form speech: workable for live captions, far from conversational timing.
Track a single end-to-end number: the time from the end of the user’s speech to the first audio sample played back, at the 95th percentile, with as many concurrent sessions as you expect at peak.
Sizing hardware for voice workloads
Published requirements for the speech components are modest next to those of a language model (verified October 2026):
| Component | Published requirement | Source |
|---|---|---|
| Whisper turbo | About 6 GB of VRAM | OpenAI’s Whisper repository |
| Whisper large | About 10 GB of VRAM | OpenAI’s Whisper repository |
| faster-whisper, large-v2 with 8-bit weights | 2,926 MB of VRAM in the project’s 13-minute benchmark | faster-whisper |
| whisper.cpp, large model | About 3.9 GB of memory | whisper.cpp |
| Parakeet TDT 0.6B v3 | At least 2 GB of RAM to load | NVIDIA model card |
| Canary-1B-v2 | At least 6 GB of RAM to load | NVIDIA model card |
| Azure Neural TTS container | 6 cores and 12 GB minimum; 8 cores and 16 GB recommended | Microsoft Learn |
| Moshi, PyTorch version | A GPU with 24 GB of memory | Kyutai |
The language model normally sets the budget. Size it first with the method in our GPU estimation guide, then add recognition and synthesis capacity for the sessions open at the same moment. A live voice session keeps audio flowing in both directions for as long as the conversation lasts, whereas a chat request frees the GPU between messages, so plan from concurrent conversations at peak rather than from named users. Our AI server guide covers chassis, GPU and power choices once those numbers are known.
Deployment checklist for on-premise voice
- Pick the shape first. Dictation, turn-based assistant or full-duplex agent: the latency target and the hardware follow from that choice.
- Clear every licence. Check code and weights separately. XTTS-v2 is non-commercial, Piper’s engine is GPL-3.0 with per-voice dataset licences, and NVIDIA’s CC-BY-4.0 models require attribution.
- Test recognition on your own audio. Accents, headsets, speakerphones, background noise and domain vocabulary all move accuracy, so measure word error rate before choosing a model.
- Put voice activity detection in front of recognition. Silero VAD’s maintainers report that a 30 ms chunk takes under a millisecond on one CPU thread, so the cost of adding it is small.
- Stream every stage. Partial transcripts, token streaming and sentence-level synthesis are what make a reply feel prompt.
- Set retention for audio and transcripts. Decide how long raw audio is kept once a reviewed transcript exists, and who may replay it.
- Disclose and mark. Tell people they are talking to an AI system where Article 50(1) applies, and mark synthetic audio if you are its provider.
- Control cloned voices. Keep the consent record with the voice model and restrict who can generate speech with it.
- Log each turn. Keep what was heard, what was transcribed, which model answered and what was spoken, under the same access controls as the audio.
- Load-test at peak concurrency. Watch the gap from end of speech to first audio at the 95th percentile, not the average.
- Mirror models for offline updates. Pull new weights on a connected staging host, verify checksums and promote them like any other release.
How VDF AI fits
VDF AI’s voice capabilities are transcription, synthesis and dictation. Agents can be assigned a speech-to-text tool for short clips, an audio transcription tool for long recordings with timestamps and speaker segments, a video transcription tool for the audio track of recordings, and a text-to-speech tool with voice, speed and MP3 or WAV output. Each can run on a local model inside your perimeter, so the audio and the text produced from it stay on infrastructure you control.
In VDF AI Chat, workspaces with voice input enabled let people dictate prompts instead of typing them, and the documentation asks users to review the transcribed text before sending. Most agents accept spoken input the same way.
VDF AI does not include telephony or a real-time voice-agent product. Its voice tools work on dictated and recorded speech.
Sources
- EDPB, Guidelines 02/2021 on virtual voice assistants
- 45 CFR 160.103, HIPAA definitions (Cornell LII)
- 18 U.S.C. § 2511 (Cornell LII)
- California Penal Code § 632
- AI Act, Regulation (EU) 2024/1689 (EUR-Lex)
- Stivers et al., turn-taking across ten languages (PNAS, 2009)
- OpenAI Whisper repository
- Whisper large-v3-turbo model card
- faster-whisper
- whisper.cpp
- whisper_streaming
- NVIDIA Parakeet TDT 0.6B v3 model card
- NVIDIA Canary-1B-v2 model card
- Kokoro-82M model card
- Piper, OHF-Voice repository
- Piper voices on Hugging Face
- Coqui XTTS-v2 model card and licence
- Silero VAD
- pyannote.audio
- NVIDIA Riva, getting started
- Microsoft Learn, Speech containers overview
- Microsoft Learn, install and run Speech containers
- Microsoft Learn, containers in disconnected environments
- Pipecat
- LiveKit Agents
- Kyutai Moshi