AI Infrastructure

On-Premise Voice AI: Speech Recognition, LLMs and Text to Speech on Your Own Servers

How an on-premise voice AI stack fits together in 2026: the open speech recognition and text-to-speech models worth shortlisting, their licences, the latency budget for a spoken turn, hardware sizing and a deployment checklist.

On-premise voice AI runs speech recognition, a language model and text to speech on servers you control, so recordings, transcripts and generated speech never pass through a cloud speech API. A workable 2026 stack pairs Whisper or NVIDIA Parakeet for recognition with a local LLM and Kokoro or Piper for synthesis, with voice activity detection judging when a speaker has finished.

What an on-premise voice AI stack contains

Voice projects come in a few shapes, and each one loads the same components differently:

  • Dictation. A person speaks, the system writes, the person edits. Recognition accuracy decides success, and a second of delay barely matters.
  • Turn-based voice assistants. The user asks aloud and hears the answer aloud. Every stage sits on the path of every reply, so latency becomes the main design constraint.
  • Full-duplex voice agents. The system keeps listening while it talks and copes with interruptions. That needs streaming at every stage, and some teams replace the chain with a single speech-to-speech model.

Transcribing recorded calls, interviews and meetings is a batch job with different trade-offs, covered in our self-hosted transcription guide.

StageJob in the pipelineOpen-source options (verified October 2026)
Voice activity detectionFinds speech in the stream and decides when a turn has endedSilero VAD (MIT)
Speech recognition (ASR)Converts speech to text, ideally with partial results while the user is still talkingWhisper and its runtimes, NVIDIA Parakeet and Canary
Language modelInterprets the request, calls tools and retrieval, drafts the replyAn open-weight LLM served on your own GPUs
Text to speech (TTS)Speaks the reply, ideally one sentence at a timeKokoro, Piper, Coqui XTTS-v2
Speaker diarizationLabels who spoke when in multi-speaker audiopyannote.audio
OrchestrationMoves audio between stages and handles interruptions and transportPipecat, LiveKit Agents

Why voice data stays inside the perimeter

A recording holds more than the words in it. The European Data Protection Board’s guidelines on virtual voice assistants point out that voice data carries personal data both in what is said and in meta-information such as the speaker’s sex or age, and they recall that voice data is inherently biometric personal data. Where it is processed to uniquely identify someone, the controller needs a derogation under Article 9 of the GDPR as well as an Article 6 legal basis. The same guidelines note that assistants can capture people who never meant to use them.

Clinical dictation brings in HIPAA. Its definition of health information covers information “whether oral or recorded in any form or medium”, so a physician’s dictated note about a patient is protected health information from the moment it is recorded, whichever system transcribes it.

Recording rules vary by jurisdiction. US federal law permits a recording when one party to the conversation consents, under 18 U.S.C. § 2511(2)(d), while California Penal Code § 632 requires the consent of all parties to a confidential communication. Running capture and transcription in-house leaves those consent duties in place, but it keeps the recordings on systems governed by your own retention and access rules.

Customer-facing voices also fall under the EU AI Act. Since 2 August 2026, Article 50(1) has required providers to design systems that interact directly with people so that those people are told they are dealing with an AI system, unless that is obvious, and Article 50(2) requires synthetic audio to be marked in a machine-readable, detectable way. Our Article 50 explainer covers the marking grace period and who counts as the provider.

Speech recognition and text-to-speech models compared

For recognition, three families cover most on-premise projects. Sizes and licences come from each project’s own repository or model card (verified October 2026):

ModelParametersLanguagesLicenceWorth knowing
Whisper large-v3-turbo809MMultilingualMITLarge-v3 with the decoder cut from 32 layers to 4; about 6 GB of VRAM; not trained for translation
NVIDIA Parakeet TDT 0.6B v3600M25 European languagesCC-BY-4.0Punctuation and word timestamps; up to 24 minutes of audio per pass with full attention
NVIDIA Canary-1B-v2978M25 European languagesCC-BY-4.0Also translates into and out of English; runs on NVIDIA’s NeMo toolkit

Runtimes, hallucination fixes and word error rate testing for Whisper are covered in the transcription guide linked above.

For text to speech, licence terms separate the options more sharply than voice quality does:

OptionLanguages and voicesLicenceWatch for
Kokoro-82M8 languages, 54 voices, 82M parametersApache 2.0Its card lists the training audio, which includes synthetic speech generated by closed TTS models from large providers
PiperVoices published per language, each with its own model cardEngine GPL-3.0 (the archived original repository was MIT); each voice links its dataset’s licenceEmbeds espeak-ng; the Open Home Foundation is looking for maintainers
Coqui XTTS-v217 languages; clones a voice from a 6-second clipCoqui Public Model License, non-commercial purposes onlyRuled out for commercial products under that licence
Azure Neural TTS containerVoices and locales listed in Microsoft’s container registryCommercial, metered through an Azure resource6 cores and 12 GB of memory at minimum
NVIDIA RivaGPU-accelerated speech synthesis alongside recognition and translationProduction use through NVIDIA AI EnterpriseShips as containers for Docker or Kubernetes

Read Piper’s voice cards before a commercial launch. The US English lessac voice, for example, links the licence of the Blizzard 2013 dataset it was trained on, while the British alba voice links CC BY 4.0. Voice cloning raises a separate question: record whose voice was cloned and on what permission before anyone generates speech with it.

Serving and orchestration choices

Models are half the build. Something has to stream audio in, hand partial transcripts to the language model, start speaking before the model has finished, and stop speaking when the user interrupts.

Open frameworks. Pipecat, licensed BSD-2-Clause, is an open-source framework for voice agents and multimodal apps, and it lists local services such as Whisper, Piper and Ollama among its integrations. LiveKit Agents is Apache 2.0 and can run with LiveKit’s open-source WebRTC media server on your own machines. Its semantic turn detection uses a transformer model to judge when a user has finished speaking, which LiveKit says helps reduce interruptions.

Packaged speech services. NVIDIA Riva is a GPU-accelerated SDK for speech recognition, synthesis and translation that deploys with Docker or Helm charts, on premises or in the cloud. NVIDIA sells ongoing production use as part of NVIDIA AI Enterprise and offers a 90-day trial licence. Microsoft publishes Docker containers for Azure Speech in Foundry Tools covering speech to text, custom speech to text and neural text to speech, plus two containers in preview. They report usage to Azure in standard mode, and only approved customers on a commitment plan can run them disconnected; the transcription guide walks through those conditions.

Speech-to-speech models. Kyutai’s Moshi is a full-duplex alternative to the chain. Its repository reports a theoretical latency of 160 ms and a practical latency as low as 200 ms on an L4 GPU, with model weights under CC-BY 4.0. The trade-off is architectural: the text hand-off between stages is also where most teams attach retrieval, tool calls, policy checks and audit logging.

A latency budget for one spoken turn

People take turns quickly. A 2009 PNAS study of conversations in ten languages found that the most common gap between one speaker finishing and the next starting fell between 0 and 200 milliseconds in every language, and the mean across the whole dataset was 208 milliseconds. A cascaded assistant has to notice that the turn is over, transcribe, generate and synthesise before any sound comes back, so the work is to shrink each stage and overlap them.

StageWhat sets the delayHow to shorten it
End-of-turn detectionHow long the system waits in silence before treating the turn as finishedTune the silence threshold on real recordings, or add a semantic turn detector
Final transcriptModel size, decoding settings and whether partial results streamStreaming recognition, smaller or turbo models, models kept loaded on the GPU
First token from the LLMPrompt length, model size and queueing on a busy GPUShort prompts, smaller models for routine turns, capacity reserved for voice traffic
First audio from TTSSynthesis speed and how much text the engine waits forSynthesise sentence by sentence while the LLM is still writing
PlaybackAudio buffers and network hopsKeep media servers close to users and avoid extra relays

Whisper itself reads audio in 30-second windows, which is why streaming wrappers exist. One of them, whisper_streaming, reports 3.3 seconds of latency on unsegmented long-form speech: workable for live captions, far from conversational timing.

Track a single end-to-end number: the time from the end of the user’s speech to the first audio sample played back, at the 95th percentile, with as many concurrent sessions as you expect at peak.

Sizing hardware for voice workloads

Published requirements for the speech components are modest next to those of a language model (verified October 2026):

ComponentPublished requirementSource
Whisper turboAbout 6 GB of VRAMOpenAI’s Whisper repository
Whisper largeAbout 10 GB of VRAMOpenAI’s Whisper repository
faster-whisper, large-v2 with 8-bit weights2,926 MB of VRAM in the project’s 13-minute benchmarkfaster-whisper
whisper.cpp, large modelAbout 3.9 GB of memorywhisper.cpp
Parakeet TDT 0.6B v3At least 2 GB of RAM to loadNVIDIA model card
Canary-1B-v2At least 6 GB of RAM to loadNVIDIA model card
Azure Neural TTS container6 cores and 12 GB minimum; 8 cores and 16 GB recommendedMicrosoft Learn
Moshi, PyTorch versionA GPU with 24 GB of memoryKyutai

The language model normally sets the budget. Size it first with the method in our GPU estimation guide, then add recognition and synthesis capacity for the sessions open at the same moment. A live voice session keeps audio flowing in both directions for as long as the conversation lasts, whereas a chat request frees the GPU between messages, so plan from concurrent conversations at peak rather than from named users. Our AI server guide covers chassis, GPU and power choices once those numbers are known.

Deployment checklist for on-premise voice

  1. Pick the shape first. Dictation, turn-based assistant or full-duplex agent: the latency target and the hardware follow from that choice.
  2. Clear every licence. Check code and weights separately. XTTS-v2 is non-commercial, Piper’s engine is GPL-3.0 with per-voice dataset licences, and NVIDIA’s CC-BY-4.0 models require attribution.
  3. Test recognition on your own audio. Accents, headsets, speakerphones, background noise and domain vocabulary all move accuracy, so measure word error rate before choosing a model.
  4. Put voice activity detection in front of recognition. Silero VAD’s maintainers report that a 30 ms chunk takes under a millisecond on one CPU thread, so the cost of adding it is small.
  5. Stream every stage. Partial transcripts, token streaming and sentence-level synthesis are what make a reply feel prompt.
  6. Set retention for audio and transcripts. Decide how long raw audio is kept once a reviewed transcript exists, and who may replay it.
  7. Disclose and mark. Tell people they are talking to an AI system where Article 50(1) applies, and mark synthetic audio if you are its provider.
  8. Control cloned voices. Keep the consent record with the voice model and restrict who can generate speech with it.
  9. Log each turn. Keep what was heard, what was transcribed, which model answered and what was spoken, under the same access controls as the audio.
  10. Load-test at peak concurrency. Watch the gap from end of speech to first audio at the 95th percentile, not the average.
  11. Mirror models for offline updates. Pull new weights on a connected staging host, verify checksums and promote them like any other release.

How VDF AI fits

VDF AI’s voice capabilities are transcription, synthesis and dictation. Agents can be assigned a speech-to-text tool for short clips, an audio transcription tool for long recordings with timestamps and speaker segments, a video transcription tool for the audio track of recordings, and a text-to-speech tool with voice, speed and MP3 or WAV output. Each can run on a local model inside your perimeter, so the audio and the text produced from it stay on infrastructure you control.

In VDF AI Chat, workspaces with voice input enabled let people dictate prompts instead of typing them, and the documentation asks users to review the transcribed text before sending. Most agents accept spoken input the same way.

VDF AI does not include telephony or a real-time voice-agent product. Its voice tools work on dictated and recorded speech.

Sources

Frequently asked questions

Can voice AI run completely on premise?

Yes. Each stage has an open model that runs on your own hardware: Whisper or NVIDIA Parakeet for speech recognition, an open-weight language model for the reasoning step, and Kokoro or Piper for text to speech, with Silero VAD spotting where speech starts and stops. Frameworks such as Pipecat and LiveKit Agents, both open source and self-hostable, connect the stages. The hard work is latency tuning, licence review and GPU sizing rather than finding components. Containers from Microsoft and NVIDIA are the commercial route.

What is the best open-source text to speech for on-premise use?

It depends on licence and languages as much as on how the voices sound. Kokoro-82M is Apache 2.0, has 82 million parameters and covers 8 languages with 54 voices, which makes it easy to approve and cheap to run. Piper is fast and local, but its maintained engine is GPL-3.0 and every voice points to its own dataset licence. Coqui XTTS-v2 clones a voice from a 6-second clip in 17 languages, yet the Coqui Public Model License restricts it to non-commercial use.

How fast does a voice assistant need to respond?

As close to human turn-taking as you can get it. A 2009 PNAS study of conversations in ten languages found that the most common gap between turns fell between 0 and 200 milliseconds in every language, with a mean of 208 milliseconds across the dataset. A cascaded assistant must detect the end of the turn, transcribe, generate and synthesise before any sound comes back, so stream every stage, start speaking on the first sentence and measure the gap from the end of speech to the first audio.

Is a voice recording personal data under the GDPR?

Usually, because a recording of someone's voice relates to an identifiable person. The European Data Protection Board's guidelines on virtual voice assistants go further and recall that voice data is inherently biometric personal data. When it is processed to uniquely identify a speaker, the controller needs a derogation under Article 9 of the GDPR in addition to a legal basis under Article 6. Recordings can also capture health details and the voices of bystanders, which is why many organisations keep them on infrastructure they control.

Can Azure text to speech run on premise?

Partly. Microsoft publishes a Neural text to speech container for Azure Speech in Foundry Tools that runs on your own Docker hosts, with a documented minimum of 6 CPU cores and 12 GB of memory. In standard mode the container must reach Azure to report usage for billing. Running it fully disconnected requires an application, Microsoft's approval and a commitment-tier plan, and access is aimed at strategic customers with offline or strictly regulated environments. Check the current terms on Microsoft Learn before planning around it.

What hardware does an on-premise voice assistant need?

Less for speech than for the language model behind it. OpenAI lists about 6 GB of VRAM for Whisper's turbo model, faster-whisper ran large-v2 in under 3 GB with 8-bit weights in its published benchmark, and Kokoro has only 82 million parameters. The language model normally sets the GPU budget, so size it first and then add speech capacity for the conversations running at the same moment. Speech-to-speech models are heavier: Kyutai's Moshi asks for a 24 GB GPU in its PyTorch version.

Filed under
on-premises AIlocal AI infrastructureopen-weight modelsdata sovereigntyregulated AIAI infrastructure
On-Prem AI

Plan your on-prem AI deployment

Book an architecture call and we will scope a private, on-prem AI deployment for your environment — integrations, hardware, and governance included.

Keep reading