A Fully Local Voice Assistant: Home Assistant Voice PE with Whisper, Piper and Ollama
Purpose
Voice control with zero cloud dependencies. Every step of the pipeline, wake word, speech-to-text, understanding, and text-to-speech, runs locally. No request leaves the LAN, no cloud provider touches the audio, and nothing relies on an internet connection to function.
The setup uses a Home Assistant Voice PE satellite as the hardware interface, with Wyoming protocol services handling the heavy lifting on a dedicated GPU Docker host and Ollama providing local LLM capabilities for open-ended conversations.
Architecture
Home Assistant runs in k3s with host networking so it sits directly on the LAN for device discovery. The Wyoming protocol services, Whisper, Piper, and OpenWakeWord, run as Docker containers on the GPU host alongside Ollama. Wyoming uses raw TCP rather than HTTP, so everything bypasses the reverse proxy entirely.
Pipeline
The voice flow works like this:
- Hardware interface: A Voice PE satellite handles audio capture and playback
- Wake word detection: Runs on the device itself using microWakeWord on an ESP32, so no audio is sent to the network until a wake word triggers
- Speech-to-text: Wyoming Whisper converts audio to text using the
base-int8model for English processing - Understanding layer: Routes between two paths:
- Home Assistant’s built-in intent engine handles simple, deterministic requests like “what time is it” and doesn’t call the LLM at all
- Ollama with a conversation model handles open-ended requests
- Text-to-speech: Wyoming Piper converts responses to audio using
en_US-lessac-mediumvoice
The system keeps an OpenWakeWord container deployed but idle, since on-device wake word detection works well and the software fallback remains available for future satellites that might need it.
Pipeline Diagram
sequenceDiagram
participant User
participant VoicePE as Voice PE Satellite
participant HA as Home Assistant
participant Whisper
participant Intent as Built-in Intent Engine
participant Ollama
participant Piper
User->>VoicePE: Speaks command
VoicePE->>VoicePE: Wake word detected (microWakeWord/local)
VoicePE->>HA: Audio stream triggered
HA->>Whisper: STT request (Wyoming TCP)
Whisper-->>HA: Transcript
HA->>Intent: Check simple intent match
alt Deterministic request
Note over Intent: "what time is it", etc.
Intent-->>HA: Response generated
HA->>Piper: TTS request (Wyoming TCP)
Piper-->>HA: Audio response
HA-->>VoicePE: Audio to speak
VoicePE-->>User: Spoken reply
else
Note over Ollama: "Explain how a heat pump works", creative tasks
HA->>Ollama: LLM call via Assist API
Ollama-->>HA: Model response
HA->>Piper: TTS request (Wyoming TCP)
Piper-->>HA: Audio response
HA-->>VoicePE: Audio to speak
VoicePE-->>User: Spoken reply
end
The hardware satellite sits on the LAN alongside the k3s cluster node running Home Assistant. The GPU host has an Intel Arc Pro B60 with 24 GB of VRAM, running Ollama’s Vulkan backend to serve both the voice assistant model and a separate 27B agent model used in other workflows.
Trade-offs
Running the conversation model locally introduces latency during initial loads compared to cloud services. The GPU serves dual purposes, with about 20 GB of the 24 GB going to the 27B agent model, so both that model and the 14B voice model cannot stay loaded at the same time. When they compete for VRAM, Ollama manages swapping by unloading one to load the other, meaning a long agent run followed by voice usage incurs a reload delay before the first voice request processes.
The choice of qwen2.5:14b-instruct over reasoning-focused models like deepseek-r1:14b comes down to response time. Thinking adds latency that doesn’t suit fast, conversational tool-calling where quick responses matter. The voice model uses keep_alive -1, which keeps it loaded until another model needs the memory.
Observability covers all the voice containers logging through Docker’s gelf driver to Graylog for centralized monitoring. See the Graylog logging post for details on the logging setup.
A Mesa driver upgrade (version 26.2, September 2026) enabled cooperative-matrix support on this GPU architecture, making generation roughly 1.5 to 2 times faster. The voice model reached about 47 tokens/second after updating, fast enough for natural conversation flow without noticeable delays between questions and responses.
Notes
Simple requests don’t go through the LLM at all. Home Assistant’s intent engine matches patterns like time queries without generating tokens. See the Home Assistant post for more on how Home Assistant fits into the homelab, and the Open WebUI post for the Ollama setup used in other contexts.
The Wyoming protocol’s raw TCP design means all inter-service communication sits outside Traefik’s reverse proxy. These services don’t need HTTP health endpoints or TLS termination. If you plan to add voice to your own setup, expect GPU contention to be the constraint if other AI workloads run on the same hardware.