Post

A Fully Local Voice Assistant: Home Assistant Voice PE with Whisper, Piper and Ollama

A Fully Local Voice Assistant: Home Assistant Voice PE with Whisper, Piper and Ollama

Purpose

Voice control with zero cloud dependencies. Every step of the pipeline, wake word, speech-to-text, understanding, and text-to-speech, runs locally. No request leaves the LAN, no cloud provider touches the audio, and nothing relies on an internet connection to function.

The setup uses a Home Assistant Voice PE satellite as the hardware interface, with Wyoming protocol services handling the heavy lifting on a dedicated GPU Docker host and Ollama providing local LLM capabilities for open-ended conversations.

Architecture

Home Assistant runs in k3s with host networking so it sits directly on the LAN for device discovery. The Wyoming protocol services, Whisper, Piper, and OpenWakeWord, run as Docker containers on the GPU host alongside Ollama. Wyoming uses raw TCP rather than HTTP, so everything bypasses the reverse proxy entirely.

Pipeline

The voice flow works like this:

  1. Hardware interface: A Voice PE satellite handles audio capture and playback
  2. Wake word detection: Runs on the device itself using microWakeWord on an ESP32, so no audio is sent to the network until a wake word triggers
  3. Speech-to-text: Wyoming Whisper converts audio to text using the base-int8 model for English processing
  4. Understanding layer: Routes between two paths:
    • Home Assistant’s built-in intent engine handles simple, deterministic requests like “what time is it” and doesn’t call the LLM at all
    • Ollama with a conversation model handles open-ended requests
  5. Text-to-speech: Wyoming Piper converts responses to audio using en_US-lessac-medium voice

The system keeps an OpenWakeWord container deployed but idle, since on-device wake word detection works well and the software fallback remains available for future satellites that might need it.

Pipeline Diagram

sequenceDiagram
    participant User
    participant VoicePE as Voice PE Satellite
    participant HA as Home Assistant
    participant Whisper
    participant Intent as Built-in Intent Engine
    participant Ollama
    participant Piper

    User->>VoicePE: Speaks command
    VoicePE->>VoicePE: Wake word detected (microWakeWord/local)

    VoicePE->>HA: Audio stream triggered
    HA->>Whisper: STT request (Wyoming TCP)
    Whisper-->>HA: Transcript
    HA->>Intent: Check simple intent match

    alt Deterministic request
        Note over Intent: "what time is it", etc.
        Intent-->>HA: Response generated
        HA->>Piper: TTS request (Wyoming TCP)
        Piper-->>HA: Audio response
        HA-->>VoicePE: Audio to speak
        VoicePE-->>User: Spoken reply
    else
        Note over Ollama: "Explain how a heat pump works", creative tasks
        HA->>Ollama: LLM call via Assist API
        Ollama-->>HA: Model response
        HA->>Piper: TTS request (Wyoming TCP)
        Piper-->>HA: Audio response
        HA-->>VoicePE: Audio to speak
        VoicePE-->>User: Spoken reply
    end

The hardware satellite sits on the LAN alongside the k3s cluster node running Home Assistant. The GPU host has an Intel Arc Pro B60 with 24 GB of VRAM, running Ollama’s Vulkan backend to serve both the voice assistant model and a separate 27B agent model used in other workflows.

Trade-offs

Running the conversation model locally introduces latency during initial loads compared to cloud services. The GPU serves dual purposes, with about 20 GB of the 24 GB going to the 27B agent model, so both that model and the 14B voice model cannot stay loaded at the same time. When they compete for VRAM, Ollama manages swapping by unloading one to load the other, meaning a long agent run followed by voice usage incurs a reload delay before the first voice request processes.

The choice of qwen2.5:14b-instruct over reasoning-focused models like deepseek-r1:14b comes down to response time. Thinking adds latency that doesn’t suit fast, conversational tool-calling where quick responses matter. The voice model uses keep_alive -1, which keeps it loaded until another model needs the memory.

Observability covers all the voice containers logging through Docker’s gelf driver to Graylog for centralized monitoring. See the Graylog logging post for details on the logging setup.

A Mesa driver upgrade (version 26.2, September 2026) enabled cooperative-matrix support on this GPU architecture, making generation roughly 1.5 to 2 times faster. The voice model reached about 47 tokens/second after updating, fast enough for natural conversation flow without noticeable delays between questions and responses.

Notes

Simple requests don’t go through the LLM at all. Home Assistant’s intent engine matches patterns like time queries without generating tokens. See the Home Assistant post for more on how Home Assistant fits into the homelab, and the Open WebUI post for the Ollama setup used in other contexts.

The Wyoming protocol’s raw TCP design means all inter-service communication sits outside Traefik’s reverse proxy. These services don’t need HTTP health endpoints or TLS termination. If you plan to add voice to your own setup, expect GPU contention to be the constraint if other AI workloads run on the same hardware.

This post is licensed under CC BY 4.0 by the author.