Home Assistant Voice Stack¶
The voice assistant on the HA box ("Hey Jarvis") is a three-leg pipeline: speech-to-text, an LLM brain, and text-to-speech back out. This page is the wiring and the testing playbook. The story of the bug that took weeks to find is in the blog post The one-byte bug.
Topology¶
Satellite (mic/speaker)
│ Wyoming protocol
▼
Home Assistant (HAOS VM)
│ dials OUT to
├──> Wyoming bridge on the 3060 box (192.168.4.45)
│ STT: whisper.cpp large-v3-turbo (port 10300)
│ TTS: Kokoro, 54 voices (port 10210)
│
└──> native llama_cpp integration ──> slot-proxy on the 5090 box (192.168.4.29)
HA is the orchestrator. The satellite is a dumb mic/speaker that speaks Wyoming. HA dials out to the Wyoming bridge, not the other way around.
Voice I/O was relocated in August 2026: STT and TTS moved from the 5090 to the 3060 box, which freed about 5.2 GB of VRAM and left the 5090 as a pure brain box. Measured end-to-end after the move: STT 0.42 s, total around 3.2 s. The pipeline is Qwen Voice — stt.whisper_gpu_2 + tts.kokoro_2 (voice af_sarah) + conversation.qwen3_8_27b, wake word "Hey Jarvis".
The slot-proxy priority model¶
The GPU runs a slot-proxy in front of llama.cpp. Voice beats the agent on purpose. A voice request needs a fast, low-latency answer; an agent task can wait for a slot. The proxy is what makes that ordering real.
The one-byte framing rule¶
Every binary frame HA receives over the websocket must be prefixed with the 1-byte stt_binary_handler_id from the run-start event's runner_data.stt_binary_handler_id (usually 1). HA reads the first byte as the handler id and the rest as audio. Send raw PCM with no prefix and HA silently drops every chunk.
The correct send:
for i in range(0, len(pcm), CHUNK):
await ws.send(bytes([HID]) + pcm[i:i+CHUNK])
await ws.send(bytes([HID])) # 1-byte terminator = end of audio (NOT b"")
An empty terminator b"" raises a Disconnect (len < 1).
Do not send a session-init message
Wyoming 1.10.0 has no session-init message. Sending one makes the server silently ignore the client and the read hangs. This mimics a "bridge hang" but is purely a client bug. The correct ASR flow is: transcribe → audio-start → audio-chunk* → audio-stop → transcript.
Testing playbook¶
Isolate a leg before touching the full harness.
- REST STT call —
POST /api/stt/stt.whisper_gpuwith the right header. If this answers in under a second, the engine is healthy and the problem is upstream in the client. - Read the bridge log — "transcribing N bytes" is the proof audio actually arrived. A bare "client disconnected" with zero audio means the framing is wrong, not the network.
- Then the full websocket pipeline.
The full bug narrative is in The one-byte bug.
2026.8.3 renames¶
The 2026.8.3 release renamed a bunch of tool calls: pipeline_id, stt.* state semantics, and /api/tts is gone. If your automation or voice tools break after that version, check the renames first.
Voice reminders¶
Alexa-style timed reminders: a chime plus a push. The naming problem to watch for is Donetick-versus-HA-todo-list — the voice brain needs to know which list a reminder lands in.
The timer routing bug worth knowing about: the satellite runs on-device timers just fine (the LED ring spins down as they count), but the Qwen Voice pipeline had prefer_local_intents=False, so timer phrases went to the 27B brain — which has no timer tool — and it over-claimed the set while being unable to read them back. Setting prefer_local_intents=True routes timer set/cancel/read to the native intent handler, and reads work. One catch on assist_pipeline/pipeline/update: it's a full-replace, every field required, including wake_word_entity and wake_word_id.
And the satellite's firmware timers are not HA timer.* entities, so timer.cancel via REST 400s on them — you can't cancel a firmware timer through the API.
See also¶
- Donetick — the task/chores side of the voice tools
- Blog post: The one-byte bug