Voice (STT → LLM → TTS)
The voice: DSL projects a full voice conversation channel into every app that
declares one. There are two modes: a pipeline (server-side
STT → LLM → TTS, the default) and realtime (a provider-direct WebSocket
relay).
Endpoints
GET /api/voice— the WebSocket conversation endpoint.GET /voice— a self-contained browser widget (mic + transcript + audio).- The upgrade answers
503when a stage's env key is unset; the route stays mounted so the app still builds and runs.
Pipeline mode (default)
sequenceDiagram
participant B as Browser
participant S as /api/voice (WS)
participant STT as STT (OpenAI-compat)
participant LLM as LLM
participant K as knowledge (search_docs)
participant TTS as TTS
B->>S: audio frames (16 kHz PCM16, b64)
S->>S: server VAD (RMS endpointing)
S->>STT: transcribe (WAV)
STT-->>S: transcript
S->>LLM: turn (+ search_docs tool loop)
LLM->>K: search_docs(query, top_k)
K-->>LLM: chunks
LLM-->>S: answer
S->>TTS: per-sentence synthesis
TTS-->>S: PCM16 @24 kHz
S-->>B: audio frames (b64)
- Endpointing. Server-side VAD: an RMS
UtteranceEndpointerends a turn onturn_silence_msof silence (default700). - STT. OpenAI-compatible multipart WAV; default model
whisper-1. - LLM turn. A single (non-streamed) completion with a server-side
search_docstool loop (max 4 iterations) over the app'svoice.site_search.kbknowledge base — the app's grounding is enforced server-side. - TTS. Per-sentence synthesis to
PCM16 @24 kHz(defaultgpt-4o-mini-tts, voicealloy). - Barge-in. Voiced audio aborts the in-flight turn.
Wire protocol
- In:
welcome | audio | text | stop. - Out:
speech | transcript | audio | tool | turn | error | goodbye.
Realtime mode
A provider-direct WebSocket relay with per-provider codecs:
- OpenAI Realtime (default) or Gemini Live.
- The server still executes the
search_docstool loop server-side, so grounding is preserved in realtime too.
Example
voice:
mode: pipeline
turn_silence_ms: 700
site_search:
kb: kb
top_k: 5
llm: { model: gpt-4o-mini }
stt: { model: whisper-1 }
tts: { model: gpt-4o-mini-tts, voice: alloy }
Configuration
| Env | Meaning |
|---|---|
MOAIC_VOICE_LLM_{API_KEY,BASE_URL,MODEL} | the voice LLM stage |
MOAIC_VOICE_STT_{API_KEY,BASE_URL,MODEL} | the STT stage |
MOAIC_VOICE_TTS_{API_KEY,BASE_URL,MODEL} | the TTS stage |
MOAIC_VOICE_TTS_VOICE | the TTS voice |
MOAIC_VOICE_REALTIME_* | the realtime relay |
Honest gaps
- One
voice:surface per app (app lens); no multi-voice/tenant variants. - The pipeline LLM turn is non-streamed (streaming is at the audio/TTS frame level); no local/self-hosted STT or TTS; voice sessions have no cross-session memory; voice is deliberately not an MCP tool.