Mosaic one model, many lenses

Voice (STT → LLM → TTS)

The voice: DSL projects a full voice conversation channel into every app that declares one. There are two modes: a pipeline (server-side STT → LLM → TTS, the default) and realtime (a provider-direct WebSocket relay).

Endpoints

  • GET /api/voice — the WebSocket conversation endpoint.
  • GET /voice — a self-contained browser widget (mic + transcript + audio).
  • The upgrade answers 503 when a stage's env key is unset; the route stays mounted so the app still builds and runs.

Pipeline mode (default)

sequenceDiagram
  participant B as Browser
  participant S as /api/voice (WS)
  participant STT as STT (OpenAI-compat)
  participant LLM as LLM
  participant K as knowledge (search_docs)
  participant TTS as TTS
  B->>S: audio frames (16 kHz PCM16, b64)
  S->>S: server VAD (RMS endpointing)
  S->>STT: transcribe (WAV)
  STT-->>S: transcript
  S->>LLM: turn (+ search_docs tool loop)
  LLM->>K: search_docs(query, top_k)
  K-->>LLM: chunks
  LLM-->>S: answer
  S->>TTS: per-sentence synthesis
  TTS-->>S: PCM16 @24 kHz
  S-->>B: audio frames (b64)
  • Endpointing. Server-side VAD: an RMS UtteranceEndpointer ends a turn on turn_silence_ms of silence (default 700).
  • STT. OpenAI-compatible multipart WAV; default model whisper-1.
  • LLM turn. A single (non-streamed) completion with a server-side search_docs tool loop (max 4 iterations) over the app's voice.site_search.kb knowledge base — the app's grounding is enforced server-side.
  • TTS. Per-sentence synthesis to PCM16 @24 kHz (default gpt-4o-mini-tts, voice alloy).
  • Barge-in. Voiced audio aborts the in-flight turn.

Wire protocol

  • In: welcome | audio | text | stop.
  • Out: speech | transcript | audio | tool | turn | error | goodbye.

Realtime mode

A provider-direct WebSocket relay with per-provider codecs:

  • OpenAI Realtime (default) or Gemini Live.
  • The server still executes the search_docs tool loop server-side, so grounding is preserved in realtime too.

Example

voice:
  mode: pipeline
  turn_silence_ms: 700
  site_search:
    kb: kb
    top_k: 5
  llm: { model: gpt-4o-mini }
  stt: { model: whisper-1 }
  tts: { model: gpt-4o-mini-tts, voice: alloy }

Configuration

EnvMeaning
MOAIC_VOICE_LLM_{API_KEY,BASE_URL,MODEL}the voice LLM stage
MOAIC_VOICE_STT_{API_KEY,BASE_URL,MODEL}the STT stage
MOAIC_VOICE_TTS_{API_KEY,BASE_URL,MODEL}the TTS stage
MOAIC_VOICE_TTS_VOICEthe TTS voice
MOAIC_VOICE_REALTIME_*the realtime relay

Honest gaps

  • One voice: surface per app (app lens); no multi-voice/tenant variants.
  • The pipeline LLM turn is non-streamed (streaming is at the audio/TTS frame level); no local/self-hosted STT or TTS; voice sessions have no cross-session memory; voice is deliberately not an MCP tool.