Mosaic one model, many lenses

ADR 0033 — Browser voice (Web Speech) on the built-in chat

This ADR closes the last "partial" row in the sync/ee.md voice feature map: EE's browser voice on the chat scene (blocks/ai/ee-chat/ui/src/voice.rs) — Web Speech dictation into the chat input plus speechSynthesis playback of replies, with zero backend. It ports that surface onto Mosaic's built-in /chat page and, where Mosaic can do more for free, does.

Context

Mosaic already ships two voice surfaces, both on the server side:

  • platform.voice (V1/V2) — the WSS /api/voice conversation endpoint (pipeline: server VAD → STT → LLM + search tool loop → per-sentence TTS; realtime: provider-direct relay) and the /voice widget. This is a full voice conversation channel and requires provider env keys.
  • Browser voice (this ADR) — the other half of EE's voice story, from ee-chat/ui/src/voice.rs: the browser's built-in Web Speech engines (Chromium SpeechRecognition for dictation, speechSynthesis for playback) wired into the ordinary text chat. No backend, no provider key, no new transport — a transcript flows into the same input field as typed text and is sent through the same /api/chat call.

EE implements it as a Rust module of hand-written wasm_bindgen extern blocks (the Web Speech API is not in web-sys), cfg-gated so the SSR half compiles to no-ops, with capability probed at interaction time so SSR and client markup never carry a capability bit.

Mosaic's built-in chat surface (/chat + /api/chat, emitted by mosaic-render::chat in the full/islands hydration modes) is a single static HTML page with a vanilla-JS client — there is no Leptos chat island to host a Rust voice module in. The Web Speech API is natively available to that page's JavaScript, so the port is the same behavior expressed one layer closer to the browser: plain JS, no wasm, no bindings, SSR-safe by construction (one static document; the server renders no capability bit).

Decision

The generated /chat page now carries the browser-voice client:

  • Dictation (STT) — a mic button in the chat form (hidden when SpeechRecognition/webkitSpeechRecognition is absent). One utterance per press; the button is a toggle (press again to stop). interimResults stream the growing phrase into the input live; final results commit into the input, which the user reviews and sends like typed text — the transcript never takes a different path. Recognition is locale-aware (navigator.language). While recording the button pulses in the --destructive color; onend (any reason) resets it; a denied microphone (not-allowed/service-not-allowed) surfaces a one-line hint in the log instead of failing silently.
  • Playback (TTS) — a "speak replies" toggle above the log, persisted in localStorage (mosaicChatTts), hidden when speechSynthesis is absent. When on, each assistant reply is spoken on arrival (cancel() + speak(), one utterance at a time, mirroring EE's speak()). Additionally, every bot message gets a small "speak" link that reads that message aloud on demand, independent of the toggle — the "more and better" part (EE only auto-speaks replies).
  • Capability at client time — the probes run in the page's own script (buttons start hidden; the client unhides what the browser supports). The served document stays a single static byte string, so there is no SSR/client markup divergence to mismatch, which is the property EE's interaction-time probing buys for Leptos.

No DSL key: like EE, browser voice is part of the built-in chat surface, always compiled in. Apps that own /chat (app_owns_chat, ADR 0025) keep their own page and are untouched. The platform.voice WSS surface is orthogonal and unchanged.

Files

  • crates/mosaic-render/src/chat.rs — chat_page_rs (mic button, TTS toggle, the Web Speech client) and chat_css (controls, recording pulse, "speak" link, sys hint).
  • crates/mosaic-render/src/web.rs — chat_page_carries_browser_voice render test.
  • Re-pinned goldens: full-stack web/src/chat_page.rs + styles.css in the four web lenses; removed stale U1-era chat.rs/chat_page.rs files from the CSR goldens (data-sync, orders-app) that predate the hydration-mode split and were no longer part of the rendered set.

Consequences

  • Browser voice now needs no provider keys and no WSS: a Chromium browser gets dictation into the assistant chat out of the box; every modern browser gets reply playback. Non-Chromium browsers simply hide the mic.
  • Dictation is per-utterance and manual (press to start, the transcript lands in the input, the user sends) — the same review-before-send posture as EE. A "push-to-talk auto-send" variant is a client-only tweak, not an API change, if ever wanted.
  • The generated page grows ~100 lines of static JS; the /api/chat contract is unchanged (voice output is just more text in message).