ADR 0033 — Browser voice (Web Speech) on the built-in chat
This ADR closes the last "partial" row in the sync/ee.md voice feature
map: EE's browser voice on the chat scene (blocks/ai/ee-chat/ui/src/voice.rs)
— Web Speech dictation into the chat input plus speechSynthesis playback of
replies, with zero backend. It ports that surface onto Mosaic's built-in
/chat page and, where Mosaic can do more for free, does.
Context
Mosaic already ships two voice surfaces, both on the server side:
platform.voice(V1/V2) — the WSS/api/voiceconversation endpoint (pipeline: server VAD → STT → LLM + search tool loop → per-sentence TTS; realtime: provider-direct relay) and the/voicewidget. This is a full voice conversation channel and requires provider env keys.- Browser voice (this ADR) — the other half of EE's voice story, from
ee-chat/ui/src/voice.rs: the browser's built-in Web Speech engines (ChromiumSpeechRecognitionfor dictation,speechSynthesisfor playback) wired into the ordinary text chat. No backend, no provider key, no new transport — a transcript flows into the same input field as typed text and is sent through the same/api/chatcall.
EE implements it as a Rust module of hand-written wasm_bindgen extern
blocks (the Web Speech API is not in web-sys), cfg-gated so the SSR half
compiles to no-ops, with capability probed at interaction time so SSR and
client markup never carry a capability bit.
Mosaic's built-in chat surface (/chat + /api/chat, emitted by
mosaic-render::chat in the full/islands hydration modes) is a single
static HTML page with a vanilla-JS client — there is no Leptos chat island
to host a Rust voice module in. The Web Speech API is natively available to
that page's JavaScript, so the port is the same behavior expressed one
layer closer to the browser: plain JS, no wasm, no bindings, SSR-safe by
construction (one static document; the server renders no capability bit).
Decision
The generated /chat page now carries the browser-voice client:
- Dictation (STT) — a mic button in the chat form (hidden when
SpeechRecognition/webkitSpeechRecognitionis absent). One utterance per press; the button is a toggle (press again to stop).interimResultsstream the growing phrase into the input live; final results commit into the input, which the user reviews and sends like typed text — the transcript never takes a different path. Recognition is locale-aware (navigator.language). While recording the button pulses in the--destructivecolor;onend(any reason) resets it; a denied microphone (not-allowed/service-not-allowed) surfaces a one-line hint in the log instead of failing silently. - Playback (TTS) — a "speak replies" toggle above the log, persisted in
localStorage(mosaicChatTts), hidden whenspeechSynthesisis absent. When on, each assistant reply is spoken on arrival (cancel()+speak(), one utterance at a time, mirroring EE'sspeak()). Additionally, every bot message gets a small "speak" link that reads that message aloud on demand, independent of the toggle — the "more and better" part (EE only auto-speaks replies). - Capability at client time — the probes run in the page's own script
(buttons start
hidden; the client unhides what the browser supports). The served document stays a single static byte string, so there is no SSR/client markup divergence to mismatch, which is the property EE's interaction-time probing buys for Leptos.
No DSL key: like EE, browser voice is part of the built-in chat surface,
always compiled in. Apps that own /chat (app_owns_chat, ADR 0025) keep
their own page and are untouched. The platform.voice WSS surface is
orthogonal and unchanged.
Files
crates/mosaic-render/src/chat.rs—chat_page_rs(mic button, TTS toggle, the Web Speech client) andchat_css(controls, recording pulse, "speak" link, sys hint).crates/mosaic-render/src/web.rs—chat_page_carries_browser_voicerender test.- Re-pinned goldens: full-stack
web/src/chat_page.rs+styles.cssin the four web lenses; removed stale U1-erachat.rs/chat_page.rsfiles from the CSR goldens (data-sync, orders-app) that predate the hydration-mode split and were no longer part of the rendered set.
Consequences
- Browser voice now needs no provider keys and no WSS: a Chromium browser gets dictation into the assistant chat out of the box; every modern browser gets reply playback. Non-Chromium browsers simply hide the mic.
- Dictation is per-utterance and manual (press to start, the transcript lands in the input, the user sends) — the same review-before-send posture as EE. A "push-to-talk auto-send" variant is a client-only tweak, not an API change, if ever wanted.
- The generated page grows ~100 lines of static JS; the
/api/chatcontract is unchanged (voice output is just more text inmessage).