AI overview
Mosaic ships a full AI stack that every generated app gets by projection —
no hand-written glue. Four surfaces compose: a RAG assistant (with a branded
auto-derived "Mosaic Aide"), a voice conversation channel
(STT → LLM → TTS), a knowledge engine (hybrid + Graph-RAG over text, code,
and books), and MCP servers/clients that expose all of it to agents.
flowchart LR
subgraph Surfaces
A["Assistant / Aide\nPOST /api/assistant/chat\nSSE"]
V["Voice\nGET /api/voice (WS)"]
K["Knowledge engine\nhybrid + Graph-RAG"]
M["MCP servers\napp (stdio) + knowledge (HTTP)"]
end
subgraph App
T["tessera.yaml\nvector_dbs / voice / app"]
end
T -->|projects| A
T -->|projects| V
T -->|projects| K
T -->|projects| M
A --> K
V -->|search_docs tool| K
M -->|23 knowledge tools| K
The surfaces
| Surface | Entry point | Page |
|---|---|---|
| Assistant / Mosaic Aide | POST /api/assistant/chat (SSE), /aide | assistant |
| Web chat | /chat, POST /api/chat (web lens) | assistant |
| Voice | GET /api/voice (WebSocket), /voice | voice |
| Knowledge / Graph-RAG | /knowledge, POST /api/vectordb/{kb}/search | knowledge |
| MCP | cargo run -- mcp (stdio), POST /mcp (knowledge) | mcp |
The seams are env, not code
Every AI surface is configured through environment variables (secret-backed in the deploy spec), so the generated app stays dependency-light and deterministic. No provider is hard-wired:
MOAIC_CHAT_*— the chat/assistant LLM (base URL, API key, model).MOAIC_EMBED_*— embeddings (MOAIC_EMBED_LOCAL=1for the deterministic local 256-dim embedder;MOAIC_EMBED_ONNX_MODELfor an ONNX sentence model).MOAIC_RERANK_*— an optional cross-encoder re-rank seam.MOAIC_VOICE_{LLM,STT,TTS}_*— the voice pipeline stages.MOAIC_MCP_<NAME>_URL— external MCP servers an app consumes.
Without the relevant keys the surfaces degrade gracefully (grounded no-LLM
answers, or a 503 on the voice upgrade) — the routes stay mounted and the app
still builds and runs.
The doctrine
- The engine is pure Rust and vendored.
mosaic-knowledgehas no network deps; it isinclude_str!-vendored into every generated app that declares avector_dbs. LLM/embedding calls happen at the app's seams, never in the engine. - Retrieval is grounded and deterministic. Hybrid BM25 + vector retrieval, graph-augmented re-ranking (Personalized PageRank), and citation/URL rewriting so answers point at real sources.
- Declarative floor, runtime ceiling. Declared knowledge sources are the floor; sources and typed entities can be added at runtime (ADR 0014, ADR 0047).