Mosaic one model, many lenses

AI overview

Mosaic ships a full AI stack that every generated app gets by projection — no hand-written glue. Four surfaces compose: a RAG assistant (with a branded auto-derived "Mosaic Aide"), a voice conversation channel (STT → LLM → TTS), a knowledge engine (hybrid + Graph-RAG over text, code, and books), and MCP servers/clients that expose all of it to agents.

flowchart LR
  subgraph Surfaces
    A["Assistant / Aide\nPOST /api/assistant/chat\nSSE"]
    V["Voice\nGET /api/voice (WS)"]
    K["Knowledge engine\nhybrid + Graph-RAG"]
    M["MCP servers\napp (stdio) + knowledge (HTTP)"]
  end
  subgraph App
    T["tessera.yaml\nvector_dbs / voice / app"]
  end
  T -->|projects| A
  T -->|projects| V
  T -->|projects| K
  T -->|projects| M
  A --> K
  V -->|search_docs tool| K
  M -->|23 knowledge tools| K

The surfaces

SurfaceEntry pointPage
Assistant / Mosaic AidePOST /api/assistant/chat (SSE), /aideassistant
Web chat/chat, POST /api/chat (web lens)assistant
VoiceGET /api/voice (WebSocket), /voicevoice
Knowledge / Graph-RAG/knowledge, POST /api/vectordb/{kb}/searchknowledge
MCPcargo run -- mcp (stdio), POST /mcp (knowledge)mcp

The seams are env, not code

Every AI surface is configured through environment variables (secret-backed in the deploy spec), so the generated app stays dependency-light and deterministic. No provider is hard-wired:

  • MOAIC_CHAT_* — the chat/assistant LLM (base URL, API key, model).
  • MOAIC_EMBED_* — embeddings (MOAIC_EMBED_LOCAL=1 for the deterministic local 256-dim embedder; MOAIC_EMBED_ONNX_MODEL for an ONNX sentence model).
  • MOAIC_RERANK_* — an optional cross-encoder re-rank seam.
  • MOAIC_VOICE_{LLM,STT,TTS}_* — the voice pipeline stages.
  • MOAIC_MCP_<NAME>_URL — external MCP servers an app consumes.

Without the relevant keys the surfaces degrade gracefully (grounded no-LLM answers, or a 503 on the voice upgrade) — the routes stay mounted and the app still builds and runs.

The doctrine

  • The engine is pure Rust and vendored. mosaic-knowledge has no network deps; it is include_str!-vendored into every generated app that declares a vector_dbs. LLM/embedding calls happen at the app's seams, never in the engine.
  • Retrieval is grounded and deterministic. Hybrid BM25 + vector retrieval, graph-augmented re-ranking (Personalized PageRank), and citation/URL rewriting so answers point at real sources.
  • Declarative floor, runtime ceiling. Declared knowledge sources are the floor; sources and typed entities can be added at runtime (ADR 0014, ADR 0047).