Mosaic one model, many lenses

ADR 0017: model registry + runtime model configurability

  • Status: accepted
  • Date: 2026-10-01
  • Supersedes: nothing (activates the dead vector_dbs[].model key; adds the model_registry: top-level declaration; extends the env-based model config with a declared floor and a runtime overlay)

Context

Which model does what is today env-only:

  • workflow llm nodes: spec.provider → MOAIC_{PROVIDER}_{API_KEY, BASE_URL,MODEL}, falling back to the global MOAIC_CHAT_* family; the model name falls back to spec.name, then gpt-4o-mini;
  • assistant chat: MOAIC_CHAT_{API_KEY,BASE_URL,MODEL} (key falls back to MOAIC_EMBED_API_KEY);
  • embeddings: MOAIC_EMBED_{API_KEY (falls back to CHAT_API_KEY), BASE_URL, MODEL} (default text-embedding-3-small);
  • rerank: MOAIC_RERANK_{API_KEY,BASE_URL,MODEL} (default rerank-english-v3.0, fail-soft).

And vector_dbs[].model ({provider, name}) is parsed and shown on the /kb/ site page but never consumed — a dead key.

There is no declared inventory of the models an app uses, no way to see (at runtime) which model/provider is effective, and no way to switch the effective model at runtime or to list the models a provider actually offers — the requirement: models for embeddings/LLM/etc. configurable, integrated with runtime configurability (provider + available models), with good software patterns.

Decision

The registry (declared floor)

A new top-level declaration — model_registry: (the models: key is already the IAM role-model vocabulary):

model_registry:
  - name: fast
    about: "Cheap model for high-volume calls."
    provider: groq        # env family: MOAIC_GROQ_{API_KEY,BASE_URL,MODEL}
    model: llama-3.1-8b   # default model name
    base_url: "https://api.groq.com/openai/v1"

Three built-in roles are always registered (a declared entry with the same name overrides its defaults — the declared-floor posture of ADR 0014/0016):

nameenv familydefault modelnotes
chatCHATgpt-4o-miniassistant chat + llm node fallback
embedEMBEDtext-embedding-3-smallkey falls back to MOAIC_CHAT_API_KEY
rerankRERANKrerank-english-v3.0fail-soft, no default base: absent key or absent base = no rerank

Resolution order per name: runtime overlay → env MOAIC_{PROVIDER}_{MODEL,BASE_URL} → declared model/base_url → built-in default. The API key is always env-only (MOAIC_{PROVIDER}_API_KEY; secrets never inlined, the Z7/ADR 0014/0016 posture).

Consumers rewired through one resolver

The generated app gains the OOP shape: a const registry (MODELS: &[ModelConfig]), a pure resolver (resolve_model(name) -> Option<EffectiveModel>: model + base_url + the key's env var name), and a runtime overlay. Every consumer resolves through it:

  • assistant chat → chat;
  • workflow llm node: a new spec.model = registry name (the node's provider/name legacy pair keeps working unchanged);
  • embeddings → embed, per collection: the dead vector_dbs[].model becomes live — its name (a registered model) selects that collection's embedding model (embed_text(text, collection));
  • rerank → rerank.

Runtime configurability

  • GET /api/models — the effective registry: per entry {name, about, provider, model, base_url, key: "set"|"unset" (never the value), origin: "declared"|"runtime", built_in};
  • POST /api/models/{name} {model?, base_url?} — a runtime override, persisted to <base>/.runtime/models.json (survives restarts; the ADR 0014 registry pattern — scratch state for e2e, reset per run). The name must exist in the registry (declared or built-in) → fail-closed named error;
  • DELETE /api/models/{name} — clear the override (back to declared);
  • GET /api/models/available?provider=X — list the models the provider offers: GET {base_url}/models with the provider key (OpenAI- compatible). Fail-closed: key unset → 400 named error; provider failure → 502 named error (the listing is best-effort information, never guessed);
  • MCP: list_models tool (one core, many bridges);
  • Site: build-time /models/ page (zola + html backends, nav entry) projecting the declared registry (the ADR 0008 projection — env values and runtime state are app data, never static content).

Proof

  • Unit tests (plan + codegen): built-ins + declared entries + override precedence; unknown-name fail-closed; the dead vector_dbs[].model selects the per-collection embed model; GET /api/models (+override + available) routes/handlers, list_models MCP tool, /models/ site page emit.
  • e2e (full-stack): GET /api/models → 200 (lists chat/embed/rerank); POST /api/models/chat {model} → 200 and GET reflects the runtime origin; POST /api/models/nope → 400; GET /api/models/available without a key → 400 (fail-closed, hermetic — no network in CI).
  • Conformance goldens re-pinned; workspace tests + clippy -D warnings + fmt clean.

Consequences

  • An operator can see exactly which model does what (declared + effective, key masked), change the effective model at runtime (persisted, restart-proof), and discover what a provider offers — before pointing a registry entry at a model id that does not exist.
  • All model config flows through one resolver with one resolution order; the env stays the secret channel (Z7), the tessera stays declarable, and the runtime overlay is the operator channel — the same declared-floor / runtime-overlay / env-for-secrets posture as the knowledge sources (ADR 0014/0015) and the part config (ADR 0016).
  • vector_dbs[].model stops being dead weight; the per-collection embed model is the first per-instance model binding.
  • Zero behavior change for apps without a model_registry: (the built-ins resolve to exactly today's env behavior, including the MOAIC_CHAT_* fallbacks and the fail-soft rerank).