ADR 0017: model registry + runtime model configurability
- Status: accepted
- Date: 2026-10-01
- Supersedes: nothing (activates the dead
vector_dbs[].modelkey; adds themodel_registry:top-level declaration; extends the env-based model config with a declared floor and a runtime overlay)
Context
Which model does what is today env-only:
- workflow
llmnodes:spec.provider→MOAIC_{PROVIDER}_{API_KEY, BASE_URL,MODEL}, falling back to the globalMOAIC_CHAT_*family; the model name falls back tospec.name, thengpt-4o-mini; - assistant chat:
MOAIC_CHAT_{API_KEY,BASE_URL,MODEL}(key falls back toMOAIC_EMBED_API_KEY); - embeddings:
MOAIC_EMBED_{API_KEY (falls back to CHAT_API_KEY), BASE_URL, MODEL}(defaulttext-embedding-3-small); - rerank:
MOAIC_RERANK_{API_KEY,BASE_URL,MODEL}(defaultrerank-english-v3.0, fail-soft).
And vector_dbs[].model ({provider, name}) is parsed and shown on the
/kb/ site page but never consumed — a dead key.
There is no declared inventory of the models an app uses, no way to see (at runtime) which model/provider is effective, and no way to switch the effective model at runtime or to list the models a provider actually offers — the requirement: models for embeddings/LLM/etc. configurable, integrated with runtime configurability (provider + available models), with good software patterns.
Decision
The registry (declared floor)
A new top-level declaration — model_registry: (the models: key is
already the IAM role-model vocabulary):
model_registry:
- name: fast
about: "Cheap model for high-volume calls."
provider: groq # env family: MOAIC_GROQ_{API_KEY,BASE_URL,MODEL}
model: llama-3.1-8b # default model name
base_url: "https://api.groq.com/openai/v1"
Three built-in roles are always registered (a declared entry with the same name overrides its defaults — the declared-floor posture of ADR 0014/0016):
| name | env family | default model | notes |
|---|---|---|---|
chat | CHAT | gpt-4o-mini | assistant chat + llm node fallback |
embed | EMBED | text-embedding-3-small | key falls back to MOAIC_CHAT_API_KEY |
rerank | RERANK | rerank-english-v3.0 | fail-soft, no default base: absent key or absent base = no rerank |
Resolution order per name: runtime overlay → env
MOAIC_{PROVIDER}_{MODEL,BASE_URL} → declared model/base_url → built-in
default. The API key is always env-only (MOAIC_{PROVIDER}_API_KEY;
secrets never inlined, the Z7/ADR 0014/0016 posture).
Consumers rewired through one resolver
The generated app gains the OOP shape: a const registry
(MODELS: &[ModelConfig]), a pure resolver
(resolve_model(name) -> Option<EffectiveModel>: model + base_url + the
key's env var name), and a runtime overlay. Every consumer resolves through
it:
- assistant chat →
chat; - workflow
llmnode: a newspec.model= registry name (the node'sprovider/namelegacy pair keeps working unchanged); - embeddings →
embed, per collection: the deadvector_dbs[].modelbecomes live — itsname(a registered model) selects that collection's embedding model (embed_text(text, collection)); - rerank →
rerank.
Runtime configurability
GET /api/models— the effective registry: per entry{name, about, provider, model, base_url, key: "set"|"unset" (never the value), origin: "declared"|"runtime", built_in};POST /api/models/{name}{model?, base_url?}— a runtime override, persisted to<base>/.runtime/models.json(survives restarts; the ADR 0014 registry pattern — scratch state for e2e, reset per run). The name must exist in the registry (declared or built-in) → fail-closed named error;DELETE /api/models/{name}— clear the override (back to declared);GET /api/models/available?provider=X— list the models the provider offers:GET {base_url}/modelswith the provider key (OpenAI- compatible). Fail-closed: key unset → 400 named error; provider failure → 502 named error (the listing is best-effort information, never guessed);- MCP:
list_modelstool (one core, many bridges); - Site: build-time
/models/page (zola + html backends, nav entry) projecting the declared registry (the ADR 0008 projection — env values and runtime state are app data, never static content).
Proof
- Unit tests (plan + codegen): built-ins + declared entries + override
precedence; unknown-name fail-closed; the dead
vector_dbs[].modelselects the per-collection embed model;GET /api/models(+override + available) routes/handlers,list_modelsMCP tool,/models/site page emit. - e2e (full-stack):
GET /api/models→ 200 (listschat/embed/rerank);POST /api/models/chat {model}→ 200 andGETreflects the runtime origin;POST /api/models/nope→ 400;GET /api/models/availablewithout a key → 400 (fail-closed, hermetic — no network in CI). - Conformance goldens re-pinned; workspace tests + clippy
-D warnings+ fmt clean.
Consequences
- An operator can see exactly which model does what (declared + effective, key masked), change the effective model at runtime (persisted, restart-proof), and discover what a provider offers — before pointing a registry entry at a model id that does not exist.
- All model config flows through one resolver with one resolution order; the env stays the secret channel (Z7), the tessera stays declarable, and the runtime overlay is the operator channel — the same declared-floor / runtime-overlay / env-for-secrets posture as the knowledge sources (ADR 0014/0015) and the part config (ADR 0016).
vector_dbs[].modelstops being dead weight; the per-collection embed model is the first per-instance model binding.- Zero behavior change for apps without a
model_registry:(the built-ins resolve to exactly today's env behavior, including theMOAIC_CHAT_*fallbacks and the fail-soft rerank).