Mosaic Handbook

A single-file, print-ready companion to the Mosaic documentation — every guide, the AI manual, and all 47 ADRs. Use your browser's print-to-PDF to export.

ADRs

The Laws of Mosaic

Mosaic is not a framework and it is not a code generator with opinions. It is a small compiler. Six laws make it what it is. Every law is enforced by an executable artifact — a test, a CI job, or a type signature — not by goodwill. If a law cannot be enforced, it is a slogan and it does not belong here.

Law 1 — Compiler Input

The entire input to mosaic is a finite, declarative set of files:

  • mosaic.yaml — the app (which tessera, which deployments),
  • <tessera>/tessera.yaml — the unit (kind, enums, entities, ops),
  • <tessera>/src/*.rs — hand-written behavior, copied verbatim,
  • deploy/*.yaml — deployments (or inline deploy: sugar in mosaic.yaml).

No plugins, no build scripts, no configuration that depends on the environment. Every key is declared, unknown keys are errors (deny_unknown_fields), and the JSON schema of each file kind is machine-printable: mosaic schema tessera|mosaic|deploy.

Enforced by: the grammar types in crates/mosaic-grammar (serde deny_unknown_fields on every spec struct) and the CI schemas are valid JSON step.

Law 2 — Zero Tax

A capability you do not declare costs you nothing. The type vocabulary is closed (nine primitives, declared entities/enums, one level of lists); the middleware vocabulary is closed; the lens set is closed. There is no expression language and therefore nothing to learn, nothing to sandbox, and nothing to optimize. When a need appears that the vocabulary does not cover, the vocabulary is extended in a deliberate ADR — it is never escaped from.

Enforced by: ADR 0001 and its proofs (mosaic-grammar::ty::tests::rejects_option_and_nested_arrays, accepts_both_list_spellings).

Law 3 — One Model

There is exactly one model: the tessera. Every artifact — REST routes, the client CLI, the OpenAPI document, the Dockerfile, the k8s manifests — is a projection of the same resolved plan. Two lenses can never drift apart, because neither of them owns the truth.

Enforced by: ADR 0002 and its proof (mosaic-conformance::tests::render_is_a_pure_function_of_the_plan).

Law 4 — Thin Renderers

A renderer prints. It does not decide. Anything that requires choosing a fact — a route, a default, a name, a derived op — happens exactly once, in mosaic-core::resolve. Renderers take a ResolvedPlan and nothing else; that signature is the law.

Enforced by: ADR 0004 and its proof (mosaic-render integration test core_never_depends_on_render, which fails the build if core ever depends on render).

Law 5 — Proven Proof

Every ADR names the test or CI step that proves its decision, and that artifact runs on every push. A decision without a proof is a rumor.

Enforced by: the Proof field in docs/adr/*.md and the CI test job, which runs all of them.

Law 6 — Copilot Law

A copilot must be able to work in this repo without tribal knowledge. Everything it needs is in the repository: the JSON schemas (mosaic schema), the golden outputs (every rendering is committed and diff-able), the laws and ADRs (short enough to fit in one context window), and examples that build. If a fact is not written down, it is not a fact.

Enforced by: the examples/ conformance suite — if the committed goldens stop matching the code, CI fails, so documentation can never silently rot.

Deployment testing

How mosaic apps are tested against real clusters and real clouds (ADR 0010: one app, many targets).

The three workflows

workflow runner what it proves trigger
e2e-k8s.yml GitHub-hosted (k3d in Docker) full data-mesh pipeline (backfill, CDC, poll→model, poll→file) against in-cluster peers + the rendered Helm chart every push/PR + manual
e2e-onprem.yml self-hosted (mosaic-onprem) the same harness against your k3s — own-infra production path manual
deploy-onprem.yml self-hosted (mosaic-onprem) builds a selected k8s deployment spec, imports the image into li7's k3s, and installs its Helm chart manual
e2e-cloud.yml GitHub-hosted app deployed with the rendered chart against managed stores (RDS/Cloud SQL + OpenSearch/qdrant) on AWS/GCP manual, secrets-gated

All three use the same entry point:

scripts/e2e/harness.sh <helm-chart-dir> <app-image> [namespace] [extra helm args...]

The harness applies the peer fixtures (scripts/e2e/peers/: postgres CDC source, mysql poll source, clickhouse sink), seeds them, installs the chart, port-forwards the app, and asserts:

  1. /health responds;
  2. GET /api/sync lists all three mirrors;
  3. backfill lands the 2 non-draft orders in ClickHouse (where clause honored) and the 3 snapshot order_items;
  4. CDC: a new row inserted into the postgres peer appears in ClickHouse;
  5. POST /api/sync/{mirror}/once is accepted;
  6. poll→file: the JSONL archive contains the CRM customers;
  7. poll→model: GET /api/order contains the derived customers.

On any failure it dumps pods, events, and the app's logs.

Running the k3d e2e locally (optional)

The CI job does it for you; to iterate locally (needs Docker + k3d):

curl -sSL https://raw.githubusercontent.com/k3d-io/k3d/main/install.sh | TAG=5.6.0 sh
k3d cluster create e2e --image docker.io/rancher/k3s:v1.30.5-k3s1

cargo run -q -p mosaic-cli -- build examples/data-sync -o /tmp/out-ds
(cd /tmp/out-ds && docker build -f deploy/e2e-k8s/Dockerfile -t mosaic/data-sync-e2e:local app/)
docker save mosaic/data-sync-e2e:local | k3d image import -c e2e -

bash scripts/e2e/harness.sh /tmp/out-ds/deploy/e2e-k8s/helm mosaic/data-sync-e2e:local mosaic-e2e --set image.pullPolicy=Never

k3d cluster delete e2e

Registering the on-prem runner

The on-prem workflow needs a GitHub Actions runner labeled mosaic-onprem with, on the box:

  • docker (for the image build),
  • kubectl configured against your k3s cluster (KUBECONFIG in the runner's env),
  • helm,
  • the k3s CLI (image import via k3s ctr images import),
  • rustup (the render step) — or a prebuilt mosaic-cli.

Register:

curl -sSfL https://raw.githubusercontent.com/actions/runner/main/docs/scripts/add-self-hosted-runner.sh | bash -s <URL> <TOKEN> mosaic-onprem

Then dispatch e2e-onprem from the Actions tab. The harness reuses the mosaic-e2e namespace on your cluster — peers and the app are (re)installed there; the image is tagged mosaic/data-sync-e2e:local, so prune old ones with k3s ctr images rm as you like.

Deploying to the on-prem cluster

Public-app infrastructure

The public-app target is the single-node k3s cluster on li7 (192.168.11.63). It is the real on-prem hosting environment for public Mosaic applications, separate from the disposable k3d cluster used by e2e-k8s.yml on GitHub-hosted runners.

The network path is:

<hostname>.eisler-systems.de
  -> Netcup wildcard DNS / FRITZ!Box public address
  -> FRITZ!Box port forwarding (80/443)
  -> li7 k3s Traefik
  -> Kubernetes Ingress
  -> Mosaic Service

Use a public eisler-systems.de hostname for an application Ingress and set the class to Traefik. cert-manager is installed in the cluster and the letsencrypt-prod ClusterIssuer obtains and renews certificates through the Traefik HTTP-01 path. The deployment workflow therefore enables TLS and sets the host, secret name, and ClusterIssuer without storing certificates in the repository.

The persistent eugeis/mosaic self-hosted runner on li7 is labeled mosaic-onprem and has Docker, Rust, Helm, kubectl, and the k3s CLI. Its kubeconfig is local to the runner at /home/ee/.kube/config; workflows do not upload a kubeconfig secret to GitHub-hosted runners or expose the Kubernetes API publicly.

Dispatch deploy-onprem.yml to deploy a rendered k8s spec to the local k3s cluster. The workflow runs only on the mosaic-onprem self-hosted runner, which has the cluster kubeconfig installed locally. It does not use a GitHub kubeconfig secret or expose the Kubernetes API to GitHub-hosted runners.

The workflow renders the selected app, builds its generated Dockerfile, imports the image through k3s ctr images import, and runs helm upgrade --install. The default inputs deploy examples/orders using its helm deployment spec to the mosaic namespace at app.eisler-systems.de. The host input overrides the generated chart's ingress host; cert-manager and the cluster ingress controller remain responsible for TLS.

The self-hosted runner is trusted with cluster-admin access. Keep this workflow manual and restrict repository write access accordingly: a workflow executed on this runner can change the local cluster.

Cloud e2e (AWS / GCP)

The cloud job assumes the spec's Terraform was applied once (it creates the cluster, the registry, the managed stores, the namespace, and the DSN secrets the chart reads). Per cloud:

AWS (deploy/prod-aws/)

cd examples/full-stack   # after: cargo run -q -p mosaic-cli -- build examples/full-stack -o /tmp/out-fs
cd /tmp/out-fs/deploy/prod-aws/infra
terraform init && terraform apply     # EKS + ECR + RDS + OpenSearch + secrets

Repo configuration (Settings → Secrets and variables → Actions):

name kind value
E2E_AWS variable true
AWS_ACCESS_KEY_ID / AWS_SECRET_ACCESS_KEY secret a key allowed to push to ECR and kubectl against the cluster
AWS_E2E_REGION variable us-east-1 (the spec's region)
AWS_E2E_REGISTRY variable the ECR repository URL (terraform output ecr_repository_uri)
AWS_E2E_KUBECONFIG secret the kubeconfig for the EKS cluster (single blob; aws eks update-kubeconfig --name <cluster> renders one)

The job builds the image, pushes it to ECR, helm installs deploy/prod-aws/helm (db + vector mode: external — the DSN secrets come from Terraform), waits for readiness, and smokes /health plus a write through the managed RDS event store, then uninstalls.

GCP (deploy/prod-gcp/)

cd /tmp/out-fs/deploy/prod-gcp/infra
terraform init && terraform apply     # GKE + Artifact Registry + Cloud SQL + db secret
name kind value
E2E_GCP variable true
GCP_SA_JSON secret a service-account key with container.admin + artifactregistry.writer on the project
GCP_E2E_PROJECT variable the project id
GCP_E2E_CLUSTER variable the GKE cluster name
GCP_E2E_REGION variable europe-west1 (the spec's region)

Note: GCP has no first-party managed Qdrant with a stable Terraform resource, so prod-gcp runs qdrant in-cluster on GKE (managed: false); only the postgres store is managed (Cloud SQL). The plan layer rejects managed combinations without a stable URL resource (deploy-managed-unavailable) — see the matrix in ADR 0010.

Azure / Alibaba

azure-aks and alibaba-ack generate the same infra/ + helm/ + Dockerfile triplet (AKS module + ACR; ACK module + ApsaraDB RDS / Redis / ClickHouse). There is no automated cloud job for them yet — run the AWS job manually as a template: apply the terraform, build/push the image from deploy/<spec>/Dockerfile, helm install deploy/<spec>/helm, smoke, uninstall.

Cost notes

  • e2e-k8s (k3d) runs on a shared runner: a few minutes of CPU, no bill.
  • The on-prem run uses your existing cluster.
  • Cloud jobs are manual only and only touch resources you provisioned; the job uninstalls the Helm release but never destroys Terraform state. terraform destroy when you are done.

Troubleshooting

  • App pod stuck ContainerCreating on a cloud target: the DSN secret is missing — kubectl -n <ns> get secret <app>-db <app>-vector. They are created by Terraform, not the chart.
  • ImagePullBackOff on k3d: the image import step failed or the tag drifted — the caller passes --set image.pullPolicy=Never precisely so a preloaded image wins.
  • CDC never lands: the peer postgres must run wal_level=logical (the fixture sets it); check the app logs for replication errors.

Re-implementation playbook

How an agent (or a human) re-implements a legacy system from its tessera + knowledge base. Everything here is deterministic and LLM-free — the LLM does the reading and the writing; mosaic supplies the ground truth.

The app's MCP server (/mcp, streamable HTTP) exposes the tools below. REST equivalents live under /api/vectordb/<kb>/….

Indexing a legacy repo (K8)

The legacy system does not have to be on the app's disk. A knowledge source can be a git repository:

vector_dbs:
  - name: kb
    sources:
      - git:
          url: "https://git.internal/legacy/billing.git"   # or a local path / .bundle file
          ref: "release-2026"                              # branch / tag / rev (optional)

The app clones it through the git CLI at boot (shallow --depth 1 for remote URLs) into a hidden cache (<knowledge-base>/.gitcache/<label>) and indexes it exactly like a local source — symbols, endpoints, git facts, hybrid search, stats, MCP tools all see it. A git bundle (git bundle create repo.bundle main) is a single committable file git clone accepts — the way to keep a repo fixture hermetic in a repo or CI. Existing checkouts are reused (delete the cache dir to force a fresh clone). mosaic knowledge-report <tessera-dir> renders the same profile offline — the re-implementation brief for a repo the app has never seen.

The loop

  1. repo_profile (kb) — the re-implementation profile in one call: the app's own DSL entity counts (aggregates, commands, events, workflows, endpoints, cli, schedules, flags), the source inventory (files, lines/pages, git facts), LOC per file (top 10), the language mix, the symbol + endpoint census, dependency manifests (parsed from Cargo.toml, package.json, go.mod, requirements.txt under the sources), the book library, and citation totals. Start every re-implementation here.
  2. search_knowledge (kb, query) — hybrid retrieval (BM25 + vector, RRF fused; MMR de-duplication; optional rerank) over the indexed chunks. Hits carry the exact locator (file:lines, book.pdf:3, doc#n) and, when a hit is a split child, its parent section for context.
  3. list_symbols / read_chunk — the code index: every function/method/ class/struct/enum with file:line locators and detected HTTP endpoints (kind: "endpoint"), and verbatim chunk reads.
  4. cite (kb, id, from?) — record a quote. Pass from (your notes section) to create a cites edge in the knowledge graph; the count feeds the stats and the report's Sources section.
  5. query_graph (kb, node?, depth?) — the deterministic graph: symbol → symbol uses-edges, file/symbol/endpoint containment, book → chapter structure, report → cited entries. Expand one node's neighborhood, or fetch the whole graph.
  6. make_report (kb, title, sections[{heading, ids[]}]) — export the notes as markdown: each section quotes its chunks verbatim (split children expand to their parent section) with the locator, followed by a Sources section built from the recorded citation counts.

CLI

The same profile renders as a sync-style markdown file without a running app:

mosaic knowledge-report <tessera-dir> [--collection <kb>] [--out sync/<name>.knowledge.md]

It ingests the tessera's declared sources with the workspace engine and writes the report (app entities, sources, top files, languages, symbols, endpoints, deps, books, git) to --out (default: stdout). Wire it into your sync pipeline to keep the re-implementation ledger current as the legacy tree changes.

Rules for the notes

  • Quote, don't paraphrase: make_report sections must reference chunk ids from search_knowledge / list_symbols / read_chunk.
  • One cite per quote, with from = the notes section heading — the graph and the report's Sources section depend on it.
  • Keep the profile fresh: re-run repo_profile (or mosaic knowledge-report) after touching the legacy tree; reindex_knowledge re-walks the sources.

Thesis / book-library graph-RAG with Mosaic

How to put a book library into a Mosaic app's knowledge base and query it as a typed-entity graph — persons, places, topics, doctrines, councils, written works, sermons — for thesis writing and research. This is the ADR 0047 P47–P52 feature set (graph-augmented retrieval + code-RAG + book concepts + typed entities).

The model, in one line: every book chapter is a graph node; the chapter's typed entities (your curated gazetteer) and topics (auto-extracted salient terms) are nodes too; chapter → entity mentions edges connect them; Personalized PageRank (P48) then resurfaces the load-bearing entities/chapters around any query.


1. What you get

Surface What it answers
GET /api/vectordb/{kb}/entity?name=X&type= "Which chapters discuss entity X?" (optionally within one type)
GET /api/vectordb/{kb}/entities?type= "List all persons / councils / doctrines … with how often each is mentioned"
GET /api/vectordb/{kb}/concept?name=X Same, but for auto-extracted topics (P51; find_concept = find_entity(X, "topic"))
GET /api/vectordb/{kb}/graph?node=…&depth=… Expand any entity/chapter into its neighborhood (who mentions it, what it co-mentions)
GET /api/vectordb/{kb}/search?q=… Hybrid (BM25 + vector) retrieval, re-ranked by PPR when retrieval.graph: true
GET /api/vectordb/{kb}/gazetteer The full typed-entity gazetteer with provenance (tessera floor + runtime layer)
POST /api/vectordb/{kb}/gazetteer/add · DELETE …/gazetteer/remove Curate the gazetteer at runtime (no redeploy, no reindex, persists across reboots)
MCP find_entity / list_entities / find_concept / get_graph / search_knowledge / make_report The same, for an LLM agent
MCP list_gazetteer / add_entity / remove_entity Curate the typed entities live, for an LLM agent
POST /api/vectordb/{kb}/report / MCP make_report Export a thesis section quoting the chunks verbatim, with locators + a Sources list

Entity types are free-form — declare whatever your domain needs. The recommended standard set (ENTITY_TYPES): person, location, topic, doctrine, council, book, sermon, institution, event, concept, other. For a theology thesis you'd typically use: person (theologians, popes, reformers), location, topic, doctrine (doctrine statements), council (ecumenical/local councils — "congresses"), book (written works), sermon (preaching), institution (churches, orders, universities), event.


2. Add the library

  1. Copy the books (EPUB / DOCX / PDF) into the app's knowledge area. For the mosaic-pages-style site that is the books/ folder (declared as a source):

yaml vector_dbs: - name: docs sources: - path: books # EPUB / DOCX / PDF chapter model kind: auto strategy: books: true # parse EPUB (OPF spine) / DOCX (Heading1) / PDF chapters

PDFs are best-effort (chapter markers via a heading heuristic; set strategy.ocr_cmd / MOAIC_OCR_CMD for scanned pages). EPUB and DOCX give the cleanest chapter model.

  1. On li7, after copying the files, force a full reindex (delta ingest keeps unchanged entries, so a new book set needs a full rebuild):

sh curl -u <user>:<pass> -X POST https://<host>/api/vectordb/docs/reindex

3. Declare the gazetteer (the important part)

The typed entities come from a small curated dictionary in the tessera — the researcher (you) knows the key entities of the thesis. Per entry: a type, a canonical name (names[0]), and alias surface forms (the rest):

vector_dbs:
  - name: docs
    retrieval:
      graph: true          # P48: PPR re-ranking (resurfaces load-bearing entities)
      graph_weight: 0.5
    entities:
      - type: person
        names: [John Calvin, Calvin, Jean Calvin]
      - type: person
        names: [Martin Luther, Luther]
      - type: council
        names: [Council of Trent, Trent, the Council of Trent]
      - type: council
        names: [Council of Nicaea, Nicaea, First Council of Nicaea]
      - type: location
        names: [Geneva]
      - type: location
        names: [Wittenberg]
      - type: doctrine
        names: [Sola Fide, sola fide, faith alone]
      - type: doctrine
        names: [Imputed Righteousness, imputation of righteousness]
      - type: book
        names: [Institutes of the Christian Religion, the Institutes]
      - type: sermon
        names: [Bannerman's Lectures, Lectures on the Church]

Rules of the matcher (deterministic, no LLM):

  • case-insensitive; word-boundary-anchored (Calvin does not match Calvinism / Calvinist);
  • longest form wins per span (Council of Trent is not double-counted by its Trent alias);
  • a surface form may be a multi-word phrase (Imputed Righteousness);
  • the same surface form in two types yields two entities (council:Trent and location:Trent both) — curate deliberately.

The auto topics (P51) need no gazetteer: each chapter's salient terms become type = "topic" entities, so even an uncurated library gets a typed topic graph.

## 3b. Curate the gazetteer at runtime (no tessera edit, no redeploy)

The tessera gazetteer is the immutable floor. On top of it there is a runtime registry — you (or an LLM agent via MCP) add/remove typed entities on the live server as you read, and they are reflected in the graph immediately (no reindex) and persisted across reboots. This is the fast loop for thesis work: read a chapter, spot a person/council/doctrine you hadn't declared, add it, and the next find_entity / make_report already uses it.

```sh # Add a typed entity (names[0] = canonical, the rest = aliases): curl -u … -X POST "…/api/vectordb/docs/gazetteer/add" \ -H 'Content-Type: application/json' \ -d '{"type":"council","names":["Council of Nicaea","Nicaea"]}'

# …or a single surface form: curl -u … -X POST "…/api/vectordb/docs/gazetteer/add" \ -d '{"type":"person","name":"Thomas Aquinas"}'

# List the whole gazetteer, with provenance (origin: "declared" | "runtime"): curl -u … "…/api/vectordb/docs/gazetteer"

# Remove a runtime entity (declared/floor entities are immutable — edit the tessera): curl -u … -X DELETE "…/api/vectordb/docs/gazetteer/remove?type=council&name=Council%20of%20Nicaea" ```

MCP equivalents: add_entity / remove_entity / list_gazetteer.

Rules (same determinism as the DSL): a runtime add whose node id is already in the tessera floor is rejected (the tessera stays the source of truth for declared entities); re-adding a runtime entity is an idempotent no-op; remove only affects the runtime layer. Because the graph is derived on demand from the effective gazetteer (floor + runtime), no reindex is needed — the change is live on the next query and survives a restart. When an entity graduates from "runtime" to "core to the thesis", promote it into the tessera entities: block so it becomes part of the immutable floor.

4. Query / research loop

# What is the corpus about, by type?
curl -u … "…/api/vectordb/docs/entities?type=person"
curl -u … "…/api/vectordb/docs/entities?type=council"

# Which chapters discuss a person, and how centrally?
curl -u … "…/api/vectordb/docs/entity?name=John%20Calvin"

# Expand the neighborhood (co-mentions, shared chapters):
curl -u … "…/api/vectordb/docs/graph?node=entity:person:John%20Calvin&depth=2"

# A grounded retrieval, PPR-re-ranked:
curl -u … -X POST "…/api/vectordb/docs/search" -d '{"query":"predestination in Calvin"}'

Then export a draft section that quotes the chapters verbatim (with locators) — the citation-ready thesis artifact:

curl -u … -X POST "…/api/vectordb/docs/report" \
  -d '{"title":"Predestination in Calvin","sections":[{"heading":"The doctrine","ids":["<chunk-id>", …]}]}'

An LLM agent gets the same through the MCP server (find_entity, list_entities, get_graph, search_knowledge, make_report) — have it enumerate the entities, expand the ones that matter, and draft from the cited chunks.


5. Best practices (libraries / thesis / research)

  • Curate the gazetteer iteratively. Start with the ~20–50 load-bearing entities (the thesis's named persons, councils, doctrines, key works). Query list_entities (all types) to see what the auto topics surface, then promote the important ones to curated typed entities with proper types + aliases.
  • Use types to disambiguate + to slice. list_entities?type=person vs ?type=doctrine is how a researcher navigates; the type is also the query filter, so "Trent" the council ≠ "Trent" the town.
  • Keep aliases. People and places have many surface forms (Latin/Greek names, abbreviations, "the Council of Trent" vs "Trent"). List them all — the matcher handles multi-word + case.
  • Lean on PPR, don't just keyword-search. With retrieval.graph: true, a query that touches a central entity resurfaces the chapters clustered around it (HippoRAG-style), which beats flat keyword ranking for "where is this doctrine developed across the corpus?"
  • Cite from the graph, not from memory. Use make_report / POST /report to build sections from real chunk ids — every claim carries a verbatim quote + a locator (book/chapter or file:line), and a Sources list is generated from the citation counts.
  • One collection per corpus, but many sources. The books, your notes, and any git repo of drafts all feed the same docs collection and the same graph, so PPR spans "the library + my notes."
  • Deterministic = reproducible. No LLM in the engine: the same library + gazetteer always yields the same graph. Re-index is idempotent.

6. What's here vs. what's a later slice

  • Here (P47–P52): hybrid retrieval + PPR re-rank; code symbol/uses/calls/ imports graph; book chapter graph; auto topic concepts (P51); typed-entity gazetteer (P52) with find_entity / list_entities + a runtime entity registry (add/remove/list typed entities on the live server, no redeploy, persists across reboots); report/citation export.
  • Later slices (not yet built):
  • Entity relations — edges between entities (is-a, authored, attended-council, co-occurrence) for a true thesis entity graph, on top of the chapter → entity mentions substrate.
  • LLM-assisted entity discovery — an agent-layer pass (an MCP tool + the LLM seam) that proposes typed entities/relations the deterministic gazetteer can't know in advance, to seed your gazetteer. The vendored engine stays LLM-free.
  • Cross-book alignment + disambiguation — the same surface form across books / homographs resolved to the right typed entity automatically.

Assistant & Mosaic Aide

Every app that declares a vector_dbs gets a RAG assistant for free. When the app is a normal (non-proxy) app, the assistant is also given a stable brand and its own surfaces — Mosaic Aide (ADR 0029).

The assistant

  • POST /api/assistant/chat — the core RAG chat. It streams over SSE: sources → page_sources → delta (token chunks) → done ({answer, sources, page_sources, llm, conversation_id}).
  • Grounding. Each turn runs chat_with_context: the conversation history window + the viewed page's context are assembled, retrieval happens first, and the final answer is post-processed to rewrite raw doc references into real page URLs (deterministic citation rewrite).
  • History. Multi-turn memory with a MOAIC_ASSISTANT_HISTORY_* / MOAIC_CHAT_HISTORY_* budget (turns + max chars) and optional summarization of dropped turns. Memory is in-process (single replica).
curl -N localhost:8080/api/assistant/chat \
  -H 'content-type: application/json' \
  -d '{"message":"what is an event-sourced ledger?","stream":true}'

Mosaic Aide

For a normal app, the same assistant is promoted to a branded helper:

  • GET /api/aide/profile — the Aide's identity + capabilities.
  • POST /api/aide/chat — an alias of the assistant chat.
  • GET /aide — a self-contained helper page.
  • aide_chat — an MCP tool so agents can use the Aide too.

Aide is auto-derived (there is no assistant: DSL key): it appears whenever ≥1 vector_db is declared, so the capability and the brand can't drift apart.

Web chat (web lens)

The web lens adds a built-in /chat page + POST /api/chat:

  • App facts (its commands, models, workflows) are ingested into a dependency-free memvid-core memory file; retrieval is lexical, with an optional LLM via MOAIC_CHAT_* and a grounded no-LLM fallback.
  • Browser voice (ADR 0033): dictation into the input via the Web Speech API (SpeechRecognition) and "speak reply" via speechSynthesis — client-side only, no backend, no provider key.
  • An app that mounts its own /api/chat or /chat opts out of the built-in (app_owns_chat).

Configuration

Env Meaning
MOAIC_CHAT_BASE_URL / MOAIC_CHAT_API_KEY / MOAIC_CHAT_MODEL the assistant LLM
MOAIC_EMBED_* embeddings (see overview)
MOAIC_RERANK_* optional re-rank seam
MOAIC_ASSISTANT_HISTORY_TURNS / ..._MAX_CHARS history budget

Honest gaps

  • Aide is gated on a vector_db and has no dedicated assistant: DSL block.
  • Conversation memory is in-process (multi-replica needs an external store).
  • The chat path has no tool-calling loop (the voice and workflow agent node do); Aide is not yet page-grounded, and there are no attachments or guardrails/content-check.

Knowledge & Graph-RAG

mosaic-knowledge is a pure-Rust, dependency-free engine that is include_str!-vendored into every generated app that declares a vector_dbs. It ingests text, code, and books; builds a knowledge graph; and answers with grounded, graph-augmented retrieval. Provider calls (embeddings) happen only at the app's env seams, never inside the engine.

Ingestion tiers

Tier What it indexes Notes
book / pdf EPUB / DOCX / PDF pure-Rust parsing, chapter markers, OCR seam (MOAIC_OCR_CMD)
code source files + Cargo.toml/package.json/… symbol extraction, call sites, import edges
docs markdown / text tier chunker (parent–child chunks)
git sources repositories ADR 0013; delta re-indexing (ADR 0019); latest-changes analysis (K11)

Sources are the declared floor, runtime ceiling: runtime sources can be added without a redeploy (ADR 0014), and persisted to any backend through the OpenDAL seam (ADR 0015).

Retrieval

  1. Hybrid. BM25 (keyword) + vector (cosine) fused by reciprocal rank fusion.
  2. Graph-augmented re-rank (ADR 0047). Personalized PageRank over the persisted document graph, with Louvain communities + modularity; optional cross-encoder re-rank seam (MOAIC_RERANK_*).
  3. Citation. Every answer carries its sources; the assistant/voice rewrite raw references into real page URLs.

Embeddings

  • Default: a deterministic local 256-dim embedder (no network).
  • MOAIC_EMBED_LOCAL=1 forces local; MOAIC_EMBED_ONNX_MODEL loads an ONNX sentence model; MOAIC_EMBED_* points at a hosted embedding API.

Book concepts & typed entities (the thesis layer)

  • Concepts (ADR 0047 P51). Book text is turned into a concept/"topic" graph.
  • Typed entities (ADR 0047 P52). A gazetteer maps typed entities (person, location, topic, doctrine, council, book, sermon, …) to graph nodes. The DSL declares a floor; the runtime registry adds/removes entities live (add_entity / remove_entity) — persisted and restored at boot, reflected in the graph without re-indexing.

Reports & profile

  • make_report renders a markdown report over a graph neighborhood.
  • repo_profile (K6) + knowledge_stats + recent_changes (K11) describe a git source.
  • CLI: mosaic knowledge-report.

Surfaces

  • /knowledge — the knowledge sources page (declared floor + runtime).
  • POST /api/vectordb/{kb}/search — hybrid search.
  • /api/vectordb/{kb}/entity?name=…, /entities, /gazetteer — typed-entity lookup + runtime curation.
  • MCP — 23 knowledge tools (see MCP).

The 23 knowledge MCP tools

search_knowledge, read_chunk, list_sources, list_knowledge, list_symbols, list_books, cite, query_graph, find_symbol, find_callers, get_call_graph, find_similar_implementation, find_concept, find_entity, list_entities, list_gazetteer, add_entity, remove_entity, make_report, repo_profile, knowledge_stats, recent_changes, reindex_knowledge.

MCP

Mosaic generates a full MCP (Model Context Protocol) surface: two servers (one per app, one per knowledge base) and a client for consuming external servers. The tools are auto-derived from the tessera — "tools derive, they aren't hand-authored".

The app MCP server (stdio)

  • Run as cargo run -- mcp (generated main.rs → mcp subcommand).
  • One tool per exposed CQRS command (ADR 0027: zero unexposed commands), plus workflow tools, ruleset/command-group tools, and meta tools (list_parts, list_models, aide_chat).
  • tools/call is method-aware (path-parameter substitution, ADR 0029).
  • Secret-backed fields marked hidden are kept off the schema.

This is the seam the Reimplementation Playbook uses to drive an agent to re-implement an app from its MCP tools alone.

The knowledge MCP server (HTTP)

  • Mounted at POST /mcp (app lens) using rmcp (streamable-HTTP, stateless).
  • Exposes the 23 knowledge tools from the knowledge engine: search, read, cite, graph query, code-RAG symbol/call-graph, book concepts, typed entities (incl. runtime add_entity/remove_entity), reports, and re-indexing.

The MCP client

An app can consume external MCP servers:

  • DSL: app.mcp.servers — each resolved from MOAIC_MCP_<NAME>_URL.
  • Workflow nodes: an mcp node (mcp_call / mcp_list_tools) and an agent node's tools.
  • Surface: GET /api/mcp/{server}/tools.
app:
  mcp:
    servers:
      - name: external
        # resolved from MOAIC_MCP_EXTERNAL_URL

Why it matters

The MCP surface is what makes an app agent-native: the same commands a human drives from the CLI or web app are the same tools an LLM/agent can call. The knowledge MCP tools turn the RAG engine into a queryable context for any agent, and the app MCP tools let agents operate the app end-to-end.

AI overview

Mosaic ships a full AI stack that every generated app gets by projection — no hand-written glue. Four surfaces compose: a RAG assistant (with a branded auto-derived "Mosaic Aide"), a voice conversation channel (STT → LLM → TTS), a knowledge engine (hybrid + Graph-RAG over text, code, and books), and MCP servers/clients that expose all of it to agents.

flowchart LR
  subgraph Surfaces
    A["Assistant / Aide\nPOST /api/assistant/chat\nSSE"]
    V["Voice\nGET /api/voice (WS)"]
    K["Knowledge engine\nhybrid + Graph-RAG"]
    M["MCP servers\napp (stdio) + knowledge (HTTP)"]
  end
  subgraph App
    T["tessera.yaml\nvector_dbs / voice / app"]
  end
  T -->|projects| A
  T -->|projects| V
  T -->|projects| K
  T -->|projects| M
  A --> K
  V -->|search_docs tool| K
  M -->|23 knowledge tools| K

The surfaces

Surface Entry point Page
Assistant / Mosaic Aide POST /api/assistant/chat (SSE), /aide assistant
Web chat /chat, POST /api/chat (web lens) assistant
Voice GET /api/voice (WebSocket), /voice voice
Knowledge / Graph-RAG /knowledge, POST /api/vectordb/{kb}/search knowledge
MCP cargo run -- mcp (stdio), POST /mcp (knowledge) mcp

The seams are env, not code

Every AI surface is configured through environment variables (secret-backed in the deploy spec), so the generated app stays dependency-light and deterministic. No provider is hard-wired:

  • MOAIC_CHAT_* — the chat/assistant LLM (base URL, API key, model).
  • MOAIC_EMBED_* — embeddings (MOAIC_EMBED_LOCAL=1 for the deterministic local 256-dim embedder; MOAIC_EMBED_ONNX_MODEL for an ONNX sentence model).
  • MOAIC_RERANK_* — an optional cross-encoder re-rank seam.
  • MOAIC_VOICE_{LLM,STT,TTS}_* — the voice pipeline stages.
  • MOAIC_MCP_<NAME>_URL — external MCP servers an app consumes.

Without the relevant keys the surfaces degrade gracefully (grounded no-LLM answers, or a 503 on the voice upgrade) — the routes stay mounted and the app still builds and runs.

The doctrine

  • The engine is pure Rust and vendored. mosaic-knowledge has no network deps; it is include_str!-vendored into every generated app that declares a vector_dbs. LLM/embedding calls happen at the app's seams, never in the engine.
  • Retrieval is grounded and deterministic. Hybrid BM25 + vector retrieval, graph-augmented re-ranking (Personalized PageRank), and citation/URL rewriting so answers point at real sources.
  • Declarative floor, runtime ceiling. Declared knowledge sources are the floor; sources and typed entities can be added at runtime (ADR 0014, ADR 0047).

Voice (STT → LLM → TTS)

The voice: DSL projects a full voice conversation channel into every app that declares one. There are two modes: a pipeline (server-side STT → LLM → TTS, the default) and realtime (a provider-direct WebSocket relay).

Endpoints

  • GET /api/voice — the WebSocket conversation endpoint.
  • GET /voice — a self-contained browser widget (mic + transcript + audio).
  • The upgrade answers 503 when a stage's env key is unset; the route stays mounted so the app still builds and runs.

Pipeline mode (default)

sequenceDiagram
  participant B as Browser
  participant S as /api/voice (WS)
  participant STT as STT (OpenAI-compat)
  participant LLM as LLM
  participant K as knowledge (search_docs)
  participant TTS as TTS
  B->>S: audio frames (16 kHz PCM16, b64)
  S->>S: server VAD (RMS endpointing)
  S->>STT: transcribe (WAV)
  STT-->>S: transcript
  S->>LLM: turn (+ search_docs tool loop)
  LLM->>K: search_docs(query, top_k)
  K-->>LLM: chunks
  LLM-->>S: answer
  S->>TTS: per-sentence synthesis
  TTS-->>S: PCM16 @24 kHz
  S-->>B: audio frames (b64)
  • Endpointing. Server-side VAD: an RMS UtteranceEndpointer ends a turn on turn_silence_ms of silence (default 700).
  • STT. OpenAI-compatible multipart WAV; default model whisper-1.
  • LLM turn. A single (non-streamed) completion with a server-side search_docs tool loop (max 4 iterations) over the app's voice.site_search.kb knowledge base — the app's grounding is enforced server-side.
  • TTS. Per-sentence synthesis to PCM16 @24 kHz (default gpt-4o-mini-tts, voice alloy).
  • Barge-in. Voiced audio aborts the in-flight turn.

Wire protocol

  • In: welcome | audio | text | stop.
  • Out: speech | transcript | audio | tool | turn | error | goodbye.

Realtime mode

A provider-direct WebSocket relay with per-provider codecs:

  • OpenAI Realtime (default) or Gemini Live.
  • The server still executes the search_docs tool loop server-side, so grounding is preserved in realtime too.

Example

voice:
  mode: pipeline
  turn_silence_ms: 700
  site_search:
    kb: kb
    top_k: 5
  llm: { model: gpt-4o-mini }
  stt: { model: whisper-1 }
  tts: { model: gpt-4o-mini-tts, voice: alloy }

Configuration

Env Meaning
MOAIC_VOICE_LLM_{API_KEY,BASE_URL,MODEL} the voice LLM stage
MOAIC_VOICE_STT_{API_KEY,BASE_URL,MODEL} the STT stage
MOAIC_VOICE_TTS_{API_KEY,BASE_URL,MODEL} the TTS stage
MOAIC_VOICE_TTS_VOICE the TTS voice
MOAIC_VOICE_REALTIME_* the realtime relay

Honest gaps

  • One voice: surface per app (app lens); no multi-voice/tenant variants.
  • The pipeline LLM turn is non-streamed (streaming is at the audio/TTS frame level); no local/self-hosted STT or TTS; voice sessions have no cross-session memory; voice is deliberately not an MCP tool.

ADR 0001: The spec is data with a closed vocabulary

  • Status: accepted
  • Date: 2026-09-23

Context

Many DSLs grew expression languages: rule expressions, policy expressions, workflow expressions — each with its own parser, evaluator, sandbox, and error model. They became second programming languages: powerful, unlintable, and impossible to reason about exhaustively.

Decision

The spec files contain no expressions. Types are a closed vocabulary of nine primitives (string, bool, i64, i32, u64, u32, f64, uuid, datetime), declared entities and enums, and one level of lists. Behavior is expressed by referencing a hand-written Rust function; middleware is a closed vocabulary (v0: cache). Anything outside the vocabulary is a parse error, not a runtime surprise.

An expression language (a small CEL-like one) is explicitly deferred: it may come in a later version, as an ADR, when a concrete need exists.

Consequences

  • The parser is serde itself; there is no grammar to maintain.
  • mosaic schema can describe the entire input precisely.
  • Features are added one vocabulary word at a time, each visible in the schema and the goldens.

Proof

cargo test -p mosaic-grammar → ty::tests::rejects_option_and_nested_arrays (optional types, nested arrays, and empty types are rejected at parse time) and ty::tests::accepts_both_list_spellings.

ADR 0002: Facts are derived exactly once, in `resolve`

  • Status: accepted
  • Date: 2026-09-23

Context

In generator systems, derivations creep into every emitter: each lens recomputes routes, defaults, and names, and they drift. Debugging "which one decided this?" stops being possible.

Decision

mosaic-core::resolve is the only place a fact is derived: CRUD ops derived from entities, route paths, op ordering, the CLI command table, the deploy merge. Its output is the single ResolvedPlan. Renderers receive the plan and print it; the plan is a pure function of the workspace, so rendering is deterministic by construction.

Consequences

  • A lens can never disagree with another lens.
  • New lenses (ui, docs, voice) are emitters only: small, reviewable, testable against goldens.
  • All validation errors carry a structured diagnostic with a stable code and a repo-relative path.

Proof

cargo test -p mosaic-conformance → render_is_a_pure_function_of_the_plan (rendering the same plan twice yields byte-identical file maps).

ADR 0003: Output is deterministic and pinned by goldens

  • Status: accepted
  • Date: 2026-09-23

Context

A code generator that changes bytes for the same input destroys trust: every release produces a diff storm, and "did my change alter the output?" has no answer.

Decision

No timestamps, no randomness, no environment leakage in generated files. Iteration is over BTreeMaps and sorted collections; the OpenAPI document is serialized from serde_json maps, which sort keys. Every example under examples/ commits its full rendered output under <example>/golden/ plus a MANIFEST with a sha256 hash per file. CI runs the conformance check on every push; authors regenerate with mosaic conformance --update.

Consequences

  • Any behavior change that alters output is visible in the golden diff — the change and its consequences are reviewed together.
  • "Did the rendering change?" is answerable in one git diff.
  • Regeneration is a deliberate act, not an accident.

Proof

cargo test -p mosaic-conformance → examples_match_goldens (runs in CI on every push; fails on any byte difference, stale golden, or manifest mismatch).

ADR 0004: Renderers are thin — they print, they do not decide

  • Status: accepted
  • Date: 2026-09-23

Context

If a renderer derives a fact (a route, a default, a name), it becomes a hidden second source of truth. The system then has as many "models" as it has lenses.

Decision

The public entry point of mosaic-render is render(plan: &ResolvedPlan, user_src: &BTreeMap<String, String>) -> Rendered. It takes the resolved plan and nothing else. The dependency graph enforces the direction: render -> core -> grammar; core never depends on render. A test parses crates/mosaic-core/Cargo.toml and fails the build if a mosaic-render dependency ever appears.

Consequences

  • Adding a lens cannot introduce a new decision point.
  • The plan type is the entire contract between the compiler and the lenses; it is small enough to read in one sitting.

Proof

cargo test -p mosaic-render --test thin_renderers → core_never_depends_on_render (dependency-graph proof) and render_entry_takes_a_resolved_plan (signature proof, compiled).

ADR 0005: The CLI lens is a client, not a twin

  • Status: accepted
  • Date: 2026-09-23

Context

Generators commonly emit an in-process CLI: the binary links the store and handlers directly, so it can run without a server. That is a second deployment of the same state — two binaries to test, two ways for the CLI to disagree with the API, and the "convenience" evaporates the moment the store is anything but trivial.

Decision

The CLI lens generates a typed HTTP client (reqwest) against the running server. mosaic <op> subcommands and the REST API are the same request. State has exactly one home: the process you started with serve. In v0 the store is in-memory, so up (build + serve + run) is the workflow; a persisted store is a later ADR, and the CLI stays a client regardless.

Consequences

  • The CLI cannot drift from the API: both hit the same code path.
  • The generated app is one binary doing two thin jobs (serve, call).
  • The --url flag makes the CLI testable against any instance.

Proof

CI job example, step smoke test (proof of ADR 0005): starts the generated serve process, then drives create/list/custom-op through the generated CLI against that process, asserting the returned total.

ADR 0006: A proxy is a byte pass-through

  • Status: accepted
  • Date: 2026-09-23

Context

Cloning an upstream API with a generator tempts you to re-model the upstream: parse its types, re-emit its routes as first-class endpoints. That is a second system of record for someone else's API, and it breaks silently every time the upstream changes.

Decision

kind: proxy tesserae declare upstream endpoints and forward them byte-for-byte: method, path, query, headers, and body go through unchanged (hop-by-hop headers stripped), status and body come back unchanged, plus two response headers (x-proxied-by, x-cache). Middleware is a closed vocabulary (v0: cache — in-memory, TTL, GET-only by default, keyed by method+path+query). mosaic import openapi turns an existing OpenAPI 3 document into a proxy tessera so cloning an API is a command, not a project.

Consequences

  • The proxy can never misrepresent the upstream's semantics in v0.
  • Enrichment and transformation are out of scope by design; they are a later ADR with its own vocabulary.
  • The cache is visible (x-cache: HIT|MISS) and bounded by TTL.

Proof

cargo test -p mosaic-conformance → examples_match_goldens, which pins the full rendering of examples/proxy (proxy tessera with cache middleware) under examples/proxy/golden/.

ADR 0007: Deployments are data, not code

  • Status: accepted
  • Date: 2026-09-23

Context

Deployment logic written as scripts (build steps, kubectl calls, registry pushes) is the part of a codebase that rots first: it is environment-coupled, unauditable, and duplicated per target.

Decision

A deployment is a declarative DeploySpec (deploy/<name>.yaml, or the inline deploy: sugar in mosaic.yaml): name, target (local | docker | k8s), host, base_path, port, replicas, literal env, and secret names (values never live in the repo). The renderer turns each target into inert artifacts — run.sh, Dockerfile, or a deployment/service/ingress triple — under deploy/<name>/ in the output directory. Conflicting sources (inline deploy: and a deploy/ directory) are a hard error. local is a first-class target, not an afterthought. Config precedence, when a later version adds overrides: compiled default < mosaic config < deploy overrides < runtime env var.

Consequences

  • Deploying is diff-able: the artifacts are generated files, reviewed like anything else.
  • Adding a target is adding a vocabulary value plus a renderer branch.
  • Secrets enter the system by name; no value is ever authored.

Proof

cargo test -p mosaic-core → load::tests::inline_deploy_conflicts_with_deploy_dir (plus inline_deploy_alone_is_accepted and missing_deploy_dir_yields_implicit_local, which pin the sugar and the implicit local deployment).

ADR 0008: Pages, documentation, and i18n are projections of one DSL

  • Status: accepted
  • Date: 2026-09-26

Context

Mosaic already renders five "lenses" from one model:

  • app — Rust CQRS server + gRPC + CLI + TUI
  • web — a Leptos (SSR / wasm-CSR) web frontend with live aggregate state
  • site — a static Zola API-reference site (commands / aggregates / workflows / projections)
  • doc — authored doc → zola pages + self-contained handbook (HTML) + presentation (one slide per chapter) + GitHub Pages
  • deploy — docker / k8s / helm / aws / vps

Two gaps block treating web pages and documentation as first-class:

  1. No i18n anywhere. The only "language" is the DDD domain.language glossary (a term→definition map). Site chrome, doc content, and app strings are all single-language (<html lang="en">); there is no locale set, no per-locale content, and no Accept-Language handling.
  2. doc is chapters-in-a-docs-site. No custom layouts, no data-driven sections, one slide per whole chapter, and the generated site is a technical reference, not an authored/marketing page.

There is also a structural fork that must be stated plainly. Two input models exist today, each with its own render pipeline:

Tessera DSL YAML spec
Files *.tessera (text) mosaic.yaml + tessera.yaml + deploy/*.yaml + src/*.rs
Behavior declarative, in the DSL (commands emit events, aggregates hold state, workflows, …) hand-written Rust in src/*.rs; YAML declares data + op routing
Kinds full model (app, part, aggregate, doc, policy/IAM, vectordb, migrate, …) Service / Proxy only
Renders all five lenses (app + web + site + doc + deploy) service / proxy + openapi + deploy
Conformance goldens here (tessera-model fixtures orders-app, full-stack + the spec goldens below) here (examples/orders, examples/proxy, incl. the AWS deploy goldens)

mosaic build picks the pipeline by file: a directory containing mosaic.yaml goes through the spec pipeline (ResolvedPlan → service/proxy); a directory containing tessera.yaml goes through the tessera-model pipeline (TesseraWorkspace → app/web/site/doc). The tessera model is the rich, fully-declarative model and the home of all recent feature work; the YAML spec is the thinner service/proxy model.

The goal: author web pages, documentation, and all user-facing strings once, in the tessera DSL, and let mosaic project them to every medium (static pages, handbooks, slide decks, localized variants, dynamic/runtime docs) — reusing the existing IAM model for audience-based access.

Decision

One source of truth, many projections, one i18n model.

0. YAML is the canonical DSL format

The model is written in YAML (a single tessera.yaml). The custom .tessera text syntax is removed from the product: no production path loads it (the CLI requires tessera.yaml), and every example + conformance fixture is authored in YAML. The text parser survives only as an internal test-fixture builder (parse_file) and for the parse_event_trigger helper the planner uses. Rationale: serde_yaml is already a workspace dependency; YAML is uniform, toolable, and content-shaped — which is where mosaic is heading (pages, doc projections, i18n, content collections). The in-memory model (TesseraWorkspace) is format-agnostic: the YAML front-end feeds the same AST the text DSL produced, so all feature work is shared — "convert to YAML" was a front-end swap, not a model redo. The thin YAML spec (service/proxy) becomes a kind in the same unified model.

1. The tessera DSL is the single source of truth

All new capability lands in the tessera-DSL model (TesseraWorkspace + the app/web/site/doc renderers). The YAML spec model is out of scope for these features.

2. Unified i18n (a value type, not a feature)

  • app { locale { default: "de" languages: ["de","en","ru"] } } declares the one locale set for the whole app.
  • Any string field becomes Localizable: a bare scalar (the default locale) or a per-locale map (title: { de: "…", en: "…" }). Long markdown and large UI sets use per-locale keys or imported catalogs (i18n { de: "i18n/de.toml" }).
  • One resolution drives everything: site chrome, doc content, page UI strings, and the generated app (a t(key, locale) helper + Accept-Language / ?lang= on API responses). The same model serves apps and docs — that is the "unified DSL."

3. First-class pages

A site top-level declaration = theme + nav + pages. A page is a layout + typed sections. Sections are authored or data-driven:

  • source: <Aggregate> binds a section to tessera data (the content collections, e.g. data/*.toml events/stations, become typed, validated aggregates that pages iterate).
  • page "/stationen/{slug}" { source: Station … } is a per-instance page.
  • A small set of section blocks (hero, cards, schedule, gallery, video, prose, cta, faq, table, embed) is "page layout in the DSL."

4. Projections (the unifying mechanism)

doc, page, and the auto reference site are all projections. A projection is source → target with a format and a context:

  • Targets: site (static pages), handbook (HTML manual), presentation (reveal / pptx / google / html), api (OpenAPI).
  • context: { audience, roles, locale, env, feature } filters the projection. roles reuses the IAM model: can(role, …) decides which commands, endpoints, and sections a given audience sees. This one mechanism covers static, dynamic, and context-based documentation.

5. Presentation backends

format: on the presentation projection:

  • html (default) — current self-contained deck.
  • reveal — a reveal.js deck: per-section slides, sub-slides from ##, speaker notes, themes.
  • pptx — a real PowerPoint file emitted at build time as Office Open XML (zip + XML; dependency-light, honors the free/own-base policy).
  • google — a .pptx that imports cleanly into Google Slides + an optional opt-in Drive-API upload step (manual, like the AWS deploy).

6. Site backend: render it, or orchestrate Zola

Per-site backend: (the DSL is identical either way):

  • html (default) — mosaic renders static HTML directly (as handbook / presentation already do). Zero external deps, deterministic.
  • zola — mosaic emits a full Zola project (config with [languages]/i18n, content/, templates/, static/) and the build wiring (a zola build
  • deploy step; generalizes the existing GitHub-Pages / CI templates). This is the "mosaic supports additional dependencies and drives them" path.

7. Dynamic + context-based app docs

The web (Leptos SSR) lens serves a /docs surface at runtime, where context is live: caller role (via the existing identity stack), locale (Accept-Language), env, current feature flags, and live projection state. Static (build-time, for the site) and dynamic (runtime, on the server) share one projection definition; only the context source differs.

Consequences

Roadmap (smallest useful increments, in order):

Phase Deliverable
P0 locale decl + Localizable value type + parse + resolution (model only)
P1 i18n in the static site: [languages], i18n/*.toml, per-language content, switcher, <html lang>
P2 site decl + theme + page + section blocks + data-driven sections (backend html)
P3 projects projections + context (audience/roles/locale/env/feature) + IAM reuse
P4 presentation backends: reveal, pptx, google
P5 dynamic /docs on the web lens (runtime role/locale/env/flags/live state)
P6 backend: zola — emit full Zola project + build/deploy wiring
P7 mosaic import <zola-site> / content migrate (port existing Zola sites)

Progress: - P0 (done): locale decl + Localizable value type + parse + resolution; the YAML front-end loads app { name, about, title (scalar or per-locale map), locale } into the same Vec<TopDecl> and renders all lenses. - P1 (done): i18n in the static site — [languages], per-locale content, nav language switcher, <html lang>, localized doc chapters. - P2 (done): first-class page decls with typed sections (hero/prose/ html/cards/collection/image) rendered to Zola pages under /pages/<slug>/ + nav links. - P3 (context) (done): ProjectionContext { audience, roles, locale, env, feature } gates page sections, pages, doc chapters, and docs. roles reuses the IAM model via ws.can_see(role, ctx) (role ladder + bypass + explicit deny — the same data feeding the generated can()). The static projections render for app.audience (default: app.default_role); gated content below the audience is hidden from the site, handbook, and presentation. Authored in YAML (context: on pages/sections/docs/chapters, app.audience). - P4 (done): presentation backends via doc { presentation_format } (YAML: presentation_format:). html (default) = the self-contained deck; reveal = a reveal.js deck (title slide, one vertical group per chapter, ## sub-slides, <!-- … --> speaker notes); pptx / google = a real PowerPoint (Office Open XML) file emitted as a ZIP of XML parts via a dependency-light in-house ZIP writer (STORE + CRC-32). Binary outputs are carried in a parallel byte-map through the build/write path. - P6 (done): per-site site_backend toggle (YAML app.site_backend, text-DSL site_backend:) — zola (default) emits the full Zola project plus build wiring (site/build.sh: installs a pinned zola release if absent, then zola build); html renders the site as self-contained static HTML directly (no Zola, no build step) — per-locale home, first-class pages (title as <h1> + sections), the model collections (commands/aggregates/workflows/projections/reactors, list + detail), and the docs (TOC + chapters), all sharing one stylesheet + an embedded nav with the language switcher. The doc lens emits its "pages" projection as static HTML under html (skipping Zola content) while the handbook/presentation are unchanged. Verified on the real fecg-bs content: zola → 63 files and a clean zola build (18 pages); html → 19 self-contained files with correct localized titles/bodies (Contact / О нас) and no Zola artifacts. - P5 (done): the web (Leptos SSR) lens serves a dynamic /docs surface at runtime. The same doc projection (chapters + audience context) is emitted as data (web/src/docs.rs: per-locale titles + build-time-rendered HTML + a min role-level per chapter); the /docs handler resolves the caller's LIVE role (identity-stack-stamped x-authz-role, dev x-role, else the app's default audience) and locale (?lang=, else Accept-Language, else default) and filters chapters by the role ladder. Static (site) and dynamic (web) share the definition — only the context source differs. No docs → no module or route. Verified: generated web app compiles (cargo check); at runtime a viewer sees only the open chapter, an admin sees the gated chapter too, and ?lang=de / Accept-Language: de resolve the German titles. - P7 (done): mosaic import zola <site> ports an existing Zola content site into a tessera.yaml (app + locale set + localized pages). Reads config.toml (title per [languages.<code>], default_language) and the flat i18n content model (<name>.md = default, <name>.<lang>.md), and emits per-locale title/prose md as Localizable. To make the import faithful, Section::Prose.md is now Localizable (per-locale bodies). Verified end-to-end on the real fecg-bs site: import → mosaic build → zola build renders the de/en/ru pages with correct localized titles + bodies. Custom Zola templates/layouts are not captured (best-effort prose). - Import hardening + first production integrations (done): the importer now also (a) captures the home prose body — content/_index*.md bodies land in a new localizable app.home field — and (b) flattens one level of nested content sections (content/stationen/*.md → pages named stationen-<name>; a section's _index.md → the bare <dir> page), so no content is silently dropped. This drove two Zola 0.20 site-lens fixes: the home is the root section index (content/_index.md / _index.<lang>.md, rendered from section.title/section.content — a section page has no page.* context; per-locale content/<lang>/index.md files were wrong and are gone), and the default language's pages are unprefixed (content/pages/, → /pages/...) while other locales are content/<code>/pages/ (→ /<code>/pages/...), with locale-aware nav links (lang-switched hrefs). First production integrations landed: fecg-bs and bibelgarten-braunschweig now carry a committed tessera.yaml (5 and 11 pages, de/en/ru); mosaic build → zola build renders every page + the localized home bodies. - The YAML front-end now covers the entire decl surface — every TopDecl kind: app, pages, docs, migrations, parts, aggregates (entities/ops with expression-as-string fields), queries, resources, flags, schedules, notifies, templates, requirements, identity (jwt/users/oidc), domains (language/contexts/maps/services/events), models (role model), policies (roles/permissions/guards/attributes/boundaries/rules), schemas (structs/enums/aliases + example exprs), admins, reactors, rulesets (reactive actions + callable output), provisioning (event + claim shapes), aspects (match pointcut + advice blocks), command_groups, workflows (nodes/edges; node props as text/list/ fields/block), tests (suites/scenarios + app e2e), vector_dbs, and persistence. All land as the same TopDecls the text DSL produces, so plan + renderers are shared. Verified end-to-end: a tessera.yaml exercising every kind plans and builds cleanly. - Examples to YAML (done): both reference examples are now authored as tessera.yaml (the text .tessera files are removed; mosaic build <dir> prefers tessera.yaml when present). Porting them required a few extra YAML surfaces: app parts/realms, part contributions (dotted slots) + nested decls inside parts (a part can own its own test/ domain/workflow...), aggregate projections (upsert/update/delete/ increment on-events), and the text-DSL enabled-defaults-to-true for schedules/rulesets. Two normalization rules keep the YAML AST identical to the text one: scalar node values become Text (matching the text scalar_text), and reactor trigger/invoke raw text is re-lexed and space-joined exactly like the text tokenizer (event Order.OrderCancelled → event Order . OrderCancelled). Verified byte-for-byte: mosaic build from tessera.yaml and from the original .tessera produce identical file trees for orders-app and full-stack (73 files each), and the ported full-stack app passes its app_e2e scenarios (/health + POST /api/orders/process). - Transition completed + YAML conformance goldens (done): the text front-end is no longer loaded — the CLI (build/check/test/up) requires a tessera.yaml and errors otherwise, and the legacy load_tessera_files walker is gone. The text parser (tessera::parse) is retained only as an internal test-fixture builder and for the parse_event_trigger helper the planner uses. The conformance harness (mosaic-conformance) now pins two fixture kinds, both rendered byte-for-byte to <fixture>/golden/ + a MANIFEST hash stamp: the mosaic.yaml workspaces (orders, proxy) and the tessera-model YAML fixtures (orders-app, full-stack, via tessera_yaml + tessera_plan::build + the app/site/doc/web lenses). mosaic conformance --update regenerates; cargo test (and CI) fails if the renderers ever change the bytes for the same YAML input. The README's model example is now authored in YAML. All P-phases (P0–P7), the full model, and the examples are done.

Integrating other systems: mosaic import … then use mosaic features

mosaic import is the canonical on-ramp for bringing an external system into mosaic. The pattern is always the same three steps:

  1. Import — mosaic import <source> converts the foreign artifact into a plain tessera.yaml. Sources today: - mosaic import zola <site> — a Zola content site (config.toml + content/): app + locale set + home prose (app.home) + all content pages (flat i18n model; one level of nested sections flattened to <section>-<name> pages). - mosaic import openapi <spec> — an OpenAPI 3 document: a kind: proxy tessera (the API surface becomes mosaic aggregates/ops). The output is marked GENERATED by mosaic import … — review before building: import is best-effort (prose/markdown bodies are captured; custom templates, layouts and template-only data are not), so a human reviews the file before committing.
  2. Commit the model — the tessera.yaml lives at the root of the source repo, next to (or replacing) the foreign artifact. From this point the content is just a tessera model — there is no "imported" marker and no reduced capability.
  3. Use mosaic features — everything a natively-authored model gets works on an imported one: mosaic build (app/site/doc/web lenses), the zola and html site backends, i18n (the imported locale set drives per-locale content + nav + doc chapters), doc handbooks + presentations, audience contexts/IAM, flags, mosaic deploy, the conformance goldens, and of course editing the model by hand (adding aggregates, pages, docs…) to grow a static site into a full app.

First production integrations (2026-09-27): eugeis/fecg-bs and eugeis/bibelgarten-braunschweig (Zola church/content sites, de/en/ru). Both repos now carry a committed tessera.yaml; mosaic build → zola build renders every page with the localized titles, bodies and home prose, so the hand-rolled Zola projects can be retired in favour of the mosaic-generated site.

A site lens renders the chrome to fit the model: a content-only site (no aggregates/commands/workflows) gets no empty technical sections — the nav and the home list only collections that exist, and first-class pages appear in the nav/home under their localized title (branching on lang), never the bare page name. A technical app keeps its collection index.

A site can also carry its original design as a committed theme. app.theme: <name> points at a theme/<name>/ directory (a normal Zola theme: theme.toml, templates/, static/). mosaic build copies it to themes/<name>/, copies the repo's data/ (for load_data) and merges the repo's legacy config.toml into the generated one (the [translations] / [extra] tables the theme templates rely on; generated keys win, the legacy base_url is kept). The generated templates/base.html then only extends the theme's base.html and overrides the two model-driven chrome blocks — nav (collection + page links) and lang_switch (locale links) — while the theme keeps the head, header, footer, scripts and all styling. There is no generated site-level index.html: Zola falls back to the theme's own home template, which owns the whole home design and reads the imported home prose via section.content. Pages keep their original Zola template: (imported from the front matter, defaulting to page.html) and any [extra] front matter (icons, accent colors, media ids) so theme templates can read page.extra.*. page.in_nav: false keeps a page out of the main nav (curated navs, e.g. a station list that links from the home). This is how eugeis/fecg-bs and eugeis/bibelgarten-braunschweig keep their hand-rolled look (carousel, station cards, lightbox, dark mode) while the content, navigation and i18n become model-driven.

Deployment (Netcup git integration): the host has no toolchain and an old glibc (Debian 11 / glibc 2.31), so each site repo commits a static musl bin/mosaic (and bin/zola) and its deploy.sh runs bin/mosaic build . → bin/zola build → httpdocs/ as the Netcup post-sync action. The static-musl build needs the bin/musl-gcc-static linker wrapper: rustc's musl target links the dynamic-loader CRT (rcrt1.o), whose startup dereferences the weak _DYNAMIC symbol — 0 in a static binary — and segfaults at startup; the wrapper swaps in the static trampoline (crt1.o).

Decisions (format + model)

  1. One model, YAML canonical — decided. YAML is the canonical DSL format; the tessera-DSL model is the single model; the text .tessera syntax is deprecated (read-only during transition, then removed); service/proxy becomes a kind. Sequencing A: build the new features YAML-native on the active slice (app/locale/i18n/doc/site) now; migrate the rest of the model + examples to YAML as a follow-up; add a conformance harness for the YAML model so site/doc/i18n output is pinned by golden output.
  2. Defaults (accepted): site backend html (zola opt-in); i18n authoring inline locale-maps and imported catalogs; presentation order reveal + pptx first, google = pptx-that-imports; P3 context starts audience+locale+env (build-time), runtime added in P5.

ADR 0009: Data movement is a `mirror` of declared facts, not an embedded engine

  • Status: accepted
  • Date: 2026-09-28

Context

A tessera app increasingly needs to move data between systems: replicate a Postgres table into an analytics warehouse, stream change-data into a second database, or derive external rows into the app's own CQRS model. The obvious implementation is to embed a movement engine (a Temporal-driven orchestrator like PeerDB, a Kafka/Debezium pipeline, or a CDC client library with its own state machine) alongside the generated app.

That conflicts with the laws:

  • Law 2 (Zero Tax) / ADR 0001 (closed vocabulary): an embedded engine is an escape hatch — an open, Turing-complete data pipeline with its own configuration, not a declared fact the plan owns.
  • Law 3 (One Model) / ADR 0002 (One Resolve): if a runtime engine decides routes, checkpoints, and schemas, the "model" of the data flow lives in two places (the DSL and the engine), and they can drift.
  • Law 4 (Thin Renderers) / ADR 0004: the flow's facts (which table, which keys, which columns map to which, the mode, the batch size) must be resolved exactly once, in the plan, and the renderer only prints.

The question is how to get PeerDB-grade sync (backfill + CDC, schema mapping, resumable checkpoints) while keeping the decision in the plan and the runtime a deterministic projection of it.

Decision

Two declarations, one resolved plan, a verbatim runtime.

Data movement is two new top-level tessera declarations — peer and mirror — that resolve (fail-closed) in mosaic-core into PeerPlan and MirrorPlan. The renderer emits a small, deterministic sync engine (app/src/sync.rs) plus a verbatim runtime core (app/src/syncrt.rs, copied byte-for-byte from crates/mosaic-sync). No engine is embedded; the app is the mover.

1. peer — a named, kinded connection (the secret is an env-var name)

peers:
  - name: legacy
    kind: postgres          # postgres | mysql | clickhouse | file
    conn: MOAIC_PEER_LEGACY # an ENV VAR NAME, never a literal

A peer is a fact: name, kind, and conn. conn is the name of an environment variable that holds the connection string; the value is read at runtime. No credential ever lands in the model or the generated source (the same law as resource.auth).

2. mirror — source → target with per-table mode

mirrors:
  - name: orders_to_warehouse
    source: legacy          # a peer
    target: warehouse       # a peer, or the literal `model`
    batch_rows: 500
    tables:
      - name: orders
        keys: [id]
        mode: cdc           # cdc | snapshot | poll
        where: "status != 'draft'"
        map: { customer: customer_id }   # target_col: source_col (rename)
      - name: order_items
        keys: [id]
        mode: snapshot

A mirror names a source peer and a target (a peer or the reserved model), a batch_rows (default 1000), an optional poll interval (for mode: poll), and a set of tables. Each table declares its keys (the primary key, required for cdc), its mode, an optional where row filter, an optional map (a rename: target_col: source_col; unmapped columns pass through), and — for target: model — into / into_delete (the CQRS command the insert/update and the delete dispatch).

3. Modes are resolved to a concrete strategy in the plan

  • cdc (Postgres source): a keyset backfill to a cursor, then a pgoutput logical-replication stream (via pgwire-replication, the one external crate reused) that decodes R/I/U/D messages, normalizes each change, and checkpoints the commit LSN.
  • snapshot: a one-shot keyset backfill (no stream).
  • poll (any source): a watermark/keyset poll every poll interval, for sources without a logical decoder (MySQL, or a chosen Postgres table).

snapshot/poll targets accept any peer; cdc requires a Postgres source and a Postgres/model target. The plan rejects a cdc table without keys, an unknown source/target, a model target without into, and a non-Postgres source with mode: cdc — all fail-closed, like every other resolver (ADR 0002).

4. Sinks are projections of the resolved change

Each decoded change is a normalized, named row (Op + RowMap + key RowMap). The sink is a pure function of that row + the MirrorPlan:

  • Postgres sink — writes to a _mosaic_raw_<mirror> staging table, then a MERGE normalizes it into the destination table (TOAST-safe: an unchanged cell is absent from the payload and the MERGE keeps the stored value via CASE WHEN _row ? col). Deletes set _mosaic_deleted = TRUE (soft delete).
  • ClickHouse sink — INSERT into a ReplacingMergeTree (the _mosaic_synced_at column is the version).
  • File sink — one JSON object per line (append-only archive).
  • Model sink — the mosaic differentiator: the row is dispatched as a command onto the app's own aggregate (into / into_delete), so an external table becomes state the app already serves, guards, and projects. This is the "derive" that no external engine offers.

5. One external crate, reused byte-exactly

pgwire-replication (0.4, Apache-2.0/MIT, pure Rust, no libpq) is the single reused dependency for the cdc path. The rest of the runtime is in-house (crates/mosaic-sync), emitted verbatim into the app so the golden output pins it byte-for-byte and there is no version drift between the crate and the generated app. Candidates evaluated and rejected: deltaforge (a standalone service, heavy rdkafka/deno_core tree — not an embeddable crate), d-engine (not CDC — an embeddable Raft KV store / etcd alternative), dbmazz (ELv2-licensed, so its code cannot be reused; its Sink trait and LSN checkpoint concepts are adopted here instead).

6. State is a file the app owns

The sync engine persists its cursors, LSNs, watermarks, and per-mirror last_error to a JSON file (MOAIC_SYNC_STATE, default sync-state.json) and is resumable: on restart it resumes the backfill cursor / replication LSN rather than re-reading. GET /api/sync reports per-mirror state; POST /api/sync/{mirror}/once triggers a manual pass.

Consequences

  • The data flow is a declared fact: peer/mirror resolve once in the plan (fail-closed), and the renderer only prints the engine + the verbatim runtime. No second source of truth, no embedded orchestrator (Laws 2–4).
  • Deriving external rows into the model (target: model) is a first-class capability no external engine has: a legacy table becomes CQRS state with the app's guards, projections, REST, web, and site — for the price of a mirror.
  • The runtime is a projection, so it is golden-pinned: a byte change to crates/mosaic-sync or the sync.rs codegen fails the conformance suite until the golden is re-pinned (Law 6).
  • Secrets stay env-var names (Law 2 / the resource.auth precedent); a credential can never be committed.
  • Cost: one new crate (mosaic-sync) and two new declarations. The cdc path adds pgwire-replication + tokio-postgres to the generated app's Cargo.toml only when a Postgres peer / cdc mode is declared (Zero Tax). MySQL adds mysql_async only when a MySQL peer is declared; ClickHouse uses HTTP/JSON (no new crate); the file sink uses std.

Proof

  • mosaic-conformance::tests::examples_match_goldens pins examples/data-sync/golden/ byte-for-byte (the full sync.rs + verbatim syncrt.rs + the gated Cargo.toml). A change to the codegen or the runtime that is not re-pinned fails CI.
  • mosaic-sync unit tests (17) prove the pgoutput decoder byte-exactly: decode_real_capture_frames, decode_real_capture_delete_frame, and decode_real_capture_update_frame decode real frames captured from a live PostgreSQL 15 stream (including the full-width delete K tuple and the keyless update), plus chunk-spanning truncation and the to_changes mapping.
  • mosaic-core plan tests (data_mesh_tests, 15) prove the resolver is fail-closed: unknown peer/kind/mode, cdc without keys, a model target without into, and a non-Postgres source with mode: cdc all reject.
  • The dependency gate is proven by the golden Cargo.toml (deps appear only for the declared kinds) and by render emitting them conditionally.

ADR 0010: One app, many targets — deploy specs are the master

  • Status: accepted
  • Date: 2026-09-28
  • Supersedes: nothing (complements ADR 0007, which covers the mosaic-spec deploy block; this ADR covers the tessera deploy { name … } specs)

Context

A product is not deployed to one place. The same mosaic app ships to our customers' own Kubernetes (bare k8s), to managed clusters on the clouds their workloads already live on (aws-eks, aws-ecs, gcp-gke, azure-aks, alibaba-ack), to plain docker and local for dev. And the datastores are the same story in miniature: sometimes a managed service (RDS, Cloud SQL, ApsaraDB RDS, Azure Database for PostgreSQL), sometimes a self-hosted engine in the cluster (postgres, mysql, redis, clickhouse, qdrant, weaviate, opensearch, milvus).

The old shape — one deploy block, one target, terraform that also rendered the k8s workload inline — could not express "same app, prod on GCP with Cloud SQL, staging on our own k3s with an in-cluster postgres".

Decision

An app declares any number of deployment specs:

deploys:
  - name: prod-gcp          # -> deploy/prod-gcp/
    target: gcp-gke         # local | docker | k8s | aws-eks | aws-ecs |
    region: europe-west1    #   gcp-gke | azure-aks | alibaba-ack (closed vocab)
    replicas: 2
    db:
      name: shop
      engine: postgres      # postgres|mysql|mariadb|redis|dynamodb|clickhouse
      managed: true         # default: cloud targets managed, own-infra not
    vector:
      name: shop-vectors
      engine: qdrant        # opensearch|qdrant|weaviate|milvus|pinecone
  - name: staging-k8s
    target: k8s             # own infra: no terraform, chart only
    db: { name: shop, engine: postgres }   # in-cluster via the chart

Rules, all fail-closed in the plan:

  • name is required and unique (it is the deploy/<name>/ directory).
  • target is a closed vocabulary; unknown targets are an error.
  • Each spec resolves a cloud (aws|gcp|azure|alibaba|none) and a per-cloud default region.
  • Stores are closed vocabularies too. managed defaults to true on cloud targets and false on own-infra targets. dynamodb and pinecone exist only as managed services; self-hosted stores require a k8s target.
  • A managed store must have a stable first-party Terraform resource with a connection URL on that cloud (deploy-managed-unavailable, fail-closed) — otherwise the chart would reference a DSN secret Terraform cannot create:
cloud managed db engines managed vector engines
aws postgres, mysql, mariadb, redis, clickhouse, dynamodb¹ opensearch
gcp postgres, mysql, mariadb, redis, dynamodb¹ — (use managed: false)
azure postgres, mysql, redis all (map to AI Search)
alibaba postgres, mysql, mariadb, redis, clickhouse, dynamodb¹ — (use managed: false)

¹ managed document store, no DSN: the app is wired through the provider SDK, and the chart installs nothing for it (enabled: false in values).

The matrix lives in the plan (managed_store_available) and must stay in sync with tf_db_url / tf_vector_url in mosaic-render.

Each spec derives its artifacts; the app, the DSL, and the codegen are identical across specs:

  • local → deploy/<name>/run.sh; docker → deploy/<name>/Dockerfile.
  • cloud targets → deploy/<name>/infra/ Terraform: the managed cluster (EKS/GKE/AKS/ACK module), the image registry, the managed store resources (RDS/DynamoDB/ElastiCache/OpenSearch, Cloud SQL/Memorystore, Azure PostgreSQL/MySQL/Redis/AI Search, ApsaraDB RDS/Redis/ClickHouse), and — for k8s targets — the namespace + DSN secrets. Terraform never renders the workload.
  • k8s targets → deploy/<name>/Dockerfile + deploy/<name>/helm/ chart: the image is the bare binary (the chart passes serve --bind as container args); the chart is the app Deployment/Service plus the stores. mode: cluster renders the store in-cluster (postgres/mysql/ mariadb/redis/clickhouse statefulsets or deployments; qdrant/weaviate/ opensearch/milvus workloads) and materializes the DSN into a chart-owned secret; mode: external reads the DSN secret Terraform created (<app>-db / <app>-vector) or that the operator provides.

The app reads exactly two store env vars in every mode: MOSAIC_DB_URL and MOSAIC_VECTOR_URL.

Consequences

  • "Deploy to another cloud" is a spec, not a fork: same image, same DSL, same generated app; only deploy/<name>/ differs.
  • The chart is the single k8s footprint for any cluster, managed or not — what ee-helm does for the enterprise engine, generalized over engines.
  • Cloud testing can be per-cloud: render the spec, terraform apply (or reuse the managed resources), helm install with the chart values, and the app is exercised against the real managed services.
  • Terraform output is codegen, reviewed like any generated file; per-cloud drift (which engines exist where) is a renderer concern, invisible to the DSL.
  • aws-ecs has no cluster: terraform owns the whole footprint (Fargate service + ALB), store DSNs wired into the task definition.
  • In-cluster store passwords live in the chart values (test/dev-grade); production footprints should prefer managed: true or mode: external so no credential is in the values file.

Proof

  • cargo test -p mosaic-core → tessera_plan::deploy_spec_tests::* (defaults, managed inference, cloud-only engines, managed-availability guard, name rules, alibaba/azure targets).
  • cargo test -p mosaic-render → app::terraform_generation::* (per-spec infra + helm for aws-eks and gcp-gke).
  • examples/full-stack/tessera.yaml carries three specs (prod-aws, prod-gcp, staging-k8s); cargo run -p mosaic-cli -- conformance pins the rendered deploy/<name>/ trees in the golden.
  • scripts/e2e/ — the k8s e2e harness + peer fixtures; .github/workflows/e2e-*.yml run it on a GH-hosted k3d cluster, on the self-hosted on-prem runner, and (secrets-gated) against the clouds.

ADR 0011: Durable store — append per dispatch, snapshot the whole store, replay the tail

  • Status: accepted
  • Date: 2026-09-30
  • Supersedes: nothing (extends the persistence construct introduced for the durable event log)

Context

The README roadmap names two ADR-gated items: durable stores and authorization enforcement on every route. This ADR is the first.

A generated app with persistence declared already has a durable event log (JSONL or SQLite): every command dispatch appends the events it emitted (with_react), and boot replays the log to rebuild aggregate state and materialized projections. Two gaps remain:

  1. Replay cannot rebuild everything. The replay path (replay_event) only re-applies aggregate events. State that lives on the Store but is not an aggregate fact — the instance-level IAM grants recorded by on_create auto-assigners (instance_grants) and the enterprise audit trail (audit, when the audit platform part is on) — is silently lost at every restart. An app that survives a crash with the right orders but the wrong grants and an empty audit log is not durable.
  2. Boot cost is linear in log size. Every start replays the whole log, including the part a snapshot would make redundant.

The Store is already serde::Serialize (the state structs derive both Serialize and Deserialize), so persisting it whole is a serialization question, not a modeling question.

Decision

Four rules, all in the generated app (the engine crates stay pure):

  1. One persistence chokepoint. server::react_and_persist(store, since) runs the event reactors for the new events, appends them to the log, and snapshots on the cadence. Every path that appends events to the live store goes through it: CQRS command dispatch (with_react), workflow runs (/api/workflows/{slug}/run streamed and plain, workflow-source endpoints, triggers, schedules, MCP tools) and the HITL resume. This closes the gap that made the log durable only for direct command dispatches — workflow-emitted events were previously persisted only at bootstrap/shutdown, so a kill between them lost them.
  2. Append per dispatch (kept). with_react keeps appending the events it emitted to the log before returning, so the log is never behind the store by more than one in-flight dispatch.
  3. Snapshot the whole store, on a cadence. New DSL key persistence.snapshot_every: <n> (events; default 256 when persistence is declared, 0 disables). A snapshot is the complete Store serialized as JSON plus an events watermark (store.events.len() at snapshot time): - jsonl backend: sidecar file <data_file>.snap.json - sqlite backend: kv(key TEXT PRIMARY KEY, value TEXT) table, key snapshot — same file as the log, so log and snapshot advance together Writes happen (a) in with_react once n events have accumulated since the last snapshot, and (b) on the graceful-shutdown path after the final append. The store carries a #[serde(skip)] snapshot_at: usize watermark so the cadence survives restarts and a failed write retries on the next dispatch. Snapshot writes are best-effort: a failure warns, never fails the request (the same fail-soft contract as the log append).
  4. Boot: snapshot first, replay the tail. Boot reads the log, then tries the snapshot: if it parses and its watermark is <= the log length, the store is adopted from it and only log[watermark..] is replayed through the existing replay_event. No snapshot, an unparseable one, or a watermark past the end of the log (log truncated out from under us) all degrade to a full replay with one warning — the pre-ADR behavior, so a broken snapshot can never lose more than a fresh replay would.

The ordering invariant that makes this safe: the snapshot is only ever written after the log append for the same events (in react_and_persist, reactors + append first, snapshot last; same at shutdown). So a crash can leave the log ahead of the snapshot (tail replay repairs the replayable state) or both behind the last in-flight dispatch (at-most-once per dispatch — the existing contract, unchanged), but never the snapshot ahead of the log. The known window: side effects a replay cannot rebuild (an on_create instance grant) are only restored when a snapshot post-dates the event that caused them — snapshot_every bounds that window.

Consequences

  • Restart preserves instance_grants and the audit trail, not just aggregates; workflow-emitted events are persisted per run, not only at shutdown; boot skips the replayed prefix.
  • One more file (jsonl) or one more table (sqlite); no new dependency (serde is already in the generated app).
  • snapshot_every: 0 reproduces today's replay-only behavior exactly, so existing apps that add persistence are unaffected unless they ask for the cadence.
  • The proof (ADR gate) is an e2e pair in the full-stack example: test 1 runs the full scenario set (workflow runs + CQRS commands, ending in a ticket creation whose on_create auto-assigner records an instance grant) and is SIGKILLed by the harness (no graceful shutdown); test 2 boots the same data file on a new port and asserts the aggregate state and the instance grant (which replay alone cannot rebuild) survived — via the mid-run snapshots written at snapshot_every: 1.

ADR 0012: Authorization enforcement on every route

  • Status: accepted
  • Date: 2026-09-30
  • Supersedes: nothing (extends the identity/IAM surface; the per-endpoint auth/min_role/authz/where keys keep working and keep winning)

Context

The README roadmap's second item. Today an app with identity declared has authentication (JWT/OIDC verification, which stamps the verified role/user onto the request) and per-endpoint opt-in authorization: a contributed REST endpoint is checked only when the author wrote auth: required (authn), min_role (RBAC level), authz (can(role, action, resource)), or where (ABAC). Everything else is open.

The gap, found by auditing the generated router: in a full-stack app the derived and platform routes carry no authorization decision at all — GET /api/{aggregate}[/{id}], projections, workflow run/preview/ runs/pauses/resume, the generic callable-ruleset routes, the whole knowledge/vectordb surface (incl. ingest/delete mutations), /api/flags (incl. set), /api/audit, /api/events (SSE leaks event payloads), /api/events/list, /metrics, /api/iam, /api/sync (incl. once), the MCP server, and voice. An app that declares users and roles still answers a bare GET /api/orders with its data. That is not "authorization enforcement on every route."

The decision vocabulary already exists in the generated app: the role ladder, ROLE_GRANTS/ROLE_DENIES (explicit deny wins over bypass), can(role, action, resource_type), the dynamic resource types (aggregates + domains), and the verified-identity headers. What is missing is a decision for every route.

Decision

Enforcement is on when identity is declared (the app has authentication); an app without identity is byte-identical (an internal tool with no identity model has no principal to authorize — no middleware, no table is emitted at all).

One middleware (authz_middleware, via from_fn_with_state) runs on every route of both lenses (app and web) and makes the decision, in order:

  1. Public routes pass. route_policy(path, method) returns None for: a fixed infra set (/health, /openapi.json, /api/auth/login, the /voice widget page), contributed endpoints declared auth: none (no table row at all), and two documented families that keep their own model: triggers (their x-mosaic-token authn is their contract; machine-to-machine) and the gateway (ADR 0006 byte-passthrough; the backend enforces). Plus the app's declared public_routes — a top-level public_routes: [<prefix>…] list of route prefixes kept open (exact or /-bounded prefix match; e.g. a public order-status read). Nothing else is public by default.
  2. Authn — three ways. - A verified Bearer token: a local JWT (HS256, the secret from MOAIC_JWT_SECRET or the identity's jwt.secret_env) or, when an OIDC IdP is configured, an id_token (RS256). - The app's api key (x-api-key or Bearer, matching the api_key_env secret): a machine credential that authorizes as DEFAULT_ROLE. - Otherwise, on a route whose policy marks it anonymous-allowed (auth: optional endpoints): the caller is admitted and authorizes as guest. - No credential and no anonymous allowance → 401 (audited when the audit part is on).
  3. Authz by a generated route table. Codegen emits route_policy(path, method) -> Option<(&str, &str, bool)> — (action, resource type, anonymous allowed) — one static entry per route, built from the same inventory as the router (route_chunks): every .route(...) in the generated chain is paired with its decision in one struct, so a new route family without a policy cannot be rendered. The default mapping (the app resource type is a synthetic entry covering app-level surfaces; scoped roles must include it, *-scoped roles cover it automatically):
route family decision
aggregate list/get, projections, ruleset-evaluate sources view on the aggregate's resource type
CQRS command-delegate sources create on the aggregate's resource type
workflow-source endpoints, workflow run / resume, voice session upgrade update on app
command_group sources create on app
callable rulesets (POST /api/rulesets/{kebab}), ruleset sources view on app
workflow preview / runs / runs/{id} / pauses, /api/resources/{name}, knowledge read ops (search/stats/docs/graph/citations/profile/books/report/index), /api/assistant/chat, the MCP server, /api/flags (get), /api/events, /api/events/list, /metrics, /api/iam, /api/sync (status), /api/auth/me view on app
knowledge write ops (ingest/reindex/cite/doc delete) update on app
/api/sync/{mirror}/once, /api/flags/{name} set (PUT/POST) administrate on app
/api/audit view on audit
any other contributed-endpoint source family administrate on app — fail closed

auth: optional contributes only the third element (anonymous allowed) — the (action, resource) decision still applies to the admitted guest. The per-endpoint keys keep their meaning and win over the table: min_role/authz/where on a contributed endpoint still run in the handler (ABAC needs the body); auth: required remains authn-only and the table still applies on top. 4. Stamp the effective principal. After the decision, the middleware sets x-authz-role / x-authz-user / x-authz-verified: 1, replacing any client-declared x-role — handler-level checks (and auth_me) see the verified principal, closing the role-spoofing hole for api-key and anonymous callers. 5. Fail closed. Unknown route pattern in the table (should not happen — it is generated from the same inventory), an unknown role (no grant row), or a JS error in a where condition: deny. Denied decisions go to the audit trail (authz.deny) when the audit platform part is on.

The table is generated, not configured: the enforcement completeness check is structural (one RouteChunk per route, policy attached), and a codegen test asserts the emitted table covers the router's routes.

Testing the matrix — the e2e harness (mosaic test) gains a per-scenario token: {user, role} field: it signs a local JWT (secret read from the same e2e env: block the server gets, via the plan's jwt_secret_env) and sends it as authorization: Bearer …, so scenarios can assert the full 401/403/200 matrix against the generated middleware.

Consequences

  • An identity app denies by default: bare requests get 401, wrong-role requests 403, and every route — derived, contributed, or platform — carries a decision. Public surface is declared, never implicit.
  • public_routes is the only new DSL key (top-level path-prefix list); the e2e token field is test tooling. The rest reuses identity, can, and the endpoint keys.
  • Existing apps without identity are byte-identical (no middleware, no table — verified against the pinned goldens); apps with identity that relied on open derived reads must either add public_routes or issue tokens — intended, and visible at the first request, not in the data model. (No shipped example declared identity before this ADR.)
  • Both lenses enforce (the web server serves the same API for SSR); the generated CLI/TUI pass a Bearer token they obtain from /api/auth/login or carry as-is — client convenience, not enforcement.
  • The proof (ADR gate) is a new small example (examples/authorized): one aggregate, identity with two local users (a viewer and an editor), one auth: none endpoint, one public_routes entry, the metrics platform part. Its e2e block (CI: the e2e-app job) asserts: /health + /api/auth/login open (login 401 on bad credentials, 200 + token on good); /metrics and /api/iam 401 without a token, 200 with a viewer token; public_routes match → 200 with no token; the CQRS delegate 401 without a token, 403 for viewer (no create on the aggregate), 200 for editor; auth: none endpoint 200 with no token.

ADR 0013: Knowledge sources can be git repositories

  • Status: accepted
  • Date: 2026-10-01
  • Supersedes: nothing (extends the vector_dbs[].sources construct from the knowledge engine, K1; the git CLI seam from K2)

Context

The knowledge engine (K1–K7) indexes local paths: every vector_dbs[].sources entry is a file or directory walked at boot. The use cases in sync/knowledge.md are about legacy systems — code bases and repositories that agents must study to re-implement or to quote from. A declared source that must be cloned by hand before the app can see it breaks the "one tessera.yaml describes the app" contract: the app's knowledge is no longer a function of its declaration + its working directory.

Two concrete gaps:

  1. No way to declare a remote (or bundled) repository as a knowledge source. sources[].path is a local path; a git URL is not a path.
  2. No hermetic way to prove it. CI runs offline-ish and must not depend on a network clone of an arbitrary repo; a test fixture must be a committed file, not a nested .git directory (which git cannot track).

Git itself is already an accepted external seam: K2 derives git facts (branch / commits / last commit / contributors) through the git CLI, fail-soft, and the engine crate stays environment-pure (Z7). Cloning through the same seam is consistent, not a new dependency class.

Decision

One DSL addition, one engine addition, no new runtime dependency:

  1. sources[].git — a source can be a git repository. ```yaml vector_dbs:

    • name: kb sources:
      • path: legacyrepo # optional label hint when git is set git: { url: "knowledge/legacyrepo.bundle", ref: "main" } ```
    • git.url is a git-clone URL: https://…, ssh://…, a local path, or a git bundle file (a single committed file that git clone accepts — this is what makes a repo fixture committable and CI-hermetic).
    • git.ref (optional) = branch/tag/rev passed as --branch.
    • When git is set, path is not required; the stable per-source label (used for entry ids and stats) is the last path segment of the URL with a trailing .git / .bundle stripped, slugified (deterministic, independent of any cache location).
    • Plan validation: git present ⇒ url non-empty; git absent ⇒ path non-empty (the previous rule).
  2. The engine clones through the git CLI seam, into a hidden cache. git::prepare_git_source(url, ref, base): - cache dir = <base>/.gitcache/<label> (a dot-dir: every source walk skips dot-dirs, so a cache can never be ingested by another source); - if the cache dir is already a git work tree it is reused as-is (deterministic boot, offline-friendly; update = delete the cache dir and reindex — documented);

    • otherwise git clone <url> <cache>, adding --depth 1 for non-local URLs (a local path / bundle is cloned in full — --depth is rejected for local clones by git itself). The clone runs with -c protocol.file.allow=always: the seam must clone local paths and bundle files from a non-interactive server process (no effect on remote URLs).
    • Post-clone verification. Some git versions exit 0 for a bundle clone whose HEAD ref the bundle does not contain ("remote HEAD refers to nonexistent ref" — observed on runners whose init.defaultBranch mismatches the bundled branch), leaving an empty checkout. The clone is therefore verified (rev-parse HEAD must resolve): when no ref was pinned, the first branch (local, else remote-tracking — sorted, so the recovery is deterministic) is checked out; with a pinned ref an empty checkout is an Err. A failed clone or an unrecoverable checkout removes the cache (never a half-clone).
    • any failure (git missing, bad URL, network) is an Err for that one source: the generated app warns and continues with the remaining sources (the existing fail-soft posture of ingest_source callers).
    • strategy.git: false turns the entire git layer off, including git sources (fail-closed for that source with a named error): "with and without git" stays a flag, not a fork.
    • entry ids for a git source are <label>/<rel-in-checkout> slugs (e.g. legacyrepo-billing-c), so a repo ingested from a URL and a local tree of the same repo get distinguishable, stable ids.
    • git facts (K2) are derived from the checkout, so a cloned source reports its branch / commits / last commit like any work-tree source.
  3. The vendored engine gains no dependency and no module. The clone lives in the existing git.rs (already vendored); Source gains git_url / git_ref. The generated knowledge_sources() emits the two fields. No new REST route: git sources are declared sources, ingested at boot and on reindex like every other source — the existing /stats, /search, /profile, MCP tools all see them without change.

Proof

  • Engine unit tests (hermetic, tempfile + the git CLI, skipped when git is absent): clone from a locally created repo, reuse of an existing cache, error on a bad URL; ingest_source on a git source produces entries with the <label>/… ids.
  • Committed fixture examples/full-stack/knowledge/legacyrepo.bundle — a git bundle of a tiny legacy repo (committed as a regular file; CI-hermetic, no network). The full-stack kb declares it as a git source; e2e asserts the boot stats list it, git facts are present, and its content is hybrid-searchable (the re-implementation use case: a repo the app has never seen on disk becomes quotable, citeable and profiled).
  • Conformance goldens re-pinned (the generated knowledge_sources() gains the git fields); workspace tests + clippy -D warnings + fmt clean.

Consequences

  • A generated app can now index repositories it does not hold on disk, from one line of DSL; the re-implementation playbook's first step ("profile the legacy repo") works against a URL or a bundle.
  • Caches persist between runs under .gitcache/; operators delete them to force a fresh clone. Shallow cloning is automatic for remote URLs (--depth 1) — large remote repos stay cheap; a ref pin makes a source reproducible.
  • What is deliberately NOT built here: remote fetching/syncing on reindex (reusing the cache is the contract), credential handling (the git CLI's own credential helpers apply, as with K2's git facts), and non-git VCS.
  • Scale note (follow-up, out of scope): the in-memory BM25 rebuild is the cost ceiling for very large corpora; the opt-in persistent index (tantivy, K4 roadmap) remains the scale path.

ADR 0014: Runtime knowledge sources — declared sources are the floor, runtime additions are the ceiling

  • Status: accepted
  • Date: 2026-10-01
  • Supersedes: nothing (extends the knowledge engine, K1; git sources, ADR 0013)

Context

Knowledge sources are today build-time facts: vector_dbs[].sources declares local paths and git repositories (ADR 0013), the CLI copies the knowledge/ directory into the output tree at build time, and the generated app ingests exactly those sources at boot. The runtime surface can only ingest inline text docs (POST /api/vectordb/{name}/ingest), delete one doc, and re-read the declared sources (reindex).

The use cases in sync/knowledge.md are operational, not build-time: an operator drops a book library (a folder of PDF/EPUB/DOCX files) or a git repository into the running app's data area and expects the app to index it — without re-rendering the app, redeploying, or editing the tessera. Conversely a source that should go away must be removable without its entries lingering in the index.

Constraints to respect:

  • ADR 0003 determinism: the declared model still renders the same bytes; runtime state lives in data files under the knowledge base, never in the rendered code.
  • ADR 0008/0010 lens split: the management surface is a web-lens feature (the /knowledge page) + REST; the static site lens gets a build-time projection of the declared sources only (a static page cannot show runtime state).
  • ADR 0012: the new routes join the route_policy table like every other route — default-deny authz applies (GET = view, mutations = update).
  • The source walk skips dot-dirs; anything the app writes under knowledge/ must live in a dot-dir so it can never be ingested by a source.

Decision

1. A persisted runtime registry under the knowledge base

<base>/knowledge/.runtime/sources.json (dot-dir ⇒ invisible to every source walk). Per collection: a map label → entry:

{
  "kb": {
    "legacyrepo2": {
      "source": { "path": "", "kind": "auto", "git_url": "knowledge/legacyrepo2.bundle", "git_ref": "main", "include": [], "exclude": [] },
      "slugs": ["legacyrepo2-inventory-c"],
      "files": 1,
      "chunks": 2,
      "added": "2026-10-01T07:00:00Z"
    }
  }
}
  • label is the stable per-source identity, computed by the engine (source_label): git sources use the ADR 0013 git_label (last URL segment, .git/.bundle stripped, slugified); path sources use the slugified last path segment. Labels are collision-checked against the declared sources of the collection — a runtime source can never shadow a declared one.
  • slugs are the entry slugs the add produced (the entry id minus the #chunk suffix). Removal deletes exactly those slugs — no prefix guessing, no damage to entries that happen to share a path prefix.
  • Boot re-ingests runtime sources (fail-soft, like declared sources), so additions survive restarts. reindex covers declared and runtime sources. Operators reset runtime state by deleting the .runtime dir.

2. REST surface (three routes, per collection)

  • GET /api/vectordb/{name}/sources — the source inventory: declared sources (recomputed live, as stats does today) plus runtime sources, each row carrying label, origin ("declared" | "runtime"), kind, files, chunks (+ git facts for git sources).
  • POST /api/vectordb/{name}/sources/add — body JSON { "path"?, "kind"?, "git": { "url", "ref"? } } or form-encoded fields path, kind, git_url, git_ref (the browser form). Fail-closed validation with named errors: unknown collection; neither path nor git url; kind outside auto|code|docs|pdf; a path source whose target does not exist; a git source while strategy.git is off; a label that collides with a declared or existing runtime source. On success: ingest (fail-closed — the caller wants an error, unlike boot's fail-soft), embed what is missing, persist the store, record the entry, return {label, files, chunks, slugs}.
  • DELETE /api/vectordb/{name}/sources/{label} — removes a runtime source (declared sources are rejected: "declared source — remove it from the tessera"), deletes its slugs, persists, answers JSON (agents / MCP).
  • POST /api/vectordb/{name}/sources/{label} — the no-JS remove form (same core, SSR-safe): answers 303 See Other to /knowledge?kb={name}&removed={label} (or …&source_error=…). A plain HTML form can only POST, so the form path is a second method on the same path — no client JS, works in full (SSR) with no hydration.

Content negotiation by body shape: the add handler answers JSON to a JSON body and, to a form-encoded body, 303 See Other to /knowledge?kb={name}&added={label} (or …&source_error=…) — so the same route serves agents (JSON) and the no-JS form (HTML navigation) without a second endpoint. All query values in a 303 Location are percent-encoded (a tiny percent_encode sits next to the percent_decode used for the form body — error messages carry spaces and quotes).

3. Surfaces

  • Web /knowledge page (all hydration modes): the per-KB sources table gains an origin column and, on runtime rows, a remove control — a plain POST form (no client JS, SSR-safe, so it works in full with no hydration and in csr/islands alike). Each KB card gains an add-source form (path or git url, ref, kind) posting to the negotiated add route. SSR mode renders a flash banner from the ?added= / ?removed= / ?source_error= query params.
  • Static site lens: a build-time knowledge sources page (/kb/, zola + html backends) projecting the declared collections and their sources (name, about, model, hybrid weights, strategy flags, source paths/git urls) — the ADR 0008 projection of the declared facts, with a nav entry. Runtime state is deliberately absent from the static lens.
  • MCP / CLI: unchanged — reindex_knowledge, knowledge_stats, repo_profile and mosaic knowledge-report now include runtime sources through the same extended sources/reports path (the registry is read by the generated app; the CLI report stays declared-only, as it runs without an app instance).

4. Engine additions (two fields, one function)

  • ingest::source_label(&Source) -> String (above).
  • IngestReport.label + SourceStats.label (serde-defaulted — existing persisted data and goldens stay valid).
  • SourceStats rows gain origin in the generated stats handler (it knows which labels are runtime), not in the engine — the engine stays source-shape-agnostic.

No new engine module, no new dependency, no new DSL key (runtime sources are data, not declaration).

Proof

  • Engine unit tests: source_label (git url, bundle, local path, nested path); IngestReport/SourceStats carry the label (serde default).
  • e2e (full-stack example, new committed fixture knowledge/legacyrepo2.bundle — a second tiny legacy repo with a distinctive token lumenledger):
  • GET …/sources lists the declared sources with origin: declared;
  • POST …/sources/add (git) adds legacyrepo2, and its content is hybrid-searchable;
  • POST …/sources/add (path) adds a directory source;
  • POST …/sources/legacyrepo2 (the no-JS remove form) answers 303 to /knowledge?removed=… and its entries go away;
  • DELETE …/sources/legacyrepo2 (the JSON path) removes it, 200;
  • fail-closed: adding a nonexistent path → 400; deleting a declared source → 400; a colliding label → 400;
  • the form-encoded add answers 303 to /knowledge?added=….
  • Conformance goldens re-pinned (new app handlers + routes, web page, site page, vendored engine); workspace tests + clippy -D warnings + fmt clean.

Consequences

  • A running app can grow and shrink its knowledge — book libraries and repositories included — from the page, the API, or an MCP-driven agent; the state persists across restarts and never leaks into the source walk or the rendered bytes.
  • Declared sources remain the source of truth: they cannot be removed at runtime, they cannot be shadowed by label, and the tessera still renders deterministically (ADR 0003).
  • The form/JSON negotiation adds one branch to the add handler; every other route stays JSON.
  • What is deliberately NOT built here: runtime edit of a runtime source's filters (re-add after delete is the documented cycle), remote (S3-style) source URIs (that is the data-access seam, a separate ADR), and multi-operator concurrency on the registry file (single process by design — the ADR 0011 posture).

ADR 0015: OpenDAL data seam — knowledge reads from, and persists to, any data backend

  • Status: accepted
  • Date: 2026-10-01
  • Supersedes: nothing (extends the vector_dbs[].sources[].path and vector_dbs[].persist constructs, K1; the runtime-source registry, ADR 0014)

Context

The knowledge engine reads sources exclusively from the local filesystem (std::fs walk + read) and persists state exclusively to local JSON files (the index via Store::save/load, the ADR 0014 runtime registry next to it). The requirement: support OpenDAL in mosaic, to be able to read and write data from different data sources/targets — concretely for the knowledge feature: a book library that lives on object storage must be addable as a source (declared or at runtime, ADR 0014), and the knowledge state (index + runtime registry) must be writable to a target beyond the local disk (object storage that survives container replacement, in-RAM for tests).

Apache OpenDAL is the accepted data-access layer for exactly this: one pure-Rust API over many backends (fs, memory, s3, gcs, azblob, oss, obs, cos, hdfs, …), services feature-gated.

Three constraints shape the design:

  1. The engine is sync, and it is called from mixed contexts. ingest_source, persist I/O and the registry I/O run at boot (inside the host app's async runtime) and in tokio::task::spawn_blocking workers (no runtime context at all). OpenDAL's sync BlockingOperator must be constructed inside a runtime context (it captures the current Handle) — which the spawn_blocking contexts do not have. The seam therefore drives OpenDAL's async Operator from the engine's own dedicated tokio runtime (data::runtime(), a small multi-thread runtime created lazily once) via block_on — valid from every calling context (a separate runtime, so no nested-block_on panic; blocking the calling thread matches the engine's existing blocking-I/O posture).
  2. The engine stays env-pure (Z7). The workspace bans std::env::var in the engine/generator crates — so storage credentials are not read by the engine. The generated app (and the mosaic knowledge-report CLI) wire a StorageEnv provider over their own env reads; the engine resolves credentials through the provider and fails closed with a named error when it has none.
  3. The vendored engine is compiled into the web lens in full/islands (native SSR) only — never for wasm (csr). A native-only dependency is therefore safe to add to the engine.

Decision

One engine module (a sync seam), two URI upgrades, env-gated credentials, no new DSL key beyond reusing path/persist as URIs:

  1. A new engine module data — the sync DataBackend trait. ``rust pub trait DataBackend: Send + Sync { /// Direct children of a directory (""` = the backend root). fn list(&self, dir: &str) -> io::Result<Vec>; fn read(&self, path: &str) -> io::Result<Vec>; /// Write bytes (parents as the service supports). fn write(&self, path: &str, data: &[u8]) -> io::Result<()>; fn exists(&self, path: &str) -> bool; fn is_dir(&self, path: &str) -> bool; /// Human-readable location (stats / the /kb page). fn describe(&self) -> String; } pub struct DataEntry { pub name: String, pub is_dir: bool }

    /// Storage credentials / region / endpoint provider (see §4). pub struct StorageEnv { / Arc Option> / } `` Two implementations + a factory: -LocalBackend { root: PathBuf }—std::fs. **The default: a plain path resolves to this, byte-identical to today's behavior** (all existing engine tests stay green unchanged). -OpenDalBackend { op: Operator, desc }— any compiled-in OpenDAL service, driven by the engine's dedicated runtime (constraint 1);list= per-directoryop.list(OpenDAL lists the queried directory itself + its direct children — the walker skips the self-entry), recursion stays in the ingest walker. -data::backend_for(uri: &str, base: &Path, env: &StorageEnv) -> Result<(Box, String), String>— the backend is rooted at the URI's **parent**, the returnedreladdresses the URI's last segment (""= the whole root): sources walk fromrel, the persist target reads/writesreldirectly. - no scheme (orfile://) →LocalBackendrooted atbase(relative) or the file's parent (absolute); -s3://bucket/…,memory://…,fs://host/…→OpenDalBackend; any other scheme is a **named error** (data: scheme x is not available in this build`); a service that cannot be configured (missing credentials/region) is a named error, fail-closed (the caller wanted a verdict).

  2. Sources can be data URIs (read from any target). sources[].path accepts an OpenDAL URI (s3://bucket/books/, memory://lib, fs://host/lib) alongside a plain path and a git source (unchanged — git stays on the git-CLI seam). Ingestion resolves the source's backend via backend_for and walks with list/read instead of the fs walk. Locators/meta for a URI source carry <label>/<rel-from-source-root> — the K8 git-label pattern: label is the slugified scheme-authority-path of the URI (deterministic, independent of any cache location); entry ids are the flat slug of that path (the existing id convention). A single-file URI (s3://bucket/book.epub) ingests as one document; include/exclude filters are unchanged (substring match on the rel path). Remote PDFs are staged to a temp file for pdf-extract (it takes an fs path) and cleaned up after. ADR 0014's source_label for a URI source is that label — the runtime add/remove flow (K9) works identically for remote data: the /knowledge form's path input accepts URIs, and the fail-closed "path not found" check becomes !backend.exists(root) (the engine performs it; the app's fs pre-check skips URIs).

  3. The persist target can be a data URI (write to any target). vector_dbs[].persist (existing key, today a local path) accepts an OpenDAL URI (e.g. s3://bucket/kb.json). The engine gains byte-level Store::serialize() -> Vec<u8> / Store::deserialize(&[u8]) -> Store (the existing Persisted shape); save(Path) / load(Path) remain as LocalBackend shims so engine tests and the CLI are unchanged. The generated app resolves the persist backend from the URI and reads/writes the index and the ADR 0014 runtime registry through it; the registry sits next to the index — local: knowledge/.runtime/sources.json (unchanged), URI target: <target-dir>/.runtime/sources.json.

  4. Env-gated credentials through a StorageEnv provider — the DSL declares placement, env declares the provider (the established pattern; secrets never in tessera.yaml): MOAIC_STORAGE_{SCHEME}_{KEY} (e.g. MOAIC_STORAGE_S3_ACCESS_KEY_ID, …_SECRET_ACCESS_KEY, …_REGION, …_ENDPOINT; scheme upper-cased, -→_), falling back to the provider-standard env (AWS_* for s3). Because the engine is env-pure (constraint 2), the env reads live in the generated app (knowledge_storage_env(), a lazily-built StorageEnv::new closure over std::env::var) and in the knowledge-report CLI; the engine only calls the provider (env.lookup(mosaic_key, standard_key)). s3 requires key + secret + region (fail-closed named errors when absent — the builder validates); no credentials are read for local paths, memory://, or fs://.

  5. Dependency. opendal 0.59 (default-features = false, features services-fs, services-memory, services-s3) plus tokio (rt-multi-thread, net, time — the engine's dedicated runtime, constraint 1), in the workspace crate and the vendored knowledge-engine/Cargo.toml. Pure Rust (rustls http transport; no native linking, no downloads at build time — same posture as the K7 ort load-dynamic seam). Additional backends (gcs, azblob, oss, …) are an enable-a-feature + one build_operator match arm change.

  6. Fail-soft / fail-closed split (the established posture).

    • Boot: an unreachable URI source or persist target → a named error for that source/target; the app boots; the remaining sources/index work (fail-soft, exactly like K8 git sources).
    • K9 runtime add of a URI source → fail-closed 400 with the named error (the caller wants a verdict).
    • memory:// is per-instance RAM: each backend_for call builds a fresh, private in-memory store — it is a zero-config test/dev seam (deterministic: always empty on a fresh process), not a shared or durable target.

Proof

  • Engine unit tests (new data module tests + ingest tests, hermetic — no network):
  • memory:// round-trip through the seam: write / read / list / exists;
  • ingest from a data URI (fs:// over a temp dir — the real OpenDAL path, hermetic): directory source (files ingested, locators <label>/<rel>, content asserted) and single-file source;
  • missing URI roots fail closed (memory:// — always fresh/empty — and a nonexistent fs:// dir both → knowledge source not found);
  • Store::serialize/deserialize round-trip (entries + citations; corrupt bytes → empty store);
  • backend_for resolution: plain path → local (unchanged), unknown scheme → named error, s3 without credentials → named error, s3 without region → named error (all before any I/O), s3 with an explicit StorageEnv provider resolves;
  • every pre-existing engine test green unchanged (the local default).
  • e2e (full-stack): a declared memory://kb-remote source on kb — boot is fail-soft (the app is healthy, the other declared sources remain hybrid-searchable), the URI source is listed in /api/vectordb/kb/sources (proving declared-URI plumbing end to end), and a runtime POST …/sources/add of memory://does-not-exist is fail-closed 400 "not found" (proving the K9 URI path deterministically — memory is always fresh/empty, so no cloud is involved). Posture per K7/K8: the seam's happy path is unit-tested hermetically; CI asserts the fail-soft/fail-closed contracts; no live cloud in CI.
  • Conformance goldens re-pinned (vendored engine + generated persist plumbing); workspace tests + clippy -D warnings + fmt clean.

Consequences

  • A mosaic app can read knowledge from and write knowledge state to any OpenDAL backend — a book library on S3 is a one-line source (declared or added at runtime, ADR 0014), and the index + runtime registry can live on a target that outlives the container.
  • Zero behavior change by default: every existing plain-path source and persist resolves to LocalBackend (byte-identical I/O); no new DSL key (path/persist are simply URIs now); no new env for local apps.
  • The engine gains one small, sync, well-tested seam (a trait + two impls + a factory) — the OOP shape the platform's external seams are moving toward (git CLI, OCR, rerank, ONNX).
  • New dependency: opendal (pure Rust, feature-gated services; native-only lenses compile it — the wasm/CSR web lens never does).

ADR 0016: parts/blocks inventory + configurability on all four levels

  • Status: accepted
  • Date: 2026-10-01
  • Supersedes: nothing (activates the dormant PartDecl.config/slots AST surface; extends the app.parts / app.instances constructs)

Context

A tessera composes an app from parts (a.k.a. blocks — the user's "BBs"): parts: declarations (name, group, about, tags, depends, audience, contributions, nested domain decls) that contribute to slots (rest.endpoints, cli.commands, mcp.tools, gateway.routes, ui.scenes, ui.pages), plus the app's enablement (app.parts: [names]) and instances (app.instances: {name: {part, realm, config}}).

The requirement (in order): an inventory of groups of parts and the parts themselves, and configurability on all levels — app → part instance → part slot/port → contribution — supported to 100%, with good software patterns.

Current state (the gap):

  • No inventory surface. PartDecl.group/audience/tags are parsed and carried into PartPlan, but rendered nowhere: no page, no REST route, no MCP tool. An operator cannot see what an app is built from.
  • The config machinery is dormant. ConfigField (name, ty, default, env, feature_gate, required) and PartPlan.config/ConfigFieldPlan exist, and the plan layer consumes them (env bindings, api_key_env, base_path) — but no YAML key ever populates part.config (YamlPart has no config), and no platform part declares one. The consumption code is dead: AppPlan.env is never read by the renderer, api_key_env is always None, part_configs (instance → field → value) is only looked up inside the dead config loop.
  • Slots are implicit. The vestigial PartDecl.slots (SlotDecl {about, multi, handler}) is never declarable or validated; a contribution's dotted slot path is matched by if/else in the planner, and a typo or an undeclared slot is silently ignored (_ => {}).
  • Enablement doesn't gate contributions. rest.endpoints/ cli.commands/mcp.tools/ui.scenes/ui.pages are collected from all declared parts, while gateway.routes is collected from enabled parts only — so "app-level configurability" (enabling/disabling parts) is a no-op for most slots.
  • Contribution items have no on/off. A part's items are always generated; there is no declared way to keep a slot contribution present but disabled.

Decision

Four levels, in the required order (app → instance → slot → contribution), plus the inventory surface that makes all of them visible:

L1 — app level: enablement gates everything

app.parts: [names] is the single source of truth for which part types are on. Fix: every slot collection (rest/cli/mcp/ui, not just gateway) now iterates the enabled part types only. A declared part not in app.parts contributes nothing (still visible in the inventory, marked disabled). Unknown part types keep the existing lenient posture (warning: platform parts are provided at build time by the compiler — rest, cli, mcp, gateway, ui, db, cqrs, scheduler, …).

L2 — instance level: validated config

parts[].config becomes declarable in YAML (activating the dormant AST):

parts:
  - name: orders
    config:
      - name: base_url
        type: string
        default: "https://orders.example.com"
        env: { var: MOAIC_ORDERS_BASE, secret: false }
      - name: token
        type: string
        env: { var: MOAIC_ORDERS_TOKEN, secret: true }
        required: true

app.instances.<name>.config (existing part_configs) is now validated against that schema, fail-closed with named errors: unknown key → part-config-unknown; missing required field with no default/env → part-config-missing; value type vs type (string/number/bool/enum) mismatch → part-config-type; a secret field with a literal value → part-config-secret (secrets are never inlined, the existing Z7/ADR 0014 posture); a secret: true field without an env binding → secret-not-env (the schema rule: secrets live in env vars, never in the tessera). secret is a standalone YAML flag (an env binding can also carry secret: true; the two are OR-ed). Resolution order per field: instance literal → declared default → env binding (the generated app reads the env var at runtime; MOAIC_* convention unchanged).

The resolved config reaches the generated app as two typed consts (the OOP shape: a part instance receives its config object, secrets by reference):

  • PART_CONFIG: &[(&str instance, &str field, &str value)] — non-secret effective values (literal or default);
  • PART_ENV: &[(&str instance, &str field, &str var, bool secret)] — env bindings; the app reads var at boot and the value shows in the inventory as set/unset (never the value).

This also activates the existing consumption paths (AppPlan.env, api_key_env, base_path) with real data.

L3 — slot/port level: declared slots, fail-closed

parts[].slots becomes declarable (activating the vestigial SlotDecl):

parts:
  - name: orders
    slots:
      actions:
        about: "Order actions this part exposes."
        multi: true

Slot paths are <part>.<slot> (the existing two-segment convention — rest.endpoints = part rest, slot endpoints). Validation, fail-closed:

  • a contribution targeting a declared user part must name one of its declared slots → otherwise part-slot-undeclared (a typo can no longer be silently ignored);
  • platform-part slots (rest.endpoints, cli.commands, mcp.tools, gateway.routes, ui.scenes, ui.pages) stay built-in — no declaration needed;
  • multi: false (default): the same item name in two contributions of one slot → part-slot-duplicate; multi: true allows it.

SlotDecl.handler stays out of scope (no runtime dispatch to register; slots are compile-time composition points).

L4 — contribution level: per-item on/off

A contribution item may declare enabled: false — the item stays in the inventory (visible, documented as disabled) but is excluded from code generation (no endpoint/route/tool/command/scene/page). Default true. This is the declared, deterministic off-switch: flipping it is a declaration change, rendered output changes accordingly (conformance re-pin).

Inventory surface (declared facts — the ADR 0008 projection)

The plan gains the complete inventory (the extended PartPlan: tags, audience, slots with their items + enabled state, instances with resolved config, enabled flag), projected to three surfaces:

  • REST (app lens): GET /api/parts (and POST, for the MCP/CLI bridges — one core, many bridges) → {groups: [{group, parts: [{name, about, audience, tags, enabled, depends, config: [{name, env, secret, required, default}], instances: [{name, realm, config: {field: value|"***"|null}}]}], slots: [{part, slot, about, multi, platform, items: [{name, enabled}]}]} (secrets masked; env-bound non-secret fields show their build-time value or null). The plan carries the inventory as ws.parts (extended PartPlan: tags/audience/enabled) + ws.slot_catalog (the six platform slots always present, declared slots alongside). The generated app emits PART_CONFIG/PART_ENV consts and PARTS_INVENTORY_JSON (a &str const — serde_json::json! is not const-evaluable) parsed once into a OnceLock by crate::ws::parts_inventory().
  • MCP (app lens): list_parts tool (the stdio server's built-in inventory entry, mounting /api/parts) — the same inventory, for agents.
  • Site (zola + html backends): build-time /parts/ page (nav entry, the ADR 0008 projection of declared facts — like /kb/ for knowledge): groups → parts table (name, about, audience, tags, slots with item counts, instances with their effective config, secrets masked).

The web lens (csr/full) gets no page in this ADR — the site page + REST + MCP cover the inventory requirement (the web lens is a per-app console, not the reference projection).

Proof

  • Unit tests (inline YAML, tessera_yaml + plan layer):
  • part.config + part.slots parse (YAML → decl);
  • L1: a disabled declared part's contributions are absent from the plan; an enabled part's are present;
  • L2: unknown instance config key / missing required / type mismatch / secret-with-literal each produce the named error; literal + default + env resolution order holds; PART_CONFIG/PART_ENV consts emitted (codegen test);
  • L3: contribution to an undeclared slot of a user part → named error; platform slots still work; duplicate item name on a non-multi slot → named error;
  • L4: enabled: false item absent from endpoints, present in the inventory as disabled;
  • GET /api/parts + list_parts + /parts/ site page emit (codegen tests) with secrets masked.
  • e2e (full-stack): GET /api/parts → 200 (the full-stack part gains a config schema + a declared slot + an enabled: false item + instance config; the new scenario asserts the inventory endpoint).
  • Conformance goldens re-pinned (new consts + route + MCP tool + site page); workspace tests + clippy -D warnings + fmt clean.

Consequences

  • An operator can see exactly what an app is built from (groups → parts → slots → items → instances) on three surfaces, and configure it on every level: enable/disable part types (app), set instance config (instance, validated), declare + reference slots (slot, fail-closed), switch individual items off (contribution).
  • The dormant config machinery (ConfigField → plan → api_key_env/ base_path/env) becomes live, fail-closed, and typed — no new AST types, only YAML keys and validation.
  • Generated apps grow two consts + one route + one MCP tool + a site page; zero behavior change for apps without config/slots (all defaults: config empty, slots = platform-only, items enabled).
  • Next (ADR 0017): the model registry — models: declarations for chat/embed/rerank (provider + model, secrets env-only), wiring the dead vector_dbs[].model, and runtime configurability (effective config
  • provider "available models" listing) — reusing the declared-floor / runtime-overlay / env-for-secrets posture established here.

ADR 0017: model registry + runtime model configurability

  • Status: accepted
  • Date: 2026-10-01
  • Supersedes: nothing (activates the dead vector_dbs[].model key; adds the model_registry: top-level declaration; extends the env-based model config with a declared floor and a runtime overlay)

Context

Which model does what is today env-only:

  • workflow llm nodes: spec.provider → MOAIC_{PROVIDER}_{API_KEY, BASE_URL,MODEL}, falling back to the global MOAIC_CHAT_* family; the model name falls back to spec.name, then gpt-4o-mini;
  • assistant chat: MOAIC_CHAT_{API_KEY,BASE_URL,MODEL} (key falls back to MOAIC_EMBED_API_KEY);
  • embeddings: MOAIC_EMBED_{API_KEY (falls back to CHAT_API_KEY), BASE_URL, MODEL} (default text-embedding-3-small);
  • rerank: MOAIC_RERANK_{API_KEY,BASE_URL,MODEL} (default rerank-english-v3.0, fail-soft).

And vector_dbs[].model ({provider, name}) is parsed and shown on the /kb/ site page but never consumed — a dead key.

There is no declared inventory of the models an app uses, no way to see (at runtime) which model/provider is effective, and no way to switch the effective model at runtime or to list the models a provider actually offers — the requirement: models for embeddings/LLM/etc. configurable, integrated with runtime configurability (provider + available models), with good software patterns.

Decision

The registry (declared floor)

A new top-level declaration — model_registry: (the models: key is already the IAM role-model vocabulary):

model_registry:
  - name: fast
    about: "Cheap model for high-volume calls."
    provider: groq        # env family: MOAIC_GROQ_{API_KEY,BASE_URL,MODEL}
    model: llama-3.1-8b   # default model name
    base_url: "https://api.groq.com/openai/v1"

Three built-in roles are always registered (a declared entry with the same name overrides its defaults — the declared-floor posture of ADR 0014/0016):

name env family default model notes
chat CHAT gpt-4o-mini assistant chat + llm node fallback
embed EMBED text-embedding-3-small key falls back to MOAIC_CHAT_API_KEY
rerank RERANK rerank-english-v3.0 fail-soft, no default base: absent key or absent base = no rerank

Resolution order per name: runtime overlay → env MOAIC_{PROVIDER}_{MODEL,BASE_URL} → declared model/base_url → built-in default. The API key is always env-only (MOAIC_{PROVIDER}_API_KEY; secrets never inlined, the Z7/ADR 0014/0016 posture).

Consumers rewired through one resolver

The generated app gains the OOP shape: a const registry (MODELS: &[ModelConfig]), a pure resolver (resolve_model(name) -> Option<EffectiveModel>: model + base_url + the key's env var name), and a runtime overlay. Every consumer resolves through it:

  • assistant chat → chat;
  • workflow llm node: a new spec.model = registry name (the node's provider/name legacy pair keeps working unchanged);
  • embeddings → embed, per collection: the dead vector_dbs[].model becomes live — its name (a registered model) selects that collection's embedding model (embed_text(text, collection));
  • rerank → rerank.

Runtime configurability

  • GET /api/models — the effective registry: per entry {name, about, provider, model, base_url, key: "set"|"unset" (never the value), origin: "declared"|"runtime", built_in};
  • POST /api/models/{name} {model?, base_url?} — a runtime override, persisted to <base>/.runtime/models.json (survives restarts; the ADR 0014 registry pattern — scratch state for e2e, reset per run). The name must exist in the registry (declared or built-in) → fail-closed named error;
  • DELETE /api/models/{name} — clear the override (back to declared);
  • GET /api/models/available?provider=X — list the models the provider offers: GET {base_url}/models with the provider key (OpenAI- compatible). Fail-closed: key unset → 400 named error; provider failure → 502 named error (the listing is best-effort information, never guessed);
  • MCP: list_models tool (one core, many bridges);
  • Site: build-time /models/ page (zola + html backends, nav entry) projecting the declared registry (the ADR 0008 projection — env values and runtime state are app data, never static content).

Proof

  • Unit tests (plan + codegen): built-ins + declared entries + override precedence; unknown-name fail-closed; the dead vector_dbs[].model selects the per-collection embed model; GET /api/models (+override + available) routes/handlers, list_models MCP tool, /models/ site page emit.
  • e2e (full-stack): GET /api/models → 200 (lists chat/embed/rerank); POST /api/models/chat {model} → 200 and GET reflects the runtime origin; POST /api/models/nope → 400; GET /api/models/available without a key → 400 (fail-closed, hermetic — no network in CI).
  • Conformance goldens re-pinned; workspace tests + clippy -D warnings + fmt clean.

Consequences

  • An operator can see exactly which model does what (declared + effective, key masked), change the effective model at runtime (persisted, restart-proof), and discover what a provider offers — before pointing a registry entry at a model id that does not exist.
  • All model config flows through one resolver with one resolution order; the env stays the secret channel (Z7), the tessera stays declarable, and the runtime overlay is the operator channel — the same declared-floor / runtime-overlay / env-for-secrets posture as the knowledge sources (ADR 0014/0015) and the part config (ADR 0016).
  • vector_dbs[].model stops being dead weight; the per-collection embed model is the first per-instance model binding.
  • Zero behavior change for apps without a model_registry: (the built-ins resolve to exactly today's env behavior, including the MOAIC_CHAT_* fallbacks and the fail-soft rerank).

ADR 0019 — Delta re-indexing: per-source file manifests

Status: accepted

Context

Every reindex (and every boot re-walk) in mosaic re-ingests all files of all sources and resets every entry's embedding to None, so the embedder re-runs on the whole corpus — even when nothing changed. The ee-pages portal fixed exactly this (its doc_files manifest + delta ingest on publish): each source keeps a manifest of relpath → content-hash; on re-index only the added and changed files are chunked + embedded, removed files are deleted from the index, and unchanged files keep their entries and their embeddings. The manifest self-heals: a corrupted or missing manifest simply re-ingests the whole source once.

For a book library or a large git source this is the difference between a reindex that costs unchanged files ≈ 0 and one that costs the whole corpus of embedding calls.

Decision

1. Per-source manifests in the engine

mosaic-knowledge gains:

  • file_hash(data) -> String — sha256 hex of the file bytes (the hash is of the content, so a touched-but-identical file is unchanged).
  • SourceManifest { files: BTreeMap<String, String> } — relpath → hash, per source label, serde-shaped (it rides the same persistence seam as the store).

2. The delta path

A new engine entry point, ingest_source_delta(src, base, strategy, env, previous: Option<&SourceManifest>, store: &mut Store) -> (IngestReport, SourceManifest):

  1. Walk the source (same walk + filters as today).
  2. For each file: hash the bytes. - hash present in previous and the doc slug already in the store → skip (entry + embedding untouched). - new or changed → process_file as today, store.upsert(slug, …) (fresh entries get emb: None; the caller's embed_missing pass embeds exactly the new ones). - in previous but absent now → store.remove_slug(slug) (the file left the source).
  3. Return the new manifest (the caller persists it).

A missing/empty previous (first run, manifest lost, label renamed) is exactly today's full ingest — self-healing, no migration.

3. Manifest persistence

One file per collection, next to the runtime registry, through the same data seam (local knowledge/.runtime/manifest.json; a data-URI persist lands it under <persist-dir>/.runtime/manifest.json):

{ "<collection>": { "<source-label>": { "<relpath>": "<sha256>" } } }
  • Boot ingest (declared + runtime sources) and POST /reindex use the delta path and rewrite their labels' manifests.
  • POST /sources/add records the new label's manifest.
  • DELETE /sources/{label} drops the label's manifest (the slugs are already removed via the registry).

4. What does NOT change

  • Inline DSL docs (VECTORDB_SEEDS): idempotent by id, tiny — no manifest.
  • Deterministic ids/slugs (unchanged — that is what makes the per-file skip and remove work).
  • The store shape, the REST surface, the e2e contract.
  • The git checkout cache (knowledge/.gitcache/<label>): a reindex still reuses the work tree (a git source's files are whatever the checkout holds; the manifest decides what of that is new).

Consequences

  • Reindex cost drops from O(corpus) embedding calls to O(changed); boot with a warm index + unchanged tree re-embeds nothing.
  • store.remove_slug must stay exact (a label's manifest only ever removes that label's slugs — no cross-source damage).
  • A content change that does not change chunk boundaries still re-upserts (new emb: None → re-embedded) — correct, slightly wasteful vs a chunk-level diff; out of scope.
  • Old indexes/manifests from pre-0019 builds simply have no manifest file → one full ingest on first boot, then delta forever.

ADR 0020 — Deploy ingress (host / ingress_class / proxy_buffering) + app content_root

Ports two ee-pages/EE deployment improvements into the tessera deploy and web lenses.

Context

Ingress. ee-pages' assistant streams SSE per token, but the deployed ingress-nginx coalesces the whole response into one burst before the client sees it. EE fixed this with a per-environment deploy-DSL field proxy_buffering: "on" | "off" (commit f93a04c96): it renders the nginx.ingress.kubernetes.io/proxy-buffering annotation into the environment's values, and the chart's ingress template picks it up. Unset keeps the nginx default (buffering on, still honoring a response's X-Accel-Buffering: no). The parser rejects any value other than on / off.

Mosaic's tessera deploy specs (ADR 0010) render a per-spec Helm chart, but that chart has no ingress at all — no host, no class, no annotations. (The legacy v0 deploy lens predates tessera and already has host / ingress_class / proxy-body-size; this ADR does not touch it.)

content_root. A portal/content app whose main page is app content cannot own the deployment root under EE: with a home path set, the generated Leptos router emits a GET / redirect (and a page claiming / is routed at /), which collides with any block that mounts its own GET / (commit b2abba3e6, "content_root — let the deployment root belong to app content"). EE added an app-spec field content_root: bool (default false, byte-identical behavior when omitted): when true, the Leptos factory emits no root route at all — the root is free for the content route, and the UI keeps all its other routes and becomes a hidden entry point.

Mosaic's web lens is the equivalent surface: the generated Leptos router unconditionally claims the root (<Route path=() view=pages::Index/>), and the axum fallback serves the shell for unknown paths. A content app (e.g. mosaic-pages' doc site, P8) that wants to serve its content at / has no way to give the root up.

Decision

1. Tessera deploy specs gain three optional ingress fields (on each deploy { name: … } spec, so every environment can differ — the spec name is the environment, per ADR 0010):

app:
  deploys:
    - name: prod
      target: gcp-gke
      host: pages.example.com
      ingress_class: nginx
      proxy_buffering: off   # on | off (streaming / SSE surfaces)
  • host — the Ingress rule host. Unset → the rule is hostless (ClusterIP / port-forward usage).
  • ingress_class — spec.ingressClassName. Unset → cluster default.
  • proxy_buffering — renders the nginx.ingress.kubernetes.io/proxy-buffering annotation. Fail closed: any value other than on / off is a bad-deploy-proxy-buffering error and the spec is skipped (same hard-error discipline as the store fields).

2. The per-spec Helm chart gains an ingress template. deploy/<name>/helm/templates/ingress.yaml renders an Ingress from values.yaml:

ingress:
  host: pages.example.com      # "" when unset
  className: nginx             # absent when unset
  proxyBuffering: off          # "" when unset

The annotation renders only when proxyBuffering is set (the nginx controller key; other controllers ignore it). The chart otherwise stays minimal — TLS is out of scope (cert-manager is cluster policy, like in EE).

3. App-level content_root: bool (default false) on app::

app:
  name: pages
  content_root: true

When true, the web lens omits the root route from the generated Leptos router (path=()). Everything else is unchanged: the other pages keep their routes (the UI remains reachable, e.g. /commands, /aggregates, /docs), and the root belongs to whatever mounts it. This is the web-lens port of EE's suppression; the static site lens (zola) is unaffected (it is a separate surface with its own /), and the app lens is pure API (no root route).

The decision is stamped once in the plan (AppPlan.content_root); the render just reads it.

Consequences

  • DeploySpecPlan grows host, ingress_class, proxy_buffering; the values.yaml and the new ingress template change only for specs that set them (hostless ingress renders - with no host, exactly as the legacy v0 lens does).
  • Conformance goldens for any example with a tessera deploy spec are re-pinned (the chart gains ingress: values + the template).
  • content_root is a render-time fact: no data-model change, no migration. Apps that never set it render byte-identical output.
  • Out of scope (deliberately): TLS/cert-manager wiring, path prefixes, the EE auto-deploy trigger / mono-SHA tagging (GitLab-mono specific), and ui_prefix (unmerged upstream).

ADR 0021 — Re-author mosaic-pages as tessera.yaml

The mosaic-pages app (pages.tessera) is written in the legacy .tessera text format, which the CLI no longer loads (ADR 0008: YAML is canonical; mosaic-cli rejects the legacy format). This ADR re-authors the app as tessera.yaml, using the features P5–P7 just landed so the port doubles as their first real-world consumer.

Context

  • pages.tessera (Sep 26) predates the YAML migration and the knowledge RAG work: its vectordb docs is inline-docs-only (three seeded texts), its chat is a workflow search node, and its deploy story doesn't exist (no deploys — the tessera has no deploy specs at all).
  • The README documents three known workarounds from that era: a disabled @env api-key binding (codegen bug, since fixed), a post-render patch adding missing web-crate deps (fixed), and the note that published docs don't reach the RAG index (still true — there is no upsert-to-vectordb DSL; documented, not solved).
  • P5 (ADR 0018) gave vector_dbs[] retrieval/chunk config + source-scoped search; P6 (ADR 0019) made re-indexing delta per source; P7 (ADR 0020) gave deploy specs host/ingress_class/proxy_buffering and the app content_root.

Decision

Rewrite mosaic-pages as tessera.yaml (dropping pages.tessera), keeping the same app shape — the Doc aggregate, the RBAC policy, the SSO identity, the ask-docs workflow, the operator guide — and adding what the legacy tessera lacked:

  1. Knowledge sources, the ee-pages way. vectordb docs gets real sources alongside the inline docs: - a books/ folder (kind auto — epub/docx get the chapter model, pdfs the page model), - a git source (K8: a committed git bundle, CI-hermetic), - a data-URI source (ADR 0015: memory://docs-remote, proving the OpenDAL seam — empty on a fresh boot, fail-soft), - retrieval: {mode: hybrid, top_k: 5} and chunk: {chars: 2000, overlap: 200} (P5 — the ee-pages defaults), - strategy: {books: true} explicit. Source-scoped search is e2e-proven: a search restricted to the git source's label finds that source's text and nothing else (and an unknown label returns empty).
  2. content_root: true (P7): the portal root belongs to the content — the web lens emits no root route; the UI keeps its other pages as a hidden entry point.
  3. A deploy spec (ADR 0010 + P7): deploys: [{name: prod, target: k8s, host: pages.example.com, ingress_class: nginx, proxy_buffering: off}] — the chart gains the ingress with the SSE annotation (the workflow run endpoint streams over ?stream=true).
  4. e2e: the legacy four scenarios (health/publish/list/get) plus token-scoped authz (viewer 403 / owner 200 on publish — the permission gate the README showcases) and the source-scoped search scenarios.

The api_key part-config env binding is restored (the codegen bug the README worked around is gone) and the README is rewritten for the YAML format (no more post-render patch).

Consequences

  • mosaic-pages becomes loadable and e2e-green again; out/ stays git-ignored and regenerable.
  • The knowledge corpus (books/, the git bundle) is committed to the app repo — the same posture as examples/full-stack/knowledge/.
  • The upsert-to-vectordb gap (published docs don't reach the RAG index at runtime) remains; it is a DSL feature request, tracked separately.
  • The legacy pages.tessera is deleted, not kept (the CLI cannot read it; keeping it would be dead weight that drifts).

ADR 0022 — Knowledge graph: persisted document graph, entity resolution, graph-aware retrieval

Status: design. This ADR decides the shape of the next knowledge-engine step; it lands with no implementation. The implementation follows as separate work items (see Consequences), each gated on this design.

Context

The engine already has a deterministic knowledge graph (K5, mosaic-knowledge/src/graph.rs): nodes and edges derived from the indexed entries — symbol/file/endpoint (code tier), book/chapter (book tier), report/doc (citation tier) — with uses/contains/serves/ next/cites edges, a neighborhood(root, depth) BFS, and REST/MCP surfaces (/graph, /neighborhood, /citations, /cite, /report).

Three properties limit it:

  1. Ephemeral. Store::graph() recomputes the whole graph from the entries on every request. Fine at the current corpus size, but it makes the graph a projection with no identity of its own — no version, no incremental update, and the uses-edge walk is O(symbols × chunks) per call. The P6 manifest already makes re-indexing delta per source; the graph cannot participate.
  2. No entity resolution. The uses edges connect code symbols via identifier-token intersection. Nothing unifies concepts across sources: the same thing written "deploy spec" in a doc, DeploySpec in code, and "the deploy block" in a book is three disconnected strings. The graph cannot say "these entries are about the same thing", so retrieval cannot exploit it either.
  3. Flat retrieval. Search is BM25 + cosine over chunks (RRF, optional MMR, optional rerank — all P5). The graph structure is never consulted: a hit's related entries (same symbol, sibling chapter, citing report) are as invisible to the query as any other chunk.

Decision

1. The graph becomes a persisted, versioned artifact

  • Built at ingest/reindex time (the same batch points as the P6 delta walk) instead of per request, and persisted next to the index through the same data seam (ADR 0015): knowledge/.index/<name>.graph.json (local) or <persist-dir>/.index/<name>.graph.json (data-URI persist).
  • The artifact carries a version stamp: the collection's P6 manifest hash (the relpath → sha256 map) plus the engine's graph-schema version. A search finds a missing or stale artifact (stamp mismatch) and rebuilds on demand — the exact self-healing posture of the P6 manifest (missing manifest = full ingest). Reads are fast in steady state; a rebuild is bounded by the sources that changed.
  • Delta: the P6 walk already knows which sources changed/removed. Changed sources recompute their nodes/edges; unchanged sources keep theirs. Cross-source edges (a uses edge crossing two files, a cites edge to another source) recompute only when either endpoint moved. The rebuild is deterministic (the engine's standing rule: no LLM, every ordering total).
  • Store::graph() becomes "load the artifact, rebuild if stale" — the signature and the REST/MCP surfaces are unchanged.

2. Entity nodes and deterministic resolution

A new node kind, entity, plus a resolution pipeline that is deterministic and auditable (no LLM in v1 — the engine's standing rule; LLM coreference is a later env-gated seam, same posture as rerank):

  • Extraction (per tier, deterministic):
  • Code: the existing symbols — every symbol node gains an implicit entity (its normalized name).
  • Books/docs: heading nodes — chapter titles (existing chapter nodes) and, v2, markdown section headings become entity candidates.
  • Declared: a new collection key, entities: [{ name, about?, aliases: [...] }] — the operator's declared floor for the vocabulary (the engine never invents an entity from this list; it only resolves mentions to it). Declared entities are the seed for cross-source identity: "name: DeploySpec", "aliases: [deploy spec, deploy block]" unifies all three spellings above.
  • Resolution rules (total order, no randomness): 1. normalize surface forms (Unicode NFKC, casefold, collapse whitespace, map -/_/camel boundaries to a canonical form); 2. a mention resolves to a declared entity when its canonical form equals the entity's canonical name or an alias; 3. otherwise it resolves to an existing code symbol when the canonical form equals a symbol's canonical name; 4. otherwise it is an unassigned mention (not a node — no hallucinated entities).
  • Auditability: the resolution map (mention → entity id, with the rule that fired) is part of the artifact. GET /api/vectordb/{name}/entities returns the entities + their mentions; an operator can see exactly why two mentions unified — or why they didn't.
  • Edges: about edges from an entry to the entities its text resolved (count = mention count), and same_as edges only between a declared entity and the code symbol it subsumes (never between two inferred mentions — that would be coreference, out of scope for v1).

3. Graph-aware retrieval

The P5 retrieval block gains an optional third facet (fail-closed validation, the standing rule):

vector_dbs:
  - name: kb
    retrieval:
      mode: hybrid
      top_k: 5
      graph:            # NEW (ADR 0022); absent = today's behavior,
        expand: true    # byte-identical
        weight: 0.2     # 0..=1, the expansion boost (default 0.2)
        hops: 1         # 1 | 2 (default 1)
  • Expansion pass: after the existing retrieval (RRF/MMR/rerank untouched), each returned hit's entity nodes seed a 1–2 hop neighborhood; neighbor entries whose text is not already in the results gain weight × (edge_count / 2^hops), and the final top_k is re-selected. Bounded by construction: 1 hop, top-N neighbors (N = 2 × top_k), source-scoped when the search is source-scoped (P5).
  • Deterministic: the boost is a pure function of the (persisted) graph and the candidate set; ties break on the existing entry-order rule. A graph: block on a collection whose artifact has no entity nodes is a no-op (nothing to expand), never an error.
  • related surface: GET /api/vectordb/{name}/related/{entry_id}?depth=1..5 returns the neighborhood as answer context (nodes with their entry ids + a one-line label each) — the follow-up material the assistant/chat surfaces consume ("what else is about this?"). Backed by the same artifact, so it is a read, not a recompute.
  • MCP: two new tools beside search_knowledge / list_sources — entities (the resolved vocabulary, source-scoped) and related (neighborhood of an entry). The assistant stays global, as the search source knob is per-search (P5 posture).

Consequences

  • Persisted/artifacts gain the graph file; the data seam (local fs or URI) is the only write path — no new storage backend, no external graph store. The engine stays embeddable and CI-hermetic.
  • vector_dbs[].entities is a new DSL key (YAML + plan validation: duplicate names fail closed; aliases may overlap with symbol names — the resolution order above decides). Conformance goldens for collections that set it are re-pinned; unset collections are byte-identical.
  • Implementation is staged, each stage independently shippable: 1. v1: persisted artifact + version stamp + delta rebuild + entities declared floor + resolution map + entities/related surfaces. (No retrieval change yet — the graph is queryable but not consulted.) 2. v2: retrieval.graph expansion + markdown-heading extraction. 3. v3 (env-gated seam, optional): LLM-assisted coreference for unassigned mentions — fail-soft to the deterministic result, keyed through the P17 model registry (coreference role), off by default, and its decisions persisted for audit like the v1 map.
  • Out of scope (deliberately): a general graph DB, cross-collection edges, temporal/versioned graphs (the artifact is a snapshot per manifest state), and anything that makes retrieval nondeterministic.

ADR 0023 — On-prem deploy: TLS issuer, app-level env, workflow-shaped chart

Status: accepted

Context

mosaic has a public on-prem deployment target (AGENTS.md → DevOps infrastructure): a single-node k3s cluster (li7) behind the *.eisler-systems.de wildcard, k3s Traefik on 80/443, cert-manager with a letsencrypt-prod ClusterIssuer, and a self-hosted GitHub Actions runner on the box (deploy-onprem.yml: render → docker build → k3s ctr images import → helm upgrade --install). The tessera deploy lens (ADR 0010, ingress refined in ADR 0020) must meet that workflow where it stands. Three gaps:

  1. No TLS in the tessera front-end. A public ingress on a wildcard domain needs a certificate; the on-prem cluster issues them with cert-manager. The chart's tls values existed (added for the workflow's --set overrides) but the tessera could not declare them — a spec that wants TLS had to rely on out-of-band helm flags.
  2. No app-level env. The generated app is configured by env (MOAIC_CHAT_API_KEY, MOAIC_EMBED_BASE_URL, MOAIC_JWT_SECRET, …). Part config env bindings exist (ADR 0016), but model/API keys are not a part's concern — there was no place for app-wide bindings, so the operator had to hand-write helm --set env.* for every deploy.
  3. Chart/workflow value mismatch. The workflow overrides image.repository / image.tag / image.pullPolicy, ingress.enabled, and db.password; the tessera chart rendered image as a single string, rendered the Ingress unconditionally (a hostless spec would install an empty-host rule — catch-all on a public controller), and left secret env secretKeyRefs non-optional (a missing AI key would wedge the whole Deployment).

Decision

Deploy spec: tls_issuer + tls_secret

deploys:
  - name: li7
    target: k8s
    host: pages.eisler-systems.de
    ingress_class: traefik
    tls_issuer: letsencrypt-prod     # cert-manager ClusterIssuer
    # tls_secret: pages-tls          # optional; defaults to <host, dots→dashes>-tls
  • tls_issuer enables the Ingress TLS block: the cert-manager.io/cluster-issuer annotation + spec.tls (host → secret).
  • tls_secret overrides the secret name; the default is the host with dots replaced by dashes plus -tls (the convention the workflow already used).
  • Fail closed (spec dropped, diagnostic raised):
  • deploy-tls-without-host — a certificate is issued for a host;
  • deploy-tls-secret-without-issuer — a secret name without an issuer is meaningless.

App-level env: app.env

app:
  env:
    MOAIC_CHAT_API_KEY:
      secret: true        # never inlined; the operator provides the k8s Secret
    MOAIC_EMBED_API_KEY:
      secret: true
    MOAIC_CHAT_BASE_URL: "https://openrouter.ai/api/v1"   # scalar = plain value
    MOAIC_JWT_SECRET: "dev-jwt-secret"
  • Each entry is var -> (value, secret, required): a scalar is a non-secret literal; a map is {value?, secret?, required?}.
  • A secret with a literal value is fail-closed (app-env-secret-literal, binding dropped) — a secret in the tessera would be inlined into the chart, which is exactly what the part-config rule already forbids (ADR 0016).
  • Plan order: app env first, then part config env bindings — a part binding for the same var wins (insert semantics, same as today's part env).
  • Render: non-secret entries land in the chart's env values map; secret entries land in the secrets list and render as secretKeyRef (secret name = key = env var name, as before).

Chart: workflow-shaped values

The per-spec tessera chart (ADR 0020's deploy/<spec>/helm) now matches what deploy-onprem.yml sets:

  • image: {repository, tag: latest, pullPolicy: IfNotPresent} — the runner builds the image, imports it into k3s, and overrides all three with --set (pullPolicy Never).
  • ingress.enabled — true when the spec declares a host, false otherwise; the Ingress template is guarded by it. A hostless spec installs no Ingress: on a public controller an empty host would match every request. (k3d-internal specs, e.g. data-sync, were hostless before and gain nothing from an Ingress object.)
  • tls: {enabled, secretName, clusterIssuer} derived from the spec's tls_issuer / tls_secret (the workflow may still override via --set).
  • Secret env secretKeyRefs carry optional: true unless the binding is required — the app boots without its AI key (the AI surfaces fail soft); a required binding keeps the wedge (boot fails loudly instead of serving a degraded app).
  • In-cluster store passwords stay required in the chart templates; the workflow derives a stable per-release password (sha256("$RELEASE-db")) so a redeploy does not rotate the credential under a live StatefulSet volume.

The workflow (deploy-onprem.yml) gains a repo input (default eugeis/mosaic) — a second checkout renders app repos that live outside the mosaic repository (e.g. eugeis/mosaic-pages) — and an AI key secrets step: when /home/ee/.ssh/openrouter exists on the runner it creates the MOAIC_CHAT_API_KEY / MOAIC_EMBED_API_KEY secrets in the namespace (one key feeds both, both are OpenAI-compatible families); without the file it warns and the deploy proceeds AI-less.

Consequences

  • A public on-prem app is declarable end to end from the tessera: host + ingress_class: traefik + tls_issuer: letsencrypt-prod renders a Traefik Ingress with a cert-manager certificate — no out-of-band helm flags.
  • Model providers (chat/embed/rerank) are configured per app via app.env + the model registry (ADR 0017), not per deploy flag; the same tessera deploys AI-less anywhere the secrets are missing.
  • Re-pinned conformance goldens: every tessera chart's values.yaml now carries the image map + ingress.enabled (full-stack, data-sync).
  • The legacy v0 lens (emit_helm) is untouched — it already had the workflow's image.repository/image.tag shape.

ADR 0024 — Embedded storage tier: sqlite + lancedb, scale-based engine defaults

Ports the EE storage-tier model into the tessera authoring format, the deploy lens, and the knowledge engine: small apps run fully embedded (sqlite + lancedb), no external store workloads at all.

Context

EE's model. In EE (blocks/ai/ee-vectordb, doc/reference/dsl/deploy.md), LanceDB is an embedded, in-process vector database (like SQLite/DuckDB): no pod, no Service, no image. The deploy DSL's lancedb { } block mounts a PersistentVolumeClaim onto the app's own pod and sets VECTOR_DB_URL to the local mount path (never a network address). Backend selection is priority-based (qdrant > weaviate > pgvector > lancedb) via VECTOR_DB_BACKEND/VECTOR_DB_URL; the vectordb block picks the matching backend through a factory (store_factory.rs), lancedb feature-gated.

The agreed tiering (operator decision):

app size db vector
small sqlite lancedb
medium postgres lancedb
large (big AI) postgres qdrant

Mosaic's current state.

  • Deploy-lens engine sets (hard-validated): db: postgres | mysql | mariadb | redis | dynamodb | clickhouse, vector: opensearch | qdrant | weaviate | milvus | pinecone. No embedded options.
  • The generated app does not connect to the deployed stores (the known MOSAIC_DB_URL/MOSAIC_VECTOR_URL gap): today a db:/vector: spec is store-workload proof, and the app persists to its own persistence: {backend: jsonl | sqlite} file store (rusqlite) — the sqlite tier is half-present already.
  • The in-app vector store is the mosaic-knowledge crate (vendored into every generated app as knowledge-engine/): in-memory entries + BM25 keyword index + brute-force cosine over in-memory embeddings, RRF fusion, JSON-file persistence (Store::save/load). No backend seam, no ANN index, no lancedb.

The li7 on-prem deployment (ADR 0010/0020/0023) runs one release per external engine; a small-app tier (embedded sqlite + lancedb) has no representation there.

Decision

1. Two new engines, embedded, k8s target.

  • db.engine: sqlite — the app's own file store; no db store workload. Valid only with persistence.backend: sqlite (fail closed: the plan errors otherwise — same hard-error discipline as the store fields).
  • vector.engine: lancedb — embedded LanceDB in the app process; no vector store workload.

Both are k8s-only in this ADR (PVC-backed; aws/azure EBS variants are out of scope). For each embedded engine the deploy lens renders, on the app Deployment:

  • a PersistentVolumeClaim + volume + mount (/data),
  • env for the app:
  • sqlite: the app's existing persistence.path_env var set to /data/app.db (no new env name invented),
  • lancedb: MOSAIC_VECTOR_BACKEND=lancedb and MOSAIC_VECTOR_URL=/data/lancedb (path, never a URL).

2. Knowledge engine gains a vector-backend seam (the EE store_factory port). mosaic-knowledge Store vector operations move behind a backend trait with two impls:

  • File — today's JSON persistence (default; apps that never opt in render and run byte-identical),
  • Lance — feature-gated lancedb crate: one Lance dataset per collection, Arrow schema (id / content / embedding FixedSizeList<f32> / metadata JSON), HNSW index for ANN.

The generated app selects the backend at boot from MOSAIC_VECTOR_BACKEND (unset → File); the keyword (BM25) leg and the RRF fusion stay in Store and are backend-agnostic, so hybrid search is unchanged in shape — only the vector leg's storage/index differs.

3. app.scale: small | medium | large (opt-in; unset keeps today's byte-identical behavior — omitted db:/vector: fields still mean no store). When set, a deploy spec that omits db:/vector: gets tier defaults:

  • small → db: sqlite, vector: lancedb
  • medium → db: postgres, vector: lancedb
  • large → db: postgres, vector: qdrant

An explicit db:/vector: on the spec always overrides the tier default (the spec name is the environment, per ADR 0010).

4. Staged outcome (documented, not hidden). medium/large defaults name external stores (postgres / qdrant) before the app can talk to them — the app-lens wiring of MOSAIC_DB_URL/MOSAIC_VECTOR_URL (postgres persistence backend, qdrant client) is a follow-up ADR. Until then the tier's external store deploys and is healthy; the app still uses its embedded store. small is fully end-to-end in this ADR.

Consequences

  • tessera_yaml/plan: engine sets grow sqlite (db) and lancedb (vector); new app.scale field; per-cloud validation gains the embedded k8s-only rule; the sqlite↔persistence-backend consistency check.
  • Deploy lens: embedded engines skip the store workloads and render the PVC/volume/env on the app Deployment instead (values.yaml gains an embedded: block). Conformance goldens for affected examples re-pin.
  • mosaic-knowledge: lancedb optional dependency (feature-gated); Store vector path routed through the backend trait; the JSON File backend is the default and the public API of Store is unchanged.
  • The li7 deployment gains a 6th release (li7-lite, app.scale: small): one app pod, no store workloads, PVC-backed sqlite + lancedb — the small-app tier proven end-to-end including AI (OpenRouter embeddings written into the Lance dataset, hybrid search over it).
  • Out of scope (deliberate): app wiring to external postgres/qdrant (follow-up ADR), the EE CQRS hybrid mode (PostgreSQL event store + SQLite projection store, doc/cqrs/storage-backends.md Mode 3 — a separate follow-up ADR), aws/azure embedded volumes, lancedb server deployments (the product is embedded by design).

ADR 0025 — Admin-UI routing: `ui_prefix`, the `ui` web-binary deploy, embedding LRU, Qdrant gRPC

Follow-up to ADR 0020 (content root) and ADR 0024 (embedded storage), from reviewing the latest EE mono commits (8944e7a5b, e3dc06369, d2bc23620, 0a7089324, 480c9a3cc, b76d626f3).

1. app.ui_prefix — nest the admin UI under a hidden route

EE learn (8944e7a5b feat(ddd): ui_prefix app field): when an app also serves content at / (content root), the Leptos admin UI is nested under a hidden route prefix (EE deploys use "/_"): /_/commands, /_/aggregates, … while / and the REST bridge (/api/...) are unchanged. It is an APP field, not a deploy field (the web shell is one binary, the prefix is part of the UI's identity).

Mosaic implementation.

  • Tessera: app: { ui_prefix: "/_" } (YAML + AppDecl.ui_prefix). Validated fail-closed (bad-app-ui-prefix): must be a non-empty path segment — leading /, no trailing /, length > 1. Unset = today's routes (byte-identical).
  • Web lens: every Leptos Route gains the prefix as its first StaticSegment (path_of(&["commands"]) in main_src); the root / route (content) and the axum REST routes are untouched.
  • The scene pages' internal links keep their absolute paths (they navigate within the UI; the browser URL bar shows the prefixed route).

2. deploy: { ui: true } — ship the web shell, not just the API binary

A k8s spec with ui: true deploys the web crate's binary ({app-name}-web: the full app — API + Leptos admin UI + shell fallback) instead of the API-only binary:

  • The k8s Dockerfile builds --bin {app-name}-web; the entrypoint takes no args (the web shell has no serve subcommand).
  • The Deployment template omits the serve --bind … args and sets MOAIC_BIND_ADDR=0.0.0.0:<containerPort>; the emitted web main.rs reads it (falling back to the leptos options' dev bind, 127.0.0.1:8787) — only when the app declares a ui deploy, so dev and non-ui deploys are byte-identical.
  • values.yaml gains ui: true (only for ui specs) and the template nil-checks .Values.ui.

3. Embedding LRU cache (learned e3dc06369)

The vectordb lens' embed_text is wrapped by a process-local LRU (4096 entries, key = exact text, emitted for every vectordb app): repeated texts — re-ingests of unchanged files, repeated query chunks, the boot-time embed-what-is-missing pass — skip the HTTP round-trip. Fail-soft and side-effect-free (pure cache), so it is emitted unconditionally for vectordb apps rather than behind an opt-in.

4. Qdrant template fixes (learned d2bc23620 + 0a7089324)

  • Image tag is qdrant/qdrant:v1.12.5 — the tag with the v prefix is the real Docker Hub tag (the bare 1.12.5 does not exist; latest would also break the chart's semver check).
  • The Service + container expose gRPC 6334 alongside REST 6333: gRPC clients (tonic) connect on 6334 — 6333 alone never serves them.
  • (P25, deferred: the qdrant client itself — including EE's arbitrary-doc-id → valid-point-id mapping from 480c9a3cc.)

5. Knowledge data-seam runtime-context fix (found while verifying)

The embedded tier's smoke test exposed a pre-existing crash: the knowledge data seam (mosaic-knowledge/src/data.rs) drives OpenDAL's async API with its own dedicated tokio runtime via block_on. Runtime::block_on panics ("Cannot start a runtime from within a runtime") when the calling thread is already driving a runtime — which is exactly where a host runs the engine's synchronous boot (the app's worker / spawn_blocking threads carry the host runtime's context). Symptom: cold boot died right after ingest (or hung on a worker), in any persist backend — the file/JSON seam included, since plain paths go through the same OpenDAL seam.

Fix: run_op — when Handle::try_current() is Ok (caller inside a runtime), the future is handed to a dedicated context-free worker thread (same mpsc pattern the host uses for its LanceDB ops); otherwise block_on directly (the cheap path: plain threads, tests).

Consequences

  • Small/medium-tier apps and any spec can now serve the admin UI on the same host as their content (/_), no extra ingress or path rewrite.
  • The data seam is context-safe in every calling context (the claim its old doc-comment made but did not keep).
  • The qdrant store template matches what a gRPC-speaking client actually needs.

ADR 0026 — EE learn batch: deploy resources, token-borne realms, guard baseline, backend boot check

From a full comparison of the EE mono (DSL + codegen framework) and the generated ee-pages portal against mosaic, plus the latest EE commits (a96bb8fe3, 8f2cdb4, b0f6521, d59df40, 5fa7bc9, 41011d028, b593e7a4b, 9b7ced46d). This ADR covers the four codegen improvements landed from that review; ADR 0024/0025 cover the earlier batches.

Already in mosaic (no work): DSL-first/codegen-first doctrine (tessera is the single source of truth; every surface is a projection over it), the single-collection source-scoped search (5fa7bc9 equivalent — mosaic sources carry a source filter on one collection), DSL-driven RAG config (d59df40 equivalent — retrieval.mode/top_k + chunk in the tessera), the native RuleSet DSL (41011d028/b593e7a4b equivalent — mosaic has reactive + callable + dynamic rulesets), on_create auto-grants, and the 4-crate generated layout (app, web, web/ui, knowledge-engine — the crate-count discipline).

1. deploy.resources { cpu, memory } — pod sizing in the tessera (EE a96bb8fe3)

EE learn: ee-pages was OOMKilled mid-RAG-indexing at the chart's hardcoded 512Mi limit with no DSL path to fix it. The fix added deploy: { resources: { cpu, memory } }; a declared quantity tunes both requests and limits (Guaranteed QoS), per-field fallback to chart defaults, and omitting the block renders byte-identically.

Mosaic implementation.

  • Tessera: deploys: [{ resources: { cpu: "500m", memory: "512Mi" } }] (AppDeploy.resources: Option<ResourceSpec> — the shared spec type). Both fields required (a partial spec is a parse error, not a silent default). Fail-closed validation on Kubernetes quantities (bad-deploy-resource-quantity, the legacy is_quantity check).
  • Deploy lens: the tessera Helm chart values gain resources: {cpu, memory} (still resources: {} when omitted — byte-identical); the deployment template emits the requests+limits block under {{- if .Values.resources.cpu }}.
  • The li7 spec (the public portal) declares 500m / 2Gi — its boot-time RAG index of the books corpus is exactly the EE failure mode.

2. Realms from the verified token — never from the request (EE ScopedCaller)

EE learn (8f2cdb4 tenant-less/app-identity rework + the ScopedCaller contract): scope/tenant values are resolved from session/ JWT state, never from the request body or headers — a caller may not claim a tenancy by sending a header. Mosaic's realm plumbing read x-realm-{dim} straight from the request headers: any caller could claim any tenant.

Mosaic implementation.

  • Tessera: identity.users.<name>.realm: { dim: value } — the user's declared realm assignments (UserDecl.realm).
  • Codegen (shared auth_handler, so app + web get it):
  • the USERS table carries a per-user realm_json;
  • issue_jwt stamps a realm claim (JSON object) when the user has one — the login is the only mint site, so every local token carries it;
  • verify_token (local JWT and OIDC id_token) parses the claim into Authed.realm;
  • the default-deny middleware re-stamps x-realm-{dim} from the verified claim over any client-declared header — downstream scope resolution (caller_realm/realm_filter) is unchanged but can no longer be spoofed. Realm-less tokens (no claim) keep the header/env/seed resolution (the dev realm bar keeps working).
  • The e2e harness signs scenario tokens with the declared user's realm claim (E2ePlan.users), so token-realm behavior is testable without a live login.

realm: { tenant: "*" } is the cross-tenant principal: realm_filter treats the * value as a wildcard (no prefix narrowing).

3. AUTHZ_GUARDS.tsv — the guard baseline as a generated artifact (EE route-guard-baseline.tsv)

EE learn: every authorized route is inventoried in a machine-generated TSV (block/mechanism/method/mount/guard), checked in; a CI drift check treats losing a guard as a regression ("adding routes is fine, moving a guard is reported, losing one is the failure").

Mosaic implementation.

  • App lens: when identity is declared, the render emits AUTHZ_GUARDS.tsv at the project root — one line per (mount, method) with its default authorization decision (action/resource/anonymous, - = public), generated from the same route inventory that feeds the router and the route_policy table.
  • Conformance: the file is pinned under golden/ (a lost guard is a byte diff) and a semantic test (authz_baseline_covers_declared_endpoints) asserts every declared rest endpoint appears with its auth class — required → non-anonymous decision, optional → anonymous admitted, none → public. A codegen change that silently re-classifies a route fails the suite even when the TSV bytes happen to match.

4. MOSAIC_VECTOR_BACKEND boot check — no silent degradation (EE b0f6521)

EE learn: EE's vectordb registry rejects any backend name absent from the compiled-in factory list at boot ("unsupported vector backend qdrant — supported values in this build: …") — a deploy env typo must fail loudly, not degrade to a weaker store.

Mosaic implementation.

  • Generated check_vector_backend_env() (app + web binaries, whenever the app has a vectordb): MOSAIC_VECTOR_BACKEND must be ""/file (always) or lancedb (only in builds with the lancedb feature — the compiled-in list is printed in the error); anything else exits the boot with code 2. Previously any unknown value silently degraded to the ephemeral file store — invisible data loss across pod restarts.

5. mosaic-pages as the reference app (all features exercised)

The portal tessera now exercises the full platform surface:

  • Realms: tenant dimension (non-global root, seeded acme); Doc and DocLog are realm-scoped (realm: [tenant] — instance ids carry the tenant prefix). Users: admin (cross-tenant *), owner/viewer (acme), owner-globex (globex).
  • ABAC: Doc.Retire carries where: 'subject.role == "admin"' — owner passes the role ladder + the doc.publish permission and is still denied by the attribute condition.
  • Rulesets: reactive doc-published / doc-retired (event → dispatch → the DocLog aggregate, which has no REST surface — written only by the rulesets) and callable doc-review-policy (decision table over the request body at POST /api/rulesets/doc-review-policy).
  • Deploy resources: li7 declares 500m / 2Gi.
  • e2e (26 scenarios): tenant isolation, token-realm-beats-header, cross-tenant admin, seeded-tenant anonymous reads, the ABAC deny, rule-set audit rows, and the decision table.

Notes

  • EE's live permission model (per-request IAM lookup instead of role-in-token) and its OIDC realms are deliberately not ported: mosaic's model is role-in-token + resource-scoped grants + ABAC, which covers the portal's needs; EE's realm is an OIDC-issuer concept, mosaic's realm is the tenancy dimension.
  • EE's memvid capacity self-grant (9b7ced46d) does not apply (mosaic uses sqlite/lancedb/opensearch/qdrant/clickhouse, not memvid).

ADR 0027 — Surface auto-derivation: zero unexposed CQRS commands

Ports the core of EE's std.derive-cqrs-services + std.derive-bridges behavior into mosaic's planner: a CQRS command is a first-class surface. Once it exists in the aggregate, it is addressable as REST, MCP, and CLI unless the author explicitly opts out or has already claimed the surface.

Context

Mosaic's surfaces are projections over the tessera model. Before this ADR, a command was only reachable on the outside when an authored slot item pointed at it:

  • rest.endpoints with source: cqrs … for a REST route;
  • mcp.tools with a delegate for an MCP tool;
  • cli.commands with a delegate for a CLI verb.

That is the right model for custom mounts, custom methods, and curated tool names — but it made the common case verbose. A new command silently had no external surface until the author remembered to add the three bridge items. EE solved the same problem with derivation blocks: the CQRS service and its bridges are generated from the command declaration, and authored items are overrides, not prerequisites.

Decision

1. expose: bool on a command (default true):

aggregates:
  - name: Order
    commands:
      - name: PlaceOrder
        about: "Place an order"
        fields: …
      - name: InternalReconcile
        expose: false   # internal only — no derived surfaces
        fields: …

expose: false is EE's visibility: internal parity: the command still runs inside workflows/reactors/rule sets, but the planner derives no external surface for it. An authored slot item can still expose such a command explicitly.

2. The planner derives the three standard surfaces for every exposed command, after authored slot items have been collected:

surface derived name / mount claim rule
REST POST /api/{aggregate-kebab}/{command-kebab} skipped when an authored endpoint already owns that (method, mount)
MCP tool named {command-kebab} skipped when an authored tool already delegates to the same command; a different owner of the name is a warning
CLI command named {command-kebab} skipped when an authored CLI command already delegates to the same command; a different owner of the name is a warning

The derived items carry the command's about text and delegate to (aggregate, command), so they render through the exact same code paths as authored items (route registration, MCP schema, CLI clap command, CLI e2e test, OpenAPI, site command pages, TUI inventory).

3. Authorization rides the delegate. The derived REST endpoint is resolved with resolve_endpoint, so the command's own min_role, permission, and where (ABAC) conditions apply exactly as for an authored endpoint. There is no separate authz story for derived surfaces.

4. Collisions are visible, not silent. A cross-aggregate kebab collision (two Submit commands) warns surface-collision on the MCP/CLI name and keeps the first owner. The REST mount cannot collide (it is namespaced by the aggregate kebab). The warning tells the author to rename a command or set expose: false on one of them.

5. The MCP surface filter still wins. app.mcp.hide / app.mcp.expose apply after derivation: a derived tool can be hidden or left out of a whitelist like any authored tool.

Consequences

  • Zero unexposed commands is the default posture: declaring a command is enough to get POST /api/{agg}/{cmd} + an MCP tool + a CLI verb. The examples/full-stack and examples/data-sync goldens now contain derived routes/tools/commands with no authored bridge items.
  • Authored surfaces remain the override mechanism. A custom mount (/api/orders/place) or a renamed tool can claim the command's surface; the planner then does not duplicate it.
  • expose: false gives internal commands a clean opt-out without deleting their in-process dispatch.
  • The derived surfaces are deterministic (BTree-ordered aggregates/commands), so conformance pins them byte-for-byte.

ADR 0028 — Browser SSO (GitHub/Google) + grounded chat context

Two identity/AI improvements for generated apps:

  1. a browser SSO login flow (authorization-code, GitHub and Google) alongside the existing local users and passive OIDC id_token verification;
  2. conversation memory, page grounding, and citation rewriting for the built-in web chat and the vectordb RAG assistant.

1. Browser SSO

Context

Mosaic's identity stack already verifies local JWTs (password login) and, when declared, OIDC id_tokens against a JWKS/public key. But a public portal has no browser login path for an external IdP: no sign-in page, no code exchange, no session cookie. EE's portal uses external SSO for exactly this shape.

Decision

Tessera: identity.sso is a list of providers:

identity:
  sso:
    - provider: github          # github | google
      client_id_env: MOAIC_GITHUB_CLIENT_ID
      client_secret_env: MOAIC_GITHUB_CLIENT_SECRET
      scopes: [read:user, user:email, read:org]
      role_map:
        default: viewer         # fail-closed floor
        rules:                  # first match wins
          - role: admin
            all:
              - claim: login
                in: [eugeis]
          - role: editor
            all:
              - claim: orgs
                contains: mobility-devops
    - provider: google
      client_id_env: MOAIC_GOOGLE_CLIENT_ID
      client_secret_env: MOAIC_GOOGLE_CLIENT_SECRET
      role_map:
        default: viewer
        rules:
          - role: admin
            all:
              - claim: email
                in: [admin@example.com]
  • Supported providers are github and google; anything else is a plan error (sso-provider-unsupported). Duplicates are rejected.
  • Secrets are env-var names only (client_id_env optional, client_secret_env required). The secret itself never appears in the DSL. Missing configuration is a plan error, not a boot surprise.
  • role_map.default is required and must be a role known to the app's IAM model. Rules are evaluated in order; the first rule whose predicates all pass wins. Predicates: claim + exactly one of eq, in, suffix, domain, contains (array claims support contains). A rule may also stamp a realm: {dim: value} onto the issued token.

Generated routes (mounted only when identity.sso is non-empty; the /auth prefix is public):

route behavior
GET /auth/login provider button page
GET /auth/{provider}/login?next=/path starts the authorization-code flow (UUID state, 300 s TTL, in-memory)
GET /auth/{provider}/callback exchanges the code, fetches provider claims, resolves role/realm, issues a local JWT, sets the mosaic_token cookie, redirects to next
GET /auth/logout clears the cookie and redirects to /auth/login

Provider claim fetch:

  • GitHub: /user, verified primary email from /user/emails, org logins from /user/orgs → claims login, id, name, email, orgs[].
  • Google: OIDC /v1/userinfo → claims sub, email, name, …

The issued JWT carries sub = "{provider}:{provider-subject}", role, optional realm, email, name, and provider. The auth middleware accepts the token from Authorization: Bearer … or the mosaic_token cookie (HttpOnly, SameSite=Lax, Secure under x-forwarded-proto: https). GET /api/auth/me now returns the verified claims instead of trusting x-authz-* headers.

Fail-closed runtime states (pinned by examples/authorized e2e): unknown provider → 404, missing client id → 503, invalid/expired state → 400.

2. Grounded chat context

Context

Both chat surfaces were stateless single-turn calls. EE's agent (f2843c02a) added page grounding via a ChatContext and a configurable chat memory window. The mosaic sync note called this "no gap" because mosaic's chat is generated; this ADR revisits that decision and ports the useful parts.

Decision

Web chat (POST /api/chat, SSR surface) accepts:

{
  "message": "…",
  "conversation_id": "c-…",
  "page_id": "order",
  "citations": {"[1]": "https://example.com/a"}
}

The generated chat module keeps an in-memory per-conversation history (capped at 200 messages) and builds a window from MOAIC_CHAT_HISTORY_{TURNS,MAX_CHARS} (defaults 10 / 24 000). When MOAIC_CHAT_HISTORY_SUMMARIZE=true, dropped older turns are summarized via the configured chat model (MOAIC_CHAT_HISTORY_SUMMARY_MAX_CHARS, default 4000) and injected as an Earlier conversation summary: system message. A non-empty page_id adds a page note to the system prompt; citations are deterministic token→URL rewrites applied to the final answer (LLM and no-LLM fallback). The chat page persists a conversation_id in localStorage and sends it with each message. The response includes conversation_id.

Vectordb RAG assistant (POST /api/assistant/chat, app lens) accepts the same optional fields. Non-streamed and streamed turns now:

  • load the conversation history (MOAIC_ASSISTANT_HISTORY_*, same defaults);
  • optionally summarize dropped turns;
  • retrieve a second page_sources set for page_id and add [page N] context plus the same page note;
  • send the full message list (system + history + question) to the LLM;
  • rewrite citations in the final answer (LLM, streamed, and grounded no-LLM fallback);
  • remember the turn when a conversation_id was sent.

The streamed response emits sources, then page_sources, then delta events, and the final done event carries {answer, sources, page_sources, llm, conversation_id}.

The e2e scenario assistant_chat_accepts_conversation_page_citations pins the no-key fail-soft path: the grounded context is returned and the citation token is rewritten to the caller-supplied URL.

3. mosaic test binary path

The e2e harness spawns the server with the rendered output dir as its working directory. A relative binary path resolved against that cwd, so the harness now canonicalizes the path before spawn.

Consequences

  • Public apps get a real browser sign-in/out flow with provider-specific role mapping, while local users and OIDC id_token verification keep working.
  • All SSO secrets remain deploy env; the DSL carries only provider names, env var names, scopes, and role rules.
  • Chat memory is process-local (in-memory). A multi-replica deployment needs an external store for shared conversations — a follow-up, not part of this ADR.
  • Citation rewriting is deterministic string substitution over caller-provided tokens; it is not a semantic doc→URL resolver.
  • SSO is intentionally limited to GitHub and Google. A generic OIDC authorization-code flow is a possible extension, but the two concrete providers keep the claim-fetch and role-mapping surface small.

ADR 0029 — App variants, OpenAPI proxy apps, domain-service surfaces, and Mosaic Aide

This ADR lands the “Mosaic beyond EE” gap program:

  1. Mosaic Aide — an out-of-the-box named assistant for every app that has a vector_db (knowledge + MCP + voice support).
  2. Modern OpenAPI proxy apps — mosaic import openapi now emits a tessera app that delegates every imported operation to the upstream, with optional cache and knowledge-capture interceptors, and exposes the same operations as MCP tools.
  3. App variants / composites — app.variant + app.variants are named, composable overlays (surface, realm, feature, and part-instance composition).
  4. Domain-service surface derivation — a domain services[].delegate (Aggregate.Command) derives its own REST endpoint and MCP tool.
  5. Dashboard UI — the web lens gains a /dashboard page (KPI cards, event mix, recent activity) on top of the existing list/detail/events/ workflow-DAG pages.

1. Mosaic Aide

Context

EE’s “Eezy assistant” is a product-level assistant that wraps the app’s knowledge, tools, and voice. Mosaic already generated a strong RAG assistant (POST /api/assistant/chat, /api/aide-less page, assistant_chat MCP tool) and a separate /voice widget, but the assistant had no stable product name and no first-class profile/page surface.

Decision

When an app declares at least one vector_db, the app lens emits:

surface behavior
GET /api/aide/profile returns Mosaic Aide, the active model name, and voice: true
POST /api/aide/chat aliases the existing grounded assistant (assistant_chat)
GET /aide a self-contained assistant page (conversation id in localStorage, page context omitted for now)
aide_chat MCP tool a model-facing tool over /api/aide/chat (query, conversation/page/collection/top_k/stream)

Mosaic Aide is the assistant brand. It reuses the existing assistant runtime (history, page grounding, citations, vectordb RAG, no-LLM grounded fallback) and the existing voice surface; it is not a second engine.

The route is app-lens only (the web lens keeps its own /api/chat surface and does not emit the Aide page route).

2. Modern OpenAPI proxy apps

Context

ADR 0006 made the legacy kind: proxy tessera a byte pass-through (no enrichment). That was the right boundary for the legacy service/proxy path, but the user request asks for an imported OpenAPI service to become a Mosaic app: delegate calls, add interceptors (cache, knowledge), and expose MCP tools for every endpoint with the app’s assistant available.

Decision

mosaic import openapi now emits a modern tessera app (not a legacy kind: proxy document):

app:
  name: petstore
  parts: [openapi-proxy, metrics]
parts:
  - name: openapi-proxy
    contributions:
      - slot: rest.endpoints
        items:
          get_pet:
            mount: /pets/{petId}
            method: get
            auth: none
            source: "proxy get_pet"
            upstream: "https://upstream.example.com"
            cache: 60s
            knowledge: upstream-kb
      - slot: mcp.tools
        items:
          get_pet:
            about: "Call the upstream get_pet operation"
            source: "proxy get_pet"
vector_dbs:
  - name: upstream-kb
    about: "Responses captured by the Mosaic proxy."
    docs: []

Planner rules:

  • source: proxy <endpoint> (or bare source: proxy using the endpoint name) resolves to a ProxyEndpointPlan (upstream, upstream_path, cache_secs, knowledge).
  • upstream must be an http(s) base URL; upstream_path defaults to the endpoint mount and must start with /.
  • cache is an interval (0s disables it); it applies to GET forwards.
  • knowledge must name a declared vector_db; when set and the upstream response is successful, the response body is ingested into that collection (vector_ingest).
  • A proxy endpoint cannot be a websocket relay.
  • An MCP tool source: proxy <endpoint> must reference a declared proxy endpoint (mcp-proxy-unknown otherwise).

Generated runtime:

  • The proxy handler reads the raw axum::extract::Request, substitutes {param} path segments from the concrete request path, preserves the query string, and forwards the method/body/headers to the upstream.
  • It skips hop-by-hop / client-shaping headers (host, content-length, connection, accept-encoding) and adds x-proxied-by: mosaic + x-cache: HIT|MISS|BYPASS.
  • The TTL cache is an in-process BTreeMap keyed by METHOD upstream_path?query.
  • Knowledge capture is optional and only emitted when the app has vectordbs.

MCP tools are now method-aware:

  • Tool carries method (default POST for CQRS/workflow/ruleset tools).
  • tools/call substitutes {name} path arguments into the mount and sends remaining arguments as the query string.
  • Body is sent only for POST / PUT / PATCH.

This intentionally extends ADR 0006 for the modern tessera proxy path; the legacy kind: proxy document remains byte pass-through.

3. App variants / composites

Context

The user asked for app/composite variants covering realm variants (no realms, tenant, tenant+workspace, workspace-only) and surface variants (REST, REST+MCP, REST+CLI, UI, etc.). Mosaic already had realm_profiles / active_realm_profile, parts, part instances, MCP surface filters, UI layout / prefix — but no named way to compose them.

Decision

app.variants is a map of named overlays, and app.variant selects the active one (default: default when present, otherwise no overlay):

app:
  name: full-stack
  variant: tenant-mcp
  variants:
    base:
      parts: [shared]
    rest-only:
      extends: [base]
      exclude: [mcp-tools, ui-scenes]
    tenant-mcp:
      extends: [base]
      parts: [mcp-tools]
      realm_profile: tenant
      mcp_expose: [place_order, list_orders]
    workspace-ui:
      realm_profile: workspace
      layout: sidebar
      ui_prefix: "/_"

A variant may set:

  • extends: [name…] — composes other variants (later extends entries are lower precedence than the current variant; cycles are a plan error).
  • parts — part instances to include.
  • exclude — part instances to remove (wins over parts).
  • realm_profile — activates a declared realm_profiles entry.
  • features — extra app features.
  • mcp_hide / mcp_expose — MCP surface filter for the variant.
  • layout / ui_prefix — UI overrides.

The planner applies the merged variant before parts, realms, MCP filtering, and UI rendering. AppPlan.active_variant records the resolved name.

This gives both requested axes without changing the existing parts/realm/MCP machinery:

  • realm variants = variants that set different realm_profile values.
  • surface variants = variants that include/exclude the parts contributing the desired surfaces and/or adjust the MCP filter.

4. Domain-service surface derivation

Context

domain.services existed as documentation-grade DDD metadata (delegate: "Aggregate.Command"), but it did not derive runtime surfaces. EE derives domain-service bridges; Mosaic’s equivalent is to make an explicit service boundary a first-class REST + MCP surface.

Decision

A domain service with delegate: "Aggregate.Command" now derives:

  • POST /api/domains/{domain-kebab}/{service-kebab} (unless an authored endpoint already claims that mount).
  • An MCP tool named after the service (unless claimed; collisions warn surface-collision).

The derived endpoint resolves through the same resolve_endpoint path as authored endpoints, so the command’s authz (min_role / permission / where) applies.

5. Dashboard UI

Context

EE ships KPI/dashboard scaffolding and CQRS ops views. Mosaic had list/detail pages, workflow DAG, projection charts, and an event log, but no top-level KPI dashboard.

Decision

The web lens now emits a /dashboard route + nav entry:

  • KPI cards: aggregates, commands, events, workflows, endpoints, CLI commands, schedules, flags, knowledge bases.
  • Event mix: events by aggregate from the live event log.
  • Recent activity: the newest events.

In SSR modes the page reads the in-process store; in CSR it fetches GET /api/events/list. The KPI counts are build-time DSL facts (deterministic golden output).

Consequences

  • Every knowledge-backed app has a stable assistant brand (Mosaic Aide) with knowledge, MCP, and voice support out of the box.
  • Imported OpenAPI services become Mosaic apps: REST pass-through + MCP tools + optional cache/knowledge + the app’s assistant, instead of an opaque proxy tessera.
  • Variants make surface/realm composition declarative and composable without duplicating apps.
  • Domain services become addressable surfaces, not just docs.
  • The dashboard is a read-only KPI/ops view. The full CQRS replay scrubber and interactive aggregate-state explorer remain follow-ups (the event log and aggregate pages already exist).
  • The proxy TTL cache is process-local; multi-replica deployments need an external cache for shared state (a follow-up).
  • Variants are plan-time static; a single build picks one active variant. Runtime variant switching is out of scope.

ADR 0030 — Aspect ops surfaces: schedule next-run, notification center, CQRS explorer, vector-search playground

This ADR lands the remaining U5 aspect surfaces from the EE sync ledger:

  1. Schedule next-run tracking — the scheduler part now exposes live next / last state for every declared schedule.
  2. Notification center — the notify aspect records every dispatch in an in-process ring buffer and exposes it on a /notifications page.
  3. CQRS replay scrubber / aggregate-state explorer — a /cqrs page replays the event log to any position and shows the aggregate state at that position.
  4. Vector-search playground — a /search page with collection, query, top_k, and source controls over the existing vectordb search API.

1. Schedule next-run tracking

Context

schedules: generated tokio tasks that fired commands/workflows, but the /schedules page was a static, codegen-time table. There was no way to see when a schedule would fire next or when it last fired.

Decision

When the scheduler part is active and at least one schedule is declared, the generated server emits:

pub static SCHEDULE_STATE:
    LazyLock<Mutex<BTreeMap<String, Value>>>
  • At boot, schedule_state_seed_rs computes the initial next for every enabled schedule (now + interval, the parsed at instant, or cron_next(...) for cron schedules) and stores last: null.
  • After each fire, the fire body records last = now and advances next (interval: now + interval; at: null; cron: recompute cron_next).
  • GET /api/schedules returns the live state map.
  • The web /schedules page is now live in every hydration mode: SSR reads SCHEDULE_STATE, the hydrated/client bundle fetches /api/schedules. The table shows declared timing/action/params plus next run and last fired.

The full web lens also spawns the scheduler (it serves the same store and API as the app lens), so a ui: true deployment has the same schedule behavior as the API-only binary.

2. Notification center

Context

notifies: generated a log line or a blocking webhook POST per matched event, but there was no stored surface: no API, no UI, no way to inspect recent dispatches or whether a webhook succeeded.

Decision

When at least one notify is declared, the generated server emits:

pub static NOTIFICATIONS:
    LazyLock<Mutex<VecDeque<Value>>>

Each notify dispatch pushes one envelope (capped at 200 entries):

{
  "notify": "<name>",
  "channel": "log|webhook",
  "aggregate": "<aggregate>",
  "event": "<event>",
  "aggregate_id": "<id>",
  "body": "<interpolated body>",
  "ok": true,
  "ts": "<RFC3339>"
}
  • Webhook dispatches record ok from the HTTP send result.
  • GET /api/notifications returns the ring newest-first as { "items": [...] }.
  • The web lens gains a conditional /notifications route + nav entry with a live list (SSR reads the static; the client fetches the API).

3. CQRS replay scrubber / aggregate-state explorer

Context

The /events page shows the event log, and aggregate pages show current instance state, but there was no way to inspect what the state was at an earlier event position.

Decision

When at least one aggregate is declared, the generated server emits:

GET /api/events/state?up_to=N

The handler:

  1. locks the live store,
  2. builds a scratch Store::default(),
  3. replays the first N events through the existing crate::agg::replay_event primitive (the same path used by durable boot recovery),
  4. returns { up_to, total, aggregates: { "<Aggregate>": { "<id>": <state> } } }.

The web lens gains a /cqrs page with a range input over the event count. Moving the slider fetches /api/events/state?up_to=N and renders the aggregate instances at that position. This is a dev/ops surface; it is deliberately O(N) per request and capped by the live event log.

4. Vector-search playground

Context

The /knowledge page already had a minimal search box, but it fixed top_k to 8, had no source control, and showed only a few hit columns.

Decision

When at least one vector_db is declared, the web lens gains a /search route + nav entry:

  • collection selector (one button per declared collection),
  • query input,
  • top_k number input,
  • optional source input,
  • a hits table with score / id / locator / context.

The page POSTs to the existing POST /api/vectordb/{name}/search endpoint ({query, top_k, source?}), so no new search engine behavior is introduced. The api_post helper now returns the parsed JSON body on success (previously it discarded it), which also fixes the CSR knowledge-page search path.

5. Web-lens client runtime

The new pages are interactive in every hydration mode. To keep the generated web lens dependency-light:

  • csr continues to use gloo-net in [dependencies].
  • full / islands gain gloo-net under [target.'cfg(target_arch = "wasm32")'.dependencies] only when an interactive surface is emitted.
  • The pages module emits a small client runtime:
  • spawn_client(f) — runs a spawn_local future on wasm and is a no-op during SSR,
  • api / api_put / api_post — cfg-aware REST helpers,
  • ev_value(&Event) — reads the current value of an input/select event target via web_sys::HtmlInputElement.

SSR branches keep their in-process store/static reads; the client bundle updates the same reactive signals after hydration.

Plan fix

check_platform_parts checked notify against the triggers presence flag (parts.1) instead of the notifications flag (parts.2). This ADR fixes the off-by-one so a notifies: declaration with notifications in app.parts is accepted, while a missing part still fails closed.

Consequences

  • The U5 aspect-surface learn-todo is complete: events, schedules, flags, dashboard, notification center, vector-search playground, and the CQRS replay scrubber / aggregate-state explorer are all codegen-derived.
  • The full web lens now runs the scheduler, matching the app lens for ui: true deployments.
  • The CQRS explorer is a scratch-replay ops tool; it does not mutate the live store and does not change the durable log format.
  • The notification ring is in-process and capped; a durable notification log would be a follow-up if it needs to survive restarts.
  • The search playground is UI-only; it reuses the existing vectordb search API and does not change retrieval behavior.

ADR 0031 — Tenant-scoped Leptos web shell + CSR/vectordb compile fixes

This ADR closes the "tenant scoping not adopted" gap in the web lens and fixes two pre-existing compile bugs in the CSR + vectordb path.

1. Tenant-scoped Leptos web shell

Context

EE's app shell (SlotShell) routes and scopes the UI by tenant. Mosaic already had the backend machinery (realms, x-realm-{dim} headers, token-borne realms, and a realm/tenant bar on the static admin HTML pages) but the Leptos pages (dashboard, list/detail, projections, events, CQRS, schedules, flags, notifications, knowledge, search) did not send the realm header. So switching tenant on an admin page scoped that page, but the Leptos pages kept showing the default (seeded) tenant — an inconsistent, partially-scoped shell.

Decision

When the app declares realms, the Leptos web shell is now tenant-scoped:

  • Dimension selection — the shell scopes by the root realm dimension (e.g. tenant), else the first declared dimension, else the first aggregate's realm dimension. This mirrors how an operator thinks about "which tenant am I in?" and matches the admin realm bar's source.
  • Tenant bar — the shell's nav-actions gains a realm input + apply button (only when a realm dimension exists). The selected tenant is persisted in localStorage under mosaic-realm-{dim} and restored on load.
  • Realm header — every client fetch helper (api, api_put, api_post) stamps the x-realm-{dim} header from the stored tenant, so all Leptos pages are scoped to the selected tenant in every hydration mode (CSR, and the hydrated full/islands browser bundle).
  • Apply reloads the page so the new tenant's data is fetched fresh (the SSR branches keep reading the in-process store, which is already realm-filtered server-side).

The realm machinery is emitted only when a realm dimension exists and is #[cfg(target_arch = "wasm32")]-gated, so non-realm and server-rendered builds are byte-identical.

2. CSR + vectordb compile fixes

Two pre-existing bugs blocked any CSR app with a vector_db from compiling (no example exercised that combination until now):

  1. The knowledge page's search closure was let do_search = move || { … } (0-arg) but used as on:click=do_search (1-arg). Fixed to move |_|.
  2. The knowledge page's collection selector used <select value=kb_sel …>, but Leptos does not support a value binding on a <select>. The select now relies on native display + on:change to sync the signal that the search reads.

Both are covered by rendering a CSR + vectordb (+ realms) app and compiling its web lens for wasm32-unknown-unknown in local validation.

Consequences

  • The Leptos web shell is now consistently tenant-scoped, matching the admin pages and EE's tenant routing. This closes the "tenant scoping not adopted" item in the app-shell feature-map row.
  • The six-position SlotShell itself is still not adopted (Mosaic uses a 4-position switchable shell) — an intentional design difference, not a functional gap.
  • CSR apps with a knowledge base now compile and run the knowledge/search pages.
  • Tenant selection is per-browser (localStorage); a server-side session-scoped tenant would be a follow-up if multi-tenant SSO sessions need it.

ADR 0032 — Six-position shell slots + split/tabs page templates

This ADR closes the remaining app-shell parity gap flagged by ADR 0031 ("the six-position SlotShell itself is still not adopted") and adopts EE's page templates (split_page, tabs_page) as a DSL choice. Where Mosaic's existing surfaces already do more, they are kept and extended (the switchable shell layout, projection charts on the split page) instead of replaced.

1. The shell is a six-position slot system

Context

EE's SlotShell (blocks/ui/ee-theme) is a frame with named injectable positions — header_left, header_right, sidebar_top, header_tools, tenant_selector, sidebar_scenes, footer, events_panel, help_panel — and the app frame composes them. Mosaic's shell had a fixed content model (brand, nav, actions, footer) with no declarative way to inject content into positions.

Decision

The generated web shell is now a six-position slot system, with the switchable layout (topbar/sidebar/bottom/floating, ADR 0022-era data-layout) kept on top of it — the layout moves the nav between positions, the slots fill them:

  • header-left — the brand (declared app.ui.brand or the app name).
  • header-center — a filter navigation search box (client-side JS filters the scene links and hides empty scene groups) plus any declared app.ui.slots entries for this position.
  • header-right — the nav actions (SSO sign-in/out, the P31 realm bar, theme toggle, layout toggle).
  • menu-top / menu-bottom — declared app.ui.slots links placed at the top/bottom of the nav.
  • footer — the P31-era footer (copyright + links).
  • events/timeline slot — when the app has aggregates, a live ticker over the CQRS event log (SSR paints from the store; spawn_client refreshes via GET /api/events/list) with a timeline link into the /cqrs replay scrubber — the full timeline-toolbar parity (EE's sticky scrubber island is that page itself).

New DSL: app.ui.slots —

app:
  ui:
    slots:
      - position: header-center   # header-center | menu-top | menu-bottom
        label: knowledge
        path: /knowledge          # internal path or external URL (new tab)

Positions are validated at plan time (bad-app-ui-slot). EE's tenant_selector maps to the P31 realm bar, sidebar_scenes to the scene nav groups, and help_panel to /docs + the nav search — no separate positions are emitted for them (their content already lives in a slot).

2. split and tabs page templates

Context

EE ships page templates in blocks/ui/ee-ui/src/components/pages/: list_page (toolbar + list), detail_page (back nav + optional sidebar), dashboard_page (KPI grid), split_page (master-detail: 360px master column + detail pane), tabs_page (header + tab panels). Mosaic already generates the list/detail/dashboard shapes from ui: hints; the split and tabs shapes were missing.

Decision

ui.template on a projection selects the list page shape:

projections:
  - name: Tickets
    fields: [ … ]
    on: [ … ]
    ui:
      template: split   # split | tabs — absent → the classic list page
  • split — master-detail: the rows card (filters + sortable headers + reactive table, unchanged) becomes the left master column (380px grid track, 1fr detail track; stacks below 768px); each row is clickable (row-selectable) and selecting one renders the row inline in the right pane as a label/value card (title = the row key), so the detail view exists even without ui.detail_page. Charts stay above the split — EE's split_page has no charts; Mosaic keeps them.
  • tabs — rows / charts as tab panels (a tab bar only when charts exist; each panel is a Show over a tab signal). This is EE's tabs_page shape with Mosaic's reactive table and SVG charts as panels.

template is validated at plan time (split | tabs); both compile in every hydration mode (verified for wasm32-unknown-unknown and native SSR).

Incidental fix

A projection page whose list is off (ui.list_page.enabled: false) with no charts and no admin commands rendered an empty <Layout> — a required-children component, so the generated app failed to compile. Such pages are now suppressed, and declared charts render even when the list is off (charts are a list-page hint, not the list). No example hit this before P32 (no example declared list_page.enabled: false).

Consequences

  • The app-shell feature-map row (codegen.web in sync/ee.md) moves from "partial (slot shell not adopted)" to landed-plus-more: the six-position slot system with the events/timeline slot, tenant scoping (P31), and the four-position switchable layout on top.
  • split/tabs are opt-in per projection; the classic list page stays the default, so existing apps render unchanged.
  • help_panel remains a /docs + search combination rather than a dedicated shell position — revisit if an app needs in-shell help content.
  • The EE transport.rs ClientTransport seam (InProcess/Browser/LocalStore/ Replay) stays mapped to Mosaic's cfg-aware api/api_put/api_post fetch helpers (P30) — a single send() entry point is not worth a rename while every call site is generated code.

ADR 0033 — Browser voice (Web Speech) on the built-in chat

This ADR closes the last "partial" row in the sync/ee.md voice feature map: EE's browser voice on the chat scene (blocks/ai/ee-chat/ui/src/voice.rs) — Web Speech dictation into the chat input plus speechSynthesis playback of replies, with zero backend. It ports that surface onto Mosaic's built-in /chat page and, where Mosaic can do more for free, does.

Context

Mosaic already ships two voice surfaces, both on the server side:

  • platform.voice (V1/V2) — the WSS /api/voice conversation endpoint (pipeline: server VAD → STT → LLM + search tool loop → per-sentence TTS; realtime: provider-direct relay) and the /voice widget. This is a full voice conversation channel and requires provider env keys.
  • Browser voice (this ADR) — the other half of EE's voice story, from ee-chat/ui/src/voice.rs: the browser's built-in Web Speech engines (Chromium SpeechRecognition for dictation, speechSynthesis for playback) wired into the ordinary text chat. No backend, no provider key, no new transport — a transcript flows into the same input field as typed text and is sent through the same /api/chat call.

EE implements it as a Rust module of hand-written wasm_bindgen extern blocks (the Web Speech API is not in web-sys), cfg-gated so the SSR half compiles to no-ops, with capability probed at interaction time so SSR and client markup never carry a capability bit.

Mosaic's built-in chat surface (/chat + /api/chat, emitted by mosaic-render::chat in the full/islands hydration modes) is a single static HTML page with a vanilla-JS client — there is no Leptos chat island to host a Rust voice module in. The Web Speech API is natively available to that page's JavaScript, so the port is the same behavior expressed one layer closer to the browser: plain JS, no wasm, no bindings, SSR-safe by construction (one static document; the server renders no capability bit).

Decision

The generated /chat page now carries the browser-voice client:

  • Dictation (STT) — a mic button in the chat form (hidden when SpeechRecognition/webkitSpeechRecognition is absent). One utterance per press; the button is a toggle (press again to stop). interimResults stream the growing phrase into the input live; final results commit into the input, which the user reviews and sends like typed text — the transcript never takes a different path. Recognition is locale-aware (navigator.language). While recording the button pulses in the --destructive color; onend (any reason) resets it; a denied microphone (not-allowed/service-not-allowed) surfaces a one-line hint in the log instead of failing silently.
  • Playback (TTS) — a "speak replies" toggle above the log, persisted in localStorage (mosaicChatTts), hidden when speechSynthesis is absent. When on, each assistant reply is spoken on arrival (cancel() + speak(), one utterance at a time, mirroring EE's speak()). Additionally, every bot message gets a small "speak" link that reads that message aloud on demand, independent of the toggle — the "more and better" part (EE only auto-speaks replies).
  • Capability at client time — the probes run in the page's own script (buttons start hidden; the client unhides what the browser supports). The served document stays a single static byte string, so there is no SSR/client markup divergence to mismatch, which is the property EE's interaction-time probing buys for Leptos.

No DSL key: like EE, browser voice is part of the built-in chat surface, always compiled in. Apps that own /chat (app_owns_chat, ADR 0025) keep their own page and are untouched. The platform.voice WSS surface is orthogonal and unchanged.

Files

  • crates/mosaic-render/src/chat.rs — chat_page_rs (mic button, TTS toggle, the Web Speech client) and chat_css (controls, recording pulse, "speak" link, sys hint).
  • crates/mosaic-render/src/web.rs — chat_page_carries_browser_voice render test.
  • Re-pinned goldens: full-stack web/src/chat_page.rs + styles.css in the four web lenses; removed stale U1-era chat.rs/chat_page.rs files from the CSR goldens (data-sync, orders-app) that predate the hydration-mode split and were no longer part of the rendered set.

Consequences

  • Browser voice now needs no provider keys and no WSS: a Chromium browser gets dictation into the assistant chat out of the box; every modern browser gets reply playback. Non-Chromium browsers simply hide the mic.
  • Dictation is per-utterance and manual (press to start, the transcript lands in the input, the user sends) — the same review-before-send posture as EE. A "push-to-talk auto-send" variant is a client-only tweak, not an API change, if ever wanted.
  • The generated page grows ~100 lines of static JS; the /api/chat contract is unchanged (voice output is just more text in message).

ADR 0034 — Interactive workflow canvas

This ADR closes the "partial" row in the sync/ee.md feature map for graph/diagram editors: EE's ee-flow is a React-Flow–parity canvas+DOM hybrid (drag, zoom, undo/redo, minimap, auto-layout, context menus, collab transport), and Mosaic's U4 answer was a read-only workflow DAG — static longest-path layering at codegen rendered as SVG. This ADR makes that graph an interactive canvas and, where Mosaic's compile-time model lets it do more for free, does: a live run tracer bound to the L7 streamed run endpoint and the L5 run history.

Context

EE workflows are runtime artifacts: the WorkflowGraphEditor island (blocks/workflow/ee-workflow-ui) wraps ee_flow::FlowCanvas to author workflows visually — node palette, connections, a property panel, and save/load against the Workflow aggregate. The canvas edits the model.

Mosaic workflows are declared in the tessera (tessera.yaml, the single source of truth, ADR 0008) and compiled into the app at build time. A visual DSL editor would be a parallel authoring path — a second source of truth for workflows — against this repo's core doctrine. What the declared model does have that EE's authoring canvas does not bind to: real runtime data. Executions stream node_start / node_output / done over POST /api/workflows/{slug}/run?stream=true (L7) and every top-level run is kept in the run history (GET /api/workflows/{slug}/runs[/{id}], L5) with a full per-node trace.

So the Mosaic canvas is the interactive viewer/tracer for the declared model: every interactive affordance of ee-flow that makes sense for a fixed graph (pan/zoom, node drag, undo/redo, auto-layout, minimap, inspection) plus live execution binding on top.

Decision

Every workflow page (/workflows/{slug}) now emits its DAG as before — the codegen-time layout is still the static SVG, so the page is fully readable with no JS — but wrapped in a canvas container:

  • .wf-canvas container — carries the declared graph as a data-wf payload (nodes with their full detail for the inspector — title, about, code, input/output fields, params, exec block, free-form props — edges with conditions, and the codegen layout as initial positions). The container is enhanced in place by web/static/wf_canvas.js, a single static script (served exactly like dispatch.js/layout.js) that upgrades every .wf-canvas in the document and a MutationObserver that catches ones rendered later (CSR route navigation). The page loads the script idempotently through a generated mount_wf_canvas_js() (wasm-only, no-op in SSR).
  • Pan / zoom — drag the background to pan, wheel to zoom at the cursor (clamped 0.2×–4×), toolbar +/− and fit-to-view (0), double-click to fit. A dot-grid background pans and zooms with the world (SVG pattern — one DOM tree, no separate canvas layer).
  • Node drag — nodes reposition freely; positions persist per workflow in localStorage (mosaic-wf-<slug>) and survive reloads. Edges re-route live (same bezier + arrow as the static render).
  • Undo / redo — an operation stack over layout actions (drag, auto-layout, reset), with Ctrl+Z / Ctrl+Shift+Z / Ctrl+Y and toolbar buttons.
  • Auto-layout — the same longest-path layering the codegen computes (layers as columns, declaration order within a layer, cycle guard), ported to the client as a one-click re-arrange; "reset" returns to the build-time layout.
  • Minimap — corner overview with a viewport rectangle; click/drag to recenter.
  • Inspector — click a node for its id, type, about, run state, last output, code, inputs, outputs, params, props and exec config; click an edge for its endpoints and condition; Esc or a background click closes it.
  • Live run tracing (the "more and better" part) — a run panel with one input per start-node variable (typed: numeric/bool coercion, defaults, required) or a raw JSON body editor when the workflow takes no declared inputs. "run" POSTs the streamed run endpoint and parses the SSE frames from the fetch body reader (POST + SSE; EventSource cannot POST). While running, each node gets a state ring — amber pulse while running, blue when done, red on failure (the last started-but-unresolved node) — and per-node outputs land in the inspector. The panel also lists the ten most recent runs from the run history (✓/✗, age, duration); clicking one replays its trace onto the canvas.
  • Keyboard — +/-/0 zoom/fit, undo/redo, Esc; ignored while an input has focus.

No DSL key: like the static DAG before it, the canvas is part of the built-in workflow surface and is emitted whenever the app declares workflows. The tessera stays the only workflow definition; nothing the canvas does writes back to the model. The WSS voice and run APIs are unchanged — the canvas is a client over existing endpoints.

Files

  • crates/mosaic-render/src/wf_canvas.rs (new) — wf_canvas_js() (the static enhancer) and canvas_payload() (the per-workflow data-wf JSON from WfPlan + the codegen layout).
  • crates/mosaic-render/src/web.rs — workflow-page codegen (canvas container, toolbar, inspector, minimap, run panel, the data-wf payload const, mount_wf_canvas_js), the canvas CSS block, and the workflow_page_is_an_interactive_canvas render test.
  • Re-pinned goldens: full-stack web/src/pages.rs + new web/static/ wf_canvas.js; styles.css in the four web lenses.

Consequences

  • The workflow page goes from a static diagram to an interactive canvas with zero backend changes: every feature (pan/zoom/drag/undo/minimap/inspector) works offline against the declared model, and the run panel works against the existing L5/L7 endpoints.
  • The no-JS path is preserved: the SSR HTML is still the complete read-only SVG (longest-path layout, tooltips, conditions on hover).
  • Layout edits are per-browser (localStorage) and session-visual — they are never a definition change, so no second source of truth for workflows is introduced. A future "persist layout app-wide" would be a one-endpoint extension, not a model change.
  • The generated page grows a ~15 KB static JS file (one per app, emitted only when workflows exist) plus per-workflow payload consts.
  • The canvas deliberately does not port ee-flow's authoring features (node/edge creation, property editing, save-to-aggregate) or its collab transport: authoring the tessera is the tessera's job, and Mosaic's run data replaces the value a second authoring surface would add.

ADR 0035 — Latest-changes analysis for git knowledge sources (K11)

  • Status: accepted
  • Date: 2026-10-04
  • Slice: P35 (knowledge roadmap K11)

Context

The Eezy-parity feature set (the EE AI assistant with MCP support, git repository indexing, source-code symbol recognition, source analysis, repo statistics, MCP tools) is already landed in Mosaic across the K1–K10 knowledge series and P29 (Mosaic Aide). The one remaining capability the feature list names is analyzing the latest changes: the K2 git facts are static (branch, commit count, last short-sha + date, contributor count). Nothing exposes the recent commit history, the files each commit touched, or the diff magnitude — and the assistant (Aide) cannot answer "what changed recently in this repo?" because no change data is indexed.

The knowledge engine already has the pieces this needs: the git CLI seam (K2 — fail-soft, pure parse layer + thin command wrapper), the ingest pipeline (K1/ADR 0019 delta with per-source manifests), the per-source stats row (which already carries the live git facts), and the MCP tool conventions (K1–K6).

Decision

K11 = latest-changes analysis, as a strategy layer (default on) on every git work-tree source — declared sources[].git, a plain path that happens to be a work tree, or a runtime-added git source (K9):

  1. Engine (mosaic-knowledge): - git.rs gains ChangedFile { path, status, added, deleted }, RecentCommit { sha, date, author, subject, files }, a pure parse_recent_commits(name_status, numstat) -> Option<Vec<RecentCommit>> (two git log outputs joined by sha — --name-status for the A/M/D/R status, --numstat for the +/- magnitude; binary files are 0/0), and the fail-soft recent_commits(dir, limit) over the git CLI seam (None when not a work tree, git missing, or the log is empty). source_worktree(base, path, git_url) resolves a source's work tree WITHOUT cloning (git source → the K8 .gitcache/<label> when present; path source → the path when present) so surfaces can read history from disk live. - ingest.rs: Strategy.changes (default on — "with and without is a flag", like symbols/graph/git/books). For a work-tree source, ingest appends synthetic knowledge entries — one per recent commit (kind: git-commit, Locator::Inline { doc: "git-commit/<sha>" }, subject + author + date + per-file status and +/-) and one aggregate summary (kind: git-changes: the window's per-commit one-liners + the most-changed files) — under a dedicated doc slug <label>/git-log (upserted as a whole, so re-ingest replaces the window cleanly). The slug is tracked in SourceManifest.git_log so ADR 0014 removal deletes it with the source, and in IngestReport.changes so every stats surface carries it (the entries themselves are excluded from the LOC language mix, like code-summary). - The window is capped at CHANGES_LIMIT = 20 commits (deterministic; the cap lives in the engine so REST/MCP/stats agree).

  2. Surfaces (generated app lens): - The per-source stats row gains changes (the last ≤20 commits with their files) — flowing through GET /api/vectordb/{name}/stats, the knowledge_stats + repo_profile MCP tools, the /knowledge page, and mosaic knowledge-report with zero per-surface code. - GET /api/vectordb/{name}/changes?limit=N — LIVE (recomputed from the checkout on request, default 10, capped 50; not ingest-time data), reporting every git work-tree source of the collection (declared + runtime) with branch + commits. Fail-soft per source. - The recent_changes MCP tool (kb + optional limit) over the same live path — the agent-facing "analyze the latest changes". - The /knowledge page renders a "recent changes" table per KB from the stats row (no new fetch — the page already consumes the stats JSON).

  3. DSL: vector_dbs[].strategy.changes (default on) — the only new key. No new dependencies (the git CLI seam is unchanged), no new surfaces beyond the existing conventions.

Consequences

  • Aide (P29) can answer "what changed recently / who touched file X recently" by RAG over the indexed commit entries, and can call recent_changes for a live read — closing the Eezy code-intelligence feature list.
  • Shallow (--depth 1) remote clones (K8) have a one-commit history, so their changes window is that single commit — inherent to the K8 reuse-if-present clone contract (delete the cache to refresh; a full local path or bundle carries the full history).
  • The commit entries live in the index (embedded like any chunk); a reindex refreshes the window. The live /changes endpoint is always current.
  • Stats payloads grow by ≤20 small commit rows per git source (deterministic order, newest first).

ADR 0036 — Command validation + page actions + row states

  • Status: accepted
  • Date: 2026-10-04
  • Slice: P36 (builder-roadmap: DSL-first, codegen-first, UI-first)

Context

Two builder-parity gaps in the CQRS/UI surface:

  1. Commands cannot declare validation rules. A guard (command.where) is a single pre-dispatch condition over state/cmd that rejects the whole dispatch, and there is no way to (a) get a list of failed rules for a payload, or (b) check whether a payload would be accepted without dispatching — the "try before you commit" check every form-bound action needs.

  2. Projection list pages are read-only. The rows table renders live read-model rows with sort/filter/columns/charts, but a user cannot act on a row from the page: no per-row buttons, no page-level action buttons, no derived row status badges. Commands are reachable only via the generic command forms (admin pages) or authored REST endpoints.

Both are DSL-first: the model declares the rules and the actions; the lenses (app + web) derive the code.

Decision

1. command.validate (DSL + app codegen)

A command may declare validate: [<expr>, …] — tessera expressions over cmd.* and the aggregate's current state.*, each compiling to a boolean (the guard's expression scope/target). Semantics:

  • Dispatch path — the generated store wrapper (the public {agg}_{cmd} entry point, not _on) evaluates the rules in order before delegating to _on. A failed rule returns Err("validation failed: <rule>") (the rule's source text), which the endpoint maps to 400. Because it sits in the public wrapper, reactor/workflow/internal dispatches through _on skip validation — validation is an edge-of-system concern (the REST caller's contract), not an invariant check (that is the guard's job).
  • Validate-only path — a CQRS-delegate endpoint whose command carries rules reads the query string; ?validate=true (any value) evaluates the rules against the parsed payload without dispatching and answers 200 { "valid": bool, "errors": [ … ] } (failed rule source texts). The handler gains an axum Query extractor only when the delegate has rules (byte-identical otherwise).

The plan layer compiles each rule once (CmdPlan.validate: Vec<(source, compiled)>); the render layer emits the block, a read-only {fn}_validate(&self, cmd) -> Vec<String> store seam, and the short-circuit.

2. ui.actions + ui.row_states (DSL + plan + web codegen)

Projection ui: gains two keys (ADR 0008: YAML-only front-end — YamlUiAction / YamlUiDialogField / YamlUiRowState):

ui:
  row_states:            # badges per row, first match wins
    - when: 'row.active == true'
      label: "active"
      variant: success   # Badge variant (default info)
  actions:               # buttons bound to CQRS commands
    - label: "activate"
      command: "RuleSet.Activate"
      placement: row     # row | toolbar
      confirm: "activate?"        # two-step arm+execute (no dialog)
      dialog:                          # declared fields (optional)
        - field: note
          label: "note"
          widget: textarea            # text|textarea|number|checkbox|json
          required: true
      toast: "activated"              # success message (default "<label> ok")

Plan-time resolution (resolve_ui_page_features, after SSO validation) per action:

  • command: "Aggregate.Command" must resolve to a declared aggregate + command (fail-closed diagnostics otherwise).
  • Mount: the authored rest.endpoints item sourcing cqrs <Agg>.<Cmd> (POST) wins; else the auto-derived POST /api/{agg-kebab}/{cmd-kebab} (P27) — so actions work with zero endpoint authoring.
  • Row prefill: for placement: row, command inputs whose name matches a projection field are prefilled from the row at click time (never shown in the dialog).
  • Dialog fields: declared dialog[] fields (validated to name command inputs; Json-typed inputs become json widgets automatically) plus any input neither declared nor row-prefilled (auto-added, text widget). Toolbar actions have no row, so ALL inputs are dialog fields.
  • Row states: each when compiles with the expression compiler in the guard scope, target bool, over a row json-param (row.<field> → row.get("<field>") on the row's serde_json::Value) — so states work over any projection row without typed state.

Web codegen (the projection list page):

  • Toolbar actions render in the rows card header: a dialog action opens the page's dialog card; a confirm action is a two-step button (first click arms, second executes); neither is declared → a direct-execute button.
  • Row actions render in a trailing actions column; dialog actions carry the row into the dialog, confirm actions arm per row key (armed: Option<String> holds the key; the button label flips only for the armed row).
  • Dialogs: one #[component] per dialog action ({kebab}_{i}_dlg), rendered in a fixed modal overlay (.dlg-overlay / .dlg-panel), one field per dialog input (type-aware widgets), required-field checks before dispatch, cancel closes; submit POSTs the command body (row-prefilled fields
  • field values, JSON-encoded per type) to the mount, flashes the toast, and reloads (the fresh state renders the result; an error flashes inline). Each action's captures are owned per-closure (row_c{i}, rk{i} clones) because every move event closure in view! must own its captures.
  • Row states render as a status badge column: a {kebab}_row_state(row) -> Option<(String, BadgeVariant)> helper (one compiled when per state, first match wins) feeds a shadcn Badge with the declared variant.
  • A transient flash line (fixed bottom) carries toasts/errors; the api_post client helper is emitted whenever any page has actions (previously only for vectordb apps).

3. Generated-code constraints discovered

  • view! attribute values are parsed by syn, which (unlike rustc) rejects ;-separated match arms — generated match arms inside handlers are block-wrapped.
  • RwSignal (leptos 0.8) is the combined read/write signal: page action state uses RwSignal::new so the same signal can be .get()/.set() from handlers and passed as a signal prop to the dialog components.

Consequences

  • Forms can pre-check payloads (?validate=true → 200 + error list) and users get precise, per-rule feedback on dispatch (400 + failed rule text) — without duplicating guard logic.
  • List pages become operational surfaces: every declared command is reachable from the rows that its inputs describe, with auto-prefill, dialogs for the remainder, and confirm for destructive/no-field actions — no endpoint authoring required.
  • Guards and validation compose: a command can carry both (guard = internal invariant, validation = edge feedback); reactors are unaffected (they bypass validation by design).
  • Golden diff is small: the Query param appears only on rule-carrying delegates; the let out: Result<Value, String> annotation on cqrs handlers is explicit (needed for the short-circuit branch) and lands on all cqrs delegates uniformly.
  • E2E (full-stack example): Order.CancelOrder validates (state.status != "cancelled", len(cmd.reason) > 0); scenarios cover 200 dispatch, double-cancel 400, empty-reason 400, and validate-only true/false. RuleSetList gains row states (active/inactive badges) + row actions (activate/deactivate, two-step confirm over the row-prefilled id).

ADR 0037 — Command forms (`command.form`)

  • Status: accepted
  • Date: 2026-10-04
  • Slice: P37 (builder-roadmap: DSL-first, codegen-first, UI-first)

Context

Every command gets a command page in the Leptos web lens (/commands/<cmd-kebab>), but the page was read-only: a table of field names + Rust types and an "invoke" card naming the endpoint. To run a command a user had to leave the web UI (admin surface, curl, CLI) — the web's only command inputs were the P36 page-action dialogs (bound to a specific projection) and the workflow canvas's raw-JSON textarea.

Meanwhile the shared cmd-form dispatch handler (the static dispatch.js) already knew how to POST any form.cmd-form to its data-mount — it just treated every control value as a string, so typed inputs (ints, bools) could not round-trip.

The gap, DSL-first: a command should be able to declare its form (title, submit label, per-field hints) and the web lens should render a working, typed form for every command.

Decision

1. command.form (DSL + plan)

A command may declare

command:
  name: PlaceOrder
  fields: [ … ]
  form:
    title: "New order"      # default: the command's `about`, else its name
    submit: "Place order"   # default: "submit <Command>"
    fields:
      - field: customer     # a command input name (fail-closed)
        label: Customer     # default: the field name
        widget: text        # text | number | checkbox | textarea | select | json
        required: true      # default: the field's requiredness
        default: ""         # initial value (string literal)
        placeholder: who?

The plan resolves a CmdFormPlan for every command (declared or not):

  • Unlisted inputs are appended in declaration order with a type-derived widget: JSON/struct → json, enum → select (variants resolved from the aggregate's in-scope enums: its own + the schema enums), bool → checkbox, int/float → number, else → text.
  • Declared hints refine the auto row; an omitted widget derives from the type the same way.
  • Fail-closed at plan time: a field that is not a command input, a duplicated field, an unknown widget, select on a non-enum, number on a non-numeric, checkbox on a non-bool.

AggPlan.enums now carries the enums in scope for the aggregate's field types (its own plus the schema enums), so the P36 dialog resolution and the new form resolution agree on what an enum is.

2. The command page is a form (web codegen)

CommandPage_<Cmd> replaces the read-only fields table with a form.cmd-form card: one <div> per form field — a label (required fields carry *) and the control for the widget (<input type=text| number>, <input type=checkbox>, <select> with an empty "—" option + enum variants (the default variant selected), <textarea rows=3> (font-mono + {} placeholder for json)) — then the submit button (form submit label) and a dispatch-status line the handler writes to. The "invoke" card (endpoint + CLI hint) stays below. Non-command pages (workflow CLI commands) keep the fields table.

3. Typed dispatch (the shared dispatch.js)

The global cmd-form handler now encodes by control type: checkbox → bool, type=number → JSON number, textarea → JSON-parsed (falling back to the raw string), else → string. This also upgrades the P36 dialogs' SSR-sibling admin forms and any authored cmd-form.

4. P36 dialogs: the unified widget vocabulary

ui.actions[].dialog[].widget now accepts the same vocabulary (text|number|checkbox|textarea|select|json); an omitted widget derives from the field type (previously it had to be spelled out, and an empty value was an error). select renders an enum <select> (variants from the plan; string encoding on dispatch).

5. The web binary honors the app's CLI seam

The generated web main now reads serve --bind <addr> (the subcommand is ignored), then MOAIC_BIND_ADDR (ADR 0025), then the leptos dev site addr. The e2e harness spawns binaries with exactly that interface, so the web SSR surface is testable by the same app_e2e blocks as the app binary (the full-stack e2e gains a full_stack_web block asserting the rendered form's markup).

Consequences

  • Every command is runnable from the web with typed, pre-fillable controls — no JSON authoring for the common case; JSON fields keep a JSON textarea.
  • Form semantics are display-side: required marks the label, the server-side truth stays the guard + command.validate (P36) + deserialization; the form never weakens a check.
  • The command page SSR HTML is e2e-assertable (GET → 200 + markup), which becomes the pattern for the upcoming web slices (P38–P40).
  • AggPlan.enums widening changes the SSR admin form's selects: enum fields typed by a schema enum now render <select>s there too (previously text) — intended parity.

Verification

  • Plan tests: type-derived widgets (text/number/checkbox/select+ variants/json), declared hints (title/submit/label/placeholder/ required/default), unlisted append, and all five fail-closed errors.
  • Render tests: the command page's typed controls (incl. selected default variant, json textarea, required asterisk, dispatch-status), the dispatch.js number encoding, and the dialog select + derived widgets.
  • Full-stack e2e: a new full_stack_web block boots the web binary via the harness (serve --bind) and asserts the SSR form markup (declared title/submit, field controls, auto-derived json widget, declared default prefill).
  • Conformance re-pinned: command-page forms in all examples' web goldens, dispatch.js, the web main bind seam, and the full-stack app surface for the new PlaceOrder.urgent bool input (proto/agg/model/mcp/openapi/site).

ADR 0038 — Task center (`/tasks`)

  • Status: accepted
  • Date: 2026-10-04
  • Slice: P38 (builder-roadmap: DSL-first, codegen-first, UI-first)

Context

Workflows with a human node pause their run waiting for an answer (the L6 HITL pause/resume seam). The only surfaces for that pause were per-workflow REST endpoints — GET /api/workflows/{slug}/pauses and POST /api/workflows/{slug}/resume/{id} — which (a) require knowing the workflow slug, and (b) have no UI at all: a paused run is invisible outside curl. As more workflows gain human gates, the operator needs one place that answers the question "what is waiting on me?" across the whole app.

Decision

1. GET /api/tasks (shared server codegen)

When any workflow has a human node, the generated server (app and web lens — the same server_rs function) gains

GET /api/tasks → { "tasks": [ { workflow, pause_id, node, question, created_ms }, … ] }

— every paused run across all workflows (oldest first, the BTreeMap key order), each entry naming its workflow so the consumer can link to and resume it. The per-workflow pause/resume endpoints stay unchanged. A pub fn paused_tasks() accessor exposes the same shape to the in-crate (SSR) side; the handler wraps it.

2. The task center page (web codegen)

The web lens gains pages::TasksPage at /tasks (plus a tasks nav entry with a check-square icon), both emitted only when a human node exists:

  • SSR reads crate::server::paused_tasks() synchronously, so the paused table renders on the first paint (the SchedulesPage dual-mode pattern); the client refetches through /api/tasks after hydration.
  • Each task row shows workflow, node, and the rendered question, plus an answer input and a resume button that POSTs {"answer": …} to the workflow's resume mount — the answer is sent as JSON when it parses, as text otherwise, and empty falls back to the server default (true). Success flashes a toast and reloads, so the answered run drops off the list; an error flashes inline.
  • An empty center renders "no runs are waiting on a human answer".

Consequences

  • A paused run is always visible in one page, for every workflow, with a one-click answer path — no slug knowledge, no curl.
  • The task list is process-local (like the pause store itself): it reflects the pauses of the running server instance.
  • The page/endpoint are fail-closed on the same gate as the pause endpoints: no human node → no /api/tasks, no route, no nav entry.
  • The page's SSR HTML is e2e-assertable (GET /tasks → 200 + markup), continuing the P37 pattern for web slices.

Verification

  • Render tests: the route, nav entry, component, and GET /api/tasks are emitted for a workflow with a human node, and absent without one. The existing human_node_emits_pause_resume_infra test now also covers paused_tasks / list_tasks / the route.
  • Full-stack e2e (app block): /api/tasks is empty after a resume, lists the new pause (workflow + pause id) after a run, the answer resumes the run, and the list clears afterward.
  • Full-stack e2e (full_stack_web block): /tasks SSR-renders the empty state, then the paused run (rendered question) with its resume control after a run pauses at the human node.
  • Conformance re-pinned: full-stack only (the only example with a human node) — app + web server.rs (endpoint), web main.rs (route), web pages.rs (nav + page).

ADR 0039 — Declared dashboards + cards/kanban list views

  • Status: accepted
  • Date: 2026-10-05
  • Slice: P39 (builder-roadmap: DSL-first, codegen-first, UI-first)

Context

P29 added a single built-in /dashboard (static entity counts + live event mix) — a fixed shape that says nothing about the app's own data. Real operational views are per-app compositions of KPI numbers, charts, and fresh-rows tables over projections, and different teams want different compositions (an ops overview vs. a billing view). Meanwhile the list pages generated for projections were always one shape — a table — but many read models are naturally better shown as cards (an order, a rule set) or as a board grouped by a status-like field (a kanban column per hit_policy). Both needs were previously impossible to express in the DSL.

Decision

1. Top-level dashboards: (DSL → plan)

app gains an optional top-level dashboards list:

dashboards:
  - name: ops
    about: "Operational overview"
    cards:
      - kind: kpi        # label + projection → row count
        label: orders
        projection: Orders
      - kind: chart      # label + projection + chart: bar|line|area|donut + x + y
        label: rules by hit policy
        projection: RuleSetList
        chart: donut
        x: hit_policy
        y: rules_count
      - kind: list       # label + projection, optional limit (default 5)
        label: recent orders
        projection: Orders
        limit: 5
      - kind: link       # label + href (+ optional icon)
        label: orders
        href: /projections/orders
        icon: database

The plan layer resolves each card against the declared projections in a post-pass (resolve_dashboards, after projections/pages are settled) and fails closed: duplicate-dashboard, dashboard-unknown-projection, dashboard-chart-kind, dashboard-chart-field (x/y must be projection fields), dashboard-link-href, dashboard-card-kind. Projection cards carry the projection's rows_fn (the generated store accessor), so codegen never re-derives names. list cards render the projection's declared list_page columns (up to 4, key field first; projection fields as fallback), limit rows, and a "view all →" link to the projection page.

2. pages::DashboardPage_{Pascal} (web codegen)

Each dashboard gets a route at /dashboards/{kebab} (the route is emitted in main, the component in the pages module, plus a nav entry under a Dashboards group). The page layout:

  • KPI/link grid (.dash-grid): KPI tiles show the live row count of their projection (SSR: the in-crate store accessor; CSR: a one-shot fetch of GET /api/projections/{kebab} into a drows_* signal — one fetch per distinct projection, deduped across cards); link tiles are anchors with a scene icon, target=_blank for external http(s) hrefs.
  • List cards: full-width Cards with a For over the (limit-trimmed) rows — the trim lives in a named closure emitted before the view! (a take(n).collect::<Vec<_>>() generic cannot sit inside a view! attribute), one column per declared list_page column, fmt_cell formatting, and a header link to the projection page.
  • Chart cards: reuse the projection-chart machinery (chart_data_stmt
  • chart_card) with the card's label as the chart title; the any_charts gate (which emits ChartDatum/CHART_COLORS) now also counts dashboard chart cards.

3. ui.list_page.view: table | cards | kanban (projection list pages)

Projection ui.list_page gains an optional view (default table) and, for kanban, a required view_group_by field:

  • cards — the rows render as a .cards-grid of .proj-card tiles: the key field (linked to the detail page when detail_page is on), up to 4 non-key value rows, and the row-state badge when row_states exist.
  • kanban — the rows are grouped into .kanban-col columns by view_group_by (a BTreeMap built in a named groups closure before the view!), each column headed by the group value + row count.
  • Fail-closed at plan time: unknown view, kanban without a view_group_by, a view_group_by that is not a projection field, and cards/kanban combined with template: split (master-detail assumes the table). template: tabs composes freely.

Consequences

  • Dashboards are declared, testable data — the plan layer is the single source of truth for what a card shows, and every bad reference is a build-time diagnostic, not a runtime blank tile.
  • One page per declared dashboard, no new server endpoints: everything a card shows is already served by the existing projection store + REST (GET /api/projections/{kebab}), so SSR and CSR stay in sync by construction.
  • The list-page shape is now a per-projection choice; the table remains the default, so no existing app changes shape unless it opts in.
  • The view! constraints that shaped the codegen (no <> generics in attribute values, closures before the macro, per-closure owned clones) are the same lessons the P36–P38 web slices learned and are now encoded in the generator rather than re-learned per app.

Verification

  • Plan tests: cards resolve against projections (rows_fn + list columns
  • link icon), fail-closed card diagnostics, duplicate dashboard names, kanban validation (group field required / must be a field / split conflict, valid case).
  • Render tests: DashboardPage_Ops + route + nav for a declared dashboard, absent without one; the cards and kanban list views render their respective markup.
  • Full-stack: new Orders projection over Order (upsert on OrderPlaced, status updates on confirm/cancel) with a cards list view; RuleSetList now renders as a kanban grouped by hit_policy; a declared ops dashboard (2 KPIs + donut + list + 2 links). All e2e scenarios green, incl. seven new full_stack_web scenarios asserting the SSR markup (KPI tiles with the seeded order, list row, link tiles, card grid, kanban column heads).
  • Conformance re-pinned: full-stack only — new orders projection (agg/server/app surface + projection page + chat fact), dashboard page
  • nav + routes, kanban/cards markup, dashboard CSS.

ADR 0040 — Command palette (Ctrl/Cmd+K)

  • Status: accepted
  • Date: 2026-10-05
  • Slice: P40 (builder-roadmap: DSL-first, codegen-first, UI-first)

Context

Every generated web app grows more surfaces with each slice: projections, aggregates, workflows, commands, schedules, flags, events, CQRS explorer, task center, declared dashboards, admin pages, knowledge/search. The shell nav lists them, but (a) in the collapsible layouts most sit behind collapsed groups, (b) the header-center filter only narrows what is already visible, and (c) command pages — the app's real "do something" surface — have no shortcut path at all: an operator has to browse /commands and click through. A command palette (the Ctrl/Cmd+K jump overlay, as in the major editors and app shells) collapses the whole navigation problem to "type what you want, press Enter", and it needs no new server surface: every item is just a link to a page that already exists.

Decision

1. app.ui.palette (DSL)

app.ui gains an optional palette block with a single enabled toggle (default on):

app:
  ui:
    palette:
      enabled: false   # opts out of the overlay + the header toggle

Absent → on. There is deliberately no per-item DSL surface: the palette is a projection of the app's existing declared surface (nav entries + CLI commands), not a new list to maintain.

2. The overlay (web codegen)

The shell (Layout) emits, right after the nav:

  • a header-right toggle (search icon + a Ctrl/⌘ kbd hint — the modifier label is corrected client-side) that opens the palette;
  • a fixed overlay (.palette, hidden by default): a filter input plus a list of a.palette-item links, each with an icon, a label, and a kind badge (page / link / command):
  • pages — the same NavEntry inventory the shell nav is built from (built-ins, declared dashboards, scene contributions, role-gated entries, app.ui.nav extras); role-gated items carry the same data-roles attribute, so the advisory role filter in layout.js drops them exactly like the nav links;
  • commands — one run: {name} item per CLI command, linking to its command page (the P37 typed form);
  • external entries keep target="_blank" rel="noopener".

3. The behavior (layout.js)

The palette wiring is static JS (the same vehicle as the nav-search filter and the layout toggles), so it works in every hydration mode with no Leptos reactivity:

  • Ctrl/Cmd+K toggles the palette from anywhere (the header toggle button does the same); opening focuses the input and resets the filter;
  • input filters items by label (case-insensitive substring, the data-label attribute) and re-highlights the first visible item;
  • ↑/↓ move the active item through the visible items (wraparound); Enter navigates to the active item's href; Esc (or a backdrop click) closes and returns focus to the toggle;
  • items removed by the role filter are excluded from navigation (isConnected check), and the kbd hint shows ⌘ on macOS, Ctrl elsewhere.

Consequences

  • Jumping to any page or any command form is ≤ keystrokes: no nav hunting, no scrolling, no collapsed groups — the palette is the flat index over the whole declared surface.
  • Zero new server endpoints and zero new client state: the overlay is a static link list in the SSR HTML (e2e-assertable), and the behavior is a few dozen lines of shell JS — the pattern P32/P38 established for shell features.
  • The palette's item list is derived, not authored: adding a projection page, a dashboard, or a CLI command automatically adds its palette entry; there is nothing to keep in sync (and nothing to typo).
  • Apps with a deliberately minimal header can opt out with one boolean; the CSS stays global (hidden overlay costs nothing).

Verification

  • Render tests: the overlay, filter input, header toggle, and the layout.js wiring (shortcut, arrows, label filter) are emitted by default; a nav entry surfaces as a data-label item; a CLI command surfaces as a run: … item linked to its command page with the command kind badge; app.ui.palette: { enabled: false } emits none of it.
  • Full-stack e2e (full_stack_web block): GET / SSR-renders the palette input and a run: process_order item.
  • Conformance re-pinned: the shell change touches every example (pages overlay + toggle, layout.js wiring, palette CSS, MANIFEST) — the palette is a shell feature like the nav-search filter, not an app-specific surface.

ADR 0041 — CLI: `schema dsl`, `diff`, `explain`

  • Status: accepted
  • Date: 2026-10-05
  • Slice: P41 (builder-roadmap: DSL-first, codegen-first, UI-first)

Context

The tessera model is authored in tessera.yaml, a large typed surface (128 Yaml* structs). Three everyday tasks had no first-class tooling:

  1. Authoring aid — editors could not validate or autocomplete tessera.yaml: mosaic schema only covered the three legacy spec formats (tessera / mosaic / deploy via schemars), not the YAML surface that is now the only authoring format (ADR 0008).
  2. Change review — when a model changes (a field added to an aggregate, a projection renamed, a dashboard dropped), the only way to see what changed at model level was to diff the rendered output — noisy (generated code) and slow. A plan-level, semantic diff was missing.
  3. Error triage — mosaic check reports error[<code>]: …; the code is the stable identity of the diagnostic, but there was no way to look up what a code means and how to fix it without reading the plan-layer source.

Decision

1. mosaic schema dsl

schema gains a fourth what value, dsl: the JSON Schema for the tessera YAML surface. It is generated with schemars from the same Yaml* structs the loader deserializes, so the schema cannot drift from the parser:

  • every Yaml* struct derives JsonSchema (added alongside Debug, Deserialize);
  • fields typed serde_yaml::Value / serde_yaml::Mapping (free-form escape hatches) are schematized as serde_json::Value;
  • Localizable (an untagged string-or-map that also flattens) gets a hand-written impl: type: ["string", "object"];
  • YamlDeploy.proxy_buffering uses a small untagged YamlOnOff (string | boolean) because a strict YAML 1.1 parser reads bare off as false; the plan layer normalizes both forms to on | off (unchanged fail-closed validation).

The entry point is mosaic_core::tessera_yaml::dsl_json_schema().

2. mosaic diff <from> <to>

Two workspace roots (dirs containing tessera.yaml) are loaded and built to plan level, then compared. The comparison unit is a digest: Digest maps entity → facet → detail string, where the entity is <kind>:<name> (e.g. aggregate:Order, projection:Orders, workflow:CheckOut, dashboard:ops) and the facet is the meaningful attribute (state.<field>, command.<name>.fields, projection.<name>.on.<i>, …). tessera_diff::digest(&ws) covers the whole plan surface (app incl. ui/env/realms, aggregates, workflows, endpoints, cli, mcp, mcp_servers, dashboards, schedules, triggers, notifications, flags, rulesets, policies, admins, scenes, vectordbs, pages/docs, identity, peers, mirrors, aspects, gateway routes, schemas, tests, model registry, command groups, requirements, reactors, templates, queries, resources, domains).

diff(&old, &new) renders a deterministic report:

+ workflow:Refund.nodes.refund_call: …
~ aggregate:Order.state.priority: - -> Prim("String")
- schedule:nightly-report: <facets>

+ added entity / facet, - removed, ~ modified. Exit 0 when identical, 1 when different (like diff); parse or plan errors on either side abort with the diagnostics.

3. mosaic explain <code>

diag_docs carries a curated catalog — code → (what it means, how to fix it) — for every diagnostic code emitted by the plan layer (~124 codes: dashboard-unknown-projection, projection-unknown-event, cli-unknown-command-source, …). explain prints the entry, or says so when the code is unknown. A test scans the mosaic-core sources for every emitted code and asserts each is cataloged, so the catalog cannot silently lag the plan layer.

Consequences

  • tessera.yaml gets editor validation + autocomplete out of the box (point the YAML language server at mosaic schema dsl); the schema is derived from the deserializer, so it stays true by construction.
  • Model review is one command: snapshot the old model dir (e.g. from git or a backup) and mosaic diff old/ new/ shows exactly which entities and facets changed — without rendering.
  • Diagnostics become self-documenting: mosaic explain <code> next to mosaic check closes the "what does this error mean" loop without source diving.
  • All three are CLI-only: no render change (conformance stayed green, no re-pin), no new server surface, no DSL surface beyond the YamlOnOff normalization.
  • The diff digest is lossy by design: it compares declared model meaning, not rendered bytes. A change that only reorders generated code (never a real model change) shows nothing — which is the point.

Verification

  • schema dsl emits a valid draft-07 JSON Schema (~70 KB); all four example tessera.yaml files validate against it (Python jsonschema). The validation caught a real drift: proxy_buffering: off was typed Option<String> but is a YAML 1.1 boolean — now accepted in both forms.
  • diff on an unmodified workspace prints no differences (exit 0); adding one aggregate state field reports ~ aggregate:Order.state.priority (exit 1).
  • explain prints the curated entry for a known code and a clear fallback for an unknown one; the catalog-completeness test fails if any emitted code is uncataloged.
  • Gates: workspace tests (new core tests for diff/explain/YamlOnOff), clippy, fmt, conformance green (no re-pin), full-stack build + e2e.

ADR 0042 — `import openapi`: spec docs as knowledge sources

  • Status: accepted
  • Date: 2026-10-05
  • Slice: P42 (builder-roadmap: DSL-first, codegen-first, UI-first)

Context

mosaic import openapi <spec> <id> (P29, ADR 0029) turns an OpenAPI document into a proxy tessera: one rest.endpoints item and one mcp.tools tool per operation, all delegating to the spec's servers[0].url. The generated app declares an upstream-kb vector db that — with --knowledge — captures successful upstream responses at runtime. But the KB never contained the spec itself: the docs describing what each operation does (parameters, request bodies, responses) lived only in the source document, so the app's existing ask surfaces (aide at /aide + /api/aide/chat, aide_chat MCP tool, /search, the knowledge REST surface) could not answer questions about the imported API. And the generated tessera.yaml was only plan-checked in users' heads: nothing in the repo proved the emitted output parses and plans cleanly.

Decision

1. Spec docs → knowledge/docs/api/

The importer now emits, alongside tessera.yaml, one markdown doc per operation plus an index, under knowledge/docs/api/:

  • index.md — the API title, version, description, server, and a table of every operation (method, path, summary) linking to its doc;
  • <endpoint-id>.md — the operation: method + path, description, tags, a parameters table (name, in, required, type, description — path-level parameters merged with operation-level, operation level winning on name:in), the request body per content type, and a responses table (status, description, body schema).

Docs are deterministic (BTree-ordered iteration of the parsed spec, no timestamps) and derived purely from the document — $ref schemas render as their component name, inline objects as object (a: string, b: …), arrays as array<T>. The generated upstream-kb then declares:

vector_dbs:
  - name: upstream-kb
    sources:
      - path: knowledge/docs/api
        kind: docs
    docs: []

The build already copies <tessera>/knowledge/ into the output (copy_knowledge_dir), the runtime resolves relative source paths against the output root (MOAIC_KNOWLEDGE_BASE → CWD → exe root), and the knowledge engine is vendored for every app with a non-empty vector_dbs — so the spec docs are ingested at boot and retrievable with no new mechanism. --knowledge keeps its P29 meaning (runtime response capture) and is orthogonal: the spec docs are emitted either way.

2. Testable core + round-trip guarantee

cmd_import_openapi was split into a pure openapi_import(doc, tessera, cache, cache_secs, knowledge) -> OpenapiImport { tessera_yaml, docs, endpoints } (the CLI does only fetch/read/parse/write) plus deterministic helpers (op_docs, op_schema_type, op_schema_body, md_cell). Unit tests cover the emitted tessera (source declaration, endpoint ids), the per-operation docs (parameter merging, request body, response table, markdown escaping), the --knowledge gate, and — the key guarantee — a round-trip test: the generated tessera.yaml loads through load_tessera_yaml_str and builds through the plan layer with zero errors, so the import output is a valid tessera by test.

3. Example + codegen fixes it exposed

A committed example (examples/openapi-proxy: the pets.json fixture + its import output + conformance golden, wired into the CI web-builds list) pins the import shape as test-of-record. Building that first zero-aggregate proxy app compiled code paths that had never been compiled, exposing three latent codegen bugs, all fixed:

  • web lens: the /dashboard KPI page is always emitted and renders its recent-events row via cell_str, but the cell_str/fmt_cell helper was only emitted for apps with aggregates/formats/charts — any such app's wasm build failed with cannot find function cell_str. The helper is now always emitted.
  • app lens: the proxy knowledge-capture stamps doc ids with now_ms(), which is defined only in the workflow run-history block — a proxy app without workflows failed with cannot find function now_ms. The helper is now also emitted when no workflow exists but a proxy endpoint declares knowledge capture.
  • app lens: the proxy forwarder read headers from req after req.into_body() moved it (E0382) — the header set is now snapshotted before the body is consumed.

Consequences

  • The imported proxy app can answer about the imported API: ask the aide ("what parameters does GET /pets take?") and get spec-grounded answers from the docs it ships — the same surfaces (aide page, /api/aide/chat, aide_chat MCP, /search, knowledge REST) every vector-db app already has.
  • The import is now self-verifying: the round-trip test fails if the emitter and the plan layer ever disagree, and the example golden pins the rendered proxy app (app + knowledge engine + web lens) in CI.
  • No new DSL surface (the sources declaration already existed — K1), no new runtime mechanism (build-time knowledge copy, source resolution, and the vendored engine are reused), no new CLI flag (--knowledge semantics unchanged).
  • The three codegen fixes are behavior-preserving for every existing app: the only rendered-output change is in apps that had zero aggregates (none before this example), so no existing golden moved.

Verification

  • New unit tests (mosaic-cli): source declaration in the emitted tessera, per-operation doc contents (incl. path-level parameter override + pipe escaping), --knowledge gating, and the plan round-trip (load + build, no errors) — 9/9 green.
  • Smoke: import a fixture spec, mosaic check + mosaic build --check pass on the output (125 files render); the rendered proxy app builds both lenses (native app + wasm web, hydration: csr).
  • Gates: workspace tests (18 suites), clippy -D warnings, fmt, conformance re-pinned (only the new example), full-stack build + e2e.

ADR 0043 — Command-form `show_if`, field groups, and wizard steps

  • Status: accepted
  • Date: 2026-10-05
  • Slice: P43 (builder-roadmap: DSL-first, codegen-first, UI-first)

Context

P37 (ADR 0037) gave every command a typed web form, but the form was a flat stack: every field rendered, always visible, in declaration order. Three common form shapes were impossible to declare:

  • Conditional fields — "show the card-number field only when payment method is card", "show an invoice reference only when the method is invoice". A form that always shows every field makes the user work out which fields apply.
  • Visual grouping — a long form with a heading per section (billing / shipping / flags) reads better than one undifferentiated stack.
  • Wizard steps — a multi-part command (pick the order, then give a reason) is easier as one step at a time than a single tall page.

These are display-side concerns: the server-side truth stays the guard + command.validate (P36) + deserialization — the form never weakens a check, the same posture as P37's required. So they belong in the DSL (command.form) and the web lens, with no new runtime mechanism.

Decision

1. command.form extensions (DSL + plan)

A command form may now declare, in addition to title/submit/ fields:

form:
  title: "New order"
  submit: "Place order"
  fields:
    - field: method
    - field: amount
    - field: card_number
      show_if: 'method == "card"'   # a restricted expression over the form's fields
  # OR visual groups (mutually exclusive with `steps`):
  groups:
    - title: "Billing"
      fields: [method, amount, card_number]
    - title: "Flags"
      fields: [urgent]
  # OR wizard steps (mutually exclusive with `groups`):
  steps:
    - title: "Method"
      fields: [method, amount]
    - title: "Details"
      fields: [card_number, urgent]
  • show_if — a per-field visibility condition. It is a restricted subset of the tessera expression language: literals (string/number/bool), single-segment field references, !, and the binary ops ==, !=, <, <=, >, >=, &&, || (parenthesization comes from the parser). A field reference must name a form field; a multi-segment path or any other node (function calls, indexing, arithmetic, …) is a load-time error. crates/mosaic-core/src/showif.rs holds the validator + the compiler to a JS expression.
  • groups — visual sections. Each lists the fields it contains (each field in at most one group; an unknown field or an empty section is a plan error). Fields not listed in any group trail after the groups, without a heading. Groups are all visible at once.
  • steps — a wizard. Each step lists its fields; together the steps must cover every form field exactly once (a plan-time diagnostic otherwise). One step is visible at a time; Next/Back move between them and the submit button sits on the last step.
  • groups and steps are mutually exclusive (a plan-time error).

The plan (CmdFormPlan) resolves groups/steps to CmdFormSectionPlan (title + field names, validated against the resolved form fields) and carries each field's show_if (the validated Expr, subset-checked at load).

2. Rendering (web codegen)

CommandPage_<Cmd> now lays the form out by shape:

  • flat (neither groups nor steps) — unchanged: every field, then the submit button.
  • groups — one heading per group + its fields, then the trailing unlisted fields, then the submit button.
  • steps — one <div data-wfstep="i"> per step (step 0 visible, the rest hidden), each with its heading + fields + a nav row (Back when not first, Next when not last, the submit button on the last step).

A field with a show_if is wrapped in <div data-show-if="<js>">. The compiled JS is embedded as a Rust string literal (the {:?} keeps the inner quotes valid in the generated Leptos view!, and consistent between the SSR-rendered string and the client DOM — getAttribute returns the decoded value in both paths).

3. Live visibility + wizard nav (the shared dispatch.js)

The global cmd-form handler now also, per form:

  • formFieldValue(f, name) — reads a field's live value (bool/number/string by control type).
  • applyShowIf(f) — for each [data-show-if] wrapper, evaluates its JS against V("<field>") (loose ==, numeric </>=) and toggles hidden; re-run on every input/change.
  • initWizard(f) — shows one [data-wfstep] at a time; data-wf-next/data-wf-back move between them.

On submit, a field whose [data-show-if] wrapper is currently hidden is omitted from the POST body (a hidden conditional field doesn't apply). Fields hidden only because they sit in a non-current wizard step still submit — a wizard collects every step before the last step's submit fires.

Consequences

  • A command form can now express conditional fields, sectioned forms, and wizards — the three display shapes P37's flat form couldn't — all declared in tessera.yaml and rendered in the web lens.
  • show_if is a deliberately small, plan-validated subset: no runtime expression interpreter ships. The expression is compiled to JS at codegen and evaluated in the browser against the live values; unknown fields / unsupported nodes fail at load, not at runtime.
  • No new runtime mechanism, no new CLI flag, no new route. The change is display-side only: the server-side truth (guard, command.validate, deserialization) is untouched, and a hidden field simply doesn't appear in the payload — the same fail-closed posture as P37's form semantics.
  • The full-stack example exercises the new shapes in a real build: PlaceOrder uses show_if (items appears once an order id is entered) + groups (Order / Flags); CancelOrder is a 2-step wizard (Which order → Why). The conformance golden for its pages.rs pins the emitted markup, so the generated Leptos is compiled in CI.

Verification

  • Unit tests (mosaic-core): the show_if subset — the JS it emits for the supported operators, and rejection of unknown fields, multi-segment paths, and unsupported nodes (function calls, indexing, unary minus, arithmetic).
  • Render tests (mosaic-render): a form with show_if + groups (headings, the compiled data-show-if on the conditional field, the trailing unlisted field, group order) and a steps form (step 0 visible / step 1 hidden, data-wf-next before data-wf-back); the dispatch.js carries formFieldValue/applyShowIf/initWizard + the hidden-field omission. Plan-time negative tests: show_if over a non-field is a load error; steps that don't cover every form field is a plan diagnostic.
  • Real build: the full-stack example (a hydration: full app) renders and compiles — full-stack-web (SSR) + full-stack (app) build, confirming the Leptos view! accepts the new data-show-if / data-wfstep markup; its declared e2e scenarios (REST + the full_stack_web command-form block) all pass.
  • Gates: workspace tests, clippy -D warnings, fmt, conformance re-pinned (full-stack pages.rs + every example's dispatch.js), full-stack build + e2e.

ADR 0044 — Per-action `requires:` role gate

  • Status: accepted
  • Date: 2026-10-05
  • Slice: P44 (builder-roadmap: DSL-first, codegen-first, UI-first)

Context

P36 (ADR 0036) gave list pages row/toolbar actions bound to CQRS commands, and P40 (ADR 0040) gave the shell an advisory client-side role filter (filterNav, driven by /api/auth/me + MOSAIC_ROLE_LEVELS) that drops any [data-roles] element whose minimum role bar exceeds the caller's level. But an action could not declare a minimum role: every action button was rendered for every authenticated caller. The server-side command auth was still the real gate, so an under-privileged caller saw a button that would 403 on click — visible but unusable.

The pieces were already in place: P40's filter understands data-roles on arbitrary elements, and every app has a role ladder (app.role_levels, default mosaic:standard). What was missing was a DSL key to attach a minimum role to an action and a plan-time check that the role exists.

Decision

1. ui.actions[].requires (DSL)

A page action may now declare a minimum role:

ui:
  actions:
    - label: "activate"
      command: "RuleSet.Activate"
      placement: row
      confirm: "activate?"
      requires: editor   # a role on the app's role ladder

requires names a role that must exist on the app's role ladder. It is a client-side advisory declaration: it controls whether the button is shown, not whether the command runs. The command's own auth (its min_role/ permission/where, enforced server-side at dispatch) remains the final gate.

2. Plan-time validation (fail-closed)

The action's requires role is checked against ws.app.role_levels at plan time. An unknown role is the fail-closed ui-action-requires-role diagnostic ("… is not on the app's role ladder (declared: …)") and the action is dropped. The fail-closed posture follows from the filter's semantics: an unmatched role contributes no level to the bar, so its minimum stays at max and the element is hidden for everyone — a silent, always-hidden button is worse than a loud plan error.

The validated role rides on UiActionPlan.requires (carried through the action resolution unchanged; a plan diagnostic does not abort the build, it just skips that action).

3. Rendering (data-roles on the action button)

action_roles_attr(a) returns data-roles="<role>" when the action has a requires, else "". It is spliced into every action button site — the three toolbar variants (dialog / confirm / plain) and the three row variants — so the button carries class="wf-tb" data-roles="<role>". P40's filterNav then does the rest on load (and after /api/auth/me resolves): it reads window.MOSAIC_ROLE_LEVELS, fetches the caller's level, and removes any [data-roles] element whose minimum role bar exceeds it. No new client mechanism, route, or JS ships — P44 only adds the attribute and the plan check.

Consequences

  • An action can now be role-gated client-side: an under-privileged caller never sees the button (matching the nav-item gating P40 already provided for nav entries and palette items).
  • Reuses the existing P40 filterNav + the app's role ladder — no new runtime, no new route, no new client JS. The change is one attribute on the button + one plan-time check.
  • Advisory, not authoritative. The data-roles attribute only hides the button; a caller can always dispatch the command directly via REST. The command's own server-side auth is the real gate, and a requires that is stricter than the command's own auth is redundant (the server would reject it anyway) while one that is looser just shows the button to people the server will then reject. requires is a UX affordance, not a security boundary.
  • View caveat. Action buttons render only on table-view list pages; the cards and kanban lenses render no action buttons (rows_view_card takes no actions). So in the full-stack example — whose only action-bearing projection, RuleSetList, is a kanban view — the requires is exercised at plan time (role validation in the conformance build) while the actual data-roles button render is pinned by the render test on a table-view projection. This is the same verification posture as P36's action buttons, which no example renders in a golden (all action-bearing example projections are non-table).
  • data-roles is a data-* attribute on a Leptos element — the same class as P43's data-show-if / data-wfstep, which are proven to compile in the full-stack-web SSR build.

Verification

  • Render test (action_requires_renders_data_roles_on_the_button): a table-view projection whose row action and toolbar action each declare a requires role renders class="wf-tb" data-roles="editor" and class="wf-tb" data-roles="manager"; an action with no requires renders no data-roles.
  • Plan test (action_requires_unknown_role_is_a_plan_error): a requires naming a role not on the app's (standard) ladder is the fail-closed ui-action-requires-role diagnostic.
  • Example: full-stack declares requires on the RuleSetList actions (activate → editor, deactivate → manager); it plans green in the conformance build (the roles are validated end-to-end).
  • Gates: 477 workspace tests, clippy -D warnings, cargo fmt --check, and conformance (green — no golden regression, since the only action-bearing example projection is kanban and renders no action buttons, so the emitted pages.rs is byte-identical to the pre-P44 golden).

ADR 0045 — Use-case-first web nav (CQRS model demoted to a developer console)

  • Status: accepted
  • Date: 2026-10-05
  • Slice: P45 (builder-roadmap: UI-first — the shell reads as an app, not a model view)

Context

The web shell's default navigation (the branch taken when ui.scenes is empty — i.e. every app that does not hand-author a scene nav) was organized around the CQRS model pages: commands, aggregates, projections, workflows, events, cqrs — all flat and always visible at the top. The use-case list pages (projections with a ui.page hint) were emitted last and sat inside nav groups. Scene groups are already collapsible — collapsed by default in the sidebar/bottom/floating layouts (layout.js applyGroups adds collapsed to every .scene-group-wrap) — so the use-case pages were hidden while the CQRS pages were prominent.

The result: the shell read as a CQRS technical representation, not an app. But CQRS (aggregates / commands / events / projections) is the backend model — how we implement things. The UI/UX should be organized around use cases: list and present things, and trigger/create/update things — without the user needing to know the word "projection" or "command".

Decision

1. Reorder the default nav to be use-case-first

The default (CQRS-derived) nav is now built in use-case-first order:

  1. Primary (always visible / open by default) — dashboard, the declared dashboards (a Dashboards group), the use-case list pages (projections with a ui.page hint — business labels, in their declared groups), and the platform/ops pages (tasks, schedules, flags, notifications).
  2. Developer group (collapsed by default) — the CQRS model pages (commands, aggregates, projections, workflows, events, cqrs) plus the admin: consoles. These are the developer/technical views of the model: still reachable by expanding the group, but no longer the app's front page.
  3. Unconditional entries (knowledge, search, app.ui.nav links) follow.

The app now leads with what the user works on (the use-case pages and their in-page actions — the "trigger/create/update" surface) and treats the CQRS model as a collapsible developer console.

2. Open-by-default groups

A group wrap can be marked open-by-default with a data-group-open="1" attribute. The shell's applyGroups collapses every group by default except those carrying data-group-open. Use-case groups and the Dashboards group are emitted open-by-default (so the app's working pages are visible without a click); the Developer group is not (so it starts collapsed). This reuses the existing collapse mechanism — one attribute plus one hasAttribute check in the shell JS.

3. Example

The full-stack reference app declares business page: hints on the Orders and Tickets projections (RuleSetList already had one), so its primary nav is dashboard / Dashboards / Orders / Tickets / Rule sets, with the CQRS model pages under the collapsed Developer console.

Consequences

  • The shell reads as a use-case app: it leads with business pages (list/present things) and their actions (trigger/create/update), and the CQRS model pages are demoted to a collapsed "Developer" console.
  • The CQRS pages are not removed — they remain reachable by expanding "Developer", preserving the reference app's value for inspecting the model.
  • Reuses the existing group-collapse machinery (one attribute + one JS condition). No new page, route, or backend change; ui.scenes (the explicit scene nav) is untouched — apps that hand-author scenes keep their scene-driven nav.
  • The command palette (P40) shares the nav inventory, so it lists the use-case pages too (a jump-to-any-page tool; it still includes the developer pages).
  • Every app that uses the default nav gets this reframe for free; apps wanting the old layout can still declare ui.scenes.

Verification

  • Goldens (all 5 examples re-pinned): the new nav order + the data-group-open attribute are pinned in each pages.rs; the layout.js applyGroups change is pinned in each shell.
  • Nav order (full-stack): dashboard / Dashboards(ops) / Orders / Tickets / data(rule sets) / tasks / schedules / flags / notifications / Developer(collapsed: commands, aggregates, projections, workflows, events, cqrs) / knowledge / search / docs / github.
  • Real build: full-stack builds (native app + SSR web) and every declared e2e scenario passes — including the command palette (now listing the new use-case pages) and the P39 use-case list views.
  • Gates: 477 workspace tests, clippy -D warnings, cargo fmt --check, conformance re-pinned.

ADR 0046 — Use-case-first pages + rendered action dialogs

  • Status: accepted
  • Date: 2026-10-05
  • Slice: P46 (builder-roadmap: UI-first — the app reads as a set of use cases)

Context

P45 demoted the CQRS model pages to a collapsed "Developer" console, so the shell no longer starts from CQRS. But the list pages themselves still read as CQRS: every projection's list page rendered a card titled rows with the description "live read model — replayed from the event log". And the "trigger / create / update" surface (the per-page actions introduced in P36) was never actually rendered in any committed app — no golden example carried a table-view action (the flagship full-stack actions were on a kanban projection, which does not render action buttons). The action/dialog codegen was therefore never compiled, and a set of latent bugs shipped.

Goal: make a generated app read as a set of use cases — pages that list and present things and let the caller create / update / trigger them — framed in business terms, with the CQRS vocabulary confined to the developer console.

Decision

1. The list page is a use-case page, not a "read model" (compiler)

The list-page card no longer injects CQRS language:

  • The CardTitle is the projection's ui.page business label (falling back to the projection name) — e.g. Documents, Tickets — threaded from the page hint into rows_card / rows_view_card.
  • The CardDescription is the projection's about (a business sentence the author writes), shown only when present. The hardcoded "live read model — replayed from the event log" is gone. ProjPlan now carries about for this.

So the page title + description are the author's business framing; the compiler adds none of its own.

2. Neutral command-form submit (compiler)

The default command-form submit label was submit <Command> (the raw command name in a button). It is now a neutral Submit. (The command page title already defaulted to the command's about — business — so the create/update form reads as an action, not a CQRS command.)

3. The reference app is a set of named use-case pages (mosaic-pages)

The pages app (eugeis/mosaic-pages, a docs portal) now declares its two projections as use-case pages:

  • DocList → Documents (ui.page label, business about) with the create/update surface: a toolbar New document action (a dialog over Doc.Publish: id / app / title / body) and a row Retire action (two-step confirm over Doc.Retire).
  • LogList → Activity (the tenant's audit trail, read-only).

These are the app's front pages (above the Developer console, per P45).

4. The rendered action / dialog codegen (compiler — bug fixes)

Because no golden exercised a rendered table-view action, the action dialog codegen was broken in four independent ways. All are fixed and now pinned in the full-stack golden (a new Open ticket toolbar action on the split Tickets page):

  • The dialog cancel button emitted unquoted text (>cancel</button> with a stray "), which broke the view! token stream and left the component with an unclosed delimiter. It is now >"cancel"</button>.
  • The dialog component was named snake_case (doc_list_0_dlg); a lowercase tag in view! resolves as an HTML element. It is now PascalCase (DocList0Dlg), definition and reference in lockstep.
  • The dialog field signals used create_signal("") → &str, which does not satisfy bind:value's IntoSplitSignal. They are now create_rw_signal(String) / create_rw_signal(bool) (read-write String/bool signals).
  • The dialog's row prop is a plain value; it is now named row (a bare _row is not getter-wrapped by #[component]) and passed as a value (the dialog is created only while its Show is open, so it is current then). A let _ = &row; suppresses the unused warning for toolbar dialogs.
  • The per-action execute block is now a ;-terminated statement (not a bare tail block), and required-field guards are ;-terminated — view!'s syn parser rejects a bare block / bare if when another statement follows it in an on:click closure.

Consequences

  • Generated apps read as use cases: business-titled pages with a business description and an in-context create/update surface; CQRS confined to the Developer console.
  • The full-stack golden now pins a rendered table-view action, so a web-rendering regression in the action/dialog codegen fails conformance (and the example web is compiled in CI / local verification).
  • The mosaic-pages reference app ships named use-case pages (Documents, Activity) with a working publish/retire surface.
  • ProjPlan.about is new public surface (plan); the list-page card title now derives from the ui.page label, so a projection with a page hint shows that label on its list page (not just in the nav).

ADR 0047 — Graph-augmented retrieval (Personalized PageRank re-ranking)

  • Status: accepted
  • Date: 2026-10-06
  • Slice: P47 (engine core) + P48 (DSL + codegen) + P49a (code-RAG tools) + P49b (precise structural edges) — part of the EE knowledge/GraphRAG parity program

Context

The K5 knowledge graph (mosaic_knowledge::graph) is a deterministic, in-process graph derived from the indexed entries: symbol / file / endpoint nodes from code (K2), book / chapter nodes from the book library (K3), and report nodes + cites edges from citation counts. It has uses / contains / serves / next / cites edges. But retrieval never used it — Store::hybrid fuses BM25 + cosine with RRF and returns; the graph was only ever rendered (the /knowledge graph page) or profiled. So a globally central entity (a hub symbol that many others call, a load-bearing chapter) was not resurfaced by search even when it was the most structurally important answer.

The EE reference (ee commit ef03fe0, "GraphRAG upgrade") ships exactly this signal: Personalized PageRank over the knowledge graph, wired as an opt-in re-ranking on top of the existing retrieval (HippoRAG-style). EE gates the storage of the graph behind an optional Apache AGE backend; Mosaic already has the graph in-process, so we take the retrieval idea without the external graph-DB.

Goal: let a generated app resurface globally-central entities by re-ranking the hybrid hits with PPR over the existing knowledge graph — opt-in per collection, graceful when the graph is empty.

Decision

1. Engine core: mosaic_knowledge::rank (P47)

A new pure-Rust module with four functions (no new dependencies):

  • undirected_weighted_adjacency(&Graph) -> BTreeMap<String, Vec<(String, f64)>> — the symmetric, weight-count adjacency the rank algorithms operate on.
  • personalized_pagerank(&Graph, seeds, damping, max_iter) -> BTreeMap<String, f64> — power-iteration PPR with a teleport vector uniform over the distinct seeds; dangling mass redistributed to the seeds. Deterministic (BTree iteration).
  • modularity(&Graph, community, r) -> f64 — Newman modularity of a community assignment on the undirected graph.
  • louvain_communities(&Graph, r) -> BTreeMap<String, u32> — greedy Louvain (one-pass, per-node best-gain moves to fixed point); returns a deterministic node→community-id map (ids assigned by first-seen order).

All are unit-tested (tests/engine.rs, 10 rank tests): PPR personalization, seed-vs-far-end ordering, mass conservation, modularity sanity, Louvain two-clique recovery + determinism, and the no-edge singleton case.

2. Retrieval: Store::graph_rerank (P48)

Store::graph_rerank(hits, graph_weight) -> Vec<Hit> seeds PPR from the graph nodes that back the hits' entries (GraphNode.entry links a node to its chunk), then re-orders the hits by a weighted blend:

score' = (1 - w) * score + w * (ppr(node) / max_ppr)     where w = graph_weight

Hits whose entry has no graph node keep score' = (1-w)*score (their PPR term is 0), so they rank after the graph-backed ones at a given w. Ties break by original position. It is a no-op when w == 0, the graph has no edges, or no hit maps to a graph node (a pure-prose corpus) — so it is always safe to call.

3. DSL: vector_dbs[].retrieval.graph + graph_weight (P48)

Two new keys under the existing retrieval: block (ADR 0018):

vector_dbs:
  - name: docs
    retrieval:
      mode: hybrid      # unchanged
      top_k: 8          # unchanged
      graph: true       # NEW: enable PPR re-ranking (default false)
      graph_weight: 0.5 # NEW: blend strength in [0,1] (default 0.5)

graph_weight is fail-closed-validated (bad-knowledge-graph-weight, must be finite and in 0..=1). The model decl VectorRetrievalDecl carries graph: bool + graph_weight: f64 (defaulted at parse: false / 0.5).

This is distinct from strategy.graph (ADR 0018), which gates building the graph layer at index time — retrieval.graph gates using it at query time.

4. Codegen (P48)

knowledge_retrieval(name) now returns the 8-tuple (vw, bw, mmr, rerank, mode, top_k, graph, graph_weight) and vector_search applies the re-ranking right after the fused hits are produced:

let hits = if graph && hits.len() > 1 { store.graph_rerank(hits, graph_weight) } else { hits };

Because it runs per collection from the collection's own config, every surface that calls vector_search — the /api/vectordb/{name}/search REST endpoint, the search_knowledge MCP tool, and the Aide/voice site-search tool loop — inherits the graph re-ranking automatically when the collection opts in. No per-surface plumbing.

Why not port AGE / a graph-DB backend

EE's headline feature is the optional Apache AGE GraphStorage (a per-KB backend: "age" that stores the graph in Postgres+AGE and runs Cypher). Mosaic already keeps the graph in-process (deterministic, cheap, no server), and its retrieval + code-RAG needs only PPR + community structure — not Cypher. Porting AGE would add a Postgres/AGE deployment dependency for a signal that the in-process graph already provides. We therefore match EE on the retrieval idea (PPR re-ranking) and defer an external graph store until a real persistence/Cypher need appears. The graph itself is already persisted transitively via the entries (the JSON knowledge store + LanceDB).

Is AGE the best OSS graph store? It depends on the posture, and Mosaic is embedded-first (SQLite for state, LanceDB for vectors — no server). On that axis AGE is the wrong shape: it is a Postgres extension, so adopting it means running Postgres. For the embedded scenario (the LanceDB analog for graphs) the best OSS option is Kuzu — a C++ embedded graph DB that persists to a single file, speaks Cypher, needs no server, and links the same way LanceDB links. So the decision ladder is: (1) in-process K5 graph (default — already persists via the entries, zero extra deps, PPR/Louvain in-crate); (2) Kuzu as the first external backend if a server-less Cypher store is ever required; (3) AGE only if the app already runs Postgres (reuse the DB, get Cypher). We do not proliferate backends — the in-process graph is the default and the others are opt-in escape hatches, added only when a real need appears.

P49a — the four bounded code-RAG tools

The EE ee-pages demo exposes four code-search tools. We re-implement them over the existing in-process K5 graph + K2 symbol census (no tree-sitter, no new deps, no LLM):

Tool Engine Surface
find_symbol codegraph::find_symbol — code symbols by case-insensitive name (file/kind filters) GET /api/vectordb/{name}/symbol?name=&file=&kind= + MCP
find_callers codegraph::find_callers — reverse of the deterministic uses edges, aggregated per caller GET /api/vectordb/{name}/callers?symbol= + MCP
get_call_graph codegraph::call_graph — bounded neighborhood (uses/contains/serves), depth 1..=5 GET /api/vectordb/{name}/call-graph?node=&depth= + MCP
find_similar_implementation Store::find_similar — embed the target chunk + vector-search code entries (keyword fallback), exclude the target GET /api/vectordb/{name}/similar?symbol=&top_k= + MCP

New module mosaic_knowledge::codegraph (pure-Rust, vendored into every generated app) holds the first three + is_code_entry; find_similar is a Store method (it needs the store's vector/keyword legs). The tools are bounded by construction (name lookup, reverse-edge aggregation, depth-clamped BFS, top_k-capped search) — the same "bounded by depth/node caps" posture EE takes. All four are unit-tested (tests/engine.rs, p49_*): name/file/kind matching, reverse-uses correctness, bounded neighborhood, and the code-only/target-excluded similarity filter.

This makes the graph usable for code-RAG over the deterministic K2 token-intersection uses-graph; P49b (below) makes it precise.

P49b — precise structural edges (calls + imports)

The K2 uses edges are a token-intersection: any indexed symbol name that appears in a chunk's text creates an edge. That is a coarse reference signal — a name in a comment, a doc line, or a string literal still counts, so X is "called by" any chunk that merely mentions X. P49b adds the precise structural edges derived from the source itself, on top of uses (not replacing it — the token-match still captures prose→symbol references).

New module mosaic_knowledge::code_ast (pure Rust, no parser dependency — the generated app's engine stays dependency-free, so no C FFI in every generated app). It exposes two deterministic extractors:

  • call_sites(text) -> BTreeSet<String> — identifiers actually invoked (an identifier immediately followed by (), excluding definitions (fn/def) and a Rust+Python keyword blocklist. A name in a comment or string is NOT a call site.
  • import_names(text) -> BTreeSet<String> — names pulled in by use a::b::c (Rust, the last :: segment) and import a / from a import b (Python) — the module dependencies a body token-match misses.

build_graph (graph.rs) resolves these to symbol node ids (same same-file-first / unique / first resolution as uses) and emits two new edge kinds, calls and imports. find_callers (codegraph.rs) now prefers the precise calls edges and falls back to the token-match uses only when a symbol has no calls edges — so a comment-only mention is no longer reported as a caller. get_call_graph / the graph page already walk all edge kinds, so they pick up calls / imports automatically.

File-level imports. import_names (and the per-chunk imports edges) read each chunk's text, so a module-level use/import in the file header is not in any symbol chunk and would be missed. code_ast::file_imports therefore extracts only the column-0 (module-level) import names — disjoint from the function-local ones — and the K2 ingest records them on the code-summary chunk's meta.imports. build_graph turns each cross-file one into a file → symbol imports edge (a same-file name is not a real import), so module dependencies (use crate::report::{build_report, …} → file:serve.rs → build_report) are in the graph too. Brace groups (use a::b::{x, y}), aliases (Read as R → the imported Read), and glob imports are handled.

Live check on the portal's code corpus: parse_date is mentioned in a comment inside handle_summary but only called by build_report + normalize_date. The coarse uses edge handle_summary → parse_date is still present (the mention), but the precise calls graph omits it, so find_callers(parse_date) returns exactly build_report + normalize_date.

Rust/Python are the tuned targets; other languages degrade to the name( call-site rule.

P51 — the book concept graph ("symbols for a thesis")

The code tier gets its "symbols" from the K2 census (real AST-ish functions), but the book tier only had a chapter hierarchy (book → chapter → next) + citations — no way to ask "which chapters discuss X?" the way a thesis writer needs. P51 adds a concept layer to the book graph, so a book is indexed with the same "symbol" discipline as code.

New module mosaic_knowledge::concepts (pure Rust, vendored, no deps). extract_concepts(chapter_text, heading) -> BTreeMap<String, u32> deterministically scores the salient terms of a chapter:

  • term frequency — how often a word appears (stopwords + single chars dropped);
  • capitalization signal — a mid-sentence capitalized word is a proper noun / technical term ("Projections", "Event Sourcing") and is weighted higher (a sentence-start capital is not a signal — that's just the grammar);
  • heading terms — the chapter title's words are always concepts (the title names the chapter's subject), weighted higher.

The result is ranked by (capitalized, frequency, name) and capped at MAX_CONCEPTS (25) per chapter, so the graph stays bounded. build_graph (book tier) turns each chapter's concepts into topic entities + chapter → entity mentions edges (weight = salience). find_concept(graph, name) is the thesis query: it returns the concept node id (case-insensitive) + the chapters that mention it (reverse of mentions), so you can then get_graph(node=<concept>) to expand the neighborhood or search the chapters that mention it. (P52 — below — unifies these auto concepts as type = "topic" entities, so the node id is now entity:topic:{name}, kind = entity; find_concept is a back-compat alias for find_entity(name, "topic").)

Exposed like the P49a tools: GET /api/vectordb/{name}/concept?name= + the find_concept MCP tool. With retrieval.graph: true (P48), a PPR re-rank now propagates through these concept nodes too, so a central concept (a "load-bearing" term in the thesis) resurfaces the chapters that cluster around it.

This is the what's-missing piece for the book-library / thesis use case: the deterministic, citation-ready concept graph over prose, mirroring the code symbol/uses graph. (Entity relations between concepts — "X is-a Y", "X used-by Z" — are a later slice; the concept→chapter mentions edges are the substrate.)

P52 — the typed-entity gazetteer ("entity types on all important levels")

P51's concepts are untyped topics. A thesis needs typed entities on the levels that matter — persons, locations, topics, doctrines, councils, written works, sermons/preaching, institutions, events. P52 adds a typed-entity gazetteer: the researcher curates a small dictionary (per entity: a type, a canonical name, and alias surface forms), and the engine matches it against the book chapters deterministically.

New module mosaic_knowledge::entities (pure Rust, vendored, no deps). extract_entities(text, gazetteer) -> BTreeMap<String, u32> matches each entity's surface forms against the chapter text, case-insensitive, word-boundary- anchored (so "Calvin" does not match "Calvinism"), and per entity the longest form wins per span (so "Council of Trent" is not double-counted by its "Trent" alias). It returns entity:{type}:{canonical} → mention count (only entities actually mentioned).

DSL: vector_dbs[].entities — a list of typed gazetteer entries:

entities:
  - type: person
    names: [John Calvin, Calvin, Jean Calvin]   # names[0] = canonical
  - type: council
    names: [Council of Trent, Trent]
  - type: location
    names: [Geneva]
  - type: doctrine
    names: [Sola Fide]

Types are free-form (the researcher declares whatever the domain needs); ENTITY_TYPES is the recommended standard set for a research/thesis corpus (esp. theology): person, location, topic, doctrine, council, book, sermon, institution, event, concept, other. build_graph (book tier) turns each matched entity into an entity:{type}:{name} node (kind=entity) + a chapter → entity mentions edge (weight = mention count). P51's auto topics fold in as type = "topic" entities, so a corpus gets a fully typed concept graph with zero gazetteer, and the researcher's curated entities layer on top.

find_entity(graph, name, type?) returns the entity node id + the chapters that mention it (the type filter is optional); list_entities(graph, type?) enumerates every typed entity with its mention salience, most-mentioned first. Exposed like the P49a tools: GET /api/vectordb/{name}/entity?name=&type=, GET /api/vectordb/{name}/entities?type=, + MCP find_entity / list_entities. The gazetteer is set on the Store at boot (idempotent) + persisted, so it survives reboots. With retrieval.graph: true (P48), PPR now propagates through the typed entities too — a central person/council/doctrine resurfaces the chapters that cluster around it.

This is the "symbols for a book" with types: the deterministic, citation-ready typed-entity graph over prose, mirroring the code symbol/uses graph, tuned for the library/thesis/research use case. (Gazetteer relations between entities — "X is-a Y", "X attended council Z" — and LLM-assisted entity discovery are later slices; the chapter → entity mentions substrate is what lands here.)

### P52 (runtime) — the runtime entity registry

The DSL gazetteer is a floor, set at boot from the tessera — but a researcher curating a thesis on a live server shouldn't have to edit the tessera + redeploy for every new person/council/doctrine they discover. So the Store carries a second, additive runtime_entities layer (persisted alongside the entries in the store file, restored at boot, never clobbered by the boot-time set_gazetteer of the DSL floor). Store::effective_gazetteer() is the DSL floor + the runtime layer de-duplicated by node id (declared wins), and that is what graph() matches against — so a runtime add is reflected in the graph immediately (no reindex) and survives a reboot.

Semantics mirror the K9 runtime-source registry: - add_entity appends to the runtime layer. A node id already in the DSL floor is rejected (the tessera is the source of truth for declared entities — they are immutable at runtime); re-adding an existing runtime entity is an idempotent no-op. - remove_entity(type, name) removes from the runtime layer only; removing a declared entity is rejected (edit the tessera). - gazetteer_listing() returns every row with provenance — origin = "declared" (tessera) or "runtime" (API-added) — for the listing surface.

Exposed like the rest of P52: GET /api/vectordb/{name}/gazetteer (list with provenance), POST /api/vectordb/{name}/gazetteer/add ({"type":"…","names":[…]} or {"type":"…","name":"…"}), DELETE /api/vectordb/{name}/gazetteer/remove?type=&name=, + MCP list_gazetteer / add_entity / remove_entity. Because the graph is derived on demand from the effective gazetteer, an add/remove needs no reindex — the next find_entity / list_entities / make_report already sees the change, and it persists across restarts. This is what lets the researcher (or an LLM agent via MCP) curate the entity types live as they read, without touching the tessera.

Non-goals (later slices)

  • Entity relations — edges between typed entities (is-a / used-by / attended / authored / co-occurrence) for a true thesis entity graph, on top of the P52 chapter → entity mentions substrate.
  • LLM-assisted entity discovery — an agent-layer pass (an MCP tool + the LLM seam) that extracts typed entities/relations from a chapter the deterministic gazetteer can't know in advance, feeding the researcher's gazetteer. The engine stays LLM-free (the LLM is the app/agent's, not the vendored engine's).
  • Cross-book concept alignment + disambiguation — the same surface form across books / homographs ("Trent" the council vs. the town; "Apple" the company vs. the fruit) resolved to the right typed entity.
  • tree-sitter precision upgrade — replace the pure-Rust code_ast call-site/import extractors with real tree-sitter ASTs (Rust/Python calls/imports/implements) for full-precision edges (method/qualified calls, implements). The dependency-free code_ast already delivers the high-value precise calls edges; tree-sitter would be a build-time (compiler) pass shipping pre-computed edges, keeping the app engine dependency-free.
  • P50 — wire graph-RAG (+ code-RAG) into the mosaic-pages site and deploy to li7.
  • Community-augmented retrieval (Louvain as a second re-ranking signal) and a local-vs-global retrieval mode split — the Louvain/modularity core is in (P47) for when those land.

Consequences

  • Opt-in and backward-compatible: default graph: false means existing apps and goldens are byte-identical unless they opt in.
  • The graph becomes load-bearing in retrieval, not just a rendered diagram.
  • Pure-prose collections are unaffected (no-op), so enabling it is safe.