Mosaic one model, many lenses

ADR 0022 — Knowledge graph: persisted document graph, entity resolution, graph-aware retrieval

Status: design. This ADR decides the shape of the next knowledge-engine step; it lands with no implementation. The implementation follows as separate work items (see Consequences), each gated on this design.

Context

The engine already has a deterministic knowledge graph (K5, mosaic-knowledge/src/graph.rs): nodes and edges derived from the indexed entries — symbol/file/endpoint (code tier), book/chapter (book tier), report/doc (citation tier) — with uses/contains/serves/ next/cites edges, a neighborhood(root, depth) BFS, and REST/MCP surfaces (/graph, /neighborhood, /citations, /cite, /report).

Three properties limit it:

  1. Ephemeral. Store::graph() recomputes the whole graph from the entries on every request. Fine at the current corpus size, but it makes the graph a projection with no identity of its own — no version, no incremental update, and the uses-edge walk is O(symbols × chunks) per call. The P6 manifest already makes re-indexing delta per source; the graph cannot participate.
  2. No entity resolution. The uses edges connect code symbols via identifier-token intersection. Nothing unifies concepts across sources: the same thing written "deploy spec" in a doc, DeploySpec in code, and "the deploy block" in a book is three disconnected strings. The graph cannot say "these entries are about the same thing", so retrieval cannot exploit it either.
  3. Flat retrieval. Search is BM25 + cosine over chunks (RRF, optional MMR, optional rerank — all P5). The graph structure is never consulted: a hit's related entries (same symbol, sibling chapter, citing report) are as invisible to the query as any other chunk.

Decision

1. The graph becomes a persisted, versioned artifact

  • Built at ingest/reindex time (the same batch points as the P6 delta walk) instead of per request, and persisted next to the index through the same data seam (ADR 0015): knowledge/.index/<name>.graph.json (local) or <persist-dir>/.index/<name>.graph.json (data-URI persist).
  • The artifact carries a version stamp: the collection's P6 manifest hash (the relpath → sha256 map) plus the engine's graph-schema version. A search finds a missing or stale artifact (stamp mismatch) and rebuilds on demand — the exact self-healing posture of the P6 manifest (missing manifest = full ingest). Reads are fast in steady state; a rebuild is bounded by the sources that changed.
  • Delta: the P6 walk already knows which sources changed/removed. Changed sources recompute their nodes/edges; unchanged sources keep theirs. Cross-source edges (a uses edge crossing two files, a cites edge to another source) recompute only when either endpoint moved. The rebuild is deterministic (the engine's standing rule: no LLM, every ordering total).
  • Store::graph() becomes "load the artifact, rebuild if stale" — the signature and the REST/MCP surfaces are unchanged.

2. Entity nodes and deterministic resolution

A new node kind, entity, plus a resolution pipeline that is deterministic and auditable (no LLM in v1 — the engine's standing rule; LLM coreference is a later env-gated seam, same posture as rerank):

  • Extraction (per tier, deterministic):
    • Code: the existing symbols — every symbol node gains an implicit entity (its normalized name).
    • Books/docs: heading nodes — chapter titles (existing chapter nodes) and, v2, markdown section headings become entity candidates.
    • Declared: a new collection key, entities: [{ name, about?, aliases: [...] }] — the operator's declared floor for the vocabulary (the engine never invents an entity from this list; it only resolves mentions to it). Declared entities are the seed for cross-source identity: "name: DeploySpec", "aliases: [deploy spec, deploy block]" unifies all three spellings above.
  • Resolution rules (total order, no randomness):
    1. normalize surface forms (Unicode NFKC, casefold, collapse whitespace, map -/_/camel boundaries to a canonical form);
    2. a mention resolves to a declared entity when its canonical form equals the entity's canonical name or an alias;
    3. otherwise it resolves to an existing code symbol when the canonical form equals a symbol's canonical name;
    4. otherwise it is an unassigned mention (not a node — no hallucinated entities).
  • Auditability: the resolution map (mention → entity id, with the rule that fired) is part of the artifact. GET /api/vectordb/{name}/entities returns the entities + their mentions; an operator can see exactly why two mentions unified — or why they didn't.
  • Edges: about edges from an entry to the entities its text resolved (count = mention count), and same_as edges only between a declared entity and the code symbol it subsumes (never between two inferred mentions — that would be coreference, out of scope for v1).

3. Graph-aware retrieval

The P5 retrieval block gains an optional third facet (fail-closed validation, the standing rule):

vector_dbs:
  - name: kb
    retrieval:
      mode: hybrid
      top_k: 5
      graph:            # NEW (ADR 0022); absent = today's behavior,
        expand: true    # byte-identical
        weight: 0.2     # 0..=1, the expansion boost (default 0.2)
        hops: 1         # 1 | 2 (default 1)
  • Expansion pass: after the existing retrieval (RRF/MMR/rerank untouched), each returned hit's entity nodes seed a 1–2 hop neighborhood; neighbor entries whose text is not already in the results gain weight × (edge_count / 2^hops), and the final top_k is re-selected. Bounded by construction: 1 hop, top-N neighbors (N = 2 × top_k), source-scoped when the search is source-scoped (P5).
  • Deterministic: the boost is a pure function of the (persisted) graph and the candidate set; ties break on the existing entry-order rule. A graph: block on a collection whose artifact has no entity nodes is a no-op (nothing to expand), never an error.
  • related surface: GET /api/vectordb/{name}/related/{entry_id}?depth=1..5 returns the neighborhood as answer context (nodes with their entry ids + a one-line label each) — the follow-up material the assistant/chat surfaces consume ("what else is about this?"). Backed by the same artifact, so it is a read, not a recompute.
  • MCP: two new tools beside search_knowledge / list_sources — entities (the resolved vocabulary, source-scoped) and related (neighborhood of an entry). The assistant stays global, as the search source knob is per-search (P5 posture).

Consequences

  • Persisted/artifacts gain the graph file; the data seam (local fs or URI) is the only write path — no new storage backend, no external graph store. The engine stays embeddable and CI-hermetic.
  • vector_dbs[].entities is a new DSL key (YAML + plan validation: duplicate names fail closed; aliases may overlap with symbol names — the resolution order above decides). Conformance goldens for collections that set it are re-pinned; unset collections are byte-identical.
  • Implementation is staged, each stage independently shippable:
    1. v1: persisted artifact + version stamp + delta rebuild + entities declared floor + resolution map + entities/related surfaces. (No retrieval change yet — the graph is queryable but not consulted.)
    2. v2: retrieval.graph expansion + markdown-heading extraction.
    3. v3 (env-gated seam, optional): LLM-assisted coreference for unassigned mentions — fail-soft to the deterministic result, keyed through the P17 model registry (coreference role), off by default, and its decisions persisted for audit like the v1 map.
  • Out of scope (deliberately): a general graph DB, cross-collection edges, temporal/versioned graphs (the artifact is a snapshot per manifest state), and anything that makes retrieval nondeterministic.