ADR 0022 — Knowledge graph: persisted document graph, entity resolution, graph-aware retrieval
Status: design. This ADR decides the shape of the next knowledge-engine step; it lands with no implementation. The implementation follows as separate work items (see Consequences), each gated on this design.
Context
The engine already has a deterministic knowledge graph (K5,
mosaic-knowledge/src/graph.rs): nodes and edges derived from the indexed
entries — symbol/file/endpoint (code tier), book/chapter (book
tier), report/doc (citation tier) — with uses/contains/serves/
next/cites edges, a neighborhood(root, depth) BFS, and REST/MCP
surfaces (/graph, /neighborhood, /citations, /cite, /report).
Three properties limit it:
- Ephemeral.
Store::graph()recomputes the whole graph from the entries on every request. Fine at the current corpus size, but it makes the graph a projection with no identity of its own — no version, no incremental update, and theuses-edge walk is O(symbols × chunks) per call. The P6 manifest already makes re-indexing delta per source; the graph cannot participate. - No entity resolution. The
usesedges connect code symbols via identifier-token intersection. Nothing unifies concepts across sources: the same thing written "deploy spec" in a doc,DeploySpecin code, and "the deploy block" in a book is three disconnected strings. The graph cannot say "these entries are about the same thing", so retrieval cannot exploit it either. - Flat retrieval. Search is BM25 + cosine over chunks (RRF, optional MMR, optional rerank — all P5). The graph structure is never consulted: a hit's related entries (same symbol, sibling chapter, citing report) are as invisible to the query as any other chunk.
Decision
1. The graph becomes a persisted, versioned artifact
- Built at ingest/reindex time (the same batch points as the P6 delta
walk) instead of per request, and persisted next to the index through
the same data seam (ADR 0015):
knowledge/.index/<name>.graph.json(local) or<persist-dir>/.index/<name>.graph.json(data-URI persist). - The artifact carries a version stamp: the collection's P6 manifest
hash (the
relpath → sha256map) plus the engine's graph-schema version. A search finds a missing or stale artifact (stamp mismatch) and rebuilds on demand — the exact self-healing posture of the P6 manifest (missing manifest = full ingest). Reads are fast in steady state; a rebuild is bounded by the sources that changed. - Delta: the P6 walk already knows which sources changed/removed.
Changed sources recompute their nodes/edges; unchanged sources keep
theirs. Cross-source edges (a
usesedge crossing two files, acitesedge to another source) recompute only when either endpoint moved. The rebuild is deterministic (the engine's standing rule: no LLM, every ordering total). Store::graph()becomes "load the artifact, rebuild if stale" — the signature and the REST/MCP surfaces are unchanged.
2. Entity nodes and deterministic resolution
A new node kind, entity, plus a resolution pipeline that is
deterministic and auditable (no LLM in v1 — the engine's standing
rule; LLM coreference is a later env-gated seam, same posture as
rerank):
- Extraction (per tier, deterministic):
- Code: the existing symbols — every symbol node gains an implicit entity (its normalized name).
- Books/docs: heading nodes — chapter titles (existing
chapternodes) and, v2, markdown section headings become entity candidates. - Declared: a new collection key,
entities: [{ name, about?, aliases: [...] }]— the operator's declared floor for the vocabulary (the engine never invents an entity from this list; it only resolves mentions to it). Declared entities are the seed for cross-source identity:"name: DeploySpec", "aliases: [deploy spec, deploy block]"unifies all three spellings above.
- Resolution rules (total order, no randomness):
- normalize surface forms (Unicode NFKC, casefold, collapse
whitespace, map
-/_/camel boundaries to a canonical form); - a mention resolves to a declared entity when its canonical form equals the entity's canonical name or an alias;
- otherwise it resolves to an existing code symbol when the canonical form equals a symbol's canonical name;
- otherwise it is an unassigned mention (not a node — no hallucinated entities).
- normalize surface forms (Unicode NFKC, casefold, collapse
whitespace, map
- Auditability: the resolution map (mention → entity id, with the
rule that fired) is part of the artifact.
GET /api/vectordb/{name}/entitiesreturns the entities + their mentions; an operator can see exactly why two mentions unified — or why they didn't. - Edges:
aboutedges from an entry to the entities its text resolved (count = mention count), andsame_asedges only between a declared entity and the code symbol it subsumes (never between two inferred mentions — that would be coreference, out of scope for v1).
3. Graph-aware retrieval
The P5 retrieval block gains an optional third facet (fail-closed
validation, the standing rule):
vector_dbs:
- name: kb
retrieval:
mode: hybrid
top_k: 5
graph: # NEW (ADR 0022); absent = today's behavior,
expand: true # byte-identical
weight: 0.2 # 0..=1, the expansion boost (default 0.2)
hops: 1 # 1 | 2 (default 1)
- Expansion pass: after the existing retrieval (RRF/MMR/rerank
untouched), each returned hit's entity nodes seed a 1–2 hop
neighborhood; neighbor entries whose text is not already in the
results gain
weight × (edge_count / 2^hops), and the final top_k is re-selected. Bounded by construction: 1 hop, top-N neighbors (N = 2 × top_k), source-scoped when the search is source-scoped (P5). - Deterministic: the boost is a pure function of the (persisted)
graph and the candidate set; ties break on the existing
entry-order rule. A
graph:block on a collection whose artifact has no entity nodes is a no-op (nothing to expand), never an error. relatedsurface:GET /api/vectordb/{name}/related/{entry_id}?depth=1..5returns the neighborhood as answer context (nodes with their entry ids + a one-line label each) — the follow-up material the assistant/chat surfaces consume ("what else is about this?"). Backed by the same artifact, so it is a read, not a recompute.- MCP: two new tools beside
search_knowledge/list_sources—entities(the resolved vocabulary, source-scoped) andrelated(neighborhood of an entry). The assistant stays global, as the search source knob is per-search (P5 posture).
Consequences
Persisted/artifacts gain the graph file; the data seam (local fs or URI) is the only write path — no new storage backend, no external graph store. The engine stays embeddable and CI-hermetic.vector_dbs[].entitiesis a new DSL key (YAML + plan validation: duplicate names fail closed; aliases may overlap with symbol names — the resolution order above decides). Conformance goldens for collections that set it are re-pinned; unset collections are byte-identical.- Implementation is staged, each stage independently shippable:
- v1: persisted artifact + version stamp + delta rebuild +
entitiesdeclared floor + resolution map +entities/relatedsurfaces. (No retrieval change yet — the graph is queryable but not consulted.) - v2:
retrieval.graphexpansion + markdown-heading extraction. - v3 (env-gated seam, optional): LLM-assisted coreference for
unassigned mentions — fail-soft to the deterministic result, keyed
through the P17 model registry (
coreferencerole), off by default, and its decisions persisted for audit like the v1 map.
- v1: persisted artifact + version stamp + delta rebuild +
- Out of scope (deliberately): a general graph DB, cross-collection edges, temporal/versioned graphs (the artifact is a snapshot per manifest state), and anything that makes retrieval nondeterministic.