Mosaic one model, many lenses

Thesis / book-library graph-RAG with Mosaic

How to put a book library into a Mosaic app's knowledge base and query it as a typed-entity graph — persons, places, topics, doctrines, councils, written works, sermons — for thesis writing and research. This is the ADR 0047 P47–P52 feature set (graph-augmented retrieval + code-RAG + book concepts + typed entities).

The model, in one line: every book chapter is a graph node; the chapter's typed entities (your curated gazetteer) and topics (auto-extracted salient terms) are nodes too; chapter → entity mentions edges connect them; Personalized PageRank (P48) then resurfaces the load-bearing entities/chapters around any query.


1. What you get

SurfaceWhat it answers
GET /api/vectordb/{kb}/entity?name=X&type="Which chapters discuss entity X?" (optionally within one type)
GET /api/vectordb/{kb}/entities?type="List all persons / councils / doctrines … with how often each is mentioned"
GET /api/vectordb/{kb}/concept?name=XSame, but for auto-extracted topics (P51; find_concept = find_entity(X, "topic"))
GET /api/vectordb/{kb}/graph?node=…&depth=…Expand any entity/chapter into its neighborhood (who mentions it, what it co-mentions)
GET /api/vectordb/{kb}/search?q=…Hybrid (BM25 + vector) retrieval, re-ranked by PPR when retrieval.graph: true
GET /api/vectordb/{kb}/gazetteerThe full typed-entity gazetteer with provenance (tessera floor + runtime layer)
POST /api/vectordb/{kb}/gazetteer/add · DELETE …/gazetteer/removeCurate the gazetteer at runtime (no redeploy, no reindex, persists across reboots)
MCP find_entity / list_entities / find_concept / get_graph / search_knowledge / make_reportThe same, for an LLM agent
MCP list_gazetteer / add_entity / remove_entityCurate the typed entities live, for an LLM agent
POST /api/vectordb/{kb}/report / MCP make_reportExport a thesis section quoting the chunks verbatim, with locators + a Sources list

Entity types are free-form — declare whatever your domain needs. The recommended standard set (ENTITY_TYPES): person, location, topic, doctrine, council, book, sermon, institution, event, concept, other. For a theology thesis you'd typically use: person (theologians, popes, reformers), location, topic, doctrine (doctrine statements), council (ecumenical/local councils — "congresses"), book (written works), sermon (preaching), institution (churches, orders, universities), event.


2. Add the library

  1. Copy the books (EPUB / DOCX / PDF) into the app's knowledge area. For the mosaic-pages-style site that is the books/ folder (declared as a source):

    vector_dbs:
      - name: docs
        sources:
          - path: books          # EPUB / DOCX / PDF chapter model
            kind: auto
        strategy:
          books: true            # parse EPUB (OPF spine) / DOCX (Heading1) / PDF chapters
    

    PDFs are best-effort (chapter markers via a heading heuristic; set strategy.ocr_cmd / MOAIC_OCR_CMD for scanned pages). EPUB and DOCX give the cleanest chapter model.

  2. On li7, after copying the files, force a full reindex (delta ingest keeps unchanged entries, so a new book set needs a full rebuild):

    curl -u <user>:<pass> -X POST https://<host>/api/vectordb/docs/reindex
    

3. Declare the gazetteer (the important part)

The typed entities come from a small curated dictionary in the tessera — the researcher (you) knows the key entities of the thesis. Per entry: a type, a canonical name (names[0]), and alias surface forms (the rest):

vector_dbs:
  - name: docs
    retrieval:
      graph: true          # P48: PPR re-ranking (resurfaces load-bearing entities)
      graph_weight: 0.5
    entities:
      - type: person
        names: [John Calvin, Calvin, Jean Calvin]
      - type: person
        names: [Martin Luther, Luther]
      - type: council
        names: [Council of Trent, Trent, the Council of Trent]
      - type: council
        names: [Council of Nicaea, Nicaea, First Council of Nicaea]
      - type: location
        names: [Geneva]
      - type: location
        names: [Wittenberg]
      - type: doctrine
        names: [Sola Fide, sola fide, faith alone]
      - type: doctrine
        names: [Imputed Righteousness, imputation of righteousness]
      - type: book
        names: [Institutes of the Christian Religion, the Institutes]
      - type: sermon
        names: [Bannerman's Lectures, Lectures on the Church]

Rules of the matcher (deterministic, no LLM):

  • case-insensitive; word-boundary-anchored (Calvin does not match Calvinism / Calvinist);
  • longest form wins per span (Council of Trent is not double-counted by its Trent alias);
  • a surface form may be a multi-word phrase (Imputed Righteousness);
  • the same surface form in two types yields two entities (council:Trent and location:Trent both) — curate deliberately.

The auto topics (P51) need no gazetteer: each chapter's salient terms become type = "topic" entities, so even an uncurated library gets a typed topic graph.

3b. Curate the gazetteer at runtime (no tessera edit, no redeploy)

The tessera gazetteer is the immutable floor. On top of it there is a runtime registry — you (or an LLM agent via MCP) add/remove typed entities on the live server as you read, and they are reflected in the graph immediately (no reindex) and persisted across reboots. This is the fast loop for thesis work: read a chapter, spot a person/council/doctrine you hadn't declared, add it, and the next find_entity / make_report already uses it.

# Add a typed entity (names[0] = canonical, the rest = aliases):
curl -u … -X POST "…/api/vectordb/docs/gazetteer/add" \
  -H 'Content-Type: application/json' \
  -d '{"type":"council","names":["Council of Nicaea","Nicaea"]}'

# …or a single surface form:
curl -u … -X POST "…/api/vectordb/docs/gazetteer/add" \
  -d '{"type":"person","name":"Thomas Aquinas"}'

# List the whole gazetteer, with provenance (origin: "declared" | "runtime"):
curl -u … "…/api/vectordb/docs/gazetteer"

# Remove a runtime entity (declared/floor entities are immutable — edit the tessera):
curl -u … -X DELETE "…/api/vectordb/docs/gazetteer/remove?type=council&name=Council%20of%20Nicaea"

MCP equivalents: add_entity / remove_entity / list_gazetteer.

Rules (same determinism as the DSL): a runtime add whose node id is already in the tessera floor is rejected (the tessera stays the source of truth for declared entities); re-adding a runtime entity is an idempotent no-op; remove only affects the runtime layer. Because the graph is derived on demand from the effective gazetteer (floor + runtime), no reindex is needed — the change is live on the next query and survives a restart. When an entity graduates from "runtime" to "core to the thesis", promote it into the tessera entities: block so it becomes part of the immutable floor.

4. Query / research loop

# What is the corpus about, by type?
curl -u … "…/api/vectordb/docs/entities?type=person"
curl -u … "…/api/vectordb/docs/entities?type=council"

# Which chapters discuss a person, and how centrally?
curl -u … "…/api/vectordb/docs/entity?name=John%20Calvin"

# Expand the neighborhood (co-mentions, shared chapters):
curl -u … "…/api/vectordb/docs/graph?node=entity:person:John%20Calvin&depth=2"

# A grounded retrieval, PPR-re-ranked:
curl -u … -X POST "…/api/vectordb/docs/search" -d '{"query":"predestination in Calvin"}'

Then export a draft section that quotes the chapters verbatim (with locators) — the citation-ready thesis artifact:

curl -u … -X POST "…/api/vectordb/docs/report" \
  -d '{"title":"Predestination in Calvin","sections":[{"heading":"The doctrine","ids":["<chunk-id>", …]}]}'

An LLM agent gets the same through the MCP server (find_entity, list_entities, get_graph, search_knowledge, make_report) — have it enumerate the entities, expand the ones that matter, and draft from the cited chunks.


5. Best practices (libraries / thesis / research)

  • Curate the gazetteer iteratively. Start with the ~20–50 load-bearing entities (the thesis's named persons, councils, doctrines, key works). Query list_entities (all types) to see what the auto topics surface, then promote the important ones to curated typed entities with proper types + aliases.
  • Use types to disambiguate + to slice. list_entities?type=person vs ?type=doctrine is how a researcher navigates; the type is also the query filter, so "Trent" the council ≠ "Trent" the town.
  • Keep aliases. People and places have many surface forms (Latin/Greek names, abbreviations, "the Council of Trent" vs "Trent"). List them all — the matcher handles multi-word + case.
  • Lean on PPR, don't just keyword-search. With retrieval.graph: true, a query that touches a central entity resurfaces the chapters clustered around it (HippoRAG-style), which beats flat keyword ranking for "where is this doctrine developed across the corpus?"
  • Cite from the graph, not from memory. Use make_report / POST /report to build sections from real chunk ids — every claim carries a verbatim quote + a locator (book/chapter or file:line), and a Sources list is generated from the citation counts.
  • One collection per corpus, but many sources. The books, your notes, and any git repo of drafts all feed the same docs collection and the same graph, so PPR spans "the library + my notes."
  • Deterministic = reproducible. No LLM in the engine: the same library + gazetteer always yields the same graph. Re-index is idempotent.

6. What's here vs. what's a later slice

  • Here (P47–P52): hybrid retrieval + PPR re-rank; code symbol/uses/calls/ imports graph; book chapter graph; auto topic concepts (P51); typed-entity gazetteer (P52) with find_entity / list_entities + a runtime entity registry (add/remove/list typed entities on the live server, no redeploy, persists across reboots); report/citation export.
  • Later slices (not yet built):
    • Entity relations — edges between entities (is-a, authored, attended-council, co-occurrence) for a true thesis entity graph, on top of the chapter → entity mentions substrate.
    • LLM-assisted entity discovery — an agent-layer pass (an MCP tool + the LLM seam) that proposes typed entities/relations the deterministic gazetteer can't know in advance, to seed your gazetteer. The vendored engine stays LLM-free.
    • Cross-book alignment + disambiguation — the same surface form across books / homographs resolved to the right typed entity automatically.