Thesis / book-library graph-RAG with Mosaic
How to put a book library into a Mosaic app's knowledge base and query it as a typed-entity graph — persons, places, topics, doctrines, councils, written works, sermons — for thesis writing and research. This is the ADR 0047 P47–P52 feature set (graph-augmented retrieval + code-RAG + book concepts + typed entities).
The model, in one line: every book chapter is a graph node; the chapter's
typed entities (your curated gazetteer) and topics (auto-extracted
salient terms) are nodes too; chapter → entity mentions edges connect them;
Personalized PageRank (P48) then resurfaces the load-bearing entities/chapters
around any query.
1. What you get
| Surface | What it answers |
|---|---|
GET /api/vectordb/{kb}/entity?name=X&type= | "Which chapters discuss entity X?" (optionally within one type) |
GET /api/vectordb/{kb}/entities?type= | "List all persons / councils / doctrines … with how often each is mentioned" |
GET /api/vectordb/{kb}/concept?name=X | Same, but for auto-extracted topics (P51; find_concept = find_entity(X, "topic")) |
GET /api/vectordb/{kb}/graph?node=…&depth=… | Expand any entity/chapter into its neighborhood (who mentions it, what it co-mentions) |
GET /api/vectordb/{kb}/search?q=… | Hybrid (BM25 + vector) retrieval, re-ranked by PPR when retrieval.graph: true |
GET /api/vectordb/{kb}/gazetteer | The full typed-entity gazetteer with provenance (tessera floor + runtime layer) |
POST /api/vectordb/{kb}/gazetteer/add · DELETE …/gazetteer/remove | Curate the gazetteer at runtime (no redeploy, no reindex, persists across reboots) |
MCP find_entity / list_entities / find_concept / get_graph / search_knowledge / make_report | The same, for an LLM agent |
MCP list_gazetteer / add_entity / remove_entity | Curate the typed entities live, for an LLM agent |
POST /api/vectordb/{kb}/report / MCP make_report | Export a thesis section quoting the chunks verbatim, with locators + a Sources list |
Entity types are free-form — declare whatever your domain needs. The
recommended standard set (ENTITY_TYPES): person, location, topic, doctrine, council, book, sermon, institution, event, concept, other. For a theology thesis
you'd typically use: person (theologians, popes, reformers), location, topic,
doctrine (doctrine statements), council (ecumenical/local councils — "congresses"),
book (written works), sermon (preaching), institution (churches, orders,
universities), event.
2. Add the library
-
Copy the books (EPUB / DOCX / PDF) into the app's knowledge area. For the
mosaic-pages-style site that is thebooks/folder (declared as a source):vector_dbs: - name: docs sources: - path: books # EPUB / DOCX / PDF chapter model kind: auto strategy: books: true # parse EPUB (OPF spine) / DOCX (Heading1) / PDF chaptersPDFs are best-effort (chapter markers via a heading heuristic; set
strategy.ocr_cmd/MOAIC_OCR_CMDfor scanned pages). EPUB and DOCX give the cleanest chapter model. -
On li7, after copying the files, force a full reindex (delta ingest keeps unchanged entries, so a new book set needs a full rebuild):
curl -u <user>:<pass> -X POST https://<host>/api/vectordb/docs/reindex
3. Declare the gazetteer (the important part)
The typed entities come from a small curated dictionary in the tessera — the
researcher (you) knows the key entities of the thesis. Per entry: a type, a
canonical name (names[0]), and alias surface forms (the rest):
vector_dbs:
- name: docs
retrieval:
graph: true # P48: PPR re-ranking (resurfaces load-bearing entities)
graph_weight: 0.5
entities:
- type: person
names: [John Calvin, Calvin, Jean Calvin]
- type: person
names: [Martin Luther, Luther]
- type: council
names: [Council of Trent, Trent, the Council of Trent]
- type: council
names: [Council of Nicaea, Nicaea, First Council of Nicaea]
- type: location
names: [Geneva]
- type: location
names: [Wittenberg]
- type: doctrine
names: [Sola Fide, sola fide, faith alone]
- type: doctrine
names: [Imputed Righteousness, imputation of righteousness]
- type: book
names: [Institutes of the Christian Religion, the Institutes]
- type: sermon
names: [Bannerman's Lectures, Lectures on the Church]
Rules of the matcher (deterministic, no LLM):
- case-insensitive; word-boundary-anchored (
Calvindoes not matchCalvinism/Calvinist); - longest form wins per span (
Council of Trentis not double-counted by itsTrentalias); - a surface form may be a multi-word phrase (
Imputed Righteousness); - the same surface form in two types yields two entities (
council:Trentandlocation:Trentboth) — curate deliberately.
The auto topics (P51) need no gazetteer: each chapter's salient terms become
type = "topic" entities, so even an uncurated library gets a typed topic graph.
3b. Curate the gazetteer at runtime (no tessera edit, no redeploy)
The tessera gazetteer is the immutable floor. On top of it there is a
runtime registry — you (or an LLM agent via MCP) add/remove typed entities on
the live server as you read, and they are reflected in the graph immediately
(no reindex) and persisted across reboots. This is the fast loop for thesis
work: read a chapter, spot a person/council/doctrine you hadn't declared, add it,
and the next find_entity / make_report already uses it.
# Add a typed entity (names[0] = canonical, the rest = aliases):
curl -u … -X POST "…/api/vectordb/docs/gazetteer/add" \
-H 'Content-Type: application/json' \
-d '{"type":"council","names":["Council of Nicaea","Nicaea"]}'
# …or a single surface form:
curl -u … -X POST "…/api/vectordb/docs/gazetteer/add" \
-d '{"type":"person","name":"Thomas Aquinas"}'
# List the whole gazetteer, with provenance (origin: "declared" | "runtime"):
curl -u … "…/api/vectordb/docs/gazetteer"
# Remove a runtime entity (declared/floor entities are immutable — edit the tessera):
curl -u … -X DELETE "…/api/vectordb/docs/gazetteer/remove?type=council&name=Council%20of%20Nicaea"
MCP equivalents: add_entity / remove_entity / list_gazetteer.
Rules (same determinism as the DSL): a runtime add whose node id is already in the
tessera floor is rejected (the tessera stays the source of truth for declared
entities); re-adding a runtime entity is an idempotent no-op; remove only affects
the runtime layer. Because the graph is derived on demand from the effective
gazetteer (floor + runtime), no reindex is needed — the change is live on the
next query and survives a restart. When an entity graduates from "runtime" to
"core to the thesis", promote it into the tessera entities: block so it becomes
part of the immutable floor.
4. Query / research loop
# What is the corpus about, by type?
curl -u … "…/api/vectordb/docs/entities?type=person"
curl -u … "…/api/vectordb/docs/entities?type=council"
# Which chapters discuss a person, and how centrally?
curl -u … "…/api/vectordb/docs/entity?name=John%20Calvin"
# Expand the neighborhood (co-mentions, shared chapters):
curl -u … "…/api/vectordb/docs/graph?node=entity:person:John%20Calvin&depth=2"
# A grounded retrieval, PPR-re-ranked:
curl -u … -X POST "…/api/vectordb/docs/search" -d '{"query":"predestination in Calvin"}'
Then export a draft section that quotes the chapters verbatim (with locators) — the citation-ready thesis artifact:
curl -u … -X POST "…/api/vectordb/docs/report" \
-d '{"title":"Predestination in Calvin","sections":[{"heading":"The doctrine","ids":["<chunk-id>", …]}]}'
An LLM agent gets the same through the MCP server (find_entity, list_entities,
get_graph, search_knowledge, make_report) — have it enumerate the entities,
expand the ones that matter, and draft from the cited chunks.
5. Best practices (libraries / thesis / research)
- Curate the gazetteer iteratively. Start with the ~20–50 load-bearing
entities (the thesis's named persons, councils, doctrines, key works). Query
list_entities(all types) to see what the auto topics surface, then promote the important ones to curated typed entities with proper types + aliases. - Use types to disambiguate + to slice.
list_entities?type=personvs?type=doctrineis how a researcher navigates; the type is also the query filter, so "Trent" the council ≠ "Trent" the town. - Keep aliases. People and places have many surface forms (Latin/Greek names, abbreviations, "the Council of Trent" vs "Trent"). List them all — the matcher handles multi-word + case.
- Lean on PPR, don't just keyword-search. With
retrieval.graph: true, a query that touches a central entity resurfaces the chapters clustered around it (HippoRAG-style), which beats flat keyword ranking for "where is this doctrine developed across the corpus?" - Cite from the graph, not from memory. Use
make_report/POST /reportto build sections from real chunk ids — every claim carries a verbatim quote + a locator (book/chapter or file:line), and a Sources list is generated from the citation counts. - One collection per corpus, but many sources. The books, your notes, and any
git repo of drafts all feed the same
docscollection and the same graph, so PPR spans "the library + my notes." - Deterministic = reproducible. No LLM in the engine: the same library + gazetteer always yields the same graph. Re-index is idempotent.
6. What's here vs. what's a later slice
- Here (P47–P52): hybrid retrieval + PPR re-rank; code symbol/uses/calls/
imports graph; book chapter graph; auto topic concepts (P51); typed-entity
gazetteer (P52) with
find_entity/list_entities+ a runtime entity registry (add/remove/list typed entities on the live server, no redeploy, persists across reboots); report/citation export. - Later slices (not yet built):
- Entity relations — edges between entities (
is-a,authored,attended-council,co-occurrence) for a true thesis entity graph, on top of thechapter → entitymentionssubstrate. - LLM-assisted entity discovery — an agent-layer pass (an MCP tool + the LLM seam) that proposes typed entities/relations the deterministic gazetteer can't know in advance, to seed your gazetteer. The vendored engine stays LLM-free.
- Cross-book alignment + disambiguation — the same surface form across books / homographs resolved to the right typed entity automatically.
- Entity relations — edges between entities (