Mosaic one model, many lenses

Re-implementation playbook

How an agent (or a human) re-implements a legacy system from its tessera + knowledge base. Everything here is deterministic and LLM-free — the LLM does the reading and the writing; mosaic supplies the ground truth.

The app's MCP server (/mcp, streamable HTTP) exposes the tools below. REST equivalents live under /api/vectordb/<kb>/….

Indexing a legacy repo (K8)

The legacy system does not have to be on the app's disk. A knowledge source can be a git repository:

vector_dbs:
  - name: kb
    sources:
      - git:
          url: "https://git.internal/legacy/billing.git"   # or a local path / .bundle file
          ref: "release-2026"                              # branch / tag / rev (optional)

The app clones it through the git CLI at boot (shallow --depth 1 for remote URLs) into a hidden cache (<knowledge-base>/.gitcache/<label>) and indexes it exactly like a local source — symbols, endpoints, git facts, hybrid search, stats, MCP tools all see it. A git bundle (git bundle create repo.bundle main) is a single committable file git clone accepts — the way to keep a repo fixture hermetic in a repo or CI. Existing checkouts are reused (delete the cache dir to force a fresh clone). mosaic knowledge-report <tessera-dir> renders the same profile offline — the re-implementation brief for a repo the app has never seen.

The loop

  1. repo_profile (kb) — the re-implementation profile in one call: the app's own DSL entity counts (aggregates, commands, events, workflows, endpoints, cli, schedules, flags), the source inventory (files, lines/pages, git facts), LOC per file (top 10), the language mix, the symbol + endpoint census, dependency manifests (parsed from Cargo.toml, package.json, go.mod, requirements.txt under the sources), the book library, and citation totals. Start every re-implementation here.
  2. search_knowledge (kb, query) — hybrid retrieval (BM25 + vector, RRF fused; MMR de-duplication; optional rerank) over the indexed chunks. Hits carry the exact locator (file:lines, book.pdf:3, doc#n) and, when a hit is a split child, its parent section for context.
  3. list_symbols / read_chunk — the code index: every function/method/ class/struct/enum with file:line locators and detected HTTP endpoints (kind: "endpoint"), and verbatim chunk reads.
  4. cite (kb, id, from?) — record a quote. Pass from (your notes section) to create a cites edge in the knowledge graph; the count feeds the stats and the report's Sources section.
  5. query_graph (kb, node?, depth?) — the deterministic graph: symbol → symbol uses-edges, file/symbol/endpoint containment, book → chapter structure, report → cited entries. Expand one node's neighborhood, or fetch the whole graph.
  6. make_report (kb, title, sections[{heading, ids[]}]) — export the notes as markdown: each section quotes its chunks verbatim (split children expand to their parent section) with the locator, followed by a Sources section built from the recorded citation counts.

CLI

The same profile renders as a sync-style markdown file without a running app:

mosaic knowledge-report <tessera-dir> [--collection <kb>] [--out sync/<name>.knowledge.md]

It ingests the tessera's declared sources with the workspace engine and writes the report (app entities, sources, top files, languages, symbols, endpoints, deps, books, git) to --out (default: stdout). Wire it into your sync pipeline to keep the re-implementation ledger current as the legacy tree changes.

Rules for the notes

  • Quote, don't paraphrase: make_report sections must reference chunk ids from search_knowledge / list_symbols / read_chunk.
  • One cite per quote, with from = the notes section heading — the graph and the report's Sources section depend on it.
  • Keep the profile fresh: re-run repo_profile (or mosaic knowledge-report) after touching the legacy tree; reindex_knowledge re-walks the sources.