Mosaic one model, many lenses

ADR 0011: Durable store — append per dispatch, snapshot the whole store, replay the tail

  • Status: accepted
  • Date: 2026-09-30
  • Supersedes: nothing (extends the persistence construct introduced for the durable event log)

Context

The README roadmap names two ADR-gated items: durable stores and authorization enforcement on every route. This ADR is the first.

A generated app with persistence declared already has a durable event log (JSONL or SQLite): every command dispatch appends the events it emitted (with_react), and boot replays the log to rebuild aggregate state and materialized projections. Two gaps remain:

  1. Replay cannot rebuild everything. The replay path (replay_event) only re-applies aggregate events. State that lives on the Store but is not an aggregate fact — the instance-level IAM grants recorded by on_create auto-assigners (instance_grants) and the enterprise audit trail (audit, when the audit platform part is on) — is silently lost at every restart. An app that survives a crash with the right orders but the wrong grants and an empty audit log is not durable.
  2. Boot cost is linear in log size. Every start replays the whole log, including the part a snapshot would make redundant.

The Store is already serde::Serialize (the state structs derive both Serialize and Deserialize), so persisting it whole is a serialization question, not a modeling question.

Decision

Four rules, all in the generated app (the engine crates stay pure):

  1. One persistence chokepoint. server::react_and_persist(store, since) runs the event reactors for the new events, appends them to the log, and snapshots on the cadence. Every path that appends events to the live store goes through it: CQRS command dispatch (with_react), workflow runs (/api/workflows/{slug}/run streamed and plain, workflow-source endpoints, triggers, schedules, MCP tools) and the HITL resume. This closes the gap that made the log durable only for direct command dispatches — workflow-emitted events were previously persisted only at bootstrap/shutdown, so a kill between them lost them.
  2. Append per dispatch (kept). with_react keeps appending the events it emitted to the log before returning, so the log is never behind the store by more than one in-flight dispatch.
  3. Snapshot the whole store, on a cadence. New DSL key persistence.snapshot_every: <n> (events; default 256 when persistence is declared, 0 disables). A snapshot is the complete Store serialized as JSON plus an events watermark (store.events.len() at snapshot time):
    • jsonl backend: sidecar file <data_file>.snap.json
    • sqlite backend: kv(key TEXT PRIMARY KEY, value TEXT) table, key snapshot — same file as the log, so log and snapshot advance together Writes happen (a) in with_react once n events have accumulated since the last snapshot, and (b) on the graceful-shutdown path after the final append. The store carries a #[serde(skip)] snapshot_at: usize watermark so the cadence survives restarts and a failed write retries on the next dispatch. Snapshot writes are best-effort: a failure warns, never fails the request (the same fail-soft contract as the log append).
  4. Boot: snapshot first, replay the tail. Boot reads the log, then tries the snapshot: if it parses and its watermark is <= the log length, the store is adopted from it and only log[watermark..] is replayed through the existing replay_event. No snapshot, an unparseable one, or a watermark past the end of the log (log truncated out from under us) all degrade to a full replay with one warning — the pre-ADR behavior, so a broken snapshot can never lose more than a fresh replay would.

The ordering invariant that makes this safe: the snapshot is only ever written after the log append for the same events (in react_and_persist, reactors + append first, snapshot last; same at shutdown). So a crash can leave the log ahead of the snapshot (tail replay repairs the replayable state) or both behind the last in-flight dispatch (at-most-once per dispatch — the existing contract, unchanged), but never the snapshot ahead of the log. The known window: side effects a replay cannot rebuild (an on_create instance grant) are only restored when a snapshot post-dates the event that caused them — snapshot_every bounds that window.

Consequences

  • Restart preserves instance_grants and the audit trail, not just aggregates; workflow-emitted events are persisted per run, not only at shutdown; boot skips the replayed prefix.
  • One more file (jsonl) or one more table (sqlite); no new dependency (serde is already in the generated app).
  • snapshot_every: 0 reproduces today's replay-only behavior exactly, so existing apps that add persistence are unaffected unless they ask for the cadence.
  • The proof (ADR gate) is an e2e pair in the full-stack example: test 1 runs the full scenario set (workflow runs + CQRS commands, ending in a ticket creation whose on_create auto-assigner records an instance grant) and is SIGKILLed by the harness (no graceful shutdown); test 2 boots the same data file on a new port and asserts the aggregate state and the instance grant (which replay alone cannot rebuild) survived — via the mid-run snapshots written at snapshot_every: 1.