ADR 0011: Durable store — append per dispatch, snapshot the whole store, replay the tail
- Status: accepted
- Date: 2026-09-30
- Supersedes: nothing (extends the
persistenceconstruct introduced for the durable event log)
Context
The README roadmap names two ADR-gated items: durable stores and authorization enforcement on every route. This ADR is the first.
A generated app with persistence declared already has a durable event
log (JSONL or SQLite): every command dispatch appends the events it emitted
(with_react), and boot replays the log to rebuild aggregate state and
materialized projections. Two gaps remain:
- Replay cannot rebuild everything. The replay path
(
replay_event) only re-applies aggregate events. State that lives on theStorebut is not an aggregate fact — the instance-level IAM grants recorded byon_createauto-assigners (instance_grants) and the enterprise audit trail (audit, when the audit platform part is on) — is silently lost at every restart. An app that survives a crash with the right orders but the wrong grants and an empty audit log is not durable. - Boot cost is linear in log size. Every start replays the whole log, including the part a snapshot would make redundant.
The Store is already serde::Serialize (the state structs derive both
Serialize and Deserialize), so persisting it whole is a serialization
question, not a modeling question.
Decision
Four rules, all in the generated app (the engine crates stay pure):
- One persistence chokepoint.
server::react_and_persist(store, since)runs the event reactors for the new events, appends them to the log, and snapshots on the cadence. Every path that appends events to the live store goes through it: CQRS command dispatch (with_react), workflow runs (/api/workflows/{slug}/runstreamed and plain, workflow-source endpoints, triggers, schedules, MCP tools) and the HITL resume. This closes the gap that made the log durable only for direct command dispatches — workflow-emitted events were previously persisted only at bootstrap/shutdown, so a kill between them lost them. - Append per dispatch (kept).
with_reactkeeps appending the events it emitted to the log before returning, so the log is never behind the store by more than one in-flight dispatch. - Snapshot the whole store, on a cadence. New DSL key
persistence.snapshot_every: <n>(events; default256whenpersistenceis declared,0disables). A snapshot is the completeStoreserialized as JSON plus aneventswatermark (store.events.len()at snapshot time):- jsonl backend: sidecar file
<data_file>.snap.json - sqlite backend:
kv(key TEXT PRIMARY KEY, value TEXT)table, keysnapshot— same file as the log, so log and snapshot advance together Writes happen (a) inwith_reactoncenevents have accumulated since the last snapshot, and (b) on the graceful-shutdown path after the final append. The store carries a#[serde(skip)] snapshot_at: usizewatermark so the cadence survives restarts and a failed write retries on the next dispatch. Snapshot writes are best-effort: a failure warns, never fails the request (the same fail-soft contract as the log append).
- jsonl backend: sidecar file
- Boot: snapshot first, replay the tail. Boot reads the log, then tries
the snapshot: if it parses and its watermark is
<=the log length, the store is adopted from it and onlylog[watermark..]is replayed through the existingreplay_event. No snapshot, an unparseable one, or a watermark past the end of the log (log truncated out from under us) all degrade to a full replay with one warning — the pre-ADR behavior, so a broken snapshot can never lose more than a fresh replay would.
The ordering invariant that makes this safe: the snapshot is only ever
written after the log append for the same events (in react_and_persist,
reactors + append first, snapshot last; same at shutdown). So a crash can
leave the log ahead of the snapshot (tail replay repairs the replayable
state) or both behind the last in-flight dispatch (at-most-once per
dispatch — the existing contract, unchanged), but never the snapshot ahead
of the log. The known window: side effects a replay cannot rebuild (an
on_create instance grant) are only restored when a snapshot post-dates
the event that caused them — snapshot_every bounds that window.
Consequences
- Restart preserves
instance_grantsand the audit trail, not just aggregates; workflow-emitted events are persisted per run, not only at shutdown; boot skips the replayed prefix. - One more file (jsonl) or one more table (sqlite); no new dependency (serde is already in the generated app).
snapshot_every: 0reproduces today's replay-only behavior exactly, so existing apps that addpersistenceare unaffected unless they ask for the cadence.- The proof (ADR gate) is an e2e pair in the full-stack example: test 1
runs the full scenario set (workflow runs + CQRS commands, ending in a
ticket creation whose
on_createauto-assigner records an instance grant) and is SIGKILLed by the harness (no graceful shutdown); test 2 boots the same data file on a new port and asserts the aggregate state and the instance grant (which replay alone cannot rebuild) survived — via the mid-run snapshots written atsnapshot_every: 1.