Relationships
#3055 Datastore rework Phase 2: one flush path, so every push is a root's flush or checkpoint (no behaviour change)
Opened by stack72 · 10/5/2026· Shipped 10/6/2026
Background (read this first)
swamp is reworking its datastore layer (tracking: swamp-club#2865). Operations stage typed changes into a unit of work, and pushes happen when a root unit ends (flush) or partway through it (checkpoint). Phase 1 is complete. Phase 2 has shipped:
- swamp-club#3025: use cases open units;
- swamp-club#3032: one root per command or request, with use-case units nested as children; a nested root given its own flush or checkpoint throws;
- swamp-club#3033: CLI commands push through roots;
- swamp-club#3034 and swamp-club#3035: serve handlers push through roots;
- swamp-club#3053: roots can
checkpoint().
Phase 2 must change no behaviour.
Where pushes still come from (main at 45184a1e). Direct .pushChanged( calls remain in about 25 production sites. Most already are a root's flush or checkpoint, written as a lambda at the call site:
- CLI:
src/cli/commands/:access_grant,access_group,access_token_mint,datastore_config_migrate,worker_prune,worker_token_create,worker_token_revoke,workflow_resume,workflow_run;src/cli/managed_config_sync.ts;- the single-phase and two-phase lock push in
src/cli/repo_context.ts.
- Serve:
pushChangedToRemoteinsrc/serve/handlers/shared.ts;executeWorkflowWithLocksinsrc/serve/deps.ts;- device auth (
device_auth_handler.ts), grant tracking (grant_write_tracking.ts), access reload (access_handlers.ts).
The rest are outside roots:
The global coordinator (
src/infrastructure/persistence/datastore_sync_coordinator.ts).requireInitializedReporegisters it (src/cli/repo_context.ts, around lines 1008 and 1072), andsrc/cli/mod.tsflushes it withflushDatastoreSync(), which pushes and releases the global lock:- in the teardown span on success (around line 2659);
- best-effort in the error path (around line 2863).
Commands that take the global lock but open no root of their own push only through this.
Model-lock flushes outside a root:
execution_service.ts(around line 1787): each workflow step's lock is pushed and released by itsstepLockHookflush;- serve's
handleModelMethodRun(model_handlers.ts, around lines 385 and 616) keepsflushLocks = lockResult.flushoutside its root.
ModelLockResult.flush()(push then release,src/cli/repo_context.ts) exists for these.Deliberate exceptions:
- the
datastore synccommand and its use case (src/cli/commands/datastore_sync.ts,src/libswamp/datastores/sync.ts); datastore setup's migration push (src/libswamp/datastores/setup.ts);- the lockfile publish (
managed_lockfile_transaction.ts), which has its own Phase 2 issue; - serve start-up and token GC (
src/cli/commands/serve.ts,src/serve/token_secret_migration.ts); - serve background GC (
bookkeeping_gc.ts,worker_gc_service.ts).
- the
"Push only on success" is implemented four times:
pushWhen: "always" | "completed"insrc/cli/command_root_unit.ts, used bydatastore_config_migrate,worker_pruneandmanaged_config_sync;- the
stagedflag insrc/serve/stage_writes_then_push.ts; - the
ranflags insrc/serve/handlers/model_handlers.ts(two sites).
The equivalence suites that prove nothing changed:
- the five
integration/usecase_sync_characterization_*_test.tsfiles, withsyncOrderand root-unit checks; integration/usecase_sync_root_unit_harness_test.ts;integration/datastore_remote_failure_test.ts;integration/serve_root_unit_test.ts;- the swamp-uat datastore suite.
Goal
Every production push is a root's flush or checkpoint, or a pinned deliberate exception, and one rule keeps it that way. The pushes themselves stay identical: the same count, order, timing relative to lock release and the reply, and failure behaviour.
Work
One "push only on success" option.
- Add
pushWhen: "always" | "completed"(default"always") torunInRootUnitOfWorkinrepo_unit_of_work.ts."completed"runs the flush only whenfnresolved. - Replace the CLI's local
pushWhenincommand_root_unit.ts(pass it through),stageWritesThenPush'sstagedflag, and bothranflags inmodel_handlers.ts. - Each must keep its exact condition. The model handlers also skip when a lock owns the push (
!flushLocks && mutating), and that part stays in their flush.
- Add
Serve model method run's lock push becomes a root flush. In
handleModelMethodRun(both sites), when a model lock is held, make its push the root's flush and release it after the root ends, as the CLI does withpush()andrelease()(#3033). Today's order is push, then release, then reply or no reply, and it must stay the same. RemoveflushLocks.Workflow step locks: decide, then do one of two things. Each step's
stepLockHookflush pushes and releases while the run's root is open, which is a mid-operation push.- Either route it through the run root's
checkpoint, with the step lock's push as the checkpoint and the release after it. This is only allowed if the order (step writes, push, release, next step) and the push count stay identical, and the domain layer (execution_service.ts) doesn't import infrastructure: the hook stays the seam. - Or, if that can't be done cleanly, keep
lockResult.flushfor step locks only, and pin it with the reason "per-step model lock owns its push and release; Phase 3 replaces model locks with leases".
Record the choice and the reason in the PR.
- Either route it through the run root's
The global coordinator.
- Map it first. List in the PR every command that takes the global lock through
requireInitializedRepoand opens no root, so it pushes only atflushDatastoreSync()in teardown. - Convert each one with a CLI helper: a root whose flush is
flushDatastoreSyncNamed(GLOBAL_LOCK_KEY), which pushes and releases as teardown did, withpushWhenmatching today's success-path and error-path behaviour. - Leave teardown in place. The
flushDatastoreSync()calls inmod.tsstay as a safety net; they do nothing once the key is flushed. - The nested-root guard. A command that already has a root must not also get a coordinator root. Nesting two roots with flushes throws (#3032), so merge them into one root, or prove the command never registers the global sync.
- Stop condition. If a command's push would move relative to its output, its lock release, or another push in a way the characterization or remote-failure suites can see, don't convert it. Pin it with the reason, and list it in the PR.
- Map it first. List in the PR every command that takes the global lock through
Hide the push implementations. Production code reaches a push only through a root.
- Move every root flush and checkpoint lambda that calls
syncService.pushChangeddirectly into a small set of push functions, for examplesrc/cli/push_paths.tsandsrc/serve/push_paths.ts: the lock push, the managed-config publish,pushChangedToRemote, the coordinator flush, the post-run push, and the token checkpoint. - Commands and handlers pass those functions as
flushorcheckpoint, and stop callingpushChangedthemselves.
- Move every root flush and checkpoint lambda that calls
Delete
ModelLockResult.flush()if item 3 removed its last caller. Otherwise keep it with a JSDoc naming its single remaining caller.One fitness rule. In
integration/datastore_write_seams_rules_test.tsorintegration/serve_root_unit_rules_test.ts, addPINNED_DIRECT_PUSHES: every production.pushChanged(call site. Each entry must be one of:- in a push-path module from item 5;
- inside the coordinator;
- a deliberate exception from Background, with a one-line reason.
It replaces
PINNED_CLI_PUSH_CALLSandPINNED_SERVE_RAW_PUSHES; fold them in. Add a self-test, like the existing scans have.Docs. Update the Phase 2 part of
design/enablers/datastores.mdto show the single flush path and the remaining exceptions.
Tests
- The new option: unit tests for
runInRootUnitOfWork'spushWhen, covering"completed"whenfnthrows, resolves, or is abandoned, plus a check thatcheckpointis unaffected by it. Also check thatcommand_root_unit,stageWritesThenPushand the model handlers behave the same after switching to it. - Equivalence:
- every characterization file, the root-unit harness,
serve_root_unit_test.tsanddatastore_remote_failure_test.tspass with no expectation changes; - for each coordinator command converted in item 4, record a characterization row or remote-failure row from the current code first, in a commit before the conversion, then convert;
- do the same for serve model method run with a model lock, and for workflow step locks if item 3 changes them.
- every characterization file, the root-unit harness,
- Rules:
PINNED_DIRECT_PUSHESpasses with its self-test, and the folded lists are gone. - These pass unchanged:
src/cli,src/libswamp,src/serve, and the swamp-uat datastore and serve suites.
Dependencies
Blocked by: nothing. swamp-club#3053 is merged. Blocks:
- removing
signalChange's hook fallback (Phase 2, not filed yet); - the lockfile publish redesign (Phase 2, not filed yet), which builds on the single flush path.
Done when
pushWhenlives only inrunInRootUnitOfWork.- Serve model method run's lock push is a root flush.
- Step locks are converted, or pinned with a reason.
- Every coordinator-only command is converted, or pinned with a reason.
- Push implementations live in the push-path modules.
PINNED_DIRECT_PUSHESis the single rule.- The suites pass unchanged, with the before-conversion rows recorded.
- Verification passes.
Out of scope
- The lockfile publish.
- Removing
signalChange's fallback. - Precise per-path marks.
- Roots for serve background GC and start-up.
- Removing the coordinator module itself.
Related
- Tracking: swamp-club#2865 (Phase 2).
- Built on: swamp-club#3032, #3033, #3034, #3035, #3053.
Shipped
Click a lifecycle step above to view its details.
Sign in to post a ripple.