Skip to main content
← Back to list
01Issue
BugOpenSwamp CLIPublic
AssigneesNone

Relationships

#2421 datastore sync: repositories mark paths dirty before writing, so a concurrent ungated push can drop the write

Opened by hammz · 9/23/2026

Description

The persistence repositories mark a path dirty before they write it:

  • YamlDefinitionRepository.save calls notifyDirty(targetPath) at yaml_definition_repository.ts:626 but writes at :755.
  • UnifiedDataRepository marks the data-name directory before allocating a version (unified_data_repository.ts:610, and the same at :1184). finalizeVersion marks the version directory before it writes metadata and the latest marker (:1225, :1288).

In swamp serve the post-run and post-resume pushes are ungated (deps.ts:515, resume_launcher.ts:245; see UNGATED_PUSH_HANDLERS in src/serve/sync_gate.ts). One of them can land between a repository's mark and its write. Its scoped walk finds the path absent and takes it as a delete (absence-on-disk rule 2), and on completion the extension clears the dirty set. The handler's own pushChanged then finds nothing dirty (fast path) and uploads nothing. The write stays local only and is lost on the next pod restart or cache rehydration.

Current exposure

  • Every serve handler that relies only on repository per-path marks (the ones PR #2459 changed) can lose a write this way when a run finishes at the same moment.
  • The OAuth login mint used to be covered by its no-path markDirty(). swamp-club#2408 replaced that with a per-path re-mark of the token's definition file and data folder, after the writes and just before pushChanged. That is a local mitigation, not the fix.

Proper fix

Record the dirty path after the write lands, inside the repositories (or mark both before and after), so no caller needs its own re-mark. This touches every repository write path and its callers.

Also needed to fully close it (datastore extensions)

markSynced in @swamp/s3-datastore and @swamp/gcs-datastore (s3_cache_sync.ts:1716-1740) clears the whole dirty set and bulkInvalidated when a push completes, including marks added while that push was running. A mark made after the write can therefore still be erased by a push that was already in flight. The extensions should clear only the paths the finishing push snapshotted at walk start, for example with a dirty-generation counter. Without that, post-write marks narrow the window but do not close it.

Evidence

Found in the swamp-club#2408 adversarial review and confirmed in code. The swamp-club#2408 unit test createDeviceAuthDeps: mintServerToken re-marks the token paths after writing them, so a concurrent push cannot drop them models the race with a sync service that drops marks for paths not yet on disk.

02Bog Flow
◉OPEN○TRIAGED○IN PROGRESS○SHIPPED

Open

9/23/2026, 5:49:48 PM

No activity in this phase yet.

03Sludge Pulse
Editable. Press Enter to edit.

sntxrr commented 9/26/2026, 5:38:49 PM

Field report that may be this race (or a #1770 regression), from a single swamp serve instance on @swamp/s3-datastore that rehydrates from the bucket on every restart.

Environment

at the traced run (09-24) now
@swamp/s3-datastore 2026.09.10.2 2026.09.24.1 (serve restarted onto it 2026-09-25 15:52Z)
@swamp/s3-datastore-bootstrap 2026.09.06.1 2026.09.06.1
swamp 20260923.231117 → 20260924.181834 (nightly promotes) 20260926.025243.0-sha.0f086bc8

Versions come from the pinned lockfile at the checkout serve was running (the 09-24 row) and from swamp extension list inside the serve container (the "now" row).

Symptom: scheduled runs that finished their work stay at stepProgress 0/N in the run index. At the next serve restart, boot reconciliation reaps them as interrupted / interrupt_reason: server_crash.

One run, traced (2026-09-24, UTC):

time event
15:30:00.150 scheduled run starts (3 steps)
15:30:00.573 step 1 writes its data version
15:31:33.645 step 3 writes its data version, which carries this workflowRunId and records the external call's HTTP 200; the downstream service logs the delivery in the same second
17:51, 18:51 serve restarts; the run is not reaped
20:52:08 serve restarts on a new build; the run is reaped as server_crash, stepProgress 0/3, along with 18 other runs

No restart, deploy or OOM between 15:30 and 17:47. The run did all its work, and only its terminal record was lost.

Scale, before and after the datastore bump. A "stuck" run is one reaped with 0 steps done more than 30 min after it started. Runs killed by a restart while still executing are counted separately; that is our own restart cadence, not this bug.

window runs started stuck killed mid-run
09-24 00:00Z → 09-25 15:53Z (s3-datastore 2026.09.10.2) 425 37 4
09-25 15:53Z → 09-26 17:26Z (s3-datastore 2026.09.24.1) 270 2 9

It is much rarer since the bump but not gone. The swamp binary changed at the same restart, so we can't attribute the drop to the extension alone.

Why it looks like #2421: the post-run push is ungated. If it lands between the mark and the write of the run record, the terminal status stays local. The next rehydration then restores the bucket's stale running copy, and boot reaps it. That fits 0 steps done, reaped at the next restart, and many runs reaped at once. The 2 stuck runs since the bump would fit the "narrowed, not closed" outcome the issue predicts if only one half is fixed.

Not verified: we did not capture the bucket object for an affected run before it was reaped, so the lost push is inferred, not observed. Run ids are available if useful.

Ask: until this is fixed, could reconciliation tell apart "no evidence the run did anything" and "the run wrote step outputs tagged with its runId"? Reporting the second case as server_crash makes a delivered run look lost.

Sign in to post a ripple.