Relationships
#2421 datastore sync: repositories mark paths dirty before writing, so a concurrent ungated push can drop the write
Opened by hammz · 9/23/2026
Description
The persistence repositories mark a path dirty before they write it:
YamlDefinitionRepository.savecallsnotifyDirty(targetPath)atyaml_definition_repository.ts:626but writes at:755.UnifiedDataRepositorymarks the data-name directory before allocating a version (unified_data_repository.ts:610, and the same at:1184).finalizeVersionmarks the version directory before it writes metadata and thelatestmarker (:1225,:1288).
In swamp serve the post-run and post-resume pushes are ungated (deps.ts:515, resume_launcher.ts:245; see UNGATED_PUSH_HANDLERS in src/serve/sync_gate.ts). One of them can land between a repository's mark and its write. Its scoped walk finds the path absent and takes it as a delete (absence-on-disk rule 2), and on completion the extension clears the dirty set. The handler's own pushChanged then finds nothing dirty (fast path) and uploads nothing. The write stays local only and is lost on the next pod restart or cache rehydration.
Current exposure
- Every serve handler that relies only on repository per-path marks (the ones PR #2459 changed) can lose a write this way when a run finishes at the same moment.
- The OAuth login mint used to be covered by its no-path
markDirty(). swamp-club#2408 replaced that with a per-path re-mark of the token's definition file and data folder, after the writes and just beforepushChanged. That is a local mitigation, not the fix.
Proper fix
Record the dirty path after the write lands, inside the repositories (or mark both before and after), so no caller needs its own re-mark. This touches every repository write path and its callers.
Also needed to fully close it (datastore extensions)
markSynced in @swamp/s3-datastore and @swamp/gcs-datastore (s3_cache_sync.ts:1716-1740) clears the whole dirty set and bulkInvalidated when a push completes, including marks added while that push was running. A mark made after the write can therefore still be erased by a push that was already in flight. The extensions should clear only the paths the finishing push snapshotted at walk start, for example with a dirty-generation counter. Without that, post-write marks narrow the window but do not close it.
Evidence
Found in the swamp-club#2408 adversarial review and confirmed in code. The swamp-club#2408 unit test createDeviceAuthDeps: mintServerToken re-marks the token paths after writing them, so a concurrent push cannot drop them models the race with a sync service that drops marks for paths not yet on disk.
Open
No activity in this phase yet.
sntxrr commented 9/26/2026, 5:38:49 PM
Field report that may be this race (or a #1770 regression), from a single swamp serve instance on @swamp/s3-datastore that rehydrates from the bucket on every restart.
Environment
| at the traced run (09-24) | now | |
|---|---|---|
@swamp/s3-datastore |
2026.09.10.2 | 2026.09.24.1 (serve restarted onto it 2026-09-25 15:52Z) |
@swamp/s3-datastore-bootstrap |
2026.09.06.1 | 2026.09.06.1 |
| swamp | 20260923.231117 → 20260924.181834 (nightly promotes) |
20260926.025243.0-sha.0f086bc8 |
Versions come from the pinned lockfile at the checkout serve was running (the 09-24 row) and from swamp extension list inside the serve container (the "now" row).
Symptom: scheduled runs that finished their work stay at stepProgress 0/N in the run index. At the next serve restart, boot reconciliation reaps them as interrupted / interrupt_reason: server_crash.
One run, traced (2026-09-24, UTC):
| time | event |
|---|---|
| 15:30:00.150 | scheduled run starts (3 steps) |
| 15:30:00.573 | step 1 writes its data version |
| 15:31:33.645 | step 3 writes its data version, which carries this workflowRunId and records the external call's HTTP 200; the downstream service logs the delivery in the same second |
| 17:51, 18:51 | serve restarts; the run is not reaped |
| 20:52:08 | serve restarts on a new build; the run is reaped as server_crash, stepProgress 0/3, along with 18 other runs |
No restart, deploy or OOM between 15:30 and 17:47. The run did all its work, and only its terminal record was lost.
Scale, before and after the datastore bump. A "stuck" run is one reaped with 0 steps done more than 30 min after it started. Runs killed by a restart while still executing are counted separately; that is our own restart cadence, not this bug.
| window | runs started | stuck | killed mid-run |
|---|---|---|---|
| 09-24 00:00Z → 09-25 15:53Z (s3-datastore 2026.09.10.2) | 425 | 37 | 4 |
| 09-25 15:53Z → 09-26 17:26Z (s3-datastore 2026.09.24.1) | 270 | 2 | 9 |
It is much rarer since the bump but not gone. The swamp binary changed at the same restart, so we can't attribute the drop to the extension alone.
Why it looks like #2421: the post-run push is ungated. If it lands between the mark and the write of the run record, the terminal status stays local. The next rehydration then restores the bucket's stale running copy, and boot reaps it. That fits 0 steps done, reaped at the next restart, and many runs reaped at once. The 2 stuck runs since the bump would fit the "narrowed, not closed" outcome the issue predicts if only one half is fixed.
Not verified: we did not capture the bucket object for an affected run before it was reaped, so the lost push is inferred, not observed. Run ids are available if useful.
Ask: until this is fixed, could reconciliation tell apart "no evidence the run did anything" and "the run wrote step outputs tagged with its runId"? Reporting the second case as server_crash makes a delivered run look lost.
Sign in to post a ripple.