Skip to main content
← Back to list
01Issue
BugOpenSwamp CLIPublic
AssigneesNone

Relationships

#3111 A nested structural swamp that outlives the run that started it keeps skipping that run's lock while the holder writes

Opened by hammz · 10/6/2026

Follow-up to swamp-club#3096, and present in the hand-offs shipped for swamp-club#2982 and swamp-club#2983.

What happens

A nested structural swamp (for example swamp data gc) skips the per-model locks named in its inherited lock list, on the assumption that the holder waits on it and writes nothing until it exits. Three paths break that assumption while the holder still holds the same lock acquisition:

  1. Cancelled dispatch. DispatchService rethrows the abort without waiting for the worker to confirm the runner is dead (src/serve/dispatch_service.ts); the worker force-kills the runner only after a grace period, and a swamp the runner started may outlive it.
  2. Lost worker. After a ChannelClosedError with no writes recorded, the step is dispatched again under the same step lock and nonce, while the first runner and its nested swamp may still be alive (for example across a network partition to a worker on another host).
  3. Killed --server client. A run requested over --server has its own abort controller (src/serve/handlers/model_handlers.ts) and keeps going when the client dies without cancelling it. The calling shell step then finishes and writes its outputs under the lock.

In each case a structural swamp still running under the abandoned run skips the holder's lock until the holder releases it, so it can work on the datastore while the holder is writing.

Notes

  • Traced by reading the code; not reproduced.
  • A stale nonce is already safe: once the holder releases the lock, a new acquisition gets a new nonce and is waited on. The window is only while the same acquisition is still held.
  • Documented under Known limits in design/enablers/datastores.md, Parent-Process Lock Awareness.

Possible directions

  • The orchestrator waits for the worker to confirm the runner and its children are dead before it releases or reuses a step lock.
  • A re-dispatch after a lost worker takes a fresh lock acquisition, so the old nonce stops matching.
  • A --server client death cancels the run it requested unless the run was explicitly detached.
02Bog Flow
◉OPEN○TRIAGED○IN PROGRESS○SHIPPED

Open

10/6/2026, 9:11:40 PM

No activity in this phase yet.

03Sludge Pulse

Sign in to post a ripple.