Skip to main content
← Back to list
01Issue
BugShippedSwamp CLIPublic
AssigneesNone

Relationships

↔ sibling #3051

#2917 Run tracker keeps some interrupted workflow rows forever: retention waits for markSettled, which several paths never call

Opened by hammz · 10/1/2026· Shipped 10/6/2026

Description

swamp-club#2896 (PR #2776) changed run-tracker retention. Startup purging (RunTrackerStore, src/infrastructure/persistence/run_tracker_store.ts) now keeps a workflow row that is interrupted with a null cancel_reason until markSettled() gives it a reason. The row is the only evidence that a run's owner died, and workflow recover and run doctor need it to settle a run record still running.

Several common paths create such rows but never call markSettled(), so those rows are never purged:

  • Graceful swamp serve shutdown with runs in flight. The shutdown path saves the run record as interrupted (server_shutdown), but leaves its tracker row running. On the next boot, reapStaleRuns / reapDeadProcessRuns mark the row interrupted with no reason. The boot reaper (reapOrphanedWorkflowRuns) then skips the run, because its record is no longer running. So nothing settles the row. Every restart with in-flight runs leaves rows behind permanently.
  • reapDeadProcessRuns later in serve (after the boot reaper has run). Rows it reaps are not settled until some later boot, and only if that boot happens to reach them.
  • Deleted or renamed workflows, or garbage-collected run records. Local swamp run doctor --fix finds the record behind an interrupted row through workflowRepo.findByName(row.workflowName). When that returns nothing, the row is skipped, and the renamed test in src/cli/commands/run_test.ts asserts this.
  • Serve-only deployments. The only general sweep is a local swamp run doctor --fix. The serve run.doctor handler only settles the runs it interrupts itself.

Impact

  • findAll(), which run history and run doctor use, loads every row, so the table and those commands grow without limit.
  • Each retained row gets a pid liveness probe on every run doctor. The probe is isProcessDead, which sends SIGCONT. With --fix, each row also gets a findByName and a findById.
  • run doctor --fix re-reads rows that are already settled, because ActiveRun does not expose cancel_reason, so they cannot be skipped.

Separately, a resume that fails before execution restores the row through complete(id, "interrupted"). That resets cancel_reason to null on an already-settled row, which un-settles it.

Suggested fix

Any of the following, ideally the first two together:

  1. Add an age cap to the retention exemption (for example 30–90 days), as a backstop so no row is kept forever.
  2. Settle rows at serve boot. After reapOrphanedWorkflowRuns, mark settled every interrupted workflow row whose record is not running.
  3. Give the row a reason at serve shutdown, for example complete(id, "interrupted", "server_shutdown"), so it never enters the exemption.
  4. Handle missing records in run doctor --fix. When the record cannot be found (workflow deleted or renamed, record GC'd), settle the row with a reason such as record_missing. Expose a settled flag on ActiveRun so settled rows are not looked up again.

Found by the adversarial review in verification of PR #2776 (commit f467dc6c), and deferred out of that PR.

02Bog Flow
✓OPEN✓TRIAGED✓IN PROGRESS✓SHIPPEDLINKED+ 9 MOREPR_MERGED+ 1 MORESESSION_SUMMARIZED

Shipped

10/6/2026, 2:08:24 PM

Click a lifecycle step above to view its details.

03Sludge Pulse
stack72 linked sibling of #305110/6/2026, 12:52:06 AM

Sign in to post a ripple.