Relationships
↔ sibling #3051#2917 Run tracker keeps some interrupted workflow rows forever: retention waits for markSettled, which several paths never call
Opened by hammz · 10/1/2026· Shipped 10/6/2026
Description
swamp-club#2896 (PR #2776) changed run-tracker retention. Startup purging (RunTrackerStore, src/infrastructure/persistence/run_tracker_store.ts) now keeps a workflow row that is interrupted with a null cancel_reason until markSettled() gives it a reason. The row is the only evidence that a run's owner died, and workflow recover and run doctor need it to settle a run record still running.
Several common paths create such rows but never call markSettled(), so those rows are never purged:
- Graceful
swamp serveshutdown with runs in flight. The shutdown path saves the run record asinterrupted(server_shutdown), but leaves its tracker rowrunning. On the next boot,reapStaleRuns/reapDeadProcessRunsmark the rowinterruptedwith no reason. The boot reaper (reapOrphanedWorkflowRuns) then skips the run, because its record is no longerrunning. So nothing settles the row. Every restart with in-flight runs leaves rows behind permanently. reapDeadProcessRunslater in serve (after the boot reaper has run). Rows it reaps are not settled until some later boot, and only if that boot happens to reach them.- Deleted or renamed workflows, or garbage-collected run records. Local
swamp run doctor --fixfinds the record behind an interrupted row throughworkflowRepo.findByName(row.workflowName). When that returns nothing, the row is skipped, and therenamedtest insrc/cli/commands/run_test.tsasserts this. - Serve-only deployments. The only general sweep is a local
swamp run doctor --fix. The serverun.doctorhandler only settles the runs it interrupts itself.
Impact
findAll(), whichrun historyandrun doctoruse, loads every row, so the table and those commands grow without limit.- Each retained row gets a pid liveness probe on every
run doctor. The probe isisProcessDead, which sendsSIGCONT. With--fix, each row also gets afindByNameand afindById. run doctor --fixre-reads rows that are already settled, becauseActiveRundoes not exposecancel_reason, so they cannot be skipped.
Separately, a resume that fails before execution restores the row through complete(id, "interrupted"). That resets cancel_reason to null on an already-settled row, which un-settles it.
Suggested fix
Any of the following, ideally the first two together:
- Add an age cap to the retention exemption (for example 30–90 days), as a backstop so no row is kept forever.
- Settle rows at serve boot. After
reapOrphanedWorkflowRuns, mark settled every interrupted workflow row whose record is notrunning. - Give the row a reason at serve shutdown, for example
complete(id, "interrupted", "server_shutdown"), so it never enters the exemption. - Handle missing records in
run doctor --fix. When the record cannot be found (workflow deleted or renamed, record GC'd), settle the row with a reason such asrecord_missing. Expose a settled flag onActiveRunso settled rows are not looked up again.
Found by the adversarial review in verification of PR #2776 (commit f467dc6c), and deferred out of that PR.
Shipped
Click a lifecycle step above to view its details.
Sign in to post a ripple.