Relationships
#2903 Docs: run doctor --fix and workflow recover settle runs left running by a force-exited process
Opened by hammz · 10/1/2026
What changed
swamp-club#2896 makes a workflow run whose owning process was force-exited (for example, a second Ctrl-C during swamp workflow run or swamp workflow resume) recoverable from the CLI:
swamp run doctortreats a tracked run on this host whose process is dead as stale straight away, instead of waiting for the 90-second heartbeat TTL.swamp run doctor --fixreaps those tracker rows and also settles the workflow run record. A run leftrunningby a dead process becomesinterrupted(interrupt_reason: owner_process_dead), and its in-flight steps becomeunknown. Log and JSON output report the orphaned workflow-run counts that--serveralready reports.swamp workflow recover --run <id>does the same reconciliation for a stuck run on its own, sodoctordoes not need to run first.swamp workflow resume --run <id>on such a run now tells you to runswamp workflow recoverinstead of "Wait for it to complete".- A step's
runningstatus is saved when it starts, so a step that was mid-flight at the force exit is recorded asunknownrather thanpending. Recovery then asks for a guard or--acknowledge-unknownbefore running it again.
Manual pages to update
content/manual/reference/operational-commands.md,swamp run doctorsection- It says that with
--fix, stale runs are "reaped (transitioned tofailed)". That is already wrong: they becomeinterrupted. - Describe the staleness rule: heartbeat TTL, or a dead pid on this host.
- Describe that
--fixalso reconciles workflow run records, and add the orphaned workflow-run fields to the JSON example.
- It says that with
content/manual/how-to/cancel-a-stuck-run.md- Add a section on a run stuck in
runningafter a force exit (Ctrl-C twice). - Recommended path:
swamp workflow recover <workflow> --run <id>(orswamp run doctor --fix), thenswamp workflow resume. - Contrast it with
swamp workflow cancel, which abandons the run.
- Add a section on a run stuck in
content/manual/reference/workflows.md, if it describes run statuses:interruptedis no longer only set whenswamp serverestarts.
Land these once the swamp-club#2896 fix has shipped.
Closed
No activity in this phase yet.
stack72 commented 10/2/2026, 1:46:15 AM
Documented in https://github.com/swamp-club/swamp-club/pull/1281. reference/operational-commands.md now gives the staleness rule (heartbeat TTL, or a dead pid on this host) and says stale runs become interrupted rather than failed. It also says --fix settles orphaned workflow runs, and the JSON example adds orphanedWorkflowRuns and orphanedReaped. how-to/cancel-a-stuck-run.md has a new section on recovering after a force exit, contrasted with cancel. clear-stuck-runs.md had the same 'failed' wording and is fixed. In reference/workflows.md, interrupted now lists a force-exited owner, with a subsection on owner_process_dead and --acknowledge-unknown. One correction to the issue text: after doctor --fix, workflow resume still refuses until workflow recover runs. The docs give doctor --fix, then recover, then resume; I checked this against a SIGKILLed run.
Sign in to post a ripple.