Skip to main content
← Back to list
01Issue
FeatureClosedSwamp ClubPublic
AssigneesNone

Relationships

#2903 Docs: run doctor --fix and workflow recover settle runs left running by a force-exited process

Opened by hammz · 10/1/2026

What changed

swamp-club#2896 makes a workflow run whose owning process was force-exited (for example, a second Ctrl-C during swamp workflow run or swamp workflow resume) recoverable from the CLI:

  • swamp run doctor treats a tracked run on this host whose process is dead as stale straight away, instead of waiting for the 90-second heartbeat TTL.
  • swamp run doctor --fix reaps those tracker rows and also settles the workflow run record. A run left running by a dead process becomes interrupted (interrupt_reason: owner_process_dead), and its in-flight steps become unknown. Log and JSON output report the orphaned workflow-run counts that --server already reports.
  • swamp workflow recover --run <id> does the same reconciliation for a stuck run on its own, so doctor does not need to run first.
  • swamp workflow resume --run <id> on such a run now tells you to run swamp workflow recover instead of "Wait for it to complete".
  • A step's running status is saved when it starts, so a step that was mid-flight at the force exit is recorded as unknown rather than pending. Recovery then asks for a guard or --acknowledge-unknown before running it again.

Manual pages to update

  1. content/manual/reference/operational-commands.md, swamp run doctor section
    • It says that with --fix, stale runs are "reaped (transitioned to failed)". That is already wrong: they become interrupted.
    • Describe the staleness rule: heartbeat TTL, or a dead pid on this host.
    • Describe that --fix also reconciles workflow run records, and add the orphaned workflow-run fields to the JSON example.
  2. content/manual/how-to/cancel-a-stuck-run.md
    • Add a section on a run stuck in running after a force exit (Ctrl-C twice).
    • Recommended path: swamp workflow recover <workflow> --run <id> (or swamp run doctor --fix), then swamp workflow resume.
    • Contrast it with swamp workflow cancel, which abandons the run.
  3. content/manual/reference/workflows.md, if it describes run statuses: interrupted is no longer only set when swamp serve restarts.

Land these once the swamp-club#2896 fix has shipped.

02Bog Flow
✓OPEN○TRIAGED○IN PROGRESS◉CLOSED

Closed

10/2/2026, 1:46:16 AM

No activity in this phase yet.

03Sludge Pulse
Editable. Press Enter to edit.

stack72 commented 10/2/2026, 1:46:15 AM

Documented in https://github.com/swamp-club/swamp-club/pull/1281. reference/operational-commands.md now gives the staleness rule (heartbeat TTL, or a dead pid on this host) and says stale runs become interrupted rather than failed. It also says --fix settles orphaned workflow runs, and the JSON example adds orphanedWorkflowRuns and orphanedReaped. how-to/cancel-a-stuck-run.md has a new section on recovering after a force exit, contrasted with cancel. clear-stuck-runs.md had the same 'failed' wording and is fixed. In reference/workflows.md, interrupted now lists a force-exited owner, with a subsection on owner_process_dead and --acknowledge-unknown. One correction to the issue text: after doctor --fix, workflow resume still refuses until workflow recover runs. The docs give doctor --fix, then recover, then resume; I checked this against a SIGKILLed run.

Sign in to post a ripple.