Relationships
#2896 A force-exited workflow run stays running with a dead pid: resume, recover and run doctor --fix cannot clear it
Opened by hammz · 10/1/2026· Shipped 10/1/2026
Description
When a swamp workflow resume (or run) is force-exited, for example by a second Ctrl-C (exit 130) while the first one's abort is still being handled, the run record is left running with the dead process's pid, and nothing in the CLI recovers it properly:
swamp workflow resume e2e-wf --run <id>:Error: Run <id> is not suspended or failed (status: running). Wait for it to complete, …, but nothing will ever complete it.swamp workflow recover e2e-wf --run <id>:Error: Interrupted run <id> not found. The run is never markedinterrupted, even though its pid is dead.swamp run doctorreports it as stale, andswamp run doctor --fixprintsReaped 2 stale run(s). Afterwardsswamp run history --activeshowsNo tracked runs., but the workflow run record is stillrunning, soresumeandrecoverstill refuse it andhistory getshowsResult: RUNNING.- Only
swamp workflow cancelclears it, and that leaves the jobsrunning(swamp-club#2895).
The record is also wrong about work done: step s had started (its command wrote to the log) but is still recorded pending.
Steps to reproduce
Repro workflow (workflows/workflow-e2e-wf.yaml; each step appends to a log so real executions can be counted):
id: <uuid>
name: e2e-wf
version: 1
concurrency: 1
jobs:
- name: a-side
steps:
- name: gate2
task: { type: manual_approval, prompt: "gate2?" }
- name: s
dependsOn: [{ step: gate2, condition: { type: succeeded } }]
task: { type: model_method, modelType: command/shell, modelName: m-s, methodName: execute,
inputs: { run: "echo s >> /tmp/exec.log; sleep 3" } }
- name: main
steps:
- name: gate
task: { type: manual_approval, prompt: "gate?" }
- name: post
dependsOn: [{ step: gate, condition: { type: succeeded } }]
task: { type: model_method, modelType: command/shell, modelName: m-post, methodName: execute,
inputs: { run: "echo post >> /tmp/exec.log" } }
- name: teardown
dependsOn: [{ job: main, condition: { type: always } }]
steps:
- name: t
task: { type: model_method, modelType: command/shell, modelName: m-t, methodName: execute,
inputs: { run: "echo t >> /tmp/exec.log" } }swamp workflow run e2e-wf, then approvegate2andgate.swamp workflow resume e2e-wf --run <id> &, then about 0.5 s later send SIGINT to it twice, 40 ms apart. It exits 130.- Run the commands above.
Reproduced in 24 of 25 attempts across SIGINT delays of 0.3–0.9 s.
Actual
run status=running
job a-side: running | gate2:succeeded s:pending (s had started)
job main: running | gate:succeeded post:pendingThe record stays like this after run doctor --fix.
Expected
A run whose owning pid is dead is recoverable. Either it is detected and marked interrupted so workflow recover works, or run doctor --fix settles the workflow run record as well as the tracker rows.
Environment: Linux x86_64. Reproduced with release 20260930.225800.0-sha.1a9b497f and with the swamp-club#2597 branch (e337ad45), so it is not a regression from that fix. Found while end-to-end validating swamp-club#2597.
Shipped
Click a lifecycle step above to view its details.
Sign in to post a ripple.