Skip to main content
← Back to list
01Issue
BugShippedSwamp CLIPublic
Assigneeshammz

Relationships

#2896 A force-exited workflow run stays running with a dead pid: resume, recover and run doctor --fix cannot clear it

Opened by hammz · 10/1/2026· Shipped 10/1/2026

Description

When a swamp workflow resume (or run) is force-exited, for example by a second Ctrl-C (exit 130) while the first one's abort is still being handled, the run record is left running with the dead process's pid, and nothing in the CLI recovers it properly:

  • swamp workflow resume e2e-wf --run <id>: Error: Run <id> is not suspended or failed (status: running). Wait for it to complete, …, but nothing will ever complete it.
  • swamp workflow recover e2e-wf --run <id>: Error: Interrupted run <id> not found. The run is never marked interrupted, even though its pid is dead.
  • swamp run doctor reports it as stale, and swamp run doctor --fix prints Reaped 2 stale run(s). Afterwards swamp run history --active shows No tracked runs., but the workflow run record is still running, so resume and recover still refuse it and history get shows Result: RUNNING.
  • Only swamp workflow cancel clears it, and that leaves the jobs running (swamp-club#2895).

The record is also wrong about work done: step s had started (its command wrote to the log) but is still recorded pending.

Steps to reproduce

Repro workflow (workflows/workflow-e2e-wf.yaml; each step appends to a log so real executions can be counted):

id: <uuid>
name: e2e-wf
version: 1
concurrency: 1
jobs:
  - name: a-side
    steps:
      - name: gate2
        task: { type: manual_approval, prompt: "gate2?" }
      - name: s
        dependsOn: [{ step: gate2, condition: { type: succeeded } }]
        task: { type: model_method, modelType: command/shell, modelName: m-s, methodName: execute,
                inputs: { run: "echo s >> /tmp/exec.log; sleep 3" } }
  - name: main
    steps:
      - name: gate
        task: { type: manual_approval, prompt: "gate?" }
      - name: post
        dependsOn: [{ step: gate, condition: { type: succeeded } }]
        task: { type: model_method, modelType: command/shell, modelName: m-post, methodName: execute,
                inputs: { run: "echo post >> /tmp/exec.log" } }
  - name: teardown
    dependsOn: [{ job: main, condition: { type: always } }]
    steps:
      - name: t
        task: { type: model_method, modelType: command/shell, modelName: m-t, methodName: execute,
                inputs: { run: "echo t >> /tmp/exec.log" } }
  1. swamp workflow run e2e-wf, then approve gate2 and gate.
  2. swamp workflow resume e2e-wf --run <id> &, then about 0.5 s later send SIGINT to it twice, 40 ms apart. It exits 130.
  3. Run the commands above.

Reproduced in 24 of 25 attempts across SIGINT delays of 0.3–0.9 s.

Actual

run status=running
  job a-side: running | gate2:succeeded s:pending   (s had started)
  job main:   running | gate:succeeded post:pending

The record stays like this after run doctor --fix.

Expected

A run whose owning pid is dead is recoverable. Either it is detected and marked interrupted so workflow recover works, or run doctor --fix settles the workflow run record as well as the tracker rows.

Environment: Linux x86_64. Reproduced with release 20260930.225800.0-sha.1a9b497f and with the swamp-club#2597 branch (e337ad45), so it is not a regression from that fix. Found while end-to-end validating swamp-club#2597.

02Bog Flow
✓OPEN✓TRIAGED✓IN PROGRESS✓SHIPPED+ 1 MOREASSIGNED+ 5 MOREREVIEW+ 19 MOREPR_MERGED+ 2 MORESESSION_SUMMARIZED

Shipped

10/1/2026, 7:08:11 PM

Click a lifecycle step above to view its details.

03Sludge Pulse
hammz assigned hammz10/1/2026, 4:32:33 PM

Sign in to post a ripple.