Skip to main content
← Back to list
01Issue
BugShippedSwamp CLIPublic
Assigneeshammz

Relationships

#2634 command/shell: aborting a step kills only sh, leaving the command's child processes orphaned

Opened by hammz · 9/28/2026· Shipped 9/28/2026

Summary

When a command/shell step is aborted, e.g. by a swamp serve shutdown, a workflow cancel, or --max-run-duration, swamp sends SIGTERM only to the direct child (sh -c <run>). Any processes the command spawned keep running after swamp is done. Swamp reports the run as cancelled, and even exits, while the real work is still running as an orphan reparented to init.

Reproduction

Found while testing the serve shutdown drain end to end (swamp-club#2484). This happens on main too; the drain branch only changes when the abort fires, not how.

  1. swamp model create command/shell sleep60 --global-arg 'run=sleep 60; echo done'
  2. Create workflow slow60 with one step: sleep60.execute.
  3. SECRET=[REDACTED-SECRET-1] swamp serve --port 19090 --webhook '/hooks/slow60:slow60:@env=SECRET' --shutdown-drain-timeout 3s
  4. Deliver a signed webhook to /hooks/slow60 and wait until pgrep -fx 'sleep 60' finds the process.
  5. kill -TERM <serve pid>

Result: serve exits after about 3.1s, and the workflow run is recorded as cancelled (workflow was cancelled). But:

$ ps -o pid,ppid,pgid,etime,cmd -p <sleep pid>
    PID    PPID    PGID     ELAPSED CMD
1845274       1 1845187       00:12 sleep 60

sleep 60 is still alive, now a child of PID 1, and runs to completion. --shutdown-drain-timeout 0 gives the same result. Anything that aborts the run's AbortSignal should behave the same way: swamp workflow cancel, --max-run-duration, or a local Ctrl-C.

Cause

src/infrastructure/process/process_executor.ts: every path (the timeout path and the streaming path, around lines 223/241 and 306/314) calls process.kill("SIGTERM") on the Deno.ChildProcess. The child is spawned without its own process group, so the signal reaches only sh. sh -c "sleep 60; echo done" dies, and its child sleep is reparented to init. The step's timeout kill has the same problem.

Expected

Aborting a step, or hitting its timeout, should terminate the whole process tree that the step started. Nothing it launched should outlive the run being marked cancelled or failed. That matters most for serve shutdown and rolling restarts: an orphaned terraform apply, deploy script or test runner keeps mutating state after swamp says the run was cancelled, and the replacement instance can start the same work at the same time.

Possible fix

On POSIX, spawn the command as its own process-group leader (e.g. through setsid), and on abort or timeout signal the group (kill(-pgid, SIGTERM)), escalating to SIGKILL after a grace period. On Windows, use a job object or taskkill /T. The escalation to SIGKILL is also missing today: a child that ignores SIGTERM is never killed.

Environment

  • swamp built from main + the #2484 branch (behaviour is unchanged on main)
  • Linux (CachyOS, kernel 7.2), command/shell typeVersion 2026.02.09.1
02Bog Flow
✓OPEN✓TRIAGED✓IN PROGRESS✓SHIPPED+ 1 MOREASSIGNED+ 5 MOREREVIEW+ 15 MOREPR_MERGED+ 2 MORESESSION_SUMMARIZED

Shipped

9/28/2026, 11:31:33 PM

Click a lifecycle step above to view its details.

03Sludge Pulse
hammz assigned hammz9/28/2026, 9:52:31 PM

Sign in to post a ripple.