Skip to main content
← Back to list
01Issue
BugShippedSwamp CLIPublic
Assigneeshammz

Relationships

#2897 workflow cancel SIGKILLs the run 2 s after SIGTERM, cutting off always/completed cleanup jobs mid-step

Opened by hammz · 10/1/2026· Shipped 10/1/2026

Description

swamp workflow cancel stops the owning process with SIGTERM and escalates to SIGKILL after 2 s (killProcessTree, src/infrastructure/process/process_kill.ts). A run that is cancelled runs its always/completed cleanup jobs under their own 30 s grace signal (CLEANUP_GRACE_TIMEOUT_MS, src/domain/workflows/execution_service.ts). Any teardown that takes more than about 2 s is therefore killed part-way, which defeats the cleanup grace period for the most common way of cancelling a run. The record is also left wrong: the teardown step had started (its command ran) but is recorded pending, and the teardown job pending.

Steps to reproduce

Repro workflow (workflows/workflow-e2e-wf.yaml; each step appends to a log so real executions can be counted):

id: <uuid>
name: e2e-wf
version: 1
concurrency: 1
jobs:
  - name: a-side
    steps:
      - name: gate2
        task: { type: manual_approval, prompt: "gate2?" }
      - name: s
        dependsOn: [{ step: gate2, condition: { type: succeeded } }]
        task: { type: model_method, modelType: command/shell, modelName: m-s, methodName: execute,
                inputs: { run: "echo s >> /tmp/exec.log; sleep 3" } }
  - name: main
    steps:
      - name: gate
        task: { type: manual_approval, prompt: "gate?" }
      - name: post
        dependsOn: [{ step: gate, condition: { type: succeeded } }]
        task: { type: model_method, modelType: command/shell, modelName: m-post, methodName: execute,
                inputs: { run: "echo post >> /tmp/exec.log" } }
  - name: teardown
    dependsOn: [{ job: main, condition: { type: always } }]
    steps:
      - name: t
        task: { type: model_method, modelType: command/shell, modelName: m-t, methodName: execute,
                inputs: { run: "echo t >> /tmp/exec.log" } }

Change teardown t to run: "echo t >> /tmp/exec.log; sleep 5" and s to sleep 30.

  1. swamp workflow run e2e-wf, approve gate2 and gate.
  2. swamp workflow resume e2e-wf --run <id> &
  3. About 1.5 s later, from a second shell: swamp workflow cancel e2e-wf --run <id>.

Actual

The resume process is killed (exit 137) about 2 s into the teardown.

run status=cancelled cancel_reason=Cancelled by user
  job a-side: failed    | gate2:succeeded s:failed
  job main:   failed    | gate:succeeded post:failed (cancelled, settledByAbort)
  job teardown: pending | t:pending          <- t started; exec log has "t" but no completion

Expected

workflow cancel gives the owner at least the cleanup grace period (or waits while it reports cleanup in progress) before escalating to SIGKILL. If it must kill, the record shows that the teardown step started and was killed rather than pending.

Environment: Linux x86_64. Reproduced with release 20260930.225800.0-sha.1a9b497f and with the swamp-club#2597 branch (e337ad45), so it is not a regression from that fix. Found while end-to-end validating swamp-club#2597.

02Bog Flow
✓OPEN✓TRIAGED✓IN PROGRESS✓SHIPPED+ 1 MOREASSIGNED+ 5 MOREREVIEW+ 14 MOREPR_MERGED+ 2 MORESESSION_SUMMARIZED

Shipped

10/1/2026, 6:22:29 PM

Click a lifecycle step above to view its details.

03Sludge Pulse
hammz assigned hammz10/1/2026, 4:32:54 PM

Sign in to post a ripple.