Relationships
#2897 workflow cancel SIGKILLs the run 2 s after SIGTERM, cutting off always/completed cleanup jobs mid-step
Opened by hammz · 10/1/2026· Shipped 10/1/2026
Description
swamp workflow cancel stops the owning process with SIGTERM and escalates to SIGKILL after 2 s (killProcessTree, src/infrastructure/process/process_kill.ts). A run that is cancelled runs its always/completed cleanup jobs under their own 30 s grace signal (CLEANUP_GRACE_TIMEOUT_MS, src/domain/workflows/execution_service.ts). Any teardown that takes more than about 2 s is therefore killed part-way, which defeats the cleanup grace period for the most common way of cancelling a run. The record is also left wrong: the teardown step had started (its command ran) but is recorded pending, and the teardown job pending.
Steps to reproduce
Repro workflow (workflows/workflow-e2e-wf.yaml; each step appends to a log so real executions can be counted):
id: <uuid>
name: e2e-wf
version: 1
concurrency: 1
jobs:
- name: a-side
steps:
- name: gate2
task: { type: manual_approval, prompt: "gate2?" }
- name: s
dependsOn: [{ step: gate2, condition: { type: succeeded } }]
task: { type: model_method, modelType: command/shell, modelName: m-s, methodName: execute,
inputs: { run: "echo s >> /tmp/exec.log; sleep 3" } }
- name: main
steps:
- name: gate
task: { type: manual_approval, prompt: "gate?" }
- name: post
dependsOn: [{ step: gate, condition: { type: succeeded } }]
task: { type: model_method, modelType: command/shell, modelName: m-post, methodName: execute,
inputs: { run: "echo post >> /tmp/exec.log" } }
- name: teardown
dependsOn: [{ job: main, condition: { type: always } }]
steps:
- name: t
task: { type: model_method, modelType: command/shell, modelName: m-t, methodName: execute,
inputs: { run: "echo t >> /tmp/exec.log" } }Change teardown t to run: "echo t >> /tmp/exec.log; sleep 5" and s to sleep 30.
swamp workflow run e2e-wf, approvegate2andgate.swamp workflow resume e2e-wf --run <id> &- About 1.5 s later, from a second shell:
swamp workflow cancel e2e-wf --run <id>.
Actual
The resume process is killed (exit 137) about 2 s into the teardown.
run status=cancelled cancel_reason=Cancelled by user
job a-side: failed | gate2:succeeded s:failed
job main: failed | gate:succeeded post:failed (cancelled, settledByAbort)
job teardown: pending | t:pending <- t started; exec log has "t" but no completionExpected
workflow cancel gives the owner at least the cleanup grace period (or waits while it reports cleanup in progress) before escalating to SIGKILL. If it must kill, the record shows that the teardown step started and was killed rather than pending.
Environment: Linux x86_64. Reproduced with release 20260930.225800.0-sha.1a9b497f and with the swamp-club#2597 branch (e337ad45), so it is not a regression from that fix. Found while end-to-end validating swamp-club#2597.
Shipped
Click a lifecycle step above to view its details.
Sign in to post a ripple.