Relationships
#2898 A job alone in its level starts after the abort, and its unstarted step is recorded as a real failure instead of settledByAbort
Opened by hammz · 10/1/2026· Shipped 10/5/2026
Description
When a workflow's abort fires before a job level starts, a job that shares its level with others is not started: its pending steps are settled as cancelled/skipped with settledByAbort, so a resume runs them again. A job that is alone in its level is started anyway, because merge() (src/infrastructure/stream/merge.ts) returns the single stream directly (if (streams.length === 1) { yield* streams[0]; return; }) before it checks signal.aborted. Its next step is then invoked with an already-aborted signal. For a command/shell step the command is spawned and killed at once, and the step is recorded as a real failure, Command exited with code -1, without settledByAbort.
Consequences:
- The same abort gives different records depending only on how many jobs share the level.
- A step that never did its work looks like a step that ran and failed, so a resume does not treat it as abort-settled work, and the command was in fact spawned. A model method that ignores the signal would run in full after the cancellation.
Steps to reproduce
Workflow with one job per level (concurrency: 1):
jobs:
- name: main
steps:
- name: gate
task: { type: manual_approval, prompt: "gate?" }
- name: post
dependsOn: [{ step: gate, condition: { type: succeeded } }]
task: { type: model_method, modelType: command/shell, modelName: m-post, methodName: execute,
inputs: { run: "echo post >> /tmp/exec.log; sleep 3" } }
- name: teardown
dependsOn: [{ job: main, condition: { type: always } }]
steps:
- name: t
task: { type: model_method, modelType: command/shell, modelName: m-t, methodName: execute,
inputs: { run: "echo t >> /tmp/exec.log" } }swamp workflow run e2e-wf,swamp workflow approve e2e-wf gate --run <id>.swamp workflow resume e2e-wf --run <id> &, and send SIGTERM to it just after its signal handler is armed (about 340 ms here).
This depends on timing: 1 of 21 attempts in a 300–700 ms sweep landed in the window. The equivalent unit-level repro is to call resume with AbortSignal.abort() on a workflow whose suspended job is alone in its level.
Actual
run status=cancelled
job main: failed | gate:succeeded post:failed err='Command exited with code -1' (no settledByAbort; "post" never reached the log)
job teardown: succeededWith a second job in the same level, the same timing gives post:failed err='cancelled' settledByAbort: true.
Expected
A job alone in its level is treated like one in a multi-job level: not started once the signal has aborted, with its pending steps settled with settledByAbort.
Environment: Linux x86_64. Reproduced with release 20260930.225800.0-sha.1a9b497f and with the swamp-club#2597 branch (e337ad45), so it is not a regression from that fix. Found while end-to-end validating swamp-club#2597.
Shipped
Click a lifecycle step above to view its details.
Sign in to post a ripple.