Skip to main content
← Back to list
01Issue
BugShippedSwamp CLIPublic
Assigneeshammz

Relationships

#2898 A job alone in its level starts after the abort, and its unstarted step is recorded as a real failure instead of settledByAbort

Opened by hammz · 10/1/2026· Shipped 10/5/2026

Description

When a workflow's abort fires before a job level starts, a job that shares its level with others is not started: its pending steps are settled as cancelled/skipped with settledByAbort, so a resume runs them again. A job that is alone in its level is started anyway, because merge() (src/infrastructure/stream/merge.ts) returns the single stream directly (if (streams.length === 1) { yield* streams[0]; return; }) before it checks signal.aborted. Its next step is then invoked with an already-aborted signal. For a command/shell step the command is spawned and killed at once, and the step is recorded as a real failure, Command exited with code -1, without settledByAbort.

Consequences:

  • The same abort gives different records depending only on how many jobs share the level.
  • A step that never did its work looks like a step that ran and failed, so a resume does not treat it as abort-settled work, and the command was in fact spawned. A model method that ignores the signal would run in full after the cancellation.

Steps to reproduce

Workflow with one job per level (concurrency: 1):

jobs:
  - name: main
    steps:
      - name: gate
        task: { type: manual_approval, prompt: "gate?" }
      - name: post
        dependsOn: [{ step: gate, condition: { type: succeeded } }]
        task: { type: model_method, modelType: command/shell, modelName: m-post, methodName: execute,
                inputs: { run: "echo post >> /tmp/exec.log; sleep 3" } }
  - name: teardown
    dependsOn: [{ job: main, condition: { type: always } }]
    steps:
      - name: t
        task: { type: model_method, modelType: command/shell, modelName: m-t, methodName: execute,
                inputs: { run: "echo t >> /tmp/exec.log" } }
  1. swamp workflow run e2e-wf, swamp workflow approve e2e-wf gate --run <id>.
  2. swamp workflow resume e2e-wf --run <id> &, and send SIGTERM to it just after its signal handler is armed (about 340 ms here).

This depends on timing: 1 of 21 attempts in a 300–700 ms sweep landed in the window. The equivalent unit-level repro is to call resume with AbortSignal.abort() on a workflow whose suspended job is alone in its level.

Actual

run status=cancelled
  job main: failed | gate:succeeded post:failed  err='Command exited with code -1'  (no settledByAbort; "post" never reached the log)
  job teardown: succeeded

With a second job in the same level, the same timing gives post:failed err='cancelled' settledByAbort: true.

Expected

A job alone in its level is treated like one in a multi-job level: not started once the signal has aborted, with its pending steps settled with settledByAbort.

Environment: Linux x86_64. Reproduced with release 20260930.225800.0-sha.1a9b497f and with the swamp-club#2597 branch (e337ad45), so it is not a regression from that fix. Found while end-to-end validating swamp-club#2597.

02Bog Flow
✓OPEN✓TRIAGED✓IN PROGRESS✓SHIPPED+ 1 MOREASSIGNED+ 2 MOREREVIEW+ 7 MOREPR_MERGED+ 2 MORESESSION_SUMMARIZED

Shipped

10/5/2026, 5:47:14 PM

Click a lifecycle step above to view its details.

03Sludge Pulse
hammz assigned hammz10/5/2026, 5:20:59 PM

Sign in to post a ripple.