Relationships
#2549 After a cancel or --timeout, jobs that share a level are abandoned mid-flight: their steps stay running and their own cleanup steps never run
Opened by hammz · 9/25/2026· Shipped 9/28/2026
Summary
When a run's signal aborts while a topological job level holds more than one job, merge() in src/infrastructure/stream/merge.ts closes its queue and the consumer in WorkflowExecutionService.run() moves on. Each job's runJob() generator keeps running in the background only until its next yield, then the drain loop fails to push and calls return() on it. runJob() never reaches its own step-level cleanup mode (swamp-club#1785), so:
- steps that were in flight stay running in the persisted run, even though the run is cancelled; the jobs stay running too;
- always-, completed- and failed-gated cleanup steps inside those jobs never run and stay pending.
A single-job level is not affected: merge() uses yield* for one stream, so runJob() runs its cleanup inline.
Steps to reproduce
One workflow, two jobs with no dependency between them, so they share a level:
- j1: step slow1 runs sleep 5; step cleanup1 has dependsOn slow1 with condition type always and echoes a marker.
- j2: step slow2 runs sleep 5.
Run swamp workflow run with --timeout 2s. Observed with 20260925.180024.0-sha.e08f1d5b: the log shows both steps start and then Cancelled workflow, with no cleanup1 line. swamp workflow history get shows the run cancelled, j1 running (slow1 running, cleanup1 pending) and j2 running (slow2 running).
Expected
As for a single-job level: in-flight steps end failed with error cancelled, and each job's cleanup steps whose conditions hold run with the 30-second cleanup grace signal. Jobs end in a terminal state.
Context
Found while planning swamp-club#2543, which covers queued steps and jobs that never started. This one is about jobs that did start.
Shipped
Click a lifecycle step above to view its details.
Sign in to post a ripple.