Relationships
#2550 workflow resume has no job-level cleanup mode: after --timeout or a cancel, always/completed cleanup jobs start with the aborted signal and fail at once
Opened by hammz · 9/25/2026· Shipped 9/28/2026
Summary
WorkflowExecutionService.run() gives job levels after an abort a fresh 30-second cleanup signal (swamp-club#1785), so always/completed/failed-gated cleanup jobs can run. The job-level loop in resume() has no cleanup mode: it passes the original, already-aborted signal to every later level. A later level with several jobs starts nothing, because mergeWithConcurrency returns at once on an aborted signal. A single-job level runs the job with the aborted signal, so its steps fail at once. Step-level cleanup inside a job is unaffected, because runJob() is shared.
Steps to reproduce
- A workflow with three jobs: gate (one manual_approval step), work (dependsOn gate succeeded; one step running sleep 5), and cleanup (dependsOn work with condition type always; one step that echoes a marker).
- swamp workflow run, then swamp workflow approve with the gate step name, then swamp workflow resume with --timeout 2s.
Observed with 20260925.180024.0-sha.e08f1d5b: work fails when the deadline passes. The cleanup job starts, its step fails 13ms later with Command exited with code -1, and the marker is never printed. The run ends cancelled with job cleanup failed.
Expected
The same as a fresh run: after an aborted job level, later levels get the cleanup grace signal and in-flight jobs and steps are marked failed with error cancelled, so the cleanup job runs.
Context
Found while planning swamp-club#2543. The resume() job loop is the one around the second mergeWithConcurrency call over job streams in src/domain/workflows/execution_service.ts.
Shipped
Click a lifecycle step above to view its details.
Sign in to post a ripple.