Relationships
#2409 workflow resume: retry the failed steps of a failed run without --from
Opened by hammz · 9/23/2026· Shipped 9/23/2026
Problem
When a workflow run fails, swamp workflow resume <workflow> refuses it unless the operator names a step with --from. The operator has to read the run history, work out which steps failed, and pick one. When steps in several jobs fail, --from selects only one template and its dependents, so the operator has to resume more than once.
This is the first step toward durable execution. Durable execution means a run keeps its completed work when something goes wrong, and the operator continues it instead of starting over. This step delivers that for runs that fail cleanly: completed step results stay in the run record, and one command retries what failed. Later steps add recovery for interrupted runs, one resume path for every run status, and recovery from a shared datastore. This step adds none of those guarantees.
Proposed solution
Let swamp workflow resume <workflow> retry the failed steps of a failed run without naming a step. The CLI and serve give the same result.
Decision: reuse the existing --from reset and execution path. Add automatic step selection for failed runs whose records are complete and unambiguous. Refuse everything else with a clear message. Do not add a run type, repository, checkpoint format, or execution engine.
Current implementation
Verified against origin/main at 379bb5f0:
| Part | Current behavior | Source |
|---|---|---|
| Resume | Without --from, accepts suspended. With --from, accepts failed. |
src/domain/workflows/execution_service.ts, resume() |
| Reset selection | Finds one template step and its dependents through step and job dependsOn edges. Matches stored forEach iterations. |
Same file, computeStepsToReset() |
| State changes | Resets selected steps and their jobs to pending. Sets the run to running. |
src/domain/workflows/workflow_run.ts, resetForResumeFrom() and resumeFromFailed() |
| Result reuse | Skips terminal steps and restores their outputs into steps.*. |
resume() and runStep() |
| Run selection | Separate suspended-run and failed-run resolvers. approve and reject use the suspended one. |
src/domain/workflows/suspended_run_resolver.ts, src/libswamp/workflows/approve.ts, src/libswamp/workflows/reject.ts |
| Serve request | WorkflowResumeRequestSchema accepts from (fixed in #2356). |
src/serve/connection.ts |
| Failure history | Records the first step with status === "failed" and allowedFailure === false. |
WorkflowRun.computeFailureInfo() |
Three facts shape the design:
- A failed run can hold unfinished steps.
WorkflowRun.complete()marks a run failed when any job is neither succeeded nor skipped. Executor error handlers call it too. An expansion error can leave a job running with pending steps. A rejected approval can leave pending steps or later approval gates. - Step names are unique only within a job. The reset helper and the
steps.*context key steps by name alone. - A template reset repeats some successful work. Resetting a
forEachtemplate resets every stored iteration. Resetting a step resets all its dependents, including successful cleanup steps.
Terms
| Term | Meaning |
|---|---|
| Failed step | A stored step with status failed and allowedFailure !== true. |
| Entry template | The workflow step a failed step came from: forEachTemplate ?? stepName. |
| Reset set | The entry templates and their transitive dependents, as stored step names. |
allowedFailure is the recorded value. It already reflects allowFailure and the assertion severity threshold in force when the run failed.
Command behavior
| Stored run status | Without --from |
With --from |
|---|---|---|
suspended |
Current approval and resume behavior. | Refuse, as today. |
failed |
Retry failed steps after the checks below. | Current explicit selection. |
| Other status | Refuse and name the applicable next action. | Refuse. |
The retry keeps the run ID and start time. --input overrides merge over stored inputs as today. An override does not reset steps that used the old value.
Select the run
Use one resolver for resume. Without --from it accepts suspended and failed runs. With --from it accepts failed only. With --run, load that run and check its status. Without --run, require exactly one candidate. Otherwise list the candidate IDs and statuses and require --run. Do not select the newest run automatically.
approve and reject keep resolveSuspendedRun unchanged.
Check the run
Before changing a failed run without --from, require all of:
- Every job and step is
succeeded,failed, orskipped. - At least one failed step exists, and every failed job contains one.
- Each failed step's entry template is a step in the current workflow.
- Step names are unique across the whole workflow, and stored step names are unique across the run.
Check 1 excludes runs with pending, running, waiting, or unknown work. Check 2 excludes jobs that failed before any step ran, such as expansion errors. Check 3 excludes older forEach records that lack forEachTemplate. Check 4 lets the change reuse the name-keyed reset helper and expression context. Supporting repeated names across jobs needs job-scoped identity in both and is separate work.
If a check fails, name the job or step, leave the run unchanged, and point the operator to workflow history and the run log. Suggest --from only when it helps: an older record without forEachTemplate, or a rejected approval. The existing --from path can reset a rejected gate; the gate then asks for a new decision.
Select and reset the steps
- Add
WorkflowRun.failedSteps(). Return the job name, step name, andforEachTemplatefor each failed step, in stored order. - Add a pure function beside
computeStepsToReset()that runs the checks above and returns the distinct entry templates. - Call
computeStepsToReset()once per entry template and union the results. - Call
resetForResumeFrom()with the union, thenresumeFromFailed(), then the existing save. Continue through the existing resume executor.
Put this in WorkflowExecutionService.resume() where fromStep is absent and the run is failed. Every caller then gets the same checks. Check for a waiting approval only when the run is suspended.
Reset clears each selected step's output references, error, approval decision, and assertion result, whatever its old status. Only jobs that contain a reset step return to pending. Steps outside the reset set keep their stored state and outputs. Guards still decide whether a reset step runs. A guard that skips it does not restore its old outputs.
Several failures produce several entry templates. Explicit --from still selects one template and its dependents.
Connect the CLI and serve
- Replace the two-resolver ternary in the CLI (
src/cli/commands/workflow_resume.ts) and both serve handler branches (src/serve/handlers/workflow_handlers.ts) with the single resolver. - Keep the existing event and JSON formats. Add no
resumedFromfield.
Explain the next action
When a run fails with a recorded failed step, print:
To retry failed steps: swamp workflow resume <workflow> --run <id>Preserve --server or repository selection in the printed command. Print no credentials. When no failed step is recorded, point to history and logs instead.
Update nextActionForStatus("failed") to the same command. Update the command description and examples. Use this --from help text:
Select the retry step in a failed run. By default, resume retries all failed steps.
Limits
These limits are part of the operator contract:
- Retry can repeat external effects. A method can change an external system and then fail. Resume does not guarantee exactly-once execution.
- Definitions and inputs must stay compatible. Resume uses the current workflow and model definitions. It does not detect definition changes or prove that stored results remain valid. If an input change affects earlier work, use
--fromor start a new run. forEachcollections must stay stable. Item identity is not preserved across collection changes. Use a new run for a changed collection.- Stored references do not guarantee data. Ephemeral data is gone after restart, and retention can remove artifacts. Resume does not reconstruct missing outputs or pin
data.latest()to its old value. - A retried nested workflow starts a new child run. It does not resume the earlier child.
- One operator per run. There is no ownership lock. Separate CLI processes or serve instances can race.
- Interrupted, cancelled, and running runs are out of scope. Interrupted runs still use
recover. This change adds no crash-recovery guarantee.
Later durable-execution steps address the last three: attempt identity and child reuse, run ownership, and checkpoints that make interrupted runs resumable.
Required tests
- Domain:
failedSteps()excludes recorded allowed failures. Entry template mapping. Each check's refusal: pending, running, waiting, and unknown steps; an expansion error; a rejection with pending work; a failed job without a failed step; a missingforEachTemplate; repeated names across jobs. A refusal saves nothing and invokes no method. - Property: for eligible runs, the reset set contains every failed step and all selected dependents. Steps outside the set are unchanged. Only containing jobs reset.
- Executor: a failed run resumes without
--from. Successful independent steps do not run again. Cover parallel failures, dependent jobs, successful cleanup steps, guards, and all iterations of a selectedforEachtemplate. - Resolver: explicit IDs, no candidates, and ambiguity across statuses. A suspended run does not make
--fromambiguous.approveandrejectignore failed runs. - Serve and presentation:
fromsurvives request validation. Both handler branches resume automatically and explicitly. Log guidance, remote targets, JSON output, and unchanged suspended-run behavior. - In-process integration: save a failed run in a real YAML repository, reload it, and resume automatically. Assert method call counts, unchanged results outside the reset set, the same run ID, and the persisted final state. Keep the existing resume provenance tests (
integration/workflow_resume_provenance_test.ts) passing. - UAT in swamp-uat: fail a workflow, run the printed command, and verify the retry. Cover log and JSON modes, locally and through
--server.
Delivery
One focused PR. Update the workflow design (design/primitives/workflows.md), CLI help, swamp skill, manual, and swamp-uat journey with the implementation. Add no new libswamp exports. In the workflow skill, tell agents to check job and step states before retrying.
Alternatives considered
- Keep requiring
--fromfor failed runs. The operator keeps doing the step selection by hand, and a run with failures in several jobs needs several resumes. - A new run type, repository, checkpoint format, or execution engine. More than this step needs. Reusing the
--fromreset and the existing resume executor keeps the change small, and later durable-execution steps can build on it. - Automatically resume the newest failed run. Rejected: with more than one candidate the resolver lists them and requires
--run, so a resume never picks the wrong run.
Related: #2153 (recover interrupted runs from durable checkpoints), #2356 (serve accepts resume --from), #1953 (durable, detached workflow execution).
Shipped
Click a lifecycle step above to view its details.
Sign in to post a ripple.