Relationships
#2825 Workflow step that times out on the model lock leaves its output and run-tracker row stuck at 'running'
Opened by hammz · 9/30/2026· Shipped 9/30/2026
Summary
When a workflow model_method step times out waiting for the per-model lock (lock_timeout), its ModelOutput stays status: running permanently and its run-tracker row stays running, even though the workflow run is marked failed and the process has exited.
Cause
In src/domain/workflows/execution_service.ts, the step does three writes before it takes the lock:
- saves the evaluated definition (
evaluatedDefRepo.save, ~line 1619); - saves the output as
running(outputRepo.saveafteroutput.markRunning, ~line 1694); - registers with the run tracker (
runTracker.register, ~line 1708).
It then calls this.stepLockHook(...) (~line 1727). That call sits outside the try/catch whose failure path calls runTracker.complete(output.id, "failed") and handleMethodFailure (which marks the output failed). A LockTimeoutError from the hook skips both. The heartbeat interval never starts either.
Steps to reproduce
swamp repo init --tool noneswamp model create command/shell web, then setmethods.execute.arguments.run: "sleep 8"in the definition.- Create a workflow with one job holding two parallel steps (
first,second), bothmodel_methodonweb/execute, nodependsOn. SWAMP_LOCK_TIMEOUT_MS=2000 swamp workflow run <workflow>
Output:
main │ step second · web · execute · start
main │ step first · web · execute · start
[WRN] Waiting for lock ".../data/command/shell/<id>/.lock" held by ...
main │ failed second in 2.0s
main │ done first in 8.1s
system │ Failed workflow ... timed out after 2002ms- Inspect what is left behind:
| Record | first |
second (lock timeout) |
|---|---|---|
Output YAML (.swamp/outputs/command/shell/execute/...) |
succeeded, has completedAt |
running, no completedAt |
| Run-tracker row | completed |
running |
| Workflow run | failed |
swamp run history --active lists second as a running command/shell/execute. swamp run doctor --json reports active: 1, stale: 0.
The default 60s timeout triggers this whenever a step on a model waits behind a method that takes over 60s: parallel steps on one model, or two workflows touching the same model.
Why it doesn't self-heal
reapStaleRuns(src/infrastructure/persistence/run_tracker_store.ts:365) only updates the tracker row. The output YAML staysrunningforever.- CLI: after the 90s stale TTL, and once the pid is dead,
run doctor --fix(or the next method run) marks the rowinterrupted. That is also wrong, because the step failed. swamp serve(from reading the code, not reproduced): the row's pid is serve's own live pid on the local host, andreapStaleRunsreaps a local run only when its process is dead. So the row is never reaped, and long-lived serve instances collect a phantom "running" row for every step that times out on a lock.
Suggested fix
Take the step lock before the evaluated-definition, output and tracker writes. Alternatively, move the stepLockHook call inside the try so the failure path marks the output and tracker row failed. Add a regression test in execution_service_test.ts with a stepLockHook that throws LockTimeoutError, asserting the output is failed and RecordingRunTracker saw complete(..., "failed").
Environment
- swamp built from source at
mainfe52cda6 - Linux, filesystem datastore
Shipped
Click a lifecycle step above to view its details.
Sign in to post a ripple.