Skip to main content
← Back to list
01Issue
BugShippedSwamp CLIPublic
Assigneeshammz

Relationships

#2825 Workflow step that times out on the model lock leaves its output and run-tracker row stuck at 'running'

Opened by hammz · 9/30/2026· Shipped 9/30/2026

Summary

When a workflow model_method step times out waiting for the per-model lock (lock_timeout), its ModelOutput stays status: running permanently and its run-tracker row stays running, even though the workflow run is marked failed and the process has exited.

Cause

In src/domain/workflows/execution_service.ts, the step does three writes before it takes the lock:

  • saves the evaluated definition (evaluatedDefRepo.save, ~line 1619);
  • saves the output as running (outputRepo.save after output.markRunning, ~line 1694);
  • registers with the run tracker (runTracker.register, ~line 1708).

It then calls this.stepLockHook(...) (~line 1727). That call sits outside the try/catch whose failure path calls runTracker.complete(output.id, "failed") and handleMethodFailure (which marks the output failed). A LockTimeoutError from the hook skips both. The heartbeat interval never starts either.

Steps to reproduce

  1. swamp repo init --tool none
  2. swamp model create command/shell web, then set methods.execute.arguments.run: "sleep 8" in the definition.
  3. Create a workflow with one job holding two parallel steps (first, second), both model_method on web / execute, no dependsOn.
  4. SWAMP_LOCK_TIMEOUT_MS=2000 swamp workflow run <workflow>

Output:

main │ step second · web · execute · start
main │ step first · web · execute · start
[WRN] Waiting for lock ".../data/command/shell/<id>/.lock" held by ...
main │ failed second in 2.0s
main │ done first in 8.1s
system │ Failed workflow ... timed out after 2002ms
  1. Inspect what is left behind:
Record first second (lock timeout)
Output YAML (.swamp/outputs/command/shell/execute/...) succeeded, has completedAt running, no completedAt
Run-tracker row completed running
Workflow run failed

swamp run history --active lists second as a running command/shell/execute. swamp run doctor --json reports active: 1, stale: 0.

The default 60s timeout triggers this whenever a step on a model waits behind a method that takes over 60s: parallel steps on one model, or two workflows touching the same model.

Why it doesn't self-heal

  • reapStaleRuns (src/infrastructure/persistence/run_tracker_store.ts:365) only updates the tracker row. The output YAML stays running forever.
  • CLI: after the 90s stale TTL, and once the pid is dead, run doctor --fix (or the next method run) marks the row interrupted. That is also wrong, because the step failed.
  • swamp serve (from reading the code, not reproduced): the row's pid is serve's own live pid on the local host, and reapStaleRuns reaps a local run only when its process is dead. So the row is never reaped, and long-lived serve instances collect a phantom "running" row for every step that times out on a lock.

Suggested fix

Take the step lock before the evaluated-definition, output and tracker writes. Alternatively, move the stepLockHook call inside the try so the failure path marks the output and tracker row failed. Add a regression test in execution_service_test.ts with a stepLockHook that throws LockTimeoutError, asserting the output is failed and RecordingRunTracker saw complete(..., "failed").

Environment

  • swamp built from source at main fe52cda6
  • Linux, filesystem datastore
02Bog Flow
✓OPEN✓TRIAGED✓IN PROGRESS✓SHIPPED+ 1 MOREASSIGNED+ 2 MOREREVIEW+ 7 MOREPR_MERGED+ 2 MORESESSION_SUMMARIZED

Shipped

9/30/2026, 6:45:34 PM

Click a lifecycle step above to view its details.

03Sludge Pulse
hammz assigned hammz9/30/2026, 6:08:36 PM

Sign in to post a ripple.