Skip to main content
← Back to list
01Issue
BugShippedSwamp CLIPublic
Assigneeshammz

Relationships

#2420 workflow resume keeps the original run's process identity (run record pid and instanceId)

Opened by hammz · 9/23/2026· Shipped 9/28/2026

Problem

A resume runs in a new process, but the run record keeps the pid and instanceId of the process that started the run. WorkflowExecutionService.resume() calls resumeFromSuspended() or resumeFromFailed(), which change only the status; run.start(pid, instanceId) runs only for a fresh run. This affects every resume: a suspended run after approval, --from on a failed run, and the automatic retry added by swamp-club#2409.

Consequences while a resume is running:

  1. swamp workflow cancel does not stop it. cancelRun() (src/cli/commands/workflow_cancel.ts) kills run.pid, which is the original, usually dead, process. The record is marked cancelled, then the live resume overwrites it on its next save. If the old pid has been reused by another swamp process, that process is killed instead (killProcessTree checks the command line contains swamp on Linux and macOS; Windows has no check).
  2. Serve-owned classification is stale. isServeOwnedRun() reads the original instanceId. A CLI resume of a run that serve started is skipped by local cancel as serve-owned, while serve's cancel endpoint has no active entry for it.
  3. Serve boot can interrupt a live resume of an old failed run. Tracker rows for finished runs are purged after 7 days (RETENTION_DAYS in run_tracker_store.ts). With no tracker row, reapOrphanedWorkflowRuns (src/cli/commands/serve.ts) falls back to the record's pid, finds it dead, and marks the live run interrupted.

swamp-club#2409 fixes the tracker half: reactivate() accepts failed rows and records the resuming pid and hostname. That covers suspended resumes and failed runs whose tracker row still exists. It does not change the run record.

Why it was not fixed in #2409

The record's pid cannot change on its own. If a resume through serve wrote serve's pid into a record with no instanceId, local workflow cancel would treat the run as local and kill the serve process (compare swamp-club#1769). The fix has to record the resuming process's pid together with serve's instance id, which means passing the instance id into resume() from the serve handler and the detached launcher. That is the run-ownership step the durable-execution plan in #2409 defers.

Expected

After a resume starts, the run record identifies the process that is running it: its pid, and serve's instance id when serve drives the resume. workflow cancel stops the live resume, and serve boot does not interrupt it. A resume of a failed run older than the tracker retention window is either tracked again or reconciled by the record's current pid.

Reproduction (Linux)

  1. swamp workflow run <wf> with a step that fails; note the run id.
  2. swamp workflow resume <wf> --run <id> --from <step> with a slow step, and leave it running.
  3. In another shell, swamp workflow cancel <wf> --run <id>: the record says cancelled, but the resume keeps running and overwrites it.
02Bog Flow
✓OPEN✓TRIAGED✓IN PROGRESS✓SHIPPED+ 1 MOREASSIGNED+ 8 MOREREVIEW+ 13 MOREPR_MERGED+ 2 MORESESSION_SUMMARIZED

Shipped

9/28/2026, 7:57:47 PM

Click a lifecycle step above to view its details.

03Sludge Pulse
hammz assigned hammz9/28/2026, 6:04:10 PM

Sign in to post a ripple.