Relationships
#2914 model cancel SIGTERMs a running swamp serve process when the method run belongs to a serve-run workflow
Opened by hammz · 10/1/2026· Shipped 10/1/2026
Description
swamp model cancel (src/cli/commands/model_cancel.ts) picks running model_method rows from the run tracker by model type only and calls killProcessTree on the row's pid. Workflow steps register those rows with ActiveRun.createModelMethodRun using pid Deno.pid and no instanceId (src/domain/workflows/execution_service.ts, src/libswamp/models/run.ts). swamp serve hands its runTracker to the WorkflowExecutionService it builds (src/serve/deps.ts) and shares the repo tracker database (RunTrackerStore.fromSwampDir).
So for a workflow that swamp serve is running, the step's tracker row carries the serve process's pid. swamp model cancel (or --all) passes isSwampProcess, sends SIGTERM to the serve process, and SIGKILLs it after the default 2 s, well under serve's 30 s shutdown drain. Every in-flight run on that server is cut off, not just the one method run.
workflow cancel already avoids this: isServeOwnedRun (src/cli/commands/workflow_cancel.ts) routes serve-owned runs through the server instead of killing a pid. model cancel has no equivalent, and the method rows carry no instanceId to tell them apart.
Found by code reading while triaging swamp-club#2910. Not reproduced end to end.
Steps to reproduce (suggested)
- swamp serve in one shell, in a repo with a workflow whose step runs a long shell method (sleep 60).
- Trigger the workflow through the server.
- From a second shell in the same repo: swamp model cancel
Expected
model cancel never signals a serve process. Serve-owned method runs are either cancelled through the server or refused with a message pointing at the server, and method rows record the instanceId of the serve that owns them.
Actual (expected from the code)
The serve process receives SIGTERM and is SIGKILLed about 2 s later.
Shipped
Click a lifecycle step above to view its details.
hammz commented 10/1/2026, 7:29:31 PM
Note from swamp-club#2910 (adversarial review): the fix for #2910 makes swamp model cancel wait longer before it SIGKILLs the owning process. A standalone method owner now gets METHOD_OWNER_STOP_GRACE_MS (10 s), and a pid that also owns a running workflow tracker row gets the workflow OWNER_STOP_GRACE_MS (40 s). For a serve-run workflow the owning pid is the serve process, so model cancel would now SIGTERM serve and wait up to 40 s before SIGKILLing it, instead of 2 s. The new log line also says to cancel again to stop immediately, which is wrong for serve: serve's shutdown handler does not set forceExitOnRepeat, so a second SIGTERM does not force an exit. Either way the root cause is unchanged: model cancel should not signal a serve-owned pid at all (route through the server, as workflow cancel does with isServeOwnedRun).
Sign in to post a ripple.