Relationships
#1738 Cancel endpoint reports status: cancelled for runs it did not actually cancel
Opened by hammz · 8/19/2026· Shipped 8/20/2026
Problem
POST /api/v1/cancel/(workflow-run|method-run)/<id> returns {"status":"cancelled"} as soon as an AbortController is aborted. Aborting is not cancelling: for in-process runs the abort is purely cooperative, so a hung or stuck run keeps executing while the caller is told it stopped. There is no later correction — the response is the only signal the client ever gets.
swamp workflow cancel --run <id> --server then prints Cancelled run <id> on server for a run that is still going.
Why the abort does not stop a stuck run
RunCancelRegistry.cancelandActiveRunRegistry.cancelonly callcontroller.abort(...)(src/serve/run_cancel_registry.ts:63-72,src/serve/active_run_registry.ts:141-146).- The workflow engine only tests
options.signal.abortedat step boundaries (src/domain/workflows/execution_service.ts:1918, 2317, 2361). A step that is already stuck never reaches one. - Extension methods execute in-process via
InProcessExecutor— same process asswamp serve. The signal is handed to the method context (src/libswamp/models/run.ts:744) but there is no subprocess to kill and no race against the signal, so a method that ignoresctx.signal(infinite loop, blocking call,fetchwithout the signal) runs to completion.
Consequences beyond the wrong status: the run stays in the registry, and because the finally that calls flushLocks() never runs, it keeps holding its model locks. The existing backstops do not help — --max-run-duration fires the same cooperative abort (src/serve/active_run_registry.ts:100-108), and stale-run reaping only rewrites tracker records. Restarting serve is the only way out.
Contrast with the two paths that do work:
- Local
swamp workflow cancel/swamp model cancelusekillProcessTree— SIGTERM, 2s poll, then SIGKILL to the process and its children (src/infrastructure/process/process_kill.ts:120-142). - Worker-dispatched steps send
runner.canceland then unconditionallychild.kill()afterRUNNER_CANCEL_GRACE_MS(src/worker/dispatch_handler.ts:258-274), which works because the runner is a separate process.
Steps to reproduce
- Start
swamp serve. - Run a workflow whose step calls a model method that ignores
ctx.signal(e.g. a loop with a plainawait new Promise(r => setTimeout(r, 1000))). POST /api/v1/cancel/workflow-run/<runId>(orswamp workflow cancel --run <id> --server).
Actual: HTTP 200 {"status":"cancelled"}; the run keeps emitting events, stays in GET /api/v1/health activeRuns, and holds its locks.
Expected: the response distinguishes "abort delivered" from "run stopped".
Suggested direction
Return a status that means what happened — e.g. cancellation_requested — and let the caller confirm termination (poll the active-run list, or have the endpoint wait a bounded grace period for the run to leave the registry before answering). A force path matching the worker's grace-then-kill behaviour would need out-of-process execution for in-process runs, so it is a larger change; at minimum the reported status should stop claiming a cancellation that did not happen.
Found while triaging #1536 (adding --server support to swamp model cancel), which would inherit this same misreporting.
Shipped
Click a lifecycle step above to view its details.