Skip to main content
← Back to list
01Issue
FeatureOpenSwamp CLIPublic
AssigneesNone

Relationships

#3108 Resume signalled workflow runs automatically under swamp serve

Opened by hammz · 10/6/2026

Problem

Split out of swamp-club#3094, which now covers only delivering a signal through swamp serve. After it, a signal can arrive over WebSocket or HTTP, but the run still sits suspended until someone resumes it by hand, also when serve started the run.

Serve has no way to learn that a signalled run can continue:

  • A signal creates an outcome record and never changes the run record (swamp-club#3093). WorkflowRun.isAwaitingResume() stays false while any signal wait is on the record, settled or not, so the derived awaitingResume is unset.
  • autoResumeAfterApproval in src/serve/resume_launcher.ts is called only from the approve handler.
  • autoResumeParentAfterChild returns false outright when the parent has a signal wait.
  • Nothing in serve sweeps for suspended runs. A lost auto-resume launch is logged and audited and never retried.

Two serve instances on one datastore could also both launch the same resume: the active-run registry is per process, and serve resumes with unclaimedRuns instead of the lock-backed run claim.

Proposal

  • Decide resumability from the run record together with the outcome records: suspended, no gate undecided, no nested run waited on, and an outcome for every wait. isAwaitingResume(store, run) in src/libswamp/workflows/signal.ts already computes this for the signal result and can move into the domain.
  • Reuse the workflow's existing autoResume setting; no second default for signals.
  • Validation requires a workflow that contains a wait and declares inputs to set autoResume explicitly, so no signal-driven run silently waits for a resume nobody knows to send.
  • A continuation claim, created with putIfAbsent per run and suspension, so that two serve instances never launch the same resume. A claim held by an instance with no live heartbeat is released and competed for again.
  • The instance that accepts a signal attempts the continuation. It may not have the run record (runRecordAvailable: false), so the sweep is the path that always works.
  • A continuation sweep at boot and on an interval: find suspended runs whose waits are all settled and whose policy allows, and launch them through the claim. Capacity, shutdown and authorization refusals leave the run suspended with a visible reason, and the next sweep retries.
  • The resume runs as the run's initiator, re-authorized when it launches. The signal's submitter is an audit actor only.
  • A parent suspended on a child that waited for a signal is auto-resumed when the child ends, as it is for a child that waited on a gate.
  • On resume takeover, pull the run record under the claim when the datastore is remote, and have serve take the lock-backed claim for this path.

Open decisions

  1. Where continuation claims live. Wait records are in the datastore's own control-plane store. Serve's other control-plane records, heartbeats included, fall back to a directory local to each repository on a filesystem datastore. Dead-holder detection needs claims and heartbeats visible to every instance.
  2. How a resume is authorized as the initiator. initiatedBy on the run is a plain string, and decideSubjectAccess refuses a subject with no token binding. The cron trigger authorizer is the nearest precedent.
  3. Whether the sweep acts on a custom datastore, where this host's copy of a run record can be missing or behind. The wait-record cleanup sweep refuses to act there for that reason.
  4. Whether the sweep also resumes approved runs whose auto-resume launch was lost. That is a fix, but it changes when such runs execute.

Acceptance criteria

  • A signal that settles a run's last wait resumes the run without a manual resume when autoResume allows, and exactly once with two serve instances on the same datastore.
  • A resume launch lost to a crash or a capacity refusal is retried by the sweep.
  • A continuation claim left by a dead instance does not strand the run.
  • A run whose initiator lost the authority to run the workflow is not resumed, and its receipt stays readable.

Tests

  • Integration tests with ordered creates, never sleeps: two instances competing for one continuation, a claim left by a dead instance, a launch refused for capacity and retried.
  • integration/workflow_run_claim_rules_test.ts updated for the resume change.

Out of scope

Settling overdue waits (swamp-club#3109), the client and dashboard display (swamp-club#3110), leases with fencing for executors, continuing one branch while another still waits.

References

swamp-club#3094 (delivery), swamp-club#3093 (outcome records), the Wait for Signal section of design/primitives/workflows.md.

02Bog Flow
◉OPEN○TRIAGED○IN PROGRESS○SHIPPED

Open

10/6/2026, 9:09:44 PM

No activity in this phase yet.

03Sludge Pulse

Sign in to post a ripple.