Relationships
#3093 Deliver workflow signals through a write-once outcome record instead of writing the run record
Opened by hammz · 10/6/2026· Shipped 10/6/2026
Problem
swamp-club#3068 added the wait_for_signal step with local delivery: workflow signal takes the run claim, re-reads the run, records the payload on the step and saves the run record. That works only when nothing else is writing the record, and three cases break it.
- Mid-level saves. The process running a run saves the record from memory, without the claim, until the level holding the wait has drained. Those saves already say suspended. A signal accepted in that window is erased by the owner's next save. #3068 guards this by reading the local run tracker (suspendedRunOwnerIsRunning) and refusing the signal as not ready, which is a heuristic that only works on the host that ran the workflow.
- Signals from another host. On a datastore shared between hosts the sender has no tracker row for the run, so the guard sees nothing. A signal can be accepted, reported as delivered, and then overwritten. This is listed as a limit in design/primitives/workflows.md.
- A serve owner that died mid-level. Its tracker row cannot be checked from the CLI, so the signal is refused as not ready until the deadline passes or the run is cancelled.
All three come from delivery writing the run record.
Proposal
Delivery stops touching the run record. Whoever settles a wait first creates one write-once record in the control-plane store with putIfAbsent, and the run record changes only inside a resume, which already holds the run claim.
Two record families:
| Key | Written by | Content |
|---|---|---|
| waits/(waitId) | The executor, when the step starts waiting | Workflow id, run id, job, step, kind, deadline, captured payload schema |
| wait-outcomes/(waitId) | putIfAbsent by whoever settles the wait first | An accepted signal (receipt and payload), timed_out, or cancelled |
- A signal, a timeout and a cancel each try to create the same outcome key. Exactly one wins, and the others read it and answer from it.
- Resume gains one step: for each waiting step, read its outcome. An accepted signal succeeds the step and copies the receipt and payload into its output; timed_out fails it with wait_timeout; no outcome means the wait is still open and the resume is refused as today.
- Copying the payload into the run record on resume keeps the record self-sufficient, so the control-plane records can be deleted once the run ends.
- Nothing user-visible changes: the wait ID, the receipt and the step output shape stay the same.
What this removes
- The not ready refusal and the run tracker readiness check in workflow signal.
- The rule that a signal must be sent from the host that ran the workflow.
- The dependence of workflow waits and workflow signal on scanning suspended runs to find a wait by ID: waits/(waitId) is the index.
To confirm first
- Whether the S3 and GCS datastore extensions implement putIfAbsent on ControlPlaneStore. The filesystem store does. This has not been checked.
- What a run does when the store lacks putIfAbsent. Proposed: validation refuses to start a run of a workflow that contains a wait, rather than degrading.
Acceptance criteria
- A signal sent while sibling steps of the wait's level are still running or queued is accepted and survives every later save by the owner.
- A signal sent from a second repository context on the same datastore is applied by the next resume.
- A signal against a timeout, and a signal against a cancel, each produce exactly one outcome; the loser is answered from the stored outcome.
- A second signal for a settled wait is refused and shown the stored receipt, as today.
- Lazy deadlines still work without serve: workflow signal refuses an expired wait, and workflow resume and workflow waits settle an expired wait as timed_out.
- Outcome and registration records are removed when their run ends or is cancelled, with a sweep for records whose run no longer exists.
- Runs suspended on a wait by a build from #3068, which hold the wait in the run record only, still accept a signal and resume.
Tests
- Property tests: one outcome per wait across random orderings of signal, timeout and cancel.
- Integration tests on a temp filesystem, ordering the creates rather than sleeping: signal during drain, signal against timeout, signal against cancel, signal from a second repository context.
- A conformance suite in packages/testing for putIfAbsent on ControlPlaneStore, since it becomes a requirement for extension authors.
- Update integration/workflow_run_claim_rules_test.ts for the changed claim use, and add rules for the new control-plane key families.
Out of scope
Serve delivery, the signal permission, the continuation and deadline sweeps (tracked separately), idempotency keys, and large or sensitive payloads.
References
swamp-club#3068 and PR 2873 in swamp-club/swamp. The design reasoning is in the Persistence section of the durable waits proposal: three write-once record families, applying an outcome, capability and cleanup.
Shipped
Click a lifecycle step above to view its details.
Sign in to post a ripple.