Relationships
#2379 Let a method inside serve start a workflow run: context.runWorkflow
Opened by hammz · 9/22/2026
Problem Statement
Split out of #2351 (reported by bixu). An extension method running inside
swamp serve has no supported way to start a workflow run. MethodContext
offers runModel, but nothing that queues a workflow through serve so that
placed steps land on workers and the caller gets a run id back.
#2351 itself is being scoped to the evaluation bug that blocked the reporter's
workaround (a nested workflow step's workflowIdOrName reading data.latest
is evaluated at run start). That fix lets one dynamic step per stage replace
the 46-step guarded matrix, but a driver method still cannot start the next
stage directly.
Proposed Solution
context.runWorkflow({ workflowIdOrName, inputs }) on MethodContext. It
queues a run through serve's run path and returns the run id.
Design questions to settle first
- Authorization. The gate precedent (
context.approveWorkflowGate/rejectWorkflowGate) was unwired in GitHub PR swamp-club/swamp#1960 (issue #1397) because model code bypassed serve's principal authorization. A method-started run must be authorized against the initiating principal'srungrant onworkflow:<name>, butMethodContextcarries noPrincipaltoday, only aninitiatedBystring insidetagOverrides. - Serve port. Serve starts runs only inside
handleWorkflowRun(src/serve/handlers/workflow_handlers.ts), which registers the run in theActiveRunRegistry(concurrency caps, cancel,run.attach), writes the control-plane active-run record, and pushes sync after the run. ArunWorkflowport would need to reuse that path rather than build an in-processWorkflowExecutionService, which would dispatch placed steps (the dispatcher is process-global) but skip all of the above. - Extension rules.
runModelgates cross-extension calls on the caller's manifestdependencies. Workflows carry no extension provenance: extension workflows are identified only by name, and a local workflow can shadow an extension workflow of the same name (CompositeWorkflowRepository). - Run id timing. The run id is minted by
WorkflowRun.createand emitted in thestartedevent after lookup, input validation and evaluation, so returning it synchronously means waiting forstartedor pre-minting it. - Remote execution. On a worker,
remote_method_context.tsstubsrunModeland the gate calls;runWorkflowwould need a stub or a new worker-to-orchestrator verb. - CLI (no serve). Decide whether
runWorkflowruns in-process outside serve, where placed steps fail at step time with no worker dispatcher.
Alternatives Considered
- In-process nested run from the method: gets worker placement for free but skips the registry, HA reconciliation, authorization and audit.
- A loopback webhook the pod signs to itself: rejected by the reporter (moves an internal call onto the network and needs a dedicated secret).
Additional Context
- Related: #2351, #1397, GitHub PR swamp-club/swamp#1960
- Reporter deployment: k3s pod,
swamp serve, managedConfig, S3 datastore, placed steps on a worker fleet
Open
No activity in this phase yet.
bixu commented 9/23/2026, 6:30:42 AM
🤖
Reporter-side answers to the design questions, from the harness that would call
runWorkflow.
Authorization. Our driver method runs inside serve for the principal that started the parent run: a webhook, a schedule, or an operator. A check of that principal's
rungrant onworkflow:<name>fits us. A webhook-started parent should carry the webhook's own principal. Then a route cannot launch workflows its grants do not cover.
Serve port. Reuse
handleWorkflowRun. We rely on the active-run registry for cancel andrun.attach, and on the post-run sync. An in-process run that skips those would break our dashboards and our retry path.
Extension rules. Name lookup is enough for us. We call only workflows our own extension bundles. A manifest
dependenciescheck for extension workflows would help, but it does not block us.
Run id timing. We need the run id to record which run owns a stage, not to wait on it. A pre-minted id returned at once suits us best. A wait for
startedalso works.
Remote execution. Our driver never runs on a worker. A stub that throws a clear error on workers is enough.
CLI without serve. A clear error outside serve is fine. We never drive propulsion from a local CLI.
Today's workaround: #2351's step-time evaluation lets us run one unrolled loop of advance steps and dynamic stage steps. It works, but a human or a webhook still starts each segment after a gate. With
runWorkflow, the approval itself could start the next segment.
Sign in to post a ripple.