Relationships
#2993 Feedback: Swamp as an approval-gated control plane for a small fleet
Opened by kklas · 10/2/2026
Swamp version: 20260930.003205.0 (self-hosted swamp serve, token auth, filesystem datastore)
Period: 2026-09-28 → 2026-10-02 (first live production runs)
How we use it
One operator, a handful of Linux servers running blockchain validators, configured with Ansible. swamp serve runs on a private network in token mode with --approve-requires-explicit-grant and deny-first grants. Workflow definitions live in a git repo; a timer fast-forwards serve's checkout, so a merged PR is live within minutes.
Two workflows so far:
- a 5-minute timer that detects upstream releases (opens a pin-bump PR) and starts the apply workflow per host as a nested run;
- the apply workflow: snapshot the commit →
ansible-playbook --check --diff→ if anything would change, Slack +manual_approval→ apply from the same snapshot → health-check steps → Slack.
The business logic is shell scripts in the repo; Swamp schedules, gates, records and authorizes. Planned next: intake from an on-call pager and a chat announcement mirror, emergency runbooks as workflows, image builds.
What worked well
- Deny-first grants +
--approve-requires-explicit-grant. "The scheduler may start this, only the operator may approve it" was easy to express, andswamp access checkmade it reviewable. - Token-gated approvals over the server, and
workflow validatecatching schema mistakes before merge. - Nested runs are independent: a waiting child survives later parent runs and is approved on its own. This is what made our design possible (see 1).
- Testable locally: we ran every path — approve, reject, expire, failed check, failed apply — against stubbed commands in a scratch repo before touching production.
swamp run gcis safe beside a live serve and never deletes suspended runs.
Friction, ranked by impact
1. Scheduled runs supersede suspended runs, with no way to opt out
Starting a run cancels any suspended run of the same workflow with matching inputs ("Superseded by new run with matching inputs"). Timer runs always carry identical inputs and the scheduler never passes noSupersede. So a timer directly on an approval-gated workflow cancels its own pending approval on the next tick. Our approval window would have been 5 minutes.
We worked around it by nesting: the timer workflow starts the gated workflow as a step, because nested starts don't supersede. It works, but it is non-obvious and has side effects (see 4).
Ask: a per-workflow supersede: false, or "scheduled runs never supersede".
2. No git/ref trigger
"Run when this branch moves" is the trigger we actually want. Without it we poll every 5 minutes and compare tree hashes ourselves, which is what forces a timer — and therefore issue 1 — in the first place.
Ask: a trigger.ref (poll a git ref in the repo dir, fire on change).
3. The approval lifecycle is silent
- Expiry is only checked when someone acts. After the timeout,
approve/rejectrefuse andapprovalshides the run, but the run stayssuspendedforever and no step can react. We wanted "on expiry, post to Slack" and had no hook for it. autoResumedefaults off at serve level, and the CLI's hint after approving says "runswamp workflow resume". Our first production approval sat suspended until we found the flag in the source.- Step outputs from before a resume are gone, so guards after an approval can't read pre-approval results; we pass everything through files instead.
Ask: an expired (and rejected) condition usable in dependsOn; autoResume on by default, or at least a CLI hint that reflects the server's setting; documented (or preserved) step outputs across resume.
4. Semantics we could only learn from the source
Each cost us a source dive; the docs didn't say:
- several
dependsOnentries are ANDed; trigger.inputsexpressions resolve for webhook runs only — scheduled runs get literals;manual_approval.timeoutis in seconds,command/shelltimeoutin milliseconds;- a nested child that suspends makes the parent step fail (we needed
allowFailure); a distinct "child waiting" status would read better; run.idvs the legacyworkflowRunId, and which expressions are deferred to step time.
Ask: a "workflow semantics" reference page covering dependencies, conditions, supersede, suspension/resume and nesting.
5. Operational papercuts
- No run-history retention in serve. A 5-minute timer leaves ~3 run records per tick (~35–50 MB/day for us).
garbageCollectionin.swamp.yamlis read by the CLI'srun gc, not applied by serve; we added an external daily timer. Serve-side retention would help. .swamp.yamlis rewritten on first use to add arepoId, dropping comments. In a read-only checkout that a timer resets, this dirties the tree and mints a new id after every reset. We now commit therepoId.SWAMP_HOMEunder the home directory gets mistaken for a repo by commands run from~.swamp serve daemon enablemapssystemctl reloadto SIGHUP, which stops serve unless--hot-reloadis set; we wrote our own unit.
6. Trust and supply chain
For a control plane that can change production servers:
- releases are frequent (~40 builds/day) and "stable" is simply the newest; artifacts carry a sha256 but no signature;
- a self-hosted serve still depends on the hosted account for authentication.
Ask: signed releases (cosign/minisign) and a slower, supported release channel; clarity on what a self-hosted serve needs from the hosted service at runtime, and what happens if it's unreachable.
Summary
Swamp's core — grants, token-gated approvals, recorded runs, definitions in git — fits this use well. Issues 1–3 shaped our design the most: supersede-on-schedule forced the nested workaround, the missing ref trigger forced polling, and the silent approval lifecycle left gaps we fill with our own scripts. Fixing those three would make Swamp a natural fit for approval-gated infrastructure changes.
Open
No activity in this phase yet.
hammz commented 10/2/2026, 10:17:52 PM
Some great feedback here! We'll consider all of it, and I'm definitely going to look into the "superseding" and think about how to solve that problem.
hammz commented 10/5/2026, 6:26:28 PM
@kklas — I tried to reproduce item 1 (scheduled runs superseding their own pending approval) and couldn't, so I'd like a bit more detail before we decide what to change.
What I tested
A scratch repo with a filesystem datastore and this workflow, under swamp serve:
name: gated
trigger:
schedule: "* * * * *"
jobs:
- name: main
steps:
- name: gate
task:
type: manual_approval
prompt: "approve?"- Serve on current main, four ticks: four suspended runs, none cancelled.
- Serve restarted, one more tick: all five runs still suspended, including the four started by the previous serve process.
- Serve built from the source of your version (20260930.003205), four ticks: four suspended runs, none cancelled.
swamp workflow run gatedtwice from the CLI: the second run cancelled the first with "Superseded ... (matching inputs)". The serve-started runs were left alone.
Why I think that is
Supersede skips any suspended run that is owned by a serve instance. Every run the scheduler starts is stamped with serve's instance id and keeps it while suspended, so the next tick doesn't cancel it. A run started by a local swamp workflow run has no instance id, so a later local run with the same inputs does supersede it. You're right that the scheduler never passes noSupersede; it just doesn't need to for runs it started itself.
Where my setup differed from yours
- I used
--auth-mode none; you run token mode with--approve-requires-explicit-grant. - My workflow suspends immediately: no steps before the gate and no
trigger.inputs. - I ran from source rather than the released binary.
What would help
- Did you see "Superseded by new run with matching inputs" on a run that serve's scheduler started, or only on runs started locally with
swamp workflow run(for example while testing in the scratch repo)? - If it was a serve-scheduled run: the output of
swamp workflow history get <workflow> --run <cancelled-run-id> --json(or the run's record), in particular itstriggerSource,instanceIdandinitiatedBy, plus the same for the run that superseded it. - Was anything else starting runs of that workflow against the same repo: a second serve instance, a
swamp workflow runfrom a timer or script, or a--servercall from another host? - Did serve restart or reload between the run suspending and being cancelled?
If it turns out to be the local CLI path only, then the gap is that the docs don't say serve-owned runs are never superseded, and we'll fix that. If you did see it on a scheduled run, the details above should let us find the path I missed.
hammz commented 10/5/2026, 6:26:46 PM
Correction to the command in point 2 above: it takes the run id directly, with no --run flag: swamp workflow history get <cancelled-run-id> --json.
Sign in to post a ripple.