Skip to main content
← Back to list
01Issue
FeatureOpenSwamp CLIPublic
AssigneesNone

Relationships

#2993 Feedback: Swamp as an approval-gated control plane for a small fleet

Opened by kklas · 10/2/2026

Swamp version: 20260930.003205.0 (self-hosted swamp serve, token auth, filesystem datastore) Period: 2026-09-28 → 2026-10-02 (first live production runs)

How we use it

One operator, a handful of Linux servers running blockchain validators, configured with Ansible. swamp serve runs on a private network in token mode with --approve-requires-explicit-grant and deny-first grants. Workflow definitions live in a git repo; a timer fast-forwards serve's checkout, so a merged PR is live within minutes.

Two workflows so far:

  • a 5-minute timer that detects upstream releases (opens a pin-bump PR) and starts the apply workflow per host as a nested run;
  • the apply workflow: snapshot the commit → ansible-playbook --check --diff → if anything would change, Slack + manual_approval → apply from the same snapshot → health-check steps → Slack.

The business logic is shell scripts in the repo; Swamp schedules, gates, records and authorizes. Planned next: intake from an on-call pager and a chat announcement mirror, emergency runbooks as workflows, image builds.

What worked well

  • Deny-first grants + --approve-requires-explicit-grant. "The scheduler may start this, only the operator may approve it" was easy to express, and swamp access check made it reviewable.
  • Token-gated approvals over the server, and workflow validate catching schema mistakes before merge.
  • Nested runs are independent: a waiting child survives later parent runs and is approved on its own. This is what made our design possible (see 1).
  • Testable locally: we ran every path — approve, reject, expire, failed check, failed apply — against stubbed commands in a scratch repo before touching production.
  • swamp run gc is safe beside a live serve and never deletes suspended runs.

Friction, ranked by impact

1. Scheduled runs supersede suspended runs, with no way to opt out

Starting a run cancels any suspended run of the same workflow with matching inputs ("Superseded by new run with matching inputs"). Timer runs always carry identical inputs and the scheduler never passes noSupersede. So a timer directly on an approval-gated workflow cancels its own pending approval on the next tick. Our approval window would have been 5 minutes.

We worked around it by nesting: the timer workflow starts the gated workflow as a step, because nested starts don't supersede. It works, but it is non-obvious and has side effects (see 4).

Ask: a per-workflow supersede: false, or "scheduled runs never supersede".

2. No git/ref trigger

"Run when this branch moves" is the trigger we actually want. Without it we poll every 5 minutes and compare tree hashes ourselves, which is what forces a timer — and therefore issue 1 — in the first place.

Ask: a trigger.ref (poll a git ref in the repo dir, fire on change).

3. The approval lifecycle is silent

  • Expiry is only checked when someone acts. After the timeout, approve/reject refuse and approvals hides the run, but the run stays suspended forever and no step can react. We wanted "on expiry, post to Slack" and had no hook for it.
  • autoResume defaults off at serve level, and the CLI's hint after approving says "run swamp workflow resume". Our first production approval sat suspended until we found the flag in the source.
  • Step outputs from before a resume are gone, so guards after an approval can't read pre-approval results; we pass everything through files instead.

Ask: an expired (and rejected) condition usable in dependsOn; autoResume on by default, or at least a CLI hint that reflects the server's setting; documented (or preserved) step outputs across resume.

4. Semantics we could only learn from the source

Each cost us a source dive; the docs didn't say:

  • several dependsOn entries are ANDed;
  • trigger.inputs expressions resolve for webhook runs only — scheduled runs get literals;
  • manual_approval.timeout is in seconds, command/shell timeout in milliseconds;
  • a nested child that suspends makes the parent step fail (we needed allowFailure); a distinct "child waiting" status would read better;
  • run.id vs the legacy workflowRunId, and which expressions are deferred to step time.

Ask: a "workflow semantics" reference page covering dependencies, conditions, supersede, suspension/resume and nesting.

5. Operational papercuts

  • No run-history retention in serve. A 5-minute timer leaves ~3 run records per tick (~35–50 MB/day for us). garbageCollection in .swamp.yaml is read by the CLI's run gc, not applied by serve; we added an external daily timer. Serve-side retention would help.
  • .swamp.yaml is rewritten on first use to add a repoId, dropping comments. In a read-only checkout that a timer resets, this dirties the tree and mints a new id after every reset. We now commit the repoId.
  • SWAMP_HOME under the home directory gets mistaken for a repo by commands run from ~.
  • swamp serve daemon enable maps systemctl reload to SIGHUP, which stops serve unless --hot-reload is set; we wrote our own unit.

6. Trust and supply chain

For a control plane that can change production servers:

  • releases are frequent (~40 builds/day) and "stable" is simply the newest; artifacts carry a sha256 but no signature;
  • a self-hosted serve still depends on the hosted account for authentication.

Ask: signed releases (cosign/minisign) and a slower, supported release channel; clarity on what a self-hosted serve needs from the hosted service at runtime, and what happens if it's unreachable.

Summary

Swamp's core — grants, token-gated approvals, recorded runs, definitions in git — fits this use well. Issues 1–3 shaped our design the most: supersede-on-schedule forced the nested workaround, the missing ref trigger forced polling, and the silent approval lifecycle left gaps we fill with our own scripts. Fixing those three would make Swamp a natural fit for approval-gated infrastructure changes.

02Bog Flow
◉OPEN○TRIAGED○IN PROGRESS○SHIPPED

Open

10/2/2026, 8:56:51 PM

No activity in this phase yet.

03Sludge Pulse
Editable. Press Enter to edit.

hammz commented 10/2/2026, 10:17:52 PM

Some great feedback here! We'll consider all of it, and I'm definitely going to look into the "superseding" and think about how to solve that problem.

hammz commented 10/5/2026, 6:26:28 PM

@kklas — I tried to reproduce item 1 (scheduled runs superseding their own pending approval) and couldn't, so I'd like a bit more detail before we decide what to change.

What I tested

A scratch repo with a filesystem datastore and this workflow, under swamp serve:

name: gated
trigger:
  schedule: "* * * * *"
jobs:
  - name: main
    steps:
      - name: gate
        task:
          type: manual_approval
          prompt: "approve?"
  • Serve on current main, four ticks: four suspended runs, none cancelled.
  • Serve restarted, one more tick: all five runs still suspended, including the four started by the previous serve process.
  • Serve built from the source of your version (20260930.003205), four ticks: four suspended runs, none cancelled.
  • swamp workflow run gated twice from the CLI: the second run cancelled the first with "Superseded ... (matching inputs)". The serve-started runs were left alone.

Why I think that is

Supersede skips any suspended run that is owned by a serve instance. Every run the scheduler starts is stamped with serve's instance id and keeps it while suspended, so the next tick doesn't cancel it. A run started by a local swamp workflow run has no instance id, so a later local run with the same inputs does supersede it. You're right that the scheduler never passes noSupersede; it just doesn't need to for runs it started itself.

Where my setup differed from yours

  • I used --auth-mode none; you run token mode with --approve-requires-explicit-grant.
  • My workflow suspends immediately: no steps before the gate and no trigger.inputs.
  • I ran from source rather than the released binary.

What would help

  1. Did you see "Superseded by new run with matching inputs" on a run that serve's scheduler started, or only on runs started locally with swamp workflow run (for example while testing in the scratch repo)?
  2. If it was a serve-scheduled run: the output of swamp workflow history get <workflow> --run <cancelled-run-id> --json (or the run's record), in particular its triggerSource, instanceId and initiatedBy, plus the same for the run that superseded it.
  3. Was anything else starting runs of that workflow against the same repo: a second serve instance, a swamp workflow run from a timer or script, or a --server call from another host?
  4. Did serve restart or reload between the run suspending and being cancelled?

If it turns out to be the local CLI path only, then the gap is that the docs don't say serve-owned runs are never superseded, and we'll fix that. If you did see it on a scheduled run, the details above should let us find the path I missed.

hammz commented 10/5/2026, 6:26:46 PM

Correction to the command in point 2 above: it takes the run id directly, with no --run flag: swamp workflow history get <cancelled-run-id> --json.

Sign in to post a ripple.