A nested structural swamp still waits on its run's lock when the lock holder is not a same-host ancestor (cross-host worker, --server into a non-ancestor serve)
A nested structural swamp that outlives the run that started it keeps skipping that run's lock while the holder writes
Show signal waits to remote clients and the dashboard
Settle expired signal waits from swamp serve
Resume signalled workflow runs automatically under swamp serve
Deliver workflow signals through a write-once outcome record instead of writing the run record
Forward held-lock identity over --server so a loopback nested swamp does not wait on its caller's lock
A nested structural swamp on a same-host remote worker waits on its own step's lock held by swamp serve
Fail fast when nested structural commands under parallel runs wait on each other's locks
Docs: remove followUpActions from the MethodResult reference in the extension model manual
signal_test: already-settled test fails by chance when a random UUID spells the sender name
Add a wait_for_signal workflow step that pauses a run for a JSON message
Deliver workflow signals through swamp serve
run doctor through serve interrupts a live run of another serve instance when no heartbeats are recorded
Flaky test: WalSink replayed segments are delivered before events written after replay leaves one WAL segment
Remove followUpActions: no model or extension can produce them
Workflow schedule watcher and workflow edit symlink lookup ignore the managed config workflows dir
main is red: repository_dirty_coverage_test fails for UnifiedData.collectGarbage(orphaned deferred write)
isProcessGone always reports alive on Linux under the test task permissions, failing 9 tests on main
Flaky test: WalSink replayed segments are delivered before events written after replay
serve: relative tls cert-file/key-file and --config resolve against the working directory, not the repository
findBySpec/findByTag lose a workflow step's output after a version is deleted, pruned or rolled back
serve HA: worker enrollment on a replica that lacks the enrollment token's definition may create a second definition (unverified)
serve HA: pull a token's definition on an auth miss so a peer accepts a new token at once
Run tracker keeps some interrupted workflow rows forever: retention waits for markSettled, which several paths never call
run doctor reads every run record on each run to rebuild the run indexes
Serve audit: two sinks with the same name share one emitter cursor, so the second never receives events
Serve audit: a durable sink write that times out is retried while the first call is still pending, so the store can hold duplicate sequences
Flaky test: ModelResolver data accessors for a renamed model return the earlier id's record from findBySpec
data query --limit returns fewer live records than exist and reports limited: false when stale catalog rows fall inside the limit
data get: help example passes --run latest, which is not a recognised run id
workflow reject leaves sibling gates waiting_approval and dependents pending in a failed, retryable run
Show expired approval gates as expired, with a Cancel action in the dashboard
digitalocean codegen leaves an orphaned security_secret.ts model file after its endpoint left the spec
workflow cancel cannot cancel a suspended run whose workflow file was deleted, and cancel --all silently skips it
A job alone in its level starts after the abort, and its unstarted step is recorded as a real failure instead of settledByAbort
workflow approve racing workflow cancel loses the cancel: the run returns to suspended after cancel reported cancelled
data query: does not follow data rename forwarding that data get follows
Unit tests share fixed /tmp catalog paths, so concurrent runs corrupt them and fail every later run
data query: renamed data items are not found by their old name
data query: model lookup by id and orphan recovery parity with data get
workflow approve on a run that is not suspended prints a fatal error with a stack trace
data query: no way to target a workflow's latest run
data query: binary content lacks contentEncoding/base64 parity with data get
data query: no fallback from spec name to data instance name
data query: no definitionHash or garbageCollection fields in query results
A nested structural swamp spawned by swamp serve skips every lock serve holds, including other runs' locks
data query: confirm catalog parity with data get for custom datastores
data query: single-result --json mode for scripts migrating from data get
Docs: document data query --single and fix the stale --limit default in reference/data.md
data get --workflow silently returns an arbitrary step's data when several steps share a data name
swamp doctor workflows misses nested and .yml extension workflows that the loader reads
auth whoami says the stored API key is no longer valid when auth.json only holds an env-key identity cache
Docs: swamp skill says cancelling a parent run cancels its nested child runs; the child is left suspended
Namespace advice in the slow-lock warning is wrong for a single repo on a local datastore
Method-run records left running by a dead owner are never reaped
Flaky: usecase_sync_characterization_data_test 'data gc (serve)' asserts markDirty order that varies under parallel load
Worker dispatch rebuilds the definition with Definition.create, so a legacy-named model cannot run on a worker
A cancelled workflow exits before its steps stop, leaving their method-run records at running
Nested swamp structural command in a shell step times out on its parent's per-model lock (SWAMP_LOCK_HOLDER_PID is stripped)
Auth gate blocks nested swamp in workflow shell steps when the credential is in SWAMP_API_KEY
swamp workflow create crashes with an [FTL] ZodError stack trace on an invalid (uppercase) workflow name
swamp model create crashes with an [FTL] ZodError stack trace on an invalid (uppercase) model name
model cancel SIGTERMs a running swamp serve process when the method run belongs to a serve-run workflow
Follow-up action retries ignore the run's abort signal and can outlast a cancel
model cancel --all on a running workflow leaves its method-run records at running, or records them failed instead of cancelled
model cancel SIGKILLs the owner 2 s after SIGTERM, before the step executor's own 3 s kill grace, so the run never records its own cancellation
workflow cancel and supersede of a suspended run leave its jobs running in the cancelled record
Tell timeouts apart from cancels in method-run records, and review the hidden 30 s fallback timer for step-called model methods
A force-exited workflow run stays running with a dead pid: resume, recover and run doctor --fix cannot clear it
A nested workflow that suspends on manual approval forwards its suspended event into the parent run's stream
workflow cancel SIGKILLs the run 2 s after SIGTERM, cutting off always/completed cleanup jobs mid-step
Warn when datastore setup keeps managedConfig but the new datastore has no config tier
Docs: run doctor --fix and workflow recover settle runs left running by a force-exited process
serve: a relative grants-dir resolves against the working directory, not the repository
workflow resume cleanup mode can run an always-gated teardown before a suspended job's approved work
Docs: serve-flags.md should state that relative grants-dir and grants-file resolve against the repository
managedConfig: extension writes update the shared lockfile from a stale cache and overwrite peers' entries
managedConfig: model/vault create and other bulk-push commands re-upload a stale extension lockfile and erase peers' entries
Cancel, reject and supersede of a parent run should settle its suspended nested child runs
datastore setup extension overwrites the shared config tier with the repo-local copy and drops managedConfig from .swamp.yaml
datastore setup can overwrite an existing remote config tier when it moves an in-repo tier into an extension datastore
serve daemon enable writes a unit that crash-loops when serve args are invalid (e.g. missing --admins)
Concurrent embedded deno extraction races: stale temp-file cleanup deletes another process's in-flight .deno.tmp file (chmod ENOENT)
Docs: manual approval gates inside nested workflows (swamp-club#2736)
Telemetry bridge drops a parent step's invocation when a nested workflow step has the same job and step names
serve daemon: unit sets SWAMP_HOME, which relocates the config dir — token/oauth daemons can't find auth.json and crash-loop
Workflow step that times out on the model lock leaves its output and run-tracker row stuck at 'running'
Docs: serve daemon enable pins SWAMP_CONFIG_DIR; document the SWAMP_CONFIG_DIR override
An extension upgrade that hits a type collision is rolled back to no version: roll back to the prior version with a commit/rollback install handle
Telemetry bridge attributes a parent run's later method invocations to its nested child run
sensitive_data_delivery_test fails where /bin/sh is bash: the final ps is exec'd and reports its own argv
build-attestation rejects every run: evaluated workflows now carry sensitiveFormat, which the provenance check treats as an uncommitted field
DuplicateTypeError names the incoming extension as the existing claimant and reports a false ghost row when it sorts first
Install extensions by stage and swap, with a journal and crash recovery
workflow resume ignores SWAMP_MAX_CONCURRENT_STEPS for job-level concurrency
Tidy up tests and helpers from the serve suspended-run cancel
Concurrent extension installs and removals can interleave in one checkout: add a pulled-extensions lock and LockfileRepository.refresh()
deno run audit prints No description available for every advisory
Add a literal() CEL function so one string can mix swamp expressions with another service template syntax
Extension archives have no size cap: cap the compressed and decompressed size on pull and push
withGeneratorSpan does not forward return() to the wrapped generator, so its finally blocks are skipped on early exit
forEach over a direct step with a shared modelName runs every iteration with one iteration's globalArgs on first creation
Split installExtension into an unlocked prepare and a repo-changing apply
workflow cancel --server shows raw JSON errors and drops --reason
Two pulled extensions providing one type + a rebuilt catalog brick every command (I-Repo-1 at startup reconcile, even rm/--repair)
Extension catalog keeps stale rows after install migration and rm; model type describe crashes with ENOENT
Flaky test: CollectiveRefreshService fallback getAccessToken test counts calls after a fixed 200ms sleep
Docs: worker-fleet guides start swamp serve without --admins and mint tokens with a removed command
model validate log output misaligns multi-entry Expression paths errors
Retry refusals print unquoted forEach template names in --from hints, which bash rejects
workflow run log output: a forEach step with a templated name labels its iterations with the raw expression, e.g. deploy-${{ self.env }}[0]
Docs: update the model validate Expression paths example in cel-expressions.md
Docker image runs swamp as PID 1 without an init, so orphaned step processes are never reaped
swamp.cli.bootstrap, configure_extension_loaders and teardown spans each start their own trace
Docs: run swamp under an init in the worker-fleet guides
Extension restore leaves a parent installed and its dependency permanently missing when the dependency install fails
serve: vault edit over --server opens an editor on the server host and ignores the managedConfig vaults dir
A suspended run started by swamp serve cannot be cancelled
command/shell: aborting a step kills only sh, leaving the command's child processes orphaned
Extension restore drops the lockfile entry's channel and leaks the parent's expectedChecksum into dependency installs
managedConfig: serve extension handlers disagree on the pulled-extensions root, and lockfile files[] are repo-relative
swamp.model.method spans are parented to the workflow run span instead of their swamp.workflow.step span
Platform certificate load failure surfaces as an uncaught error on network commands
serve: server-login can't be traced: no request spans, and gcs-datastore push is unparented with an 81s gap in child spans
workflow resume keeps the original run's process identity (run record pid and instanceId)
managedConfig: store pulled extension archives in the config tier and converge each checkout to the shared lockfile
workflow run --json: guarded-skip lines land on stdout or stderr depending on timing
workflow run and model method run reject --ca-cert, but their TLS error tells users to pass it
After a cancel or --timeout, jobs that share a level are abandoned mid-flight: their steps stay running and their own cleanup steps never run
SWAMP_CA_CERT overrides the --ca-cert flag, contrary to the documented flag-wins precedence
workflow resume --timeout records cancel_reason 'The signal has been aborted' instead of a timeout reason
InProcessExecutor mutates the process-wide TRACEPARENT, so concurrent in-process executions clobber and leak trace context
workflow resume has no job-level cleanup mode: after --timeout or a cancel, always/completed cleanup jobs start with the aborted signal and fail at once
workflow cancel on a live local run overwrites the owning process's final run record with a stale snapshot
s3-datastore: SWAMP_S3_REQUEST_TIMEOUT_MS does not appear to apply to ListObjectsV2 during pull
s3-datastore: S3Lock.acquire overshoots maxWaitMs by up to a full backoff interval
s3-datastore: a model-scoped pull still lists, walks and indexes the whole namespace
After a cancel or --timeout, a never-started dependency stays pending, so failed- and completed-gated cleanup steps are skipped
Direct-type forEach fan-out waits about 1s on the auto-definition lock when its definition is first created (missed by swamp-club#2126)
managedConfig: fix startup config-base resolution and extension list (prerequisites for swamp-club#2429)
workflow run log output: a skipped plain step's line names only the job, not the step
Docs: manual examples show the old workflow skip line without the step name
A forEach step's dependsOn condition is never evaluated, so its iterations run when the condition is false
managedConfig: the datastore extension is only found through a lockfile, so the config base is guessed (part 3c of swamp-club#2483)
model output data/logs/search can't find outputs of user extension model types locally
serve: ServerTokenGcService is never started, so expired server tokens accumulate forever
workflow resume --input key=value does not coerce the value to the declared input type (workflow run does)
Direct-execution workflow steps validate evaluated globalArgs, so the CEL concatenation workaround for colliding template names fails
workflow resume of a suspended run crashes mid-run when the workflow was edited during the approval window
Install rollback recursively deletes pre-existing skill dirs that pull merged into
Foreign CEL like GitHub Actions ${{ github.sha }} in a global argument fails validate and throws when the method reads it
model validate blames foreign {{ ... }} text for an unclosed ${{ expression
workflow resume --from: a step moved or renamed since the run failed is skipped and the run reported succeeded, or crashes with 'Step run not found'
managedConfig: extension lockfile writes lose updates, and auto-resolved installs never reach the shared lockfile
workflow resume prints unexpected errors as a fatal stack trace, and --json returns them unclassified with the stack
serve: remove the remaining no-path markDirty() calls that force full-cache datastore pushes
workflow resume: Ctrl-C leaves the run 'running' with nothing driving it, and it can no longer be resumed
workflow recover crashes on every invocation: 'Path must be a string'
workflow resume: retry the failed steps of a failed run without --from
Remote indicator prints the full --server URL, including a ?token= query credential, to stderr
serve: vault annotate --label over --server is rejected (CLI sends labels as an object, protocol expects string[])
workflow recover: a changed definition leaves the interrupted run with no way to recover or resume
serve: the device-token poll that returns the token takes 88s, making server-login to ops.swamp-club.com take ~90s
datastore sync: repositories mark paths dirty before writing, so a concurrent ungated push can drop the write
gcs-datastore: push spends almost all its time in an untraced gap between index read and first upload (81s in the #2408 login, up to 836s)
serve: after lazy hydration, scoped poller pulls never take the fast path, and the login mint waits behind them on the sync gate
datastore extensions: full pushes re-hash every pulled file on every run because index mtimes never match pulled copies
steps.<name>.outputs is never populated for model_method steps
auth server-login hangs after device-flow authorization is approved in browser; never completes or times out
Let a method inside serve start a workflow run: context.runWorkflow
s3-datastore: every scoped push on a fresh datastore warns 'Index merge failed, fast path stays unarmed' (missing .datastore-index.json) and the fast path never arms
Flaky integration tests: cel_data_access and webhook_payload_inputs race on SWAMP_TEST_2172_AUTHORED env var
serve rejects doctor.datastores, cluster.instances and serve.config requests: no zod schema
s3/gcs datastore: no safe recovery when _meta.json v2 lists a shard missing from the bucket
gcs-datastore: a failed lock read is reported as "unlocked" — #2298 in the GCS backend
Scheduled runs that suspend on an approval gate are logged as failures
swamp run history reports completed for failed workflow runs
s3/gcs-datastore: clean up data left behind by deletes that silently no-opped before the swamp-club#2249 fix
Every webhook lifecycle line is logged twice by swamp serve
gcs-datastore: verify generation-precondition support at setup and in doctor; an emulator or proxy that ignores ifGenerationMatch makes locking silently unsafe
agent-constraints: planning-conventions CI fan-out checklist points at jobs that no longer live in ci.yml
Property test for deferred bindings fails intermittently on a __proto__ key
serve: poller pull can race a handler's delete-then-push and resurrect the deleted item
s3-datastore: pullChanged ignores subdirs while advertising configRefresh, so serve pollers re-download the whole datastore
verify-build fails on the codegen-verify guard: 'No such key: stdout'
s3-datastore: data 'latest' pointer objects are never indexed, so deleting a data item leaves an orphaned latest object in S3
Server-side workflow resume fails on direct-type model steps
Data queries read unrelated artifact bodies before metadata filters reject them
data.latest() cannot see ephemeral data once the light expression context is used (breaks verify-reviews and verify-skills guards)
Expression evaluation scans all model definitions without model or file references
Repeated targeted lookups reparse unrelated model and workflow definitions
Workflow history get/logs enumerate all run contents for a full run UUID
Brief local model-lock contention incurs a roughly one-second initial retry delay
Cancel endpoint reports status: cancelled for runs it did not actually cancel
serve: cron ticks leak pending-run records, replaying every past tick on restart
Ghosts on the leaderboard 500
Relationships
#2246 s3-datastore: pullChanged ignores subdirs while advertising configRefresh, so serve pollers re-download the whole datastore
Opened by hammz · 9/17/2026· Shipped 9/17/2026
Summary
@swamp/s3-datastore (2026.09.10.2) sets configRefresh: true in its capabilities, but pullChanged never reads options.subdirs. The word subdirs does not appear anywhere in datastores/_lib/s3_cache_sync.ts. Every scoped pull therefore compares the whole remote index with the local cache and downloads any file that is missing locally.
Why it matters
swamp serve runs three pollers every 30s, and each one asks for a scoped pull:
- config poller:
subdirs: ["config"] - access data poller
- runtime data poller:
subdirs: ["data"]
Because the scope is ignored, each poll walks the entire datastore. A file removed from serve's local cache comes back from S3 on the next poll, whichever poller runs first. This made the deleted-data resurrection in swamp-club#2240 happen within about 4s instead of on the next runtime-data poll. The missing push on delete is being fixed in core under #2240. This issue covers only the extension side.
Reproduction (local MinIO, swamp 20260917.034341.0-sha.81c00b86)
- Set up the S3 datastore against MinIO (
endpoint,forcePathStyle: true), withmanagedConfig: trueandhydrationStrategyfull. - Start
swamp serve --auth-mode none --log-level debug. - Write a data item through serve, then remove its directory from serve's local cache (
swamp data delete ... --serverdoes this today without pushing). - Within seconds serve logs
Config poller: 3 file(s) updated, invalidating catalogs. The three restored files are underdata/, notconfig/.
Expected
When subdirs is given, pullChanged (including the fast path) only considers index entries under those datastore subdirectories. Alternatively, the extension should stop advertising configRefresh until it honors the scope.
Upstream repository: https://github.com/systeminit/swamp-extensions
Environment
- Extension:
@swamp/s3-datastore@2026.09.10.2 - swamp:
20260917.034341.0-sha.81c00b86 - OS:
darwin(aarch64) - Deno:
2.9.6 - Shell:
/bin/zsh
Shipped
Click a lifecycle step above to view its details.