A nested structural swamp still waits on its run's lock when the lock holder is not a same-host ancestor (cross-host worker, --server into a non-ancestor serve)
Deliver workflow signals through a write-once outcome record instead of writing the run record
Forward held-lock identity over --server so a loopback nested swamp does not wait on its caller's lock
A nested structural swamp on a same-host remote worker waits on its own step's lock held by swamp serve
Fail fast when nested structural commands under parallel runs wait on each other's locks
Docs: remove followUpActions from the MethodResult reference in the extension model manual
signal_test: already-settled test fails by chance when a random UUID spells the sender name
Add a wait_for_signal workflow step that pauses a run for a JSON message
run doctor through serve interrupts a live run of another serve instance when no heartbeats are recorded
Flaky test: WalSink replayed segments are delivered before events written after replay leaves one WAL segment
Remove followUpActions: no model or extension can produce them
Workflow schedule watcher and workflow edit symlink lookup ignore the managed config workflows dir
main is red: repository_dirty_coverage_test fails for UnifiedData.collectGarbage(orphaned deferred write)
isProcessGone always reports alive on Linux under the test task permissions, failing 9 tests on main
serve: relative tls cert-file/key-file and --config resolve against the working directory, not the repository
findBySpec/findByTag lose a workflow step's output after a version is deleted, pruned or rolled back
Run tracker keeps some interrupted workflow rows forever: retention waits for markSettled, which several paths never call
run doctor reads every run record on each run to rebuild the run indexes
Serve audit: two sinks with the same name share one emitter cursor, so the second never receives events
Serve audit: a durable sink write that times out is retried while the first call is still pending, so the store can hold duplicate sequences
Flaky test: ModelResolver data accessors for a renamed model return the earlier id's record from findBySpec
data query --limit returns fewer live records than exist and reports limited: false when stale catalog rows fall inside the limit
data get: help example passes --run latest, which is not a recognised run id
workflow reject leaves sibling gates waiting_approval and dependents pending in a failed, retryable run
workflow cancel cannot cancel a suspended run whose workflow file was deleted, and cancel --all silently skips it
A job alone in its level starts after the abort, and its unstarted step is recorded as a real failure instead of settledByAbort
workflow approve racing workflow cancel loses the cancel: the run returns to suspended after cancel reported cancelled
data query: does not follow data rename forwarding that data get follows
Unit tests share fixed /tmp catalog paths, so concurrent runs corrupt them and fail every later run
data query: renamed data items are not found by their old name
data query: model lookup by id and orphan recovery parity with data get
workflow approve on a run that is not suspended prints a fatal error with a stack trace
data query: no way to target a workflow's latest run
data query: binary content lacks contentEncoding/base64 parity with data get
data query: no fallback from spec name to data instance name
data query: no definitionHash or garbageCollection fields in query results
A nested structural swamp spawned by swamp serve skips every lock serve holds, including other runs' locks
data query: confirm catalog parity with data get for custom datastores
data query: single-result --json mode for scripts migrating from data get
data get --workflow silently returns an arbitrary step's data when several steps share a data name
swamp doctor workflows misses nested and .yml extension workflows that the loader reads
auth whoami says the stored API key is no longer valid when auth.json only holds an env-key identity cache
Docs: swamp skill says cancelling a parent run cancels its nested child runs; the child is left suspended
Namespace advice in the slow-lock warning is wrong for a single repo on a local datastore
Method-run records left running by a dead owner are never reaped
Flaky: usecase_sync_characterization_data_test 'data gc (serve)' asserts markDirty order that varies under parallel load
Worker dispatch rebuilds the definition with Definition.create, so a legacy-named model cannot run on a worker
A cancelled workflow exits before its steps stop, leaving their method-run records at running
Nested swamp structural command in a shell step times out on its parent's per-model lock (SWAMP_LOCK_HOLDER_PID is stripped)
Auth gate blocks nested swamp in workflow shell steps when the credential is in SWAMP_API_KEY
swamp workflow create crashes with an [FTL] ZodError stack trace on an invalid (uppercase) workflow name
swamp model create crashes with an [FTL] ZodError stack trace on an invalid (uppercase) model name
model cancel SIGTERMs a running swamp serve process when the method run belongs to a serve-run workflow
model cancel --all on a running workflow leaves its method-run records at running, or records them failed instead of cancelled
model cancel SIGKILLs the owner 2 s after SIGTERM, before the step executor's own 3 s kill grace, so the run never records its own cancellation
workflow cancel and supersede of a suspended run leave its jobs running in the cancelled record
A force-exited workflow run stays running with a dead pid: resume, recover and run doctor --fix cannot clear it
A nested workflow that suspends on manual approval forwards its suspended event into the parent run's stream
workflow cancel SIGKILLs the run 2 s after SIGTERM, cutting off always/completed cleanup jobs mid-step
Warn when datastore setup keeps managedConfig but the new datastore has no config tier
serve: a relative grants-dir resolves against the working directory, not the repository
workflow resume cleanup mode can run an always-gated teardown before a suspended job's approved work
managedConfig: extension writes update the shared lockfile from a stale cache and overwrite peers' entries
datastore setup extension overwrites the shared config tier with the repo-local copy and drops managedConfig from .swamp.yaml
serve daemon enable writes a unit that crash-loops when serve args are invalid (e.g. missing --admins)
Concurrent embedded deno extraction races: stale temp-file cleanup deletes another process's in-flight .deno.tmp file (chmod ENOENT)
Telemetry bridge drops a parent step's invocation when a nested workflow step has the same job and step names
serve daemon: unit sets SWAMP_HOME, which relocates the config dir — token/oauth daemons can't find auth.json and crash-loop
Workflow step that times out on the model lock leaves its output and run-tracker row stuck at 'running'
An extension upgrade that hits a type collision is rolled back to no version: roll back to the prior version with a commit/rollback install handle
Telemetry bridge attributes a parent run's later method invocations to its nested child run
sensitive_data_delivery_test fails where /bin/sh is bash: the final ps is exec'd and reports its own argv
build-attestation rejects every run: evaluated workflows now carry sensitiveFormat, which the provenance check treats as an uncommitted field
DuplicateTypeError names the incoming extension as the existing claimant and reports a false ghost row when it sorts first
Install extensions by stage and swap, with a journal and crash recovery
workflow resume ignores SWAMP_MAX_CONCURRENT_STEPS for job-level concurrency
Tidy up tests and helpers from the serve suspended-run cancel
Concurrent extension installs and removals can interleave in one checkout: add a pulled-extensions lock and LockfileRepository.refresh()
deno run audit prints No description available for every advisory
Add a literal() CEL function so one string can mix swamp expressions with another service template syntax
Extension archives have no size cap: cap the compressed and decompressed size on pull and push
withGeneratorSpan does not forward return() to the wrapped generator, so its finally blocks are skipped on early exit
forEach over a direct step with a shared modelName runs every iteration with one iteration's globalArgs on first creation
Split installExtension into an unlocked prepare and a repo-changing apply
workflow cancel --server shows raw JSON errors and drops --reason
Two pulled extensions providing one type + a rebuilt catalog brick every command (I-Repo-1 at startup reconcile, even rm/--repair)
Extension catalog keeps stale rows after install migration and rm; model type describe crashes with ENOENT
Flaky test: CollectiveRefreshService fallback getAccessToken test counts calls after a fixed 200ms sleep
model validate log output misaligns multi-entry Expression paths errors
Retry refusals print unquoted forEach template names in --from hints, which bash rejects
workflow run log output: a forEach step with a templated name labels its iterations with the raw expression, e.g. deploy-${{ self.env }}[0]
Docker image runs swamp as PID 1 without an init, so orphaned step processes are never reaped
swamp.cli.bootstrap, configure_extension_loaders and teardown spans each start their own trace
Extension restore leaves a parent installed and its dependency permanently missing when the dependency install fails
serve: vault edit over --server opens an editor on the server host and ignores the managedConfig vaults dir
A suspended run started by swamp serve cannot be cancelled
command/shell: aborting a step kills only sh, leaving the command's child processes orphaned
Extension restore drops the lockfile entry's channel and leaks the parent's expectedChecksum into dependency installs
managedConfig: serve extension handlers disagree on the pulled-extensions root, and lockfile files[] are repo-relative
swamp.model.method spans are parented to the workflow run span instead of their swamp.workflow.step span
Platform certificate load failure surfaces as an uncaught error on network commands
serve: server-login can't be traced: no request spans, and gcs-datastore push is unparented with an 81s gap in child spans
workflow resume keeps the original run's process identity (run record pid and instanceId)
workflow run --json: guarded-skip lines land on stdout or stderr depending on timing
workflow run and model method run reject --ca-cert, but their TLS error tells users to pass it
After a cancel or --timeout, jobs that share a level are abandoned mid-flight: their steps stay running and their own cleanup steps never run
SWAMP_CA_CERT overrides the --ca-cert flag, contrary to the documented flag-wins precedence
workflow resume --timeout records cancel_reason 'The signal has been aborted' instead of a timeout reason
InProcessExecutor mutates the process-wide TRACEPARENT, so concurrent in-process executions clobber and leak trace context
workflow resume has no job-level cleanup mode: after --timeout or a cancel, always/completed cleanup jobs start with the aborted signal and fail at once
workflow cancel on a live local run overwrites the owning process's final run record with a stale snapshot
After a cancel or --timeout, a never-started dependency stays pending, so failed- and completed-gated cleanup steps are skipped
Direct-type forEach fan-out waits about 1s on the auto-definition lock when its definition is first created (missed by swamp-club#2126)
managedConfig: fix startup config-base resolution and extension list (prerequisites for swamp-club#2429)
workflow run log output: a skipped plain step's line names only the job, not the step
A forEach step's dependsOn condition is never evaluated, so its iterations run when the condition is false
managedConfig: the datastore extension is only found through a lockfile, so the config base is guessed (part 3c of swamp-club#2483)
model output data/logs/search can't find outputs of user extension model types locally
Relationships
#2646 Extension restore leaves a parent installed and its dependency permanently missing when the dependency install fails
Opened by hammz · 9/28/2026· Shipped 9/29/2026
Summary
During a lockfile restore (swamp extension install, repo upgrade, serve extension.install), installExtension writes the parent's lockfile entry and extracts its files before installing dependencies (src/libswamp/extensions/pull.ts, the writeEntry call before the dependency loop). If a dependency that is missing from the lockfile then fails to install (network error, not found, integrity error), the command reports the parent as failed but the parent's files are back on disk.
On the next extension install, the parent is classified up_to_date because its lockfile entry and files are present. Restore never looks at a parent's declared dependencies once the parent is up to date, so the missing dependency is never retried. The half-installed state is permanent and silent: later runs report failed 0.
Found while triaging swamp-club#2639. There the trigger was the parent's expectedChecksum leaking into dependency installs, which #2639 fixes. Any other dependency failure still produces the same end state, and so does a hand-edited lockfile with a dependency entry removed.
Reproduction
In a scratch repo, pull @kneel/eda-factory with --channel beta. It depends on @kneel/kicad and @kneel/eda-review-lenses. Remove @kneel/kicad from the lockfile, delete the files of both the parent and kicad, and make kicad's install fail (before #2639 the checksum leak does this). Run swamp extension install: the parent is failed. Run it again: the result is installed 0, upToDate 2, failed 0, and kicad stays absent.
Possible directions
- Roll back the parent's created files and lockfile entry when a dependency install throws, so the next restore retries the whole tree.
- Make the restore's up-to-date check treat a parent as needing install when a declared dependency has no lockfile entry. This also covers hand-edited lockfiles.
Shipped
Click a lifecycle step above to view its details.