A nested structural swamp still waits on its run's lock when the lock holder is not a same-host ancestor (cross-host worker, --server into a non-ancestor serve)
A nested structural swamp that outlives the run that started it keeps skipping that run's lock while the holder writes
Show signal waits to remote clients and the dashboard
Settle expired signal waits from swamp serve
Resume signalled workflow runs automatically under swamp serve
Deliver workflow signals through a write-once outcome record instead of writing the run record
Forward held-lock identity over --server so a loopback nested swamp does not wait on its caller's lock
A nested structural swamp on a same-host remote worker waits on its own step's lock held by swamp serve
Fail fast when nested structural commands under parallel runs wait on each other's locks
Docs: remove followUpActions from the MethodResult reference in the extension model manual
signal_test: already-settled test fails by chance when a random UUID spells the sender name
Add a wait_for_signal workflow step that pauses a run for a JSON message
Deliver workflow signals through swamp serve
run doctor through serve interrupts a live run of another serve instance when no heartbeats are recorded
Flaky test: WalSink replayed segments are delivered before events written after replay leaves one WAL segment
Remove followUpActions: no model or extension can produce them
Workflow schedule watcher and workflow edit symlink lookup ignore the managed config workflows dir
main is red: repository_dirty_coverage_test fails for UnifiedData.collectGarbage(orphaned deferred write)
isProcessGone always reports alive on Linux under the test task permissions, failing 9 tests on main
Flaky test: WalSink replayed segments are delivered before events written after replay
serve: relative tls cert-file/key-file and --config resolve against the working directory, not the repository
findBySpec/findByTag lose a workflow step's output after a version is deleted, pruned or rolled back
serve HA: worker enrollment on a replica that lacks the enrollment token's definition may create a second definition (unverified)
serve HA: pull a token's definition on an auth miss so a peer accepts a new token at once
Run tracker keeps some interrupted workflow rows forever: retention waits for markSettled, which several paths never call
run doctor reads every run record on each run to rebuild the run indexes
Serve audit: two sinks with the same name share one emitter cursor, so the second never receives events
Serve audit: a durable sink write that times out is retried while the first call is still pending, so the store can hold duplicate sequences
Flaky test: ModelResolver data accessors for a renamed model return the earlier id's record from findBySpec
data query --limit returns fewer live records than exist and reports limited: false when stale catalog rows fall inside the limit
data get: help example passes --run latest, which is not a recognised run id
workflow reject leaves sibling gates waiting_approval and dependents pending in a failed, retryable run
Show expired approval gates as expired, with a Cancel action in the dashboard
digitalocean codegen leaves an orphaned security_secret.ts model file after its endpoint left the spec
workflow cancel cannot cancel a suspended run whose workflow file was deleted, and cancel --all silently skips it
A job alone in its level starts after the abort, and its unstarted step is recorded as a real failure instead of settledByAbort
workflow approve racing workflow cancel loses the cancel: the run returns to suspended after cancel reported cancelled
data query: does not follow data rename forwarding that data get follows
Unit tests share fixed /tmp catalog paths, so concurrent runs corrupt them and fail every later run
data query: renamed data items are not found by their old name
data query: model lookup by id and orphan recovery parity with data get
workflow approve on a run that is not suspended prints a fatal error with a stack trace
data query: no way to target a workflow's latest run
data query: binary content lacks contentEncoding/base64 parity with data get
data query: no fallback from spec name to data instance name
data query: no definitionHash or garbageCollection fields in query results
A nested structural swamp spawned by swamp serve skips every lock serve holds, including other runs' locks
data query: confirm catalog parity with data get for custom datastores
data query: single-result --json mode for scripts migrating from data get
Docs: document data query --single and fix the stale --limit default in reference/data.md
data get --workflow silently returns an arbitrary step's data when several steps share a data name
swamp doctor workflows misses nested and .yml extension workflows that the loader reads
auth whoami says the stored API key is no longer valid when auth.json only holds an env-key identity cache
Docs: swamp skill says cancelling a parent run cancels its nested child runs; the child is left suspended
Namespace advice in the slow-lock warning is wrong for a single repo on a local datastore
Method-run records left running by a dead owner are never reaped
Flaky: usecase_sync_characterization_data_test 'data gc (serve)' asserts markDirty order that varies under parallel load
Worker dispatch rebuilds the definition with Definition.create, so a legacy-named model cannot run on a worker
A cancelled workflow exits before its steps stop, leaving their method-run records at running
Nested swamp structural command in a shell step times out on its parent's per-model lock (SWAMP_LOCK_HOLDER_PID is stripped)
Auth gate blocks nested swamp in workflow shell steps when the credential is in SWAMP_API_KEY
swamp workflow create crashes with an [FTL] ZodError stack trace on an invalid (uppercase) workflow name
swamp model create crashes with an [FTL] ZodError stack trace on an invalid (uppercase) model name
model cancel SIGTERMs a running swamp serve process when the method run belongs to a serve-run workflow
Follow-up action retries ignore the run's abort signal and can outlast a cancel
model cancel --all on a running workflow leaves its method-run records at running, or records them failed instead of cancelled
model cancel SIGKILLs the owner 2 s after SIGTERM, before the step executor's own 3 s kill grace, so the run never records its own cancellation
workflow cancel and supersede of a suspended run leave its jobs running in the cancelled record
Tell timeouts apart from cancels in method-run records, and review the hidden 30 s fallback timer for step-called model methods
A force-exited workflow run stays running with a dead pid: resume, recover and run doctor --fix cannot clear it
A nested workflow that suspends on manual approval forwards its suspended event into the parent run's stream
workflow cancel SIGKILLs the run 2 s after SIGTERM, cutting off always/completed cleanup jobs mid-step
Warn when datastore setup keeps managedConfig but the new datastore has no config tier
Docs: run doctor --fix and workflow recover settle runs left running by a force-exited process
serve: a relative grants-dir resolves against the working directory, not the repository
workflow resume cleanup mode can run an always-gated teardown before a suspended job's approved work
Docs: serve-flags.md should state that relative grants-dir and grants-file resolve against the repository
managedConfig: extension writes update the shared lockfile from a stale cache and overwrite peers' entries
managedConfig: model/vault create and other bulk-push commands re-upload a stale extension lockfile and erase peers' entries
Cancel, reject and supersede of a parent run should settle its suspended nested child runs
datastore setup extension overwrites the shared config tier with the repo-local copy and drops managedConfig from .swamp.yaml
datastore setup can overwrite an existing remote config tier when it moves an in-repo tier into an extension datastore
serve daemon enable writes a unit that crash-loops when serve args are invalid (e.g. missing --admins)
Concurrent embedded deno extraction races: stale temp-file cleanup deletes another process's in-flight .deno.tmp file (chmod ENOENT)
Docs: manual approval gates inside nested workflows (swamp-club#2736)
Telemetry bridge drops a parent step's invocation when a nested workflow step has the same job and step names
serve daemon: unit sets SWAMP_HOME, which relocates the config dir — token/oauth daemons can't find auth.json and crash-loop
Workflow step that times out on the model lock leaves its output and run-tracker row stuck at 'running'
Docs: serve daemon enable pins SWAMP_CONFIG_DIR; document the SWAMP_CONFIG_DIR override
An extension upgrade that hits a type collision is rolled back to no version: roll back to the prior version with a commit/rollback install handle
Telemetry bridge attributes a parent run's later method invocations to its nested child run
sensitive_data_delivery_test fails where /bin/sh is bash: the final ps is exec'd and reports its own argv
build-attestation rejects every run: evaluated workflows now carry sensitiveFormat, which the provenance check treats as an uncommitted field
DuplicateTypeError names the incoming extension as the existing claimant and reports a false ghost row when it sorts first
Install extensions by stage and swap, with a journal and crash recovery
workflow resume ignores SWAMP_MAX_CONCURRENT_STEPS for job-level concurrency
Tidy up tests and helpers from the serve suspended-run cancel
Concurrent extension installs and removals can interleave in one checkout: add a pulled-extensions lock and LockfileRepository.refresh()
deno run audit prints No description available for every advisory
Add a literal() CEL function so one string can mix swamp expressions with another service template syntax
Extension archives have no size cap: cap the compressed and decompressed size on pull and push
withGeneratorSpan does not forward return() to the wrapped generator, so its finally blocks are skipped on early exit
forEach over a direct step with a shared modelName runs every iteration with one iteration's globalArgs on first creation
Split installExtension into an unlocked prepare and a repo-changing apply
workflow cancel --server shows raw JSON errors and drops --reason
Two pulled extensions providing one type + a rebuilt catalog brick every command (I-Repo-1 at startup reconcile, even rm/--repair)
Extension catalog keeps stale rows after install migration and rm; model type describe crashes with ENOENT
Flaky test: CollectiveRefreshService fallback getAccessToken test counts calls after a fixed 200ms sleep
Docs: worker-fleet guides start swamp serve without --admins and mint tokens with a removed command
model validate log output misaligns multi-entry Expression paths errors
Retry refusals print unquoted forEach template names in --from hints, which bash rejects
workflow run log output: a forEach step with a templated name labels its iterations with the raw expression, e.g. deploy-${{ self.env }}[0]
Docs: update the model validate Expression paths example in cel-expressions.md
Docker image runs swamp as PID 1 without an init, so orphaned step processes are never reaped
swamp.cli.bootstrap, configure_extension_loaders and teardown spans each start their own trace
Relationships
#2735 Telemetry bridge attributes a parent run's later method invocations to its nested child run
Opened by hammz · 9/29/2026· Shipped 9/30/2026
Summary
The libswamp telemetry bridge (src/libswamp/workflows/telemetry_bridge.ts, observe) sets workflowName and runId from every started event in the stream. A workflow step that runs a nested workflow forwards the child's started event into the parent's stream. So after a nested step starts, every method invocation the bridge records, including ones from the parent's later steps, is attributed to the child's workflow name and run id.
Context
Found while fixing swamp-club#2470. Nested started events now carry parentRunId, so the bridge can tell them apart. The fix needs a decision on attribution: record child steps under the child run (the bridge would need a per-run stack), or record everything under the top-level run.
Steps to reproduce
- Workflow child with one model_method step. Workflow parent with a workflow step running child, followed by a model_method step.
- Run parent with telemetry enabled.
- The parent's second step invocation is reported with the child's workflowName and runId.
Shipped
Click a lifecycle step above to view its details.