Relationships
#2612 managedConfig: store pulled extension archives in the config tier and converge each checkout to the shared lockfile
Opened by hammz · 9/28/2026
Summary
Under managedConfig, the design says the datastore config tier is the single source of truth for configuration (design/enablers/datastores.md, "Managed Config Deployment Architecture"). Pulled extension sources are not part of it today: they live in each checkout's in-repo pulled root, and every pod has to run swamp extension install against the registry to restore them. A peer's extension pull, update or rm changes only the shared lockfile, and nothing makes other checkouts or serve instances match it.
Split out of swamp-club#2429, which ships the prerequisites: the ConfigPoller reloads when the lockfile hash changes, and extension writes push per-path marks instead of a bulk mark. This issue carries the design that went through six adversarial review rounds on #2429 (plan v6).
Design: registry archives in the tier, and per-checkout convergence
The tier stores the registry-verified archive of every entry in the tier lockfile. Each checkout's sources stay in the in-repo pulled root as a disposable local copy, which is brought back in line with the shared lockfile.
- Tier archive store. Archives live at
<config base>/pulled-archives/<name>/<sha256>.tar.gz. They are content-addressed and immutable. The name must match SCOPED_NAME_PATTERN and the checksum must be 64 hex characters. Paths are contained by realpath. Archives are size-capped at no less than the registry publish limit, and the decompressed-size cap covers both decompression passes. Every read checks the bytes against the lockfile entry's checksum. Reads and writes are split into separate source and sink ports. - installExtension.
- An archive-source hit is keyed on the entry checksum, which also covers constraint pins. On a hit the install skips getExtension, downloadArchive and getChecksum and runs offline. The extraction guards and the safety analysis run unchanged.
- A record policy (always, if-missing-checksum, never) controls writeEntry, the archive sink, dependency recursion and orphan pruning.
- Archives are stored only when the registry verified them, and only after the lockfile entry is written.
- Stage and swap: nothing on disk changes until the archive has been verified, extracted and safety-checked. The swap is two-phase, under
<pulledRoot>/.swamp-staging/<uuid>/, with full reverse-order rollback and manifest.yaml swapped last. Nested entry roots are copied in place. files[] records final paths. - createInstallContext passes the entry's channel. Today a restore drops it (create_extension_install_deps.ts:92-103, pull.ts:1372).
- extension install restores from tier archives and falls back to the registry. It backfills a missing archive for every tier entry, using the exact version from the on-disk manifest for constraint pins. Backfill failures are recorded until the entry changes. The output gains
sourceandbackfilledfields. - update and rm delete the superseded archive after the lockfile write, marking the path for push.
- Convergence is extensionInstall in converge mode with record set to never, and it only reads archives.
- It is gated by a validated per-checkout ConvergenceState (
.swamp/pulled-extensions-state.json: config base, lockfile hash, pending entries, backfill failures, last-verified checksums). - It restores missing or version-skewed entries only. Same-version drift (local edits) gets a warning, not an overwrite.
- It never writes the shared lockfile, contacts the registry, touches skills, or deletes sources. Delisted non-datastore extensions are inert because loaders enumerate lockfile names; datastore-kind extensions keep loading from the on-disk scan, as they do today.
- It does nothing when the lockfile is absent or unparseable. A changed config base triggers a full pass.
- The tier entry wins over transitional duplicates.
- Concurrent calls share one in-flight run. A try-lock is taken at command and service boundaries.
- Failures are best-effort warnings on stderr.
- It is gated by a validated per-checkout ConvergenceState (
- Wiring.
- CLI: convergence runs inside requireInitializedRepo, requireInitializedRepoUnlocked and requireInitializedRepoReadOnly, before the pulled workflow dirs are enumerated. For the locked variant that is after the coordinator pull. A fallback hook in withLoaderContext covers commands without a repo context. The missing-files check runs after convergence. A
skipConvergenceoption exists for serve's early context. - Serve: convergence runs in pullManagedConfigAtBoot, between the pull and the workflow-dir refresh, and also on the no-syncService boot path. The ConfigPoller handler runs converge, then the workflow refresh, then performServeReload.
- CLI: convergence runs inside requireInitializedRepo, requireInitializedRepoUnlocked and requireInitializedRepoReadOnly, before the pulled workflow dirs are enumerated. For the locked variant that is after the coordinator pull. A fallback hook in withLoaderContext covers commands without a repo context. The missing-files check runs after convergence. A
- Docs. The same PR updates datastores.md, serve.md, extensions.md (including the Integrity Verification and Integrity Anchor sections, Trusted Collectives, and the wrong sha256- example), architecture.md and remote-execution.md.
What this guarantees
After any change to the tier lockfile, each checkout restores missing or version-skewed pulled sources from checksum-verified archives. In the CLI that happens on the next repo command; in serve, at boot or on the next poll. A pod needs only the bucket, except for bootstrap extensions (those shipping datastores/*.ts), which still come from the registry. The lockfile format is unchanged and old binaries keep working, so rolling upgrades need no coordination.
Trust model
A tier writer can make every checkout run code, just as tier model definitions can today. This is consistent with #2489 (the lockfile and tier are trusted input). Archives enter the tier only after registry verification.
Known warnings from the last review (resolve during implementation)
- The stale
.swamp-stagingsweep needs an age threshold and crash recovery: rename an orphanedold/back when its live counterpart is missing. Otherwise have the auto-resolver take the boundary lock, because it calls installExtension without it (auto_resolver_adapters.ts:308). - The swap rollback must undo every completed step in reverse order, with manifest.yaml swapped last.
- The first-run baseline for constraint pins: record the entry checksum as verified only when the on-disk digest equals filesChecksum; otherwise warn as drift.
- The whole-dir swap drops files that are not in the archive, which today's merge-copy keeps. Document this and add a test.
Testing
- Unit tests covering each component.
- An in-process integration test with two checkouts (with different marker tools) sharing a filesystem datastore outside both repos. It covers:
- pull, then update, then rm propagation;
- workflows of a restored extension found in the same command;
- the lockfile bytes left unchanged by convergence;
- an entry with no archive keeping its current files;
- tamper refusal;
- a missing lockfile deleting nothing;
- skills surviving convergence;
- a failure on the second rename rolling back;
- backfill.
- A serve instance on a filesystem datastore.
- The extension-backed bootstrap case.
Follow-ups to file
- Manual docs (9 swamp-club pages).
- swamp-uat coverage: there is no managedConfig coverage today.
- Serve reload does not unregister a removed extension's types.
- The S3/GCS extensions never delete locally on pull (tombstone-aware pulls).
datastore config migratestill copies the legacy pulled tree intoconfig/pulled-extensions.doctor extensionsshould report sources of delisted extensions.- The plan-v1 stale-catalog
describecrash: reproduce first.
Triaged
Click a lifecycle step above to view its details.
hammz commented 9/28/2026, 10:18:25 PM
Triage and planning record (1 of 4): overview and proposed split
Triage. Classified as a feature (high confidence). The #2429 prerequisites are merged (PR #2676): ConfigPoller reloads on the tier lockfile hash, and the CLI extension writers push by path through pushManagedConfigPathsDeferred. The code references in the issue body were checked and are accurate.
Found beyond the issue text:
- The channel is dropped on restore in two places, not one:
createExtensionInstallDeps(src/cli/create_extension_install_deps.ts:92-103) and the auto-resolver context (src/cli/auto_resolver_adapters.ts:289-302). installExtensionspreads the parent context into dependency installs (src/libswamp/extensions/pull.ts:1421), so the parent'sexpectedChecksumleaks into any dependency missing from the lockfile.extractTarGz/listTarGzEntrieshave no decompressed-size cap, and there is no archive byte-limit constant anywhere (the nearest is the safety analyzer's 10 MB source limit).SCOPED_NAME_PATTERNis duplicated in 7 places and is module-private inextension_manifest.ts.- The intro of "Managed Config Deployment Architecture" in
design/enablers/datastores.mdalready claims sources live in the tier, contradicting the rest of the section. LockfileRepositoryserves reads from a construction-time snapshot, so any lock-based design needs an explicit refresh.
Planning. An 18-step plan went through four adversarial review rounds (v1 to v4; 36 findings, all blocking ones resolved). It is too large for one change, so it is being split. Ripples 2 and 3 record the v4 design; ripple 4 records the open findings and the decisions the reviews forced.
Proposed split (each unit ships on its own and does not break existing users):
| Unit | Scope | Depends on |
|---|---|---|
| U1 | Restore-path correctness: keep the entry channel on restore (both sites) and stop leaking expectedChecksum into dependency installs |
none |
| U2 | Archive safety caps: decompressed-size cap on both tar passes, a compressed archive cap shared with push, close the tar file handles | none |
| U3 | Per-checkout pulled-extensions lock at the install/remove service boundary, LockfileRepository.refresh(), install split into unlocked prepare and locked apply |
none |
| U4 | Stage-and-swap install with journal, crash recovery and a commit/rollback handle; filesChecksum excluding nested entry roots | U3 |
| U5 | Archive key and ports, tier archive store, archive storage from every writer, superseded-archive deletion, extension install restore from archives plus backfill |
U2, U4 |
| U6 | ConvergenceState and the convergence service, wired into the CLI repo-context helpers and startup | U5 |
| U7 | Serve wiring: boot convergence, per-poll convergence, restoreGeneration-driven reloads | U6 |
Each unit carries its own design-doc updates. The seven follow-ups listed in the issue body are filed once U5 to U7 land. U1 is being filed now as its own issue.
hammz commented 9/28/2026, 10:18:27 PM
Triage and planning record (2 of 4): plan v4, groundwork (steps 1-8)
- Archive identity and interfaces. Export
SCOPED_NAME_PATTERN. Add anExtensionArchiveKeyvalue object (scoped name plus exactly 64 lowercase hex characters;create()rejects anything else). Add portsExtensionArchiveSource.read(key, signal)andExtensionArchiveSink.store/delete, each returning the paths touched. AddMAX_EXTENSION_ARCHIVE_BYTESandMAX_EXTENSION_ARCHIVE_DECOMPRESSED_BYTES; push checks the same compressed constant so they cannot drift. - Decompressed-size cap on both tar passes, as one byte-counting stream option on
listTarGzEntriesandextractTarGz. Also close the two file handles atpull.ts:929and944. - TierExtensionArchiveStore over
<config base>/pulled-archives/<name>/<sha256>.tar.gz. Paths are built only from a validated key and contained withassertSafePath. Writes are atomic and idempotent. Reads refuse oversize files and re-hash the bytes; a mismatch is a tamper warning. A local miss hydrates the one archive, which needs aHydrateFileHookvariant that forwards a signal, plus aPromise.racedeadline. Archives live underconfig/, so config pulls bring them: a fresh pod downloads everything once, and later pulls fetch only new archives. - Per-checkout pulled-extensions lock: an in-process keyed mutex plus a cross-process
FileLockon.swamp/pulled-extensions.lock, with atryAcquire(). Reentrancy uses anAsyncLocalStoragelease cleared infinally, so an auto-resolve inside a locked section runs inline, but cached promises resolved later cannot bypass the lock.installExtensionis split into an unlocked prepare (fetch, verify, extract, safety analysis) and a locked apply (refresh the lockfile snapshot, stage, swap,writeEntry, dependencies, phase 8, commit). The lock wrapsInstallExtensionService.execute, the direct-caller wrapper andRemoveExtensionService.execute, so every entry point holds it. It is never held across the interactive conflict prompt. Lock order: this lock first, the lockfile advisory.lockinnermost.LockfileRepository.refresh()is called at the start of apply. - installExtension archive source and record policy.
InstallContextgainsarchiveSource,archiveSink,record(always/if-missing-checksum/never, defaultalways) andmode: 'converge'. An archive hit skips the registry calls, and requires the manifest name to match and, for an exact pin, the version. The record policy gateswriteEntry, the sink, dependency recursion and orphan pruning. Only archives whose registry checksum came back non-null and matched may be stored. Dependencies no longer inheritexpectedChecksumor the overrides. - Stage, swap, commit. A staged swap under
<pulledRoot>/.swamp-staging/<uuid>/replaces the merge copy. Bundle folders get a sibling.swamp-staging-<uuid>at least two levels deep. A journal is written before any staging folder is created; it holds the owner id, phase, the lockfile path this install writes, the new checksum, and old/new manifest digests and paths. Phase 1 moves live roots toold/<i>; phase 2 moves new roots in,manifest.yamllast, and the phase becomesswapped. Any failure after the swap rolls back the install and its dependencies. APendingInstallhandle hascommit()(deleteold/, store the archive, delete the superseded archive) androllback()(reverse the swap, and restore the prior entry with channel andpulledAtonly ifwriteEntryran). The service commits aftersaveAll, rolls back onDuplicateTypeError, and commits then rethrows on other phase-8 errors.filesChecksumexcludes nested entry roots. The swap drops local files that are not in the archive; this is documented. - Crash recovery. The journal is Zod-validated with containment checks. Recovery runs under the exclusive lock and treats a journal whose owner id is not in this process's active set as orphaned; there are no pid heuristics. It rolls forward when the journal reached
swappedand the recorded lockfile's entry matches the new checksum; otherwise it rolls back. It never discards the only copy of a root. It runs at the start of every apply, at the start ofrm, and in the convergence gate. - Keep the channel on restore in
createExtensionInstallDepsand the auto-resolver.
hammz commented 9/28/2026, 10:18:29 PM
Triage and planning record (3 of 4): plan v4, archives and convergence (steps 9-18)
- Archive wiring for every writer. One factory,
tierArchiveStoreFor(repoDir, marker, lockfilePath), returns a store only when managedConfig is active, the base is resolved and the lockfile is the tier lockfile (the #445 exemption never stores archives). It is passed through the shared factories:extensionPull(pull, search install, doctor repair), libswampcreateInstallContext(update),createExtensionInstallDeps(install, serve install, repo upgrade), and the auto-resolver. The auto-resolver also gets the archive source, and usesif-missing-checksumfor pinned entries. Archive paths go to the publish call each caller already uses (pushManagedConfigPathsDeferred,pushManagedLockfileIfChangedDeferredwith a newextraPathsargument, serve'smarkExtensionChanges). The pinned-writer rule test checks for this. - Superseded archives are deleted in
commit()(for rm, after the lockfile write), only when the old and new checksums differ. Serve propagates the deletion. The CLI on S3/GCS does not, because its first push in a process is a full walk that skips deletion detection (#2273); the leftover is unreferenced and harmless. This is documented. - ConvergenceState in
.swamp/pulled-extensions-state.json, Zod-validated: config base, lockfile hash, pending entries with backoff (1 minute to 1 hour), unrestorable entries, backfill failures,lastVerified,restoreGeneration, and a warn-once digest. Invalid state means a full pass. A failed write is one warning. - extension install restores from archives first, then the registry. It records
alwaysfor legacy migration or missing checksums, otherwiseif-missing-checksum. It backfills missing archives (exact version from the on-disk manifest for constraint pins, verified against the entry checksum and a non-null registry checksum). The output gains a per-entrysourceand a top-levelbackfilledlist, both additive. - convergePulledExtensions reads only the tier lockfile, the archive source and the state, and never throws. The gate: recover orphaned journals; skip when the lockfile is absent or unparseable; full pass when the config base changed; skip when the hash matches and nothing is due. Converge mode uses record
never,skillsDirs: [], and no registry functions. It checks only files under<pulledRoot>/<name>, excluding skills, bundles and nested roots, and never migrates. Restore rule: files missing, OR an exact-pin version skew, OR (checksum differs fromlastVerifiedAND the digest differs fromfilesChecksum). A digest match recordslastVerified. Same-version drift is a warning. Calls in one process share a single in-flight run created outside any lock lease; across processes it uses the try-lock and reports busy. - CLI wiring in
requireInitializedRepoReadOnly,requireInitializedRepo(after the coordinator pull) andrequireInitializedRepoUnlocked, plus askipConvergenceoption. It also runs inconfigureStartupExtensionsbefore reconcile and the missing-files check, except for thin-client commands and when local file checks are off.withLoaderContextis the fallback. - Serve wiring. Convergence runs in
pullManagedConfigAtBootand on the no-sync-service boot path, before the boot hash and registry loads. The poller converges every poll. A busy result delays the hash-change reload until a non-busy pass. It reloads when the hash changed orrestoreGenerationmoved.ServeReloadResponseis unchanged. - Tests. In-process integration tests with two checkouts using different marker tools on a shared filesystem datastore; serve on a filesystem datastore; the extension-backed bootstrap case; a runtime contract test that convergence never calls the registry or
writeEntry. - Docs: datastores.md, serve.md, extensions.md (including the wrong
sha256-example), architecture.md, remote-execution.md, and the "until #2612" notes in code and tests. - File the issue's seven follow-ups.
hammz commented 9/28/2026, 10:18:30 PM
Triage and planning record (4 of 4): review decisions and open findings
Decisions forced by the adversarial reviews (carry these into the split issues):
- Lock at the install/remove service boundary and hold it through commit, not around whole commands. A command-level lock deadlocks against the auto-resolver, which
resolveManagedLockfileForWritecan trigger, and would be held across the conflict prompt. - Take the lock and the archive sink in the shared factories, so search install, doctor repair, repo upgrade and serve install are not missed.
- Converge mode must ignore skill paths:
entry.filesrecords the marker tool of whoever pulled, so checkouts on different tools would otherwise re-swap forever. It must also never run legacy migration, which deletes files. - The
if-missing-checksumrecord policy breaks legacy-layout migration (install.ts:254-262relies on the rewrittenfiles[]); usealwaysthere. - Pending entries need backoff and an unrestorable bucket, or every CLI command pays a network round trip.
- Recovery must not rely on pid liveness: containers reuse pids. Use owner ids plus the exclusive lock.
- Restoring whenever the checksum differs from
lastVerifiedoverwrites same-version edits on first run; also require the digest to differ. - The lockfile snapshot cache makes "re-read under the lock" a no-op without
refresh().
Open findings on v4 (not yet folded in; the U3, U4 and U7 issues should address them):
- ADV-32: journal lockfile containment rejects the non-managed lockfile at
extensions/models/upstream_extensions.json; accept exact matches of the resolved lockfile paths. - ADV-33: decide each root's recovery from the staging layout (old/new presence), not from the extension root's manifest verdict, or a root that was never moved can be lost.
- ADV-34: the
restoreGenerationreload trigger needs the same baseline and failure cap as the hash trigger. - ADV-35: a busy converge should defer every reload, and the generation should bump only after commit, recovery or rollback.
- ADV-36:
RemoveExtensionService.executeneedsrefresh()under the lock. - ADV-29 (low): pass the poller's stop signal into convergence so shutdown is not held by a hydrate.
- ADV-30 (low): write the journal before creating bundle staging folders.
- ADV-31 (low): the lock holder should re-check ownership before each rename phase; document the shared-volume limitation.
hammz commented 9/28/2026, 10:20:12 PM
U1 (restore-path correctness: channel dropped on restore, and the parent's expectedChecksum leaking into dependency installs) is filed as swamp-club#2639. It has no dependencies and will be triaged on its own.
hammz commented 9/29/2026, 5:44:48 PM
Finer split (2026-09-29)
U2 and U3 are now filed as three sub-issues. Each ships on its own and affects every repo, not just managedConfig ones:
| Sub-issue | Scope | Depends on |
|---|---|---|
| swamp-club#2707 (was U2) | Compressed and decompressed size caps on extension archives, shared by pull and push; check the tar handles at pull.ts:930/945 |
none |
| swamp-club#2708 (new, from U3) | Split installExtension into an unlocked prepare and a repo-changing apply; refactor with no behavior change |
none |
| swamp-club#2709 (rest of U3) | Per-checkout pulled-extensions lock at the install/remove boundary, LockfileRepository.refresh(), lock order (fixes ADV-36) |
#2708 |
Corrections to the plan found while filing:
- The dependency-override leak is already covered: U1 (#2639) clears
expectedChecksumfor dependencies atpull.ts:1430, and the other overrides do not exist yet. FsFile.readablecloses on stream end or cancel, so #2707 confirms whether the flagged handles actually leak before changing them.- Lock order: datastore global lock → pulled-extensions lock → lockfile
.lock. The auto-resolver runs inside commands that already hold the global lock, so swamp-club#2495 must take the global lock before the pulled-extensions lock, never inside apply.
Remaining units, to be filed later:
- U4 → D (stage-and-swap with journal and crash recovery) and E (commit/rollback handle across the catalog save).
- U5 → F (tier archive store plus every writer storing archives), G (delete superseded archives) and H (
extension installrestores from archives, plus backfill). - U6 → I (CLI convergence).
- U7 → J (serve boot convergence) and K (per-poll convergence and reloads).
swamp-club#2495 comes next after #2709. It should be split into (a) hydrate-before-write and scoped publish, (b) the auto-resolver writing the shared lockfile, (c) adoption (after D), (d) legacy-lockfile retirement, and (e) the datastore setup extension overwrite. Its (a) and (b) should land before I.
hammz commented 9/29/2026, 7:15:39 PM
U4 is filed as swamp-club#2723 (D: stage-and-swap install with a journal and crash recovery; covers ADV-30 to ADV-33, the staging sweep, nested roots, and the filesChecksum fix) and swamp-club#2724 (E: commit/rollback install handle across the catalog save). Order: #2708 → #2709 → #2723 → #2724. Found while filing: today a v1→v2 upgrade that hits DuplicateTypeError is rolled back to no version at all, because createdPaths includes the files that overwrote v1 and v1-only files were already pruned. Also, filesChecksum walks nested entry roots, so installing @a/b/c makes @a/b look locally edited.
Sign in to post a ripple.