Relationships
#2868 managedConfig: model/vault create and other bulk-push commands re-upload a stale extension lockfile and erase peers' entries
Opened by hammz · 10/1/2026
Description
In a managedConfig repo on an S3 datastore, commands that change the config tier can re-upload a checkout's stale cache copy of config/upstream_extensions.json. When they do, any entries that other checkouts have added since this one last synced are lost. Affected commands include model create, model edit, vault create, vault edit, vault migrate, the access group, grant and token commands, datastore namespace, the worker token commands, and worker prune.
Each of these calls syncService.markDirty() with no path. The S3 extension treats that as bulk invalidation, so the next pushChanged walks the whole cache and uploads the lockfile along with everything else. The push is last-writer-wins per file. Neither the global lock nor a prior fetch protects the lockfile here, so the stale copy overwrites the newer remote one.
This is the same lost-update problem as swamp-club#2838, but with a different writer. #2838 adds a managed lockfile transaction that fetches the lockfile under the datastore global lock, applies the change, and publishes only the lockfile. That keeps extension writes safe from each other. It does not stop these other writers from rolling the lockfile back. A pending-delta replay can't recover the lost entries either, because they were never in this checkout's delta. (swamp-club#2495 noted that pushManagedConfigChanges uses a bare markDirty(), but only for the extension write path.)
datastore sync --push behaves differently. From a clean checkout it does not re-upload an unmodified stale lockfile. After a failed lockfile publish, though, it does overwrite peers' entries.
The released binary (20260930.180631.0-sha.e34705c3) and the #2838 branch (afa40bb5) both behave this way.
Steps to reproduce
- Run MinIO and set up two checkouts, A and C. Each has its own repo and its own
SWAMP_HOMEcache, and both share one bucket and prefix:swamp datastore setup extension @swamp/s3-datastore --config '{"bucket":"b","prefix":"p","region":"us-east-1","endpoint":"http://127.0.0.1:9000","forcePathStyle":true}', thenswamp datastore config migrate. - In A, run
swamp extension pull @swamp/1password -y. - In C, run
swamp extension pull @swamp/ssh -y. The bucket lockfile now lists both extensions. - In A, run
swamp model create command/shell m1. - Read
<prefix>/config/upstream_extensions.jsonfrom the bucket.@swamp/sshis gone. The object's ETag is identical to A's publish from step 2, so A re-uploaded its unmodified stale copy.
Replacing step 4 with swamp vault create local_encryption v1 gives the same result.
Variant: datastore sync --push after a failed publish.
- A's
extension pullfails to publish, so the v1 pending record is kept. - C pulls another extension.
- A runs
datastore sync --push. C's entry is overwritten. - A runs
swamp extension installto retry. The retry replays only A's own delta, so C's entry stays lost.
Expected: C's entry survives in every case. Actual: it is permanently removed from the shared lockfile.
A scripted repro, with a scenario per case and run against both the released and branch binaries, is in the #2838 e2e harness (scenario F1 for the main case, B5 for the variant).
Environment
- swamp
20260930.180631.0-sha.e34705c3(released) and the #2838 branch atafa40bb5 @swamp/s3-datastore2026.09.24.1 against MinIO RELEASE.2025-09-07T16-13-09Z- Linux
Suggested fix
Make the managed lockfile publish only through the managed lockfile transaction. Either of these would do it:
- Exclude the managed lockfile from bulk and walk pushes.
- Have walk pushes skip any config file whose remote changed since this checkout's last sync, instead of overwriting it.
The commands above could also mark only the paths they actually wrote, rather than calling a bare markDirty().
In Progress
Click a lifecycle step above to view its details.
Sign in to post a ripple.