Skip to main content
← Back to list
01Issue
BugIn ProgressSwamp CLIPublic
Assigneeshammz

Relationships

#2868 managedConfig: model/vault create and other bulk-push commands re-upload a stale extension lockfile and erase peers' entries

Opened by hammz · 10/1/2026

Description

In a managedConfig repo on an S3 datastore, commands that change the config tier can re-upload a checkout's stale cache copy of config/upstream_extensions.json. When they do, any entries that other checkouts have added since this one last synced are lost. Affected commands include model create, model edit, vault create, vault edit, vault migrate, the access group, grant and token commands, datastore namespace, the worker token commands, and worker prune.

Each of these calls syncService.markDirty() with no path. The S3 extension treats that as bulk invalidation, so the next pushChanged walks the whole cache and uploads the lockfile along with everything else. The push is last-writer-wins per file. Neither the global lock nor a prior fetch protects the lockfile here, so the stale copy overwrites the newer remote one.

This is the same lost-update problem as swamp-club#2838, but with a different writer. #2838 adds a managed lockfile transaction that fetches the lockfile under the datastore global lock, applies the change, and publishes only the lockfile. That keeps extension writes safe from each other. It does not stop these other writers from rolling the lockfile back. A pending-delta replay can't recover the lost entries either, because they were never in this checkout's delta. (swamp-club#2495 noted that pushManagedConfigChanges uses a bare markDirty(), but only for the extension write path.)

datastore sync --push behaves differently. From a clean checkout it does not re-upload an unmodified stale lockfile. After a failed lockfile publish, though, it does overwrite peers' entries.

The released binary (20260930.180631.0-sha.e34705c3) and the #2838 branch (afa40bb5) both behave this way.

Steps to reproduce

  1. Run MinIO and set up two checkouts, A and C. Each has its own repo and its own SWAMP_HOME cache, and both share one bucket and prefix: swamp datastore setup extension @swamp/s3-datastore --config '{"bucket":"b","prefix":"p","region":"us-east-1","endpoint":"http://127.0.0.1:9000","forcePathStyle":true}', then swamp datastore config migrate.
  2. In A, run swamp extension pull @swamp/1password -y.
  3. In C, run swamp extension pull @swamp/ssh -y. The bucket lockfile now lists both extensions.
  4. In A, run swamp model create command/shell m1.
  5. Read <prefix>/config/upstream_extensions.json from the bucket. @swamp/ssh is gone. The object's ETag is identical to A's publish from step 2, so A re-uploaded its unmodified stale copy.

Replacing step 4 with swamp vault create local_encryption v1 gives the same result.

Variant: datastore sync --push after a failed publish.

  1. A's extension pull fails to publish, so the v1 pending record is kept.
  2. C pulls another extension.
  3. A runs datastore sync --push. C's entry is overwritten.
  4. A runs swamp extension install to retry. The retry replays only A's own delta, so C's entry stays lost.

Expected: C's entry survives in every case. Actual: it is permanently removed from the shared lockfile.

A scripted repro, with a scenario per case and run against both the released and branch binaries, is in the #2838 e2e harness (scenario F1 for the main case, B5 for the variant).

Environment

  • swamp 20260930.180631.0-sha.e34705c3 (released) and the #2838 branch at afa40bb5
  • @swamp/s3-datastore 2026.09.24.1 against MinIO RELEASE.2025-09-07T16-13-09Z
  • Linux

Suggested fix

Make the managed lockfile publish only through the managed lockfile transaction. Either of these would do it:

  • Exclude the managed lockfile from bulk and walk pushes.
  • Have walk pushes skip any config file whose remote changed since this checkout's last sync, instead of overwriting it.

The commands above could also mark only the paths they actually wrote, rather than calling a bare markDirty().

02Bog Flow
✓OPEN✓TRIAGED◉IN PROGRESS○SHIPPED+ 1 MOREASSIGNED+ 5 MOREREVIEW+ 8 MOREVERIFICATION_PASSED

In Progress

10/1/2026, 7:08:22 PM

Click a lifecycle step above to view its details.

03Sludge Pulse
hammz assigned hammz10/1/2026, 6:51:02 PM

Sign in to post a ripple.