Skip to main content
← Back to list
01Issue
BugOpenSwamp CLIPublic
AssigneesNone

Relationships

#2416 datastore extensions: full pushes re-hash every pulled file on every run because index mtimes never match pulled copies

Opened by hammz · 9/23/2026

Description

In the @swamp/s3-datastore and @swamp/gcs-datastore extensions, every full-walk pushChanged on a cache filled by pulling reads and SHA-256s every pulled file, and it does so on every push with no warm-up.

The cause:

  1. A full push rebuilds the index from the remote shards.
  2. Each shard entry carries the localMtime of the machine that wrote the file.
  3. A pulled copy has its download time as its mtime, so the two never match, and fileNeedsPush (s3_cache_sync.ts:3904, gcs_cache_sync.ts:3720) takes the hash branch for the file.
  4. When the hash shows the file is unchanged, its localMtime is never written back to match. The next full push hashes the same files again.

This affects every consumer of a shared datastore whose cache was hydrated by pull: swamp serve, HA peers, and CI runners.

Evidence

Reproduction on MinIO with @swamp/s3-datastore 2026.09.23.1 and swamp 20260923.142649.0-sha.379bb5f0. A writer repo seeded 25,000 files, and a serve repo hydrated them by pull. Each OAuth login on that serve did a full push (the no-path markDirty() from swamp-club#2408).

  • An offline replay of fileNeedsPush against the remote shards hashed 25,014 of 25,158 index entries. The only files skipped were ones the serve had written itself.
  • /proc/<pid>/io showed about 15.7 MB read on every login, which matches the size of the cache, on all 5 runs.
  • Hashing was about 3.4 s of a 5.5 s push at 25k files, about 60% of the push.
  • A shard entry showed localMtime 2026-09-23T16:07:28.484Z, the time the writer wrote the file. The serve's copy had a download-time mtime of about 16:09:43.
  • Full-push writeback: localHasAllRemoteEntries() stats every index entry a second time, and the 10 MB .datastore-index.json is rewritten. That took about 0.8 s at 25k.
  • Scoped (per-path) pushes still read, merge and rewrite the whole local .datastore-index.json in writeLocalIndexAfterPush. That is about 10 MB of I/O at 25k, and most of the remaining 150–250 ms of a 4-file push.
  • Shard assembly fetches one GET per partition, in sequential batches of 10. On a real bucket the cost is the number of round trips times the latency.

Expected

On a pulled cache, a file that was verified unchanged should not be hashed again on every full push. For example, the pull could record the local mtime of each downloaded file in the index, or a hash-verified entry could get its localMtime reconciled.

Environment

S3 extension verified directly. The GCS extension has the same bulk-dirty, hash-on-mtime-mismatch and second-stat logic. The platform-swamp vendored GCS adapter used by ops.swamp-club.com is worse: under bulkInvalidated it hashes every file, even when the mtimes match.

02Bog Flow
◉OPEN○TRIAGED○IN PROGRESS○SHIPPED

Open

9/23/2026, 4:35:11 PM

No activity in this phase yet.

03Sludge Pulse

Sign in to post a ripple.