Relationships
#2416 datastore extensions: full pushes re-hash every pulled file on every run because index mtimes never match pulled copies
Opened by hammz · 9/23/2026
Description
In the @swamp/s3-datastore and @swamp/gcs-datastore extensions, every full-walk pushChanged on a cache filled by pulling reads and SHA-256s every pulled file, and it does so on every push with no warm-up.
The cause:
- A full push rebuilds the index from the remote shards.
- Each shard entry carries the
localMtimeof the machine that wrote the file. - A pulled copy has its download time as its mtime, so the two never match, and
fileNeedsPush(s3_cache_sync.ts:3904,gcs_cache_sync.ts:3720) takes the hash branch for the file. - When the hash shows the file is unchanged, its
localMtimeis never written back to match. The next full push hashes the same files again.
This affects every consumer of a shared datastore whose cache was hydrated by pull: swamp serve, HA peers, and CI runners.
Evidence
Reproduction on MinIO with @swamp/s3-datastore 2026.09.23.1 and swamp 20260923.142649.0-sha.379bb5f0. A writer repo seeded 25,000 files, and a serve repo hydrated them by pull. Each OAuth login on that serve did a full push (the no-path markDirty() from swamp-club#2408).
- An offline replay of
fileNeedsPushagainst the remote shards hashed 25,014 of 25,158 index entries. The only files skipped were ones the serve had written itself. /proc/<pid>/ioshowed about 15.7 MB read on every login, which matches the size of the cache, on all 5 runs.- Hashing was about 3.4 s of a 5.5 s push at 25k files, about 60% of the push.
- A shard entry showed
localMtime2026-09-23T16:07:28.484Z, the time the writer wrote the file. The serve's copy had a download-time mtime of about 16:09:43.
Related costs seen in the same reproduction
- Full-push writeback:
localHasAllRemoteEntries()stats every index entry a second time, and the 10 MB.datastore-index.jsonis rewritten. That took about 0.8 s at 25k. - Scoped (per-path) pushes still read, merge and rewrite the whole local
.datastore-index.jsoninwriteLocalIndexAfterPush. That is about 10 MB of I/O at 25k, and most of the remaining 150–250 ms of a 4-file push. - Shard assembly fetches one GET per partition, in sequential batches of 10. On a real bucket the cost is the number of round trips times the latency.
Expected
On a pulled cache, a file that was verified unchanged should not be hashed again on every full push. For example, the pull could record the local mtime of each downloaded file in the index, or a hash-verified entry could get its localMtime reconciled.
Environment
S3 extension verified directly. The GCS extension has the same bulk-dirty, hash-on-mtime-mismatch and second-stat logic. The platform-swamp vendored GCS adapter used by ops.swamp-club.com is worse: under bulkInvalidated it hashes every file, even when the mtimes match.
Open
No activity in this phase yet.
Sign in to post a ripple.