Relationships
#2998 s3-datastore: every fast-path miss re-downloads every _index shard, so a busy serve pulls the whole index every poll (~230 GB/day S3 egress)
Opened by sntxrr · 10/3/2026
Summary
On a fast-path miss, pullChanged in @swamp/s3-datastore rebuilds the index by downloading every shard listed in _meta.json (assembleIndexFromShards), even when only one or two shards changed. Scoped pulls (the serve pollers' subdirs, fixed in #2246) narrow the files walked afterwards, but not the index read that comes first.
The fast path only hits when _meta.json's commitSeq is unchanged since the last pull. On a repo where scheduled workflows write regularly, commitSeq changes about once a minute, so the fast path almost never hits, and every poll downloads the full index.
Measured (one serve instance, one namespace)
_index/: 445 shards, 114 MB in total. The largest shards are 14 MB, 11 MB and 9.7 MB (models and workflows with long data histories)._meta.jsoncommits: 400 in 8 h, about 1 every 73 s.- serve's network intake: 9.03 GB in its first 57 min of uptime, steady at 2.2–3 MB/s, about 230 GB/day. This matches the bucket's billed
DataTransfer-Out-Bytes: 130–315 GB/day, about $484 of a $745 S3 month. - A manual
swamp datastore sync --pull --timeout 1800against an up-to-date cache (17, then 45 files changed) took 172 s and 185 s. - Because each pull takes longer than the 120 s default, serve's pollers log
Datastore pull to "runtime data poller" timed out after 120000msevery few minutes, and the next poll starts the full download again. - Raising the timeout would only let the same 114 MB download finish, so it doesn't cut the cost.
- After
swamp data gc(120,250 excess versions reclaimed), the index dropped to 34 MB and serve's intake to about 1.3 MB/s. That shows per-poll bytes scale with the size of the index, not with the amount of change. Within about 15 minutes the index was back at 48 MB and intake at 2.3 MB/s. - After upgrading serve and the CLI to
2026.10.01.1, intake measured 2.8 MB/s over 5 minutes (index 48 MB). No change, as expected from the code path below.
Code path (swamp-extensions main @ ae548a8)
datastore/s3/extensions/datastores/_lib/s3_cache_sync.ts
pullChanged(~L2395):tryFastPullChangedmisses → unscoped / subdir-scoped branch →assembleIndexFromShards(signal)(~L2466)assembleIndexFromShards(~L1483): GETs every partition in_meta.jsonv2, regardless of which ones changed
Expected
On a fast-path miss, fetch only the shards that changed since the last pull. For example:
- record each shard's ETag in the sync-state sidecar, HEAD or conditionally GET (
If-None-Match) each shard, and reuse the cached copy on 304; or - record a per-partition
commitSeqin_meta.json, so the reader can tell which partitions moved without reading them.
The pull cost would then scale with the amount of change, not with the size of the index.
Workarounds in the meantime
swamp data gcto shrink shards. This helps in proportion to how much history is dropped.- Nothing on the serve side reduces per-poll bytes. Raising
SWAMP_DATASTORE_SYNC_TIMEOUT_MSmakes the problem worse.
Related
- #2246: scoped pulls ignored
subdirs(shipped). This is the remaining half: the index read before the scope is applied. - #2865: the datastore commit-log rework would likely remove this, but it's still at the test-baseline stage.
Environment
- Extension:
@swamp/s3-datastore@2026.09.24.1, reproduced on2026.10.01.1(code path unchanged on main @ ae548a8) - swamp:
20261002.202148.0-sha.0fc928c8(serve, linux container) - AWS S3, single region, versioning enabled
Upstream repository: https://github.com/systeminit/swamp-extensions
Environment
- Extension:
@swamp/s3-datastore@2026.10.01.1 - swamp:
20261002.194016.0-sha.c0c751a1 - OS:
darwin(aarch64) - Deno:
2.9.7 - Shell:
/bin/zsh
Open
No activity in this phase yet.
system commented 10/3/2026, 12:35:59 AM
Classified automatically when this issue was filed.
- Source: Extensions
If you feel this classification is incorrect, add a ripple to tell us so.
Sign in to post a ripple.