Relationships
#2418 serve: after lazy hydration, scoped poller pulls never take the fast path, and the login mint waits behind them on the sync gate
Opened by hammz · 9/23/2026
Description
On a serve whose datastore cache was hydrated lazily, every poller pull takes the slow path indefinitely. The OAuth login mint (and every gated WebSocket mutation) waits behind those pulls on the sync gate.
- A lazy hydration pull (
metadataOnly) setslazyPullActivein the S3/GCS extension's sync state (gcs_cache_sync.ts:2507-2511). - While
lazyPullActiveis set, every pull that is notmetadataOnlyskips the commitSeq fast path (gcs_cache_sync.ts:2178). - Only an unscoped pull with no model context clears the flag (
gcs_cache_sync.ts:2512-2519). - All three serve pollers (config, access-data, runtime-data) pass
subdirs, so they are scoped and never clear the flag. Each poll every 30 s therefore does a full slow-path pull: assembling the index from every shard, then listing the subdir prefix. The runtime poller lists all ofdata/. - Each poller pull holds the one-permit FIFO sync gate (
src/serve/sync_gate.ts), for up toPOLLER_PULL_TIMEOUT_MS(120 s). The login mint (device_auth_handler.ts,withSyncGate) and the gated WebSocket mutations queue behind up to three such holds. They go ahead ungated only afterGATE_WAIT_TIMEOUT_MS(150 s).
The result: on a large lazily hydrated datastore, login and mutation latency can be dominated by the gate wait even once the mint's own push is scoped (swamp-club#2408).
Where this applies
- ops.swamp-club.com hydrates lazily through the platform-swamp
serve-lazy-hydration.patch. It currently runs core375d7eca, which predates the sync gate. So this becomes the main login delay once platform-swamp moves its core pin to current main while keeping that patch. - Current main's serve does not hydrate lazily at boot. But a cache hydrated lazily by the CLI (
hydrationStrategy: lazy) and later served byswamp servekeepslazyPullActivein its sync-state sidecar, so the same behaviour follows.
Possible directions
- Extension: let a scoped pull that covers the subdirs it was asked for use the fast path when the remote commitSeq has not moved, or track lazy hydration per subdir.
- Serve: bound the mint's gate wait separately, or shorten poller holds, for example by skipping a poll cycle while a gated request is queued.
Evidence
- Code references are from swamp-extensions at c47093989 (
datastore/gcs/extensions/datastores/_lib/gcs_cache_sync.ts). The S3 extension has the same!scopedguard. - The swamp-club#2408 MinIO repro did not hydrate lazily. Its pollers hit the fast path (
pullChanged.fastpath ~5ms hit), so gate contention was not observed there. Reproducing this issue needs a lazily hydrated serve cache.
Open
No activity in this phase yet.
bixu commented 9/24/2026, 6:17:46 AM
🤖
Production evidence from a serve with
hydrationStrategy: lazy. Setup: k3s-dev, namespacefactory-dev, swamp20260923.001003.0-sha.0c3a78cb,@swamp/s3-datastoreon S3.
The slow-path poller pulls also have a large tracing cost.
S3 getObjectspans went from 356K to 3.16M in 24 hours (about 9x).- This is 81% of the serve's span volume.
- Each poll reads all 269
factory-dev/_index/*shards again.- The steady rate is about 45K
getObjectspans per hour.
All poller spans attach to one root span for each pod lifetime. That root span never ends. One pod has 1.9M
getObjectspans under one trace ID. The root span is never exported. Soparent.*joins find nothing and the trace view is not usable.
Earlier builds (
20260922.*) showed the same pattern at a lower rate that decayed as the pod aged.20260923.001003holds the high rate.
Suggestion: in addition to the fast-path fix, start a new root span for each poll cycle. A poll that finds no change could also emit no child spans.
Our mitigation: we are moving dev to
hydrationStrategy: fulluntil a fix ships.
bixu commented 9/24/2026, 6:20:23 AM
🤖
Correction to my previous ripple. This serve runs
hydrationStrategy: full, notlazy. Dev has usedfullsince 2026-09-17. The mitigation line in my last ripple is wrong.
The sidecar on the live pod shows
lazyPullActive: false. It has nocommitSeqfield, andremoteIndexETagis empty. SotryCommitSeqFastPathreturnsnulland the ETag probe also misses. Every scoped poller pull takes the slow path.
Probable writer, in
@swamp/s3-datastore2026.09.10.2s3_cache_sync.ts: the v2 no-writeback branch ofpushChangedrebuilds the sidecar. It setscommitSeqonly whenlocalHasAllRemoteEntries()is true. Otherwise it writes a sidecar withoutcommitSeq, which disarms the fast path.
So the same symptom can occur without lazy hydration. It may deserve its own issue. The span evidence in my previous ripple still applies.
Sign in to post a ripple.