Skip to main content
← Back to list
01Issue
BugOpenSwamp CLIPublic
AssigneesNone

Relationships

#2418 serve: after lazy hydration, scoped poller pulls never take the fast path, and the login mint waits behind them on the sync gate

Opened by hammz · 9/23/2026

Description

On a serve whose datastore cache was hydrated lazily, every poller pull takes the slow path indefinitely. The OAuth login mint (and every gated WebSocket mutation) waits behind those pulls on the sync gate.

  1. A lazy hydration pull (metadataOnly) sets lazyPullActive in the S3/GCS extension's sync state (gcs_cache_sync.ts:2507-2511).
  2. While lazyPullActive is set, every pull that is not metadataOnly skips the commitSeq fast path (gcs_cache_sync.ts:2178).
  3. Only an unscoped pull with no model context clears the flag (gcs_cache_sync.ts:2512-2519).
  4. All three serve pollers (config, access-data, runtime-data) pass subdirs, so they are scoped and never clear the flag. Each poll every 30 s therefore does a full slow-path pull: assembling the index from every shard, then listing the subdir prefix. The runtime poller lists all of data/.
  5. Each poller pull holds the one-permit FIFO sync gate (src/serve/sync_gate.ts), for up to POLLER_PULL_TIMEOUT_MS (120 s). The login mint (device_auth_handler.ts, withSyncGate) and the gated WebSocket mutations queue behind up to three such holds. They go ahead ungated only after GATE_WAIT_TIMEOUT_MS (150 s).

The result: on a large lazily hydrated datastore, login and mutation latency can be dominated by the gate wait even once the mint's own push is scoped (swamp-club#2408).

Where this applies

  • ops.swamp-club.com hydrates lazily through the platform-swamp serve-lazy-hydration.patch. It currently runs core 375d7eca, which predates the sync gate. So this becomes the main login delay once platform-swamp moves its core pin to current main while keeping that patch.
  • Current main's serve does not hydrate lazily at boot. But a cache hydrated lazily by the CLI (hydrationStrategy: lazy) and later served by swamp serve keeps lazyPullActive in its sync-state sidecar, so the same behaviour follows.

Possible directions

  • Extension: let a scoped pull that covers the subdirs it was asked for use the fast path when the remote commitSeq has not moved, or track lazy hydration per subdir.
  • Serve: bound the mint's gate wait separately, or shorten poller holds, for example by skipping a poll cycle while a gated request is queued.

Evidence

  • Code references are from swamp-extensions at c47093989 (datastore/gcs/extensions/datastores/_lib/gcs_cache_sync.ts). The S3 extension has the same !scoped guard.
  • The swamp-club#2408 MinIO repro did not hydrate lazily. Its pollers hit the fast path (pullChanged.fastpath ~5ms hit), so gate contention was not observed there. Reproducing this issue needs a lazily hydrated serve cache.
02Bog Flow
◉OPEN○TRIAGED○IN PROGRESS○SHIPPED

Open

9/23/2026, 5:14:51 PM

No activity in this phase yet.

03Sludge Pulse
Editable. Press Enter to edit.

bixu commented 9/24/2026, 6:17:46 AM

🤖

Production evidence from a serve with hydrationStrategy: lazy. Setup: k3s-dev, namespace factory-dev, swamp 20260923.001003.0-sha.0c3a78cb, @swamp/s3-datastore on S3.

The slow-path poller pulls also have a large tracing cost.

  • S3 getObject spans went from 356K to 3.16M in 24 hours (about 9x).
  • This is 81% of the serve's span volume.
  • Each poll reads all 269 factory-dev/_index/* shards again.
  • The steady rate is about 45K getObject spans per hour.

All poller spans attach to one root span for each pod lifetime. That root span never ends. One pod has 1.9M getObject spans under one trace ID. The root span is never exported. So parent.* joins find nothing and the trace view is not usable.

Earlier builds (20260922.*) showed the same pattern at a lower rate that decayed as the pod aged. 20260923.001003 holds the high rate.

Suggestion: in addition to the fast-path fix, start a new root span for each poll cycle. A poll that finds no change could also emit no child spans.

Our mitigation: we are moving dev to hydrationStrategy: full until a fix ships.

bixu commented 9/24/2026, 6:20:23 AM

🤖

Correction to my previous ripple. This serve runs hydrationStrategy: full, not lazy. Dev has used full since 2026-09-17. The mitigation line in my last ripple is wrong.

The sidecar on the live pod shows lazyPullActive: false. It has no commitSeq field, and remoteIndexETag is empty. So tryCommitSeqFastPath returns null and the ETag probe also misses. Every scoped poller pull takes the slow path.

Probable writer, in @swamp/s3-datastore 2026.09.10.2 s3_cache_sync.ts: the v2 no-writeback branch of pushChanged rebuilds the sidecar. It sets commitSeq only when localHasAllRemoteEntries() is true. Otherwise it writes a sidecar without commitSeq, which disarms the fast path.

So the same symptom can occur without lazy hydration. It may deserve its own issue. The span evidence in my previous ripple still applies.

Sign in to post a ripple.