Relationships
#3073 serve HA: pull a token's definition on an auth miss so a peer accepts a new token at once
Opened by hammz · 10/6/2026
Summary
Follow-up to #2481 (fixed in swamp-club/swamp#2861). A running replica now pulls auto-definitions/ with data/ on every poll, so a server token minted on one replica is accepted by its peers within the poll interval. Until a peer's next poll it still rejects the token as unknown.
Why it matters
- The default poll interval is 30s, and under steady run load a poller can skip three cycles before it queues for the sync gate, so the window can be about four intervals.
- Device login is the case users will hit:
device_auth_handler.tsmints the token on one replica and the CLI's next connection can land on another replica straight away, behind a load balancer. The user logs in successfully and is then told authentication failed. - The same window applies to
access token mint --serverfollowed by immediate use.
Proposal
On a definition or record miss in readServerTokenRecord (src/serve/token_auth.ts), do one targeted pull of that token's definition and data before failing, then retry the lookup once.
Constraints to design for
- The miss path is reachable by unauthenticated callers, so a pull on miss lets anyone trigger datastore reads. It needs a bound: single-flight, plus a minimum interval between miss-driven pulls (per replica, and probably per token name), with a negative result remembered for that interval.
- The pull must run under the sync gate like every other pull (
gatedPullinsrc/serve/sync_gate.ts), and must not make connection upgrades wait behind a long-held gate. - Scope the pull as narrowly as the datastore extension allows (the server-token type directory under
auto-definitions/anddata/), not the whole subtree. - The client-facing error must stay generic; the reason belongs in the server log only.
Acceptance
With two replicas on a shared remote datastore, a token minted through replica A authenticates on replica B on the first attempt, without waiting for B's poll, and repeated attempts with unknown token names cause a bounded number of datastore pulls.
Reproduction environment
The two-replica MinIO setup used for #2481 is described in that issue's lifecycle record: two repos sharing one S3 datastore, each running swamp serve --auth-mode token --datastore-poll-interval 5s, mint through A with swamp access token mint <name> --principal user:<name> --server <A>, then swamp access can-i --server <B> --token-file <file> --action read --on 'workflow:*'. Read the server log for the result: can-i exits 1 for both an auth failure and an authenticated deny.
Open
No activity in this phase yet.
Sign in to post a ripple.