Skip to main content
← Back to list
01Issue
BugShippedSwamp CLIPublic
Assigneeshammz

Relationships

#2901 serve: a failed pull in acquireModelLocks leaves the per-model lock held until serve exits

Opened by stack72 · 10/1/2026· Shipped 10/6/2026

Problem

acquireModelLocks (src/cli/repo_context.ts:1489) acquires each per-model lock and registers it with the datastore sync coordinator (registerDatastoreSyncNamed at :1631, which acquires the lock) before it pulls. If that pull fails (:1727-1735), it throws without unwinding the registered entry: the lock stays held and the coordinator entry stays in the module-scope map. The caller never receives a lock result, so it cannot call flush.

In the CLI this is bounded: the runCli error path calls flushDatastoreSync (src/cli/mod.ts:2612), which releases every registered entry. Serve has no such teardown. src/serve contains no flushDatastoreSync call, and serve calls acquireModelLocks per request (src/serve/handlers/model_handlers.ts:342 and :554, src/serve/deps.ts:369 for workflow steps). One transient remote failure during the pull therefore leaves that model's lock held for the life of the serve process. Later requests for the same model wait for the lock until it times out, and other machines sharing the datastore see a held lock.

Expected

A pull failure inside acquireModelLocks releases the locks and coordinator entries it registered for this call, the way registerDatastoreSyncNamed already unwinds its own entry on a failed pull, and then throws.

Notes

Found while planning swamp-club#2859, which pins today's behaviour (one leaked coordinator entry after a pull failure) in integration/datastore_remote_failure_test.ts. When this is fixed, that pin must be updated.

02Bog Flow
✓OPEN✓TRIAGED✓IN PROGRESS✓SHIPPED+ 1 MOREASSIGNED+ 5 MOREREVIEW+ 9 MOREPR_MERGED+ 2 MORESESSION_SUMMARIZED

Shipped

10/6/2026, 3:07:29 PM

Click a lifecycle step above to view its details.

03Sludge Pulse
hammz assigned hammz10/6/2026, 1:51:28 PM

Sign in to post a ripple.