Relationships
#2901 serve: a failed pull in acquireModelLocks leaves the per-model lock held until serve exits
Opened by stack72 · 10/1/2026· Shipped 10/6/2026
Problem
acquireModelLocks (src/cli/repo_context.ts:1489) acquires each per-model lock and registers it with the datastore sync coordinator (registerDatastoreSyncNamed at :1631, which acquires the lock) before it pulls. If that pull fails (:1727-1735), it throws without unwinding the registered entry: the lock stays held and the coordinator entry stays in the module-scope map. The caller never receives a lock result, so it cannot call flush.
In the CLI this is bounded: the runCli error path calls flushDatastoreSync (src/cli/mod.ts:2612), which releases every registered entry. Serve has no such teardown. src/serve contains no flushDatastoreSync call, and serve calls acquireModelLocks per request (src/serve/handlers/model_handlers.ts:342 and :554, src/serve/deps.ts:369 for workflow steps). One transient remote failure during the pull therefore leaves that model's lock held for the life of the serve process. Later requests for the same model wait for the lock until it times out, and other machines sharing the datastore see a held lock.
Expected
A pull failure inside acquireModelLocks releases the locks and coordinator entries it registered for this call, the way registerDatastoreSyncNamed already unwinds its own entry on a failed pull, and then throws.
Notes
Found while planning swamp-club#2859, which pins today's behaviour (one leaked coordinator entry after a pull failure) in integration/datastore_remote_failure_test.ts. When this is fixed, that pin must be updated.
Shipped
Click a lifecycle step above to view its details.
Sign in to post a ripple.