Skip to main content
← Back to list
01Issue
BugShippedSwamp CLIPublic
Assigneeshammzstack72

Relationships

↔ sibling #2917

#3051 run doctor reads every run record on each run to rebuild the run indexes

Opened by hammz · 10/5/2026· Shipped 10/6/2026

Follow-up to #2518, which shipped in 20261005.222556.0.

Problem

Since #2518, swamp run doctor rebuilds every workflow's run index from the run records before it looks for orphaned runs, on every invocation, with or without --fix, locally and through serve. That is what stops a stale index entry hiding a running run from doctor, but it makes doctor's cost grow with the number of run records, and it rewrites every .runs-index.json each time even when nothing changed.

Measured

Wall-clock seconds, median of 3, old binary (20261005.174318.0) against the #2518 build, on tmpfs on a 32-core machine, records of about 18.5 KB:

  • 100 records: run doctor 0.49 before, 0.58 after
  • 1,000 records: 0.50 before, 1.04 after
  • 10,000 records: 0.53 before, 5.58 after (5.98 with all records in one workflow)

About 0.5 ms per record over a 0.5 s floor. A real disk with a cold cache will be slower. Serve boot, workflow history search and workflow runs did not change.

Two side effects of the rebuild:

  • The old doctor rewrote no index file; the new one rewrites all of them with identical content.
  • A doctor in a second process can overwrite an entry a serve process wrote during the rebuild, because index writes are only queued within one process. Under load this left about one wrong or missing entry per several hundred runs until that run was next read (one sample per configuration, so a rough rate). The old binary showed the same class of drift.

Options

  1. Verify instead of rebuild: read each record's status cheaply and rewrite an index only when an entry is wrong. Removes the unconditional rewrite and narrows the cross-process race, but still reads every record.
  2. Rebuild only with --fix, and have a plain doctor report that indexes were not verified.
  3. Bound the work: skip records whose file is older than the index and whose entry is terminal, so only records that could have changed are read.
  4. Make index writes safe across processes (a lock on the index file, or a per-run entry file instead of one JSON per workflow), which removes the reason to rebuild at all.

Code: diagnoseLocalRuns in src/cli/commands/run.ts, handleRunDoctor in src/serve/handlers/admin_handlers.ts, rebuildIndexes in src/infrastructure/persistence/yaml_workflow_run_repository.ts. Measurements and scripts from the validation run are described in the PR for #2518 (swamp-club/swamp#2849).

02Bog Flow
✓OPEN✓TRIAGED✓IN PROGRESS✓SHIPPED+ 4 MOREISSUE_LINKED+ 5 MOREREVIEW+ 17 MOREPR_MERGED+ 2 MORESESSION_SUMMARIZED

Shipped

10/6/2026, 2:08:24 PM

Click a lifecycle step above to view its details.

03Sludge Pulse
hammz assigned hammz10/5/2026, 10:39:02 PM
stack72 assigned stack7210/6/2026, 12:50:52 AM
stack72 linked sibling of #291710/6/2026, 12:52:06 AM

Sign in to post a ripple.