Skip to main content
← Back to list
01Issue
BugShippedSwamp CLIPublic
Assigneesstack72

Relationships

#3054 Serve audit: WAL replay at startup deletes segments before the store confirms them

Opened by stack72 · 10/5/2026· Shipped 10/6/2026

WalSink.replay (src/serve/audit_sinks/wal_sink.ts) runs at serve startup and, for each WAL segment left from the previous session, calls downstream.write and then deletes the segment straight away. StoreSink.write only appends events to its in-memory batch unless the batch is full, so the segment is deleted before its events are written to any store. If the store put then fails, or serve exits before the batch is flushed, those events are lost from both the WAL and the store. Live delivery no longer has this gap after swamp-club#3039/#3052: a checkpoint deletes a segment only after StoreSink.flush confirms its events. Expected: replay follows the same rule, for example by queuing replayed segments through the normal delivery loop so the next checkpoint confirms them before they are deleted. Found in agent review of the swamp-club#3039 change, which does not change replay.

02Bog Flow
✓OPEN✓TRIAGED✓IN PROGRESS✓SHIPPED+ 1 MOREASSIGNED+ 5 MOREREVIEW+ 10 MOREPR_MERGED+ 2 MORESESSION_SUMMARIZED

Shipped

10/6/2026, 12:46:21 AM

Click a lifecycle step above to view its details.

03Sludge Pulse
stack72 assigned stack7210/6/2026, 12:19:53 AM
Editable. Press Enter to edit.

stack72 commented 10/5/2026, 11:46:43 PM

When this is fixed, check the WalSink paragraph in design/enablers/serve-audit.md (from swamp-club#3039). It says the WAL holds only events the store has not confirmed, which is true for live delivery but not yet for startup replay. Once replayed segments go through the same checkpoint as live ones, that sentence holds everywhere; until then it needs a qualifier.

Sign in to post a ripple.