Skip to main content
← Back to list
01Issue
BugShippedSwamp CLIPublic
Assigneeshammz

Relationships

#2988 Serve audit: emitter breaks the durable hash chain and drops WAL events when a non-durable sink fails or lags

Opened by stack72 · 10/2/2026· Shipped 10/5/2026

Summary

AuditEmitter (src/domain/serve_audit/audit_emitter.ts) has four related defects. Each is triggered when a non-durable sink (webhook, syslog, and soon extension audit sinks from swamp-club#2112) fails repeatedly, lags, or is slow. They are a prerequisite for #2112, because a throwing or hanging extension sink makes them easy to hit.

1. Retries re-chain events, which breaks the durable store's chain

#drain takes a chain snapshot, chains every buffered event from the lowest sink cursor, and restores the snapshot only when no durable sink succeeded. When the WAL/store succeeds but a non-durable sink fails, the next drain chains the lagging sink's events again on top of the advanced chain state. The store's next event then links to the digest of a replayed copy it never stored, so swamp audit verify (verifyChain in audit_query_service.ts) reports a broken chain. The lagging sink also receives different sequence and digest values than the store recorded.

Expected: each event is chained exactly once and the chained form is cached by sequence; replays reuse it.

2. Ring-buffer wrap starves or skips the WAL

#drain reads readFrom(minCursor), but RingBuffer.readFrom clamps the start to oldestSeq once the buffer has wrapped (capacity 10k). #drain still slices each sink by sinkCursor minus minCursor, assuming the first item is minCursor plus one. Once any sink lags by more than capacity, the WAL either gets an empty slice on every drain (total durable loss while the lag persists) or skips a range and still jumps its cursor to throughSeq (permanent gap). Fail-secure mode only checks WAL fullness and dropped events, so it does not notice.

Expected: use the real start sequence from readFrom, clamp each sink cursor to it, and count and log events dropped for the lagging sink only. The durable sink never loses events because of another sink.

3. A slow or hung sink blocks the WAL path

#drain awaits sinks one after another and drains are serialized. A hung write holds every later batch, WAL writes included, for up to the 30s per-sink timeout, and a slow working sink delays every drain, so the ring buffer can overflow before the WAL sees events. The timeout also does not cancel the underlying write, so the next drain can call write() on the same sink while the previous call is still pending.

Expected: durable sinks are written first and never wait on non-durable sinks; a non-durable sink has at most one write in flight.

4. A sink that always throws hot-loops the drain

#drainSerialized re-drains immediately while highSeq is ahead of the lowest cursor. A sink whose write always rejects never advances its cursor, so the emitter re-drains back to back with no delay, re-running HMAC and hashing and logging a warning on every pass.

Expected: per-sink exponential backoff (never for durable sinks), keyed by sink identity, reset by replaceSinks, timer cleared on close.

Tests

  • capacity 4 with an always-failing non-durable sink: the durable sink receives every event in order
  • a lagging non-durable sink plus new events: verifyChain passes over the durable output
  • a hanging non-durable sink does not delay durable writes
  • an always-throwing sink: bounded write count with an injected clock

Found during adversarial review of the plan for swamp-club#2112.

02Bog Flow
✓OPEN✓TRIAGED✓IN PROGRESS✓SHIPPED+ 1 MOREASSIGNED+ 2 MOREREVIEW+ 13 MOREPR_MERGED+ 2 MORESESSION_SUMMARIZED

Shipped

10/5/2026, 7:23:06 PM

Click a lifecycle step above to view its details.

03Sludge Pulse
hammz assigned hammz10/5/2026, 6:32:47 PM

Sign in to post a ripple.