Skip to main content
← Back to list
01Issue
FeatureOpenSwamp ClubPublic
AssigneesNone

Relationships

#2102 Serve Audit Log - Phase 6: Full documentation

Opened by stack72 · 9/10/2026

Summary

The serve audit log is feature-complete through Phase 5 (#2004, #2028, #2049, #2074, #2103) but the documentation has not kept pace. The in-repo design doc is implementation-focused and incomplete. The swamp skill has no audit routing, no audit guide, and no audit reference. CLI help text for the parent audit command still describes the old local audit timeline, not the serve audit subsystem.

This issue closes every documentation gap. The implementation reference needed to write docs without exploring the codebase is in the comments below.

Documentation gaps

Location Current state Gap
design/README.md index serve-audit NOT listed in enablers table Missing row
design/enablers/serve-audit.md Implementation-focused, incomplete Missing: event schema, categories, sinks, HMAC, alerts, compliance reports, storage, failure modes, full config
design/primitives/serve.md 3-line bullet at line 383 Audit is a major subsystem, not a bullet
.claude/skills/swamp/SKILL.md No audit row or commands Agents cannot route audit questions
.claude/skills/swamp/references/serve/guide.md No audit mention Operators looking at serve docs cannot find audit
.claude/skills/swamp/references/audit/ Does not exist No dedicated audit guide or reference
swamp audit CLI desc (src/cli/commands/audit.ts) View audit timeline of swamp vs direct CLI commands Stale - does not mention log/verify/export/report/rotate-key/alerts
Troubleshooting No serve audit content WAL, sink connectivity, chain breaks, fail-secure, HMAC

What to build

1. Expand design/enablers/serve-audit.md

Authoritative reference. Include: event schema (full AuditEvent interface), 7 categories table with actions and security signal, management/data tiers, 4 audit policy levels, chain hashing, HMAC (fields hashed, rotation, opt-out), storage architecture (ring buffer, WAL, store sink, date-partitioned JSONL), durable delivery failure modes, multi-target fan-out, all 5 sinks (store, WAL, WebSocket, webhook, syslog) with config, CEF field mapping, alert rules (config, states, delivery), 5 compliance report templates, full config reference, system events, prior art (Kubernetes, Vault, CloudTrail, Temporal).

2. Update design/README.md

Add serve-audit to enablers index table.

3. Update design/primitives/serve.md

Expand audit from a 3-line bullet to a proper subsection.

4. Create .claude/skills/swamp/references/audit/guide.md

Agent-facing guide: what audit is, enabling it, querying (swamp audit log), verifying (swamp audit verify), bulk export (swamp audit export), live tailing (--follow), configuring sinks, audit policy, HMAC, alert rules, compliance reports (swamp audit report), troubleshooting.

5. Create .claude/skills/swamp/references/audit/reference.md

Deep reference: full event schema, complete config schema with defaults, CEF mapping, syslog mapping, policy evaluation, CLI flag reference for all 7 subcommands, server request types, filter shapes, alert config, compliance report details, troubleshooting recipes.

6. Update .claude/skills/swamp/SKILL.md

Routing table row: Audit - serve audit log, query, verify, export, SIEM, alerts, compliance -> references/audit/guide.md Common Commands: swamp audit log --since 24h --category secrets, swamp audit verify --since 7d, swamp audit export --since 2026-07-01 --until 2026-09-30 --format csv, swamp audit report secret-access --since 30d, swamp audit alerts

7. Update .claude/skills/swamp/references/serve/guide.md

Add audit section or link to dedicated audit guide.

8. Fix swamp audit CLI description

src/cli/commands/audit.ts - change description to cover both local audit record and serve audit subcommands.

9. Troubleshooting content

In audit guide or reference: WAL filling up, webhook errors, syslog circuit breaker, chain breaks, fail-secure lockout, HMAC key missing, empty query results, alerts not firing.

Files to create

  • .claude/skills/swamp/references/audit/guide.md
  • .claude/skills/swamp/references/audit/reference.md

Files to modify

  • design/enablers/serve-audit.md
  • design/README.md
  • design/primitives/serve.md
  • .claude/skills/swamp/SKILL.md
  • .claude/skills/swamp/references/serve/guide.md
  • src/cli/commands/audit.ts

Verification

  • design/enablers/serve-audit.md covers every section
  • design/README.md lists serve-audit
  • SKILL.md has audit routing row and common commands
  • references/audit/guide.md and reference.md exist and are complete
  • swamp audit --help shows updated description
  • No stale Phase 4 references - everything reflects Phase 5 shipped state
02Bog Flow
◉OPEN○TRIAGED○IN PROGRESS○SHIPPED

Open

9/10/2026, 6:00:00 PM

No activity in this phase yet.

03Sludge Pulse
Editable. Press Enter to edit.

stack72 commented 10/2/2026, 8:57:41 PM

Implementation Reference — Part 1: Event Schema and Categories

This comment and the ones below contain the full current implementation details. Use these to write docs without exploring the codebase.

AuditEvent interface (src/domain/serve_audit/audit_event.ts)

type AuditCategory = auth | access | execution | secrets | admin | data | system type AuditStage = request | response type AuditOutcome = success | failure | denied

interface AuditDecision { action: string; // run | read | write | admin resourceKind: string; resourceName: string; effect: allow | deny; grantId: string | null; // UUID of matched grant, or null principalGroups: readonly string[]; // groups at decision time }

interface AuditEvent { id: string; // crypto.randomUUID() timestamp: string; // ISO 8601 UTC instanceId: string; // serve instance ID category: AuditCategory; stage: AuditStage; outcome: AuditOutcome; action: string; // e.g. model.create, vault.read-secret resourceKind: string; resourceName: string; principalKind: string; // user, worker, system, or anonymous principalId: string; initiatedBy: string; // resolved display name sourceIp: string; requestId: string; methodName?: string; detail?: string; version?: number; // schema version (1) sequence?: number; // per-instance monotonic counter digest?: string; // SHA-256 chain hash decision?: AuditDecision; hmacKeyVersion?: number; }

ChainedAuditEvent = AuditEvent with version, sequence, digest required.

Categories table

auth: Login, logout, token create/revoke/rotate — identity lifecycle access: access.grant., access.check, access.can-i, access.reload, access.group. — permission changes, grant evaluations execution: workflow.run, [HOST-1], cancel, resume, approve/reject — workload execution secrets: vault.read-secret, vault.put, vault.delete, vault.annotate — secret access (highest sensitivity) admin: serve.reload, cluster., worker., extension., datastore. — system-level operator changes data: data.get, data.query, data.delete, data.gc, data.prune, model/workflow CRUD — data reads, writes, destructive ops system: instance.start, instance.stop, instance.join, instance.leave, health.transition, alert.fired — infrastructure lifecycle (automatic)

Event tiers

Management: audit.query, audit.verify, audit.subscribe, serve.reload, serve.health, all system events — always on, all sinks Data: everything else — audit store only; external sinks opt-in via tier: all

System events

instance.start — serve instance starts (with version) instance.stop — graceful shutdown instance.join — joins HA cluster instance.leave — leaves HA cluster health.transition — health state change (previous + new state in detail) alert.fired — alert rule threshold breached (rule name, window count)

All system events: principalKind=system, metadata level, management tier.

stack72 commented 10/2/2026, 8:58:12 PM

Implementation Reference — Part 2: Policy, Chain Hashing, HMAC, Storage

Audit policy — four detail levels

none: Event not recorded (health checks, probes) metadata: Who, what, when, outcome — no request/response bodies request: Metadata + request parameters requestResponse: Metadata + request + response bodies

Rules evaluated in order, first match wins. Default level: metadata.

Policy config: audit: policy: default: metadata rules: - actions: [server.version] level: none - actions: [vault.read-secret] level: metadata - categories: [execution] level: requestResponse hmac: false

Chain hashing (src/domain/serve_audit/audit_chain.ts)

Seed: 64 hex zeros (0000...0000) Canonical JSON: strips version/sequence/digest from event, sorts keys deterministically, JSON.stringify digest = SHA-256(previousDigest + canonicalJSON) Per-instance sequence counter increments from 0 verifyChain(events, startDigest?) replays and returns { valid, brokenAt? } Chain state supports snapshot/restore for rollback on durable-sink failure

HMAC (src/domain/serve_audit/audit_hmac.ts)

Fields hashed: resourceName, detail (if present), methodName (if present), decision.resourceName (if present) Algorithm: HMAC-SHA256 Key stored in vault (default: _audit/hmac-key) Each event carries hmacKeyVersion

Key rotation: swamp audit rotate-key generates a new version, stores alongside old versions in vault. Old versions retained indefinitely. HmacKeyRegistry holds all versions (Map of version to CryptoKey). currentContext() returns highest version.

Opt-out: set hmac: false on a policy rule.

Config: audit: hmac: vault: _audit key: hmac-key enabled: true

Storage architecture

Pipeline: Handler -> AuditEmitter (ring buffer, 10,000 capacity) -> Sinks

StoreSink: batched JSONL writes to remote store targets. Defaults: batchSize=100, flushIntervalMs=5000. Storage layout: events/YYYY-MM-DD/uuid.jsonl. Date-partitioned. Retention GC runs hourly, deletes entire date partitions older than retention days.

WAL (src/domain/serve_audit/audit_wal.ts): append-only JSONL segments on local disk. Default max 100MB. Per-target delivery cursors. Replays unflushed segments on restart. WalSink wraps StoreSink: append to WAL first, then try downstream. On flush success, delete delivered WAL segments.

Multi-target fan-out: writes to all configured store targets simultaneously. One target can be marked primary (for query API reads).

Durable delivery failure modes

Backend unreachable: events queue in ring buffer, spill to WAL, drain on reconnection Node crash (normal): events in ring buffer not yet in WAL are lost (one batch interval). WAL replays on restart. Chain hash detects gap. Node crash (fail-secure): WAL fsync before ack. Minimal loss window. Disk loss: entire WAL lost. Events already flushed to remote targets are safe. Chain hash detects gap.

Fail-secure mode (fail-open: false): refuses requests when no store reachable AND WAL is full.

stack72 commented 10/2/2026, 8:58:43 PM

Implementation Reference — Part 3: Sinks, CEF, Alert Rules, Compliance Reports

Sinks

StoreSink (always on, durable=true): batched JSONL to remote stores. Defaults: batchSize=100, flushIntervalMs=5000, gcIntervalMs=1hr. Retention GC deletes entire date partitions.

WalSink (always on, wraps StoreSink): WAL for durable delivery. Append first, then downstream. Replay on restart.

WebSocketSink (always on, durable=false): real-time broadcast to audit.subscribe clients. Per-subscriber filtering: categories, principals, actions, outcomes, resourceKind.

WebhookSink (config-driven, durable=false): HTTP POST in JSON or CEF. Defaults: batchSize=100, batchIntervalMs=5000, maxAttempts=3, backoffMs=1000, maxPending=10. Auth types: bearer, basic, header. Per-sink filtering: categories, tier (management/data/all), outcomes.

SyslogSink (config-driven, durable=false): RFC 5424 structured data. Transport: tcp, tcp+tls, udp. Facility: auth/access/secrets=[REDACTED-SECRET-1] admin=10, execution/data=1, system=3. Severity: denied=4, failure=3, success=6. Circuit breaker for connection failures.

CEF field mapping (src/serve/audit_sinks/cef_formatter.ts)

Format: CEF:0|SwampClub|SwampServe|1.0|action|label|severity|extensions

Severity: secrets=[REDACTED-SECRET-2] admin=7, auth/access=6, execution=5, data/system=3 suser=principalKind:principalId, dhost=resourceKind:resourceName, outcome=outcome, rt=timestamp epoch ms, cs1=namespace, cs2=grantId, cs3=instanceId, src=sourceIp

Action labels defined for: vault.read-secret, vault.put-secret, vault.delete-secret, auth.login, auth.logout, access.grant.create, access.grant.delete, instance.start, instance.stop, instance.join, instance.leave, health.transition

Alert rules (src/domain/serve_audit/audit_alerts.ts)

Config: audit: alerts: - name: brute-force-auth description: Multiple denied auth attempts match: { category: auth, outcome: denied } threshold: { count: 5, window-seconds: 60 } action: { type: log }

Match fields: category, action, outcome, principal (all optional) States: armed (watching), triggered (threshold breached), cooldown (waiting for window to pass) Delivery: type=log emits system audit event (action: alert.fired). type=webhook POSTs JSON to URL (5s timeout, catch-and-warn). Alert events are themselves audit events (category: system) flowing through all sinks. Skips system/alert events to prevent loops.

Compliance reports (src/domain/serve_audit/audit_compliance_reports.ts)

Five templates, each producing markdown + JSON:

access-review: Who has access, grouped by principal, effective permissions, last activity secret-access: vault.read-secret events grouped by principal and key change-history: All mutating actions (create/edit/delete/update/put/remove) denied-access: All denied events grouped by principal, rule, resource system-events: All system-category events (lifecycle, health, HA, alerts)

CLI: swamp audit report [name] --since ISO8601 --until ISO8601 --list

stack72 commented 10/2/2026, 8:59:21 PM

Implementation Reference — Part 4: Config, CLI, Server Requests, Troubleshooting

Full config reference (src/serve/serve_config.ts)

KNOWN_AUDIT_KEYS: stores, sinks, alerts, batch-size, flush-interval, fail-open, policy, wal, hmac

audit: stores: - target: primary type: s3 config: { bucket: acme-audit, region: us-east-1 } retention: { days: 90 } batch-size: 100 flush-interval: 5000 fail-open: true wal: directory: /var/lib/swamp/audit-wal max-size: 104857600 policy: default: metadata rules: - actions: [server.version] level: none - categories: [execution] level: requestResponse hmac: false hmac: vault: _audit key: hmac-key enabled: true sinks: - type: webhook url: https://siem.example.com/api/ingest format: cef auth: { type: bearer, token: vault:integrations/siem-token } batch: { size: 100, interval-ms: 5000 } retry: { max-attempts: 3, backoff-ms: 1000 } filter: { categories: [auth, access, secrets, admin], tier: management } max-pending: 10 - type: syslog host: syslog.example.com port: 514 transport: tcp filter: { tier: management } alerts: - name: brute-force-auth match: { category: auth, outcome: denied } threshold: { count: 5, window-seconds: 60 } action: { type: log }

CLI commands (src/cli/commands/audit.ts)

Registered subcommands: record, alerts, export, log, report, rotate-key, verify

swamp audit log: --since, --until, --principal, --category, --action, --outcome, --cursor, --limit (default 100), --follow swamp audit verify: --since, --until swamp audit export: --since (required), --until (required), --format (json|cef|csv, default json), --output, --principal, --category, --action, --outcome, --resource swamp audit rotate-key: no extra flags (uses --server/--token) swamp audit alerts: no extra flags swamp audit report [name]: --list, --since, --until swamp audit record: pre-existing local CLI audit timeline (separate system)

Server request types (src/serve/protocol.ts)

audit.query: { since?, until?, principal?, category?, action?, outcome?, resource?, limit?, cursor? } -> { events[], cursor?, total? } audit.verify: { since?, until? } -> { valid, eventsChecked, brokenAt?, hmacValid?, hmacChecked?, hmacFailed?, message } audit.subscribe: { categories?, principals?, actions?, outcomes?, resourceKind? } -> { subscriptionId } audit.unsubscribe: (none) audit.export: { from, to, format?, principal?, category?, action?, outcome?, resource? } -> { events/data, format, count, truncated?, streaming?, done? } audit.rotate-key: (none) -> { previousVersion, newVersion, message } audit.alerts: (none) -> { rules: [{ name, description?, state, windowCount, lastFiredAt? }] } audit.report: { name, from, to } -> { name, description, from, to, generatedAt, markdown, data }

Troubleshooting recipes

WAL filling up: remote store unreachable. Check store connectivity and config. WAL max default 100MB. Webhook errors: auth failure (check token resolution), connectivity (check URL), format mismatch. Syslog circuit breaker: connection refused/timeout. Check host:port, transport (tcp/udp), firewall. Chain breaks: caused by crash or WAL loss. brokenAt reports the sequence number. Chain detects gaps, does not prevent them. Fail-secure rejecting: all stores unreachable AND WAL full. Fix store connectivity first. HMAC key missing: vault not configured or key not created. Run swamp audit rotate-key to create initial key. Empty query results: check time range (max 90 days, max 50000 events loaded), filter mismatch, retention may have expired. Alerts not firing: check match pattern matches event fields exactly, threshold window, cooldown state (swamp audit alerts shows state).

Prior art to document

Kubernetes: audit policy with four detail levels, first-match rule evaluation HashiCorp Vault: fail-secure mode, HMAC-SHA256 of sensitive fields, per-device salt AWS CloudTrail: management vs data event split, hourly digest files, organization-wide trail Temporal Cloud: push-based sink model (events stream to customer infrastructure)

Sign in to post a ripple.