Relationships
#2102 Serve Audit Log - Phase 6: Full documentation
Opened by stack72 · 9/10/2026
Summary
The serve audit log is feature-complete through Phase 5 (#2004, #2028, #2049, #2074, #2103) but the documentation has not kept pace. The in-repo design doc is implementation-focused and incomplete. The swamp skill has no audit routing, no audit guide, and no audit reference. CLI help text for the parent audit command still describes the old local audit timeline, not the serve audit subsystem.
This issue closes every documentation gap. The implementation reference needed to write docs without exploring the codebase is in the comments below.
Documentation gaps
| Location | Current state | Gap |
|---|---|---|
| design/README.md index | serve-audit NOT listed in enablers table | Missing row |
| design/enablers/serve-audit.md | Implementation-focused, incomplete | Missing: event schema, categories, sinks, HMAC, alerts, compliance reports, storage, failure modes, full config |
| design/primitives/serve.md | 3-line bullet at line 383 | Audit is a major subsystem, not a bullet |
| .claude/skills/swamp/SKILL.md | No audit row or commands | Agents cannot route audit questions |
| .claude/skills/swamp/references/serve/guide.md | No audit mention | Operators looking at serve docs cannot find audit |
| .claude/skills/swamp/references/audit/ | Does not exist | No dedicated audit guide or reference |
| swamp audit CLI desc (src/cli/commands/audit.ts) | View audit timeline of swamp vs direct CLI commands | Stale - does not mention log/verify/export/report/rotate-key/alerts |
| Troubleshooting | No serve audit content | WAL, sink connectivity, chain breaks, fail-secure, HMAC |
What to build
1. Expand design/enablers/serve-audit.md
Authoritative reference. Include: event schema (full AuditEvent interface), 7 categories table with actions and security signal, management/data tiers, 4 audit policy levels, chain hashing, HMAC (fields hashed, rotation, opt-out), storage architecture (ring buffer, WAL, store sink, date-partitioned JSONL), durable delivery failure modes, multi-target fan-out, all 5 sinks (store, WAL, WebSocket, webhook, syslog) with config, CEF field mapping, alert rules (config, states, delivery), 5 compliance report templates, full config reference, system events, prior art (Kubernetes, Vault, CloudTrail, Temporal).
2. Update design/README.md
Add serve-audit to enablers index table.
3. Update design/primitives/serve.md
Expand audit from a 3-line bullet to a proper subsection.
4. Create .claude/skills/swamp/references/audit/guide.md
Agent-facing guide: what audit is, enabling it, querying (swamp audit log), verifying (swamp audit verify), bulk export (swamp audit export), live tailing (--follow), configuring sinks, audit policy, HMAC, alert rules, compliance reports (swamp audit report), troubleshooting.
5. Create .claude/skills/swamp/references/audit/reference.md
Deep reference: full event schema, complete config schema with defaults, CEF mapping, syslog mapping, policy evaluation, CLI flag reference for all 7 subcommands, server request types, filter shapes, alert config, compliance report details, troubleshooting recipes.
6. Update .claude/skills/swamp/SKILL.md
Routing table row: Audit - serve audit log, query, verify, export, SIEM, alerts, compliance -> references/audit/guide.md Common Commands: swamp audit log --since 24h --category secrets, swamp audit verify --since 7d, swamp audit export --since 2026-07-01 --until 2026-09-30 --format csv, swamp audit report secret-access --since 30d, swamp audit alerts
7. Update .claude/skills/swamp/references/serve/guide.md
Add audit section or link to dedicated audit guide.
8. Fix swamp audit CLI description
src/cli/commands/audit.ts - change description to cover both local audit record and serve audit subcommands.
9. Troubleshooting content
In audit guide or reference: WAL filling up, webhook errors, syslog circuit breaker, chain breaks, fail-secure lockout, HMAC key missing, empty query results, alerts not firing.
Files to create
- .claude/skills/swamp/references/audit/guide.md
- .claude/skills/swamp/references/audit/reference.md
Files to modify
- design/enablers/serve-audit.md
- design/README.md
- design/primitives/serve.md
- .claude/skills/swamp/SKILL.md
- .claude/skills/swamp/references/serve/guide.md
- src/cli/commands/audit.ts
Verification
- design/enablers/serve-audit.md covers every section
- design/README.md lists serve-audit
- SKILL.md has audit routing row and common commands
- references/audit/guide.md and reference.md exist and are complete
- swamp audit --help shows updated description
- No stale Phase 4 references - everything reflects Phase 5 shipped state
Open
No activity in this phase yet.
stack72 commented 10/2/2026, 8:57:41 PM
Implementation Reference — Part 1: Event Schema and Categories
This comment and the ones below contain the full current implementation details. Use these to write docs without exploring the codebase.
AuditEvent interface (src/domain/serve_audit/audit_event.ts)
type AuditCategory = auth | access | execution | secrets | admin | data | system type AuditStage = request | response type AuditOutcome = success | failure | denied
interface AuditDecision { action: string; // run | read | write | admin resourceKind: string; resourceName: string; effect: allow | deny; grantId: string | null; // UUID of matched grant, or null principalGroups: readonly string[]; // groups at decision time }
interface AuditEvent { id: string; // crypto.randomUUID() timestamp: string; // ISO 8601 UTC instanceId: string; // serve instance ID category: AuditCategory; stage: AuditStage; outcome: AuditOutcome; action: string; // e.g. model.create, vault.read-secret resourceKind: string; resourceName: string; principalKind: string; // user, worker, system, or anonymous principalId: string; initiatedBy: string; // resolved display name sourceIp: string; requestId: string; methodName?: string; detail?: string; version?: number; // schema version (1) sequence?: number; // per-instance monotonic counter digest?: string; // SHA-256 chain hash decision?: AuditDecision; hmacKeyVersion?: number; }
ChainedAuditEvent = AuditEvent with version, sequence, digest required.
Categories table
auth: Login, logout, token create/revoke/rotate — identity lifecycle access: access.grant., access.check, access.can-i, access.reload, access.group. — permission changes, grant evaluations execution: workflow.run, [HOST-1], cancel, resume, approve/reject — workload execution secrets: vault.read-secret, vault.put, vault.delete, vault.annotate — secret access (highest sensitivity) admin: serve.reload, cluster., worker., extension., datastore. — system-level operator changes data: data.get, data.query, data.delete, data.gc, data.prune, model/workflow CRUD — data reads, writes, destructive ops system: instance.start, instance.stop, instance.join, instance.leave, health.transition, alert.fired — infrastructure lifecycle (automatic)
Event tiers
Management: audit.query, audit.verify, audit.subscribe, serve.reload, serve.health, all system events — always on, all sinks Data: everything else — audit store only; external sinks opt-in via tier: all
System events
instance.start — serve instance starts (with version) instance.stop — graceful shutdown instance.join — joins HA cluster instance.leave — leaves HA cluster health.transition — health state change (previous + new state in detail) alert.fired — alert rule threshold breached (rule name, window count)
All system events: principalKind=system, metadata level, management tier.
stack72 commented 10/2/2026, 8:58:12 PM
Implementation Reference — Part 2: Policy, Chain Hashing, HMAC, Storage
Audit policy — four detail levels
none: Event not recorded (health checks, probes) metadata: Who, what, when, outcome — no request/response bodies request: Metadata + request parameters requestResponse: Metadata + request + response bodies
Rules evaluated in order, first match wins. Default level: metadata.
Policy config: audit: policy: default: metadata rules: - actions: [server.version] level: none - actions: [vault.read-secret] level: metadata - categories: [execution] level: requestResponse hmac: false
Chain hashing (src/domain/serve_audit/audit_chain.ts)
Seed: 64 hex zeros (0000...0000) Canonical JSON: strips version/sequence/digest from event, sorts keys deterministically, JSON.stringify digest = SHA-256(previousDigest + canonicalJSON) Per-instance sequence counter increments from 0 verifyChain(events, startDigest?) replays and returns { valid, brokenAt? } Chain state supports snapshot/restore for rollback on durable-sink failure
HMAC (src/domain/serve_audit/audit_hmac.ts)
Fields hashed: resourceName, detail (if present), methodName (if present), decision.resourceName (if present) Algorithm: HMAC-SHA256 Key stored in vault (default: _audit/hmac-key) Each event carries hmacKeyVersion
Key rotation: swamp audit rotate-key generates a new version, stores alongside old versions in vault. Old versions retained indefinitely. HmacKeyRegistry holds all versions (Map of version to CryptoKey). currentContext() returns highest version.
Opt-out: set hmac: false on a policy rule.
Config: audit: hmac: vault: _audit key: hmac-key enabled: true
Storage architecture
Pipeline: Handler -> AuditEmitter (ring buffer, 10,000 capacity) -> Sinks
StoreSink: batched JSONL writes to remote store targets. Defaults: batchSize=100, flushIntervalMs=5000. Storage layout: events/YYYY-MM-DD/uuid.jsonl. Date-partitioned. Retention GC runs hourly, deletes entire date partitions older than retention days.
WAL (src/domain/serve_audit/audit_wal.ts): append-only JSONL segments on local disk. Default max 100MB. Per-target delivery cursors. Replays unflushed segments on restart. WalSink wraps StoreSink: append to WAL first, then try downstream. On flush success, delete delivered WAL segments.
Multi-target fan-out: writes to all configured store targets simultaneously. One target can be marked primary (for query API reads).
Durable delivery failure modes
Backend unreachable: events queue in ring buffer, spill to WAL, drain on reconnection Node crash (normal): events in ring buffer not yet in WAL are lost (one batch interval). WAL replays on restart. Chain hash detects gap. Node crash (fail-secure): WAL fsync before ack. Minimal loss window. Disk loss: entire WAL lost. Events already flushed to remote targets are safe. Chain hash detects gap.
Fail-secure mode (fail-open: false): refuses requests when no store reachable AND WAL is full.
stack72 commented 10/2/2026, 8:58:43 PM
Implementation Reference — Part 3: Sinks, CEF, Alert Rules, Compliance Reports
Sinks
StoreSink (always on, durable=true): batched JSONL to remote stores. Defaults: batchSize=100, flushIntervalMs=5000, gcIntervalMs=1hr. Retention GC deletes entire date partitions.
WalSink (always on, wraps StoreSink): WAL for durable delivery. Append first, then downstream. Replay on restart.
WebSocketSink (always on, durable=false): real-time broadcast to audit.subscribe clients. Per-subscriber filtering: categories, principals, actions, outcomes, resourceKind.
WebhookSink (config-driven, durable=false): HTTP POST in JSON or CEF. Defaults: batchSize=100, batchIntervalMs=5000, maxAttempts=3, backoffMs=1000, maxPending=10. Auth types: bearer, basic, header. Per-sink filtering: categories, tier (management/data/all), outcomes.
SyslogSink (config-driven, durable=false): RFC 5424 structured data. Transport: tcp, tcp+tls, udp. Facility: auth/access/secrets=[REDACTED-SECRET-1] admin=10, execution/data=1, system=3. Severity: denied=4, failure=3, success=6. Circuit breaker for connection failures.
CEF field mapping (src/serve/audit_sinks/cef_formatter.ts)
Format: CEF:0|SwampClub|SwampServe|1.0|action|label|severity|extensions
Severity: secrets=[REDACTED-SECRET-2] admin=7, auth/access=6, execution=5, data/system=3 suser=principalKind:principalId, dhost=resourceKind:resourceName, outcome=outcome, rt=timestamp epoch ms, cs1=namespace, cs2=grantId, cs3=instanceId, src=sourceIp
Action labels defined for: vault.read-secret, vault.put-secret, vault.delete-secret, auth.login, auth.logout, access.grant.create, access.grant.delete, instance.start, instance.stop, instance.join, instance.leave, health.transition
Alert rules (src/domain/serve_audit/audit_alerts.ts)
Config: audit: alerts: - name: brute-force-auth description: Multiple denied auth attempts match: { category: auth, outcome: denied } threshold: { count: 5, window-seconds: 60 } action: { type: log }
Match fields: category, action, outcome, principal (all optional) States: armed (watching), triggered (threshold breached), cooldown (waiting for window to pass) Delivery: type=log emits system audit event (action: alert.fired). type=webhook POSTs JSON to URL (5s timeout, catch-and-warn). Alert events are themselves audit events (category: system) flowing through all sinks. Skips system/alert events to prevent loops.
Compliance reports (src/domain/serve_audit/audit_compliance_reports.ts)
Five templates, each producing markdown + JSON:
access-review: Who has access, grouped by principal, effective permissions, last activity secret-access: vault.read-secret events grouped by principal and key change-history: All mutating actions (create/edit/delete/update/put/remove) denied-access: All denied events grouped by principal, rule, resource system-events: All system-category events (lifecycle, health, HA, alerts)
CLI: swamp audit report [name] --since ISO8601 --until ISO8601 --list
stack72 commented 10/2/2026, 8:59:21 PM
Implementation Reference — Part 4: Config, CLI, Server Requests, Troubleshooting
Full config reference (src/serve/serve_config.ts)
KNOWN_AUDIT_KEYS: stores, sinks, alerts, batch-size, flush-interval, fail-open, policy, wal, hmac
audit: stores: - target: primary type: s3 config: { bucket: acme-audit, region: us-east-1 } retention: { days: 90 } batch-size: 100 flush-interval: 5000 fail-open: true wal: directory: /var/lib/swamp/audit-wal max-size: 104857600 policy: default: metadata rules: - actions: [server.version] level: none - categories: [execution] level: requestResponse hmac: false hmac: vault: _audit key: hmac-key enabled: true sinks: - type: webhook url: https://siem.example.com/api/ingest format: cef auth: { type: bearer, token: vault:integrations/siem-token } batch: { size: 100, interval-ms: 5000 } retry: { max-attempts: 3, backoff-ms: 1000 } filter: { categories: [auth, access, secrets, admin], tier: management } max-pending: 10 - type: syslog host: syslog.example.com port: 514 transport: tcp filter: { tier: management } alerts: - name: brute-force-auth match: { category: auth, outcome: denied } threshold: { count: 5, window-seconds: 60 } action: { type: log }
CLI commands (src/cli/commands/audit.ts)
Registered subcommands: record, alerts, export, log, report, rotate-key, verify
swamp audit log: --since, --until, --principal, --category, --action, --outcome, --cursor, --limit (default 100), --follow swamp audit verify: --since, --until swamp audit export: --since (required), --until (required), --format (json|cef|csv, default json), --output, --principal, --category, --action, --outcome, --resource swamp audit rotate-key: no extra flags (uses --server/--token) swamp audit alerts: no extra flags swamp audit report [name]: --list, --since, --until swamp audit record: pre-existing local CLI audit timeline (separate system)
Server request types (src/serve/protocol.ts)
audit.query: { since?, until?, principal?, category?, action?, outcome?, resource?, limit?, cursor? } -> { events[], cursor?, total? } audit.verify: { since?, until? } -> { valid, eventsChecked, brokenAt?, hmacValid?, hmacChecked?, hmacFailed?, message } audit.subscribe: { categories?, principals?, actions?, outcomes?, resourceKind? } -> { subscriptionId } audit.unsubscribe: (none) audit.export: { from, to, format?, principal?, category?, action?, outcome?, resource? } -> { events/data, format, count, truncated?, streaming?, done? } audit.rotate-key: (none) -> { previousVersion, newVersion, message } audit.alerts: (none) -> { rules: [{ name, description?, state, windowCount, lastFiredAt? }] } audit.report: { name, from, to } -> { name, description, from, to, generatedAt, markdown, data }
Troubleshooting recipes
WAL filling up: remote store unreachable. Check store connectivity and config. WAL max default 100MB. Webhook errors: auth failure (check token resolution), connectivity (check URL), format mismatch. Syslog circuit breaker: connection refused/timeout. Check host:port, transport (tcp/udp), firewall. Chain breaks: caused by crash or WAL loss. brokenAt reports the sequence number. Chain detects gaps, does not prevent them. Fail-secure rejecting: all stores unreachable AND WAL full. Fix store connectivity first. HMAC key missing: vault not configured or key not created. Run swamp audit rotate-key to create initial key. Empty query results: check time range (max 90 days, max 50000 events loaded), filter mismatch, retention may have expired. Alerts not firing: check match pattern matches event fields exactly, threshold window, cooldown state (swamp audit alerts shows state).
Prior art to document
Kubernetes: audit policy with four detail levels, first-match rule evaluation HashiCorp Vault: fail-secure mode, HMAC-SHA256 of sensitive fields, per-device salt AWS CloudTrail: management vs data event split, hourly digest files, organization-wide trail Temporal Cloud: push-based sink model (events stream to customer infrastructure)
Sign in to post a ripple.