EXTENSIONS
Built by operatives — models, drivers, vaults, and reports, the parts that plug into Swamp.
Filter by what you need and pull what fits.
Gcp/monitoring
Google Cloud monitoring infrastructure models
Gcp/vertex Usage
GCP Vertex AI and Gemini API usage and cost analysis. Reads token and request
Systemd Panel
Enable, disable, and inspect the status of any systemd unit through swamp — one typed, versioned model instance per unit, with enablement, activation, and last-run outcome (parsed from journalctl for timers, from systemctl show for services) all in one place. Not tied to any particular workload or hardware.
Credential Expiry
Probe the credentials a fleet actually holds and report how long each has left, distinguishing expiry from an outage in progress
Ai Usage
Unified cross-provider AI token usage monitoring — workflow, model, and
Aws/kiro Usage
AWS Kiro per-user usage and spend monitoring — queries the Cost and Usage
Aws/metrics
Query and analyze CloudWatch Metrics for operational visibility and performance monitoring.
Aws/bedrock Usage
AWS Bedrock token usage monitoring — multi-account fan-out scanning of
Backrest
Keep a Backrest server's snapshot index current for restic repositories it does not itself back up, and report how fresh each one is. Backrest indexes snapshots only for repositories it runs backups for; one that is merely configured — the normal shape when restic runs from systemd timers on each host and Backrest is only the console — is never indexed at all, and reports no error while doing so. The repository simply stays empty in the UI, which reads as `no backups` for a host whose backups are in fact current. `sync` reads the operation log and reports each repository's newest indexed snapshot. `reindex` triggers Backrest's own TASK_INDEX_SNAPSHOTS for every configured repository, waits for them to appear, and reports the same shape; it reads repositories and never creates, forgets or prunes a snapshot, so it is safe to schedule. Four server behaviours shape the implementation because each one silently produces a wrong answer if ignored. GetOperations' repoId selector does not filter — a selector matching nothing returns the ENTIRE operation log rather than an empty set, so a per-repository query makes every repository report the whole fleet's totals; operations are therefore fetched once and grouped on each operation's own repoId. A failed index task leaves no trace in the operation log, only successes being recorded, so a repository the server cannot read is indistinguishable through the API from one whose task has not run yet — both are reported as `unindexed` with the reason named as the server log rather than invented. The task queue is serial, so triggering N repositories enqueues N tasks behind each other and one whose credentials were revoked does not fail fast but retries with exponential backoff for six minutes or more while everything behind it waits, which is why the wait is a deadline over the whole set rather than a per-task timeout. And a repository with no indexed snapshot at all is a different failure from one whose snapshots have stopped advancing — the first is a credential this server holds that no longer exists, the second is a backup that has stopped running — so they are counted separately as `unindexed` and `stale` instead of being folded into one unhelpful total. Because it reaches repositories through the server's own stored credentials rather than the ones the backup hosts use, disagreement with a host-side view is informative: it means exactly one of the two credential sets has gone bad.
Datadog/synthetics
Datadog Synthetics — synthetic monitoring tests, results, and locations
Azure/openai Usage
Azure OpenAI / AI Services token usage monitoring — multi-subscription
Swamp Watch
Per-workflow observability for a swamp repo's own scheduled work.
Observability Agent
Install and configure a host-native metrics + logs agent on a remote
Cadvisor
cAdvisor container metrics for swamp — deploy a cAdvisor container over SSH,
Statuspage
Manage an Atlassian Statuspage from swamp — component and incident model types wrapping the Statuspage REST API v1. Full CRUD lifecycle plus drift-detecting sync, idempotent create/delete, and rate-limit-aware retries.
Tls Cert Expiry
Standalone TLS certificate expiry checker. Give it hostnames, get back days-until-expiry — no external service dependency.
Zabbix Retention
Zabbix database-growth diagnostics — find what fills the history tables and
Zabbix
Zabbix Monitoring — read-only integration for troubleshooting and observing monitored infrastructure via the Zabbix JSON-RPC 2.0 API. Retrieves hosts, problems, triggers, items, history, host groups, maintenance windows, events, network maps, and template linkage.
Onion
Tor onion-service observability model: bootstrap a persistent Tor daemon, probe onion endpoints, capture sha256-attributed evidence snapshots, crawl dump-inventory listings (filebrowser HTML or enumeration JSON) with PII-bearing filename masking. Network layer of the breach-verification bench.
Speedport Plus 2
Authenticated observation and explicit session management for Arcadyan Speedport Plus 2 routers.
Openobserve
Operate a self-hosted OpenObserve instance from swamp, and decide safely whether a newer upstream release should be rolled out to it. OpenObserve publishes release candidates into the same tag namespace as stable releases -- on 2026-08-29 the newest tag was v1.0.0-rc1, published eleven days AFTER the newest stable v0.92.2 -- so anything that reads the tag list and sorts it pins an RC. That is worse here than for a stateless app: OpenObserve is a log store, and an RC that migrates the on-disk schema is not undone by re-pinning the previous tag, because the previous binary can no longer read what the new one wrote. The `check_update` method reads GitHub *releases* rather than tags, excludes drafts and prereleases (by the GitHub flag OR a semver prerelease suffix, since the flag is hand-set and occasionally wrong), applies real semver precedence so a prerelease sorts before its own release, and reports the newer prereleases it skipped so a pending major stays visible without being auto-applied. It then confirms the candidate tag actually resolves in the registry before offering the update -- a GitHub release and a pushed image are separate events, and proposing a bump whose image does not exist yet fails the deploy at `compose pull`, after the running container has already been stopped. The `health` method probes /healthz, treating a refused connection as a health result rather than a model error so a down instance is distinguishable from a broken check.
Prometheus
Evidence-preserving instant and range PromQL queries against Prometheus, with explicit timestamps and optional SSH transport.
Prometheus Pushgateway
Push metrics from a scheduled swamp method into a Prometheus Pushgateway. Scheduled jobs are exactly what Pushgateway exists for: a batch that runs, computes numbers, and exits long before any scrape could reach it. The `@sntxrr/prometheus/pushgateway` model's `push` method takes a flat `{name, value, labels}` series — the shape several models already emit alongside their verdicts — renders it as text exposition format, and writes it to a grouping key. Defaults to PUT so a series that disappears from the source disappears from the gateway, rather than lingering at its last value forever. Validates metric and label names and rejects non-finite values before sending, because Pushgateway answers 400 without naming the offending series and a single NaN discards the whole batch. Handles the base64 grouping-key escape for values containing a slash or empty values, which would otherwise corrupt or collapse a path segment. Emits an optional heartbeat counter, because pushed metrics go stale silently: the gateway serves the last value forever, so a job that stops running leaves every dashboard green while nothing is being checked.
Unifi Fabric
Structural health monitoring for a UniFi fabric. The `@sntxrr/unifi-fabric/topology` model's `check` method compares a declared topology against live `/stat/device` rows and reports the failures that outcome-based monitoring cannot see: a device expected on the wire that has silently fallen back to a wireless mesh uplink, attachment to the wrong upstream device, links negotiated below their expected speed, ports carrying error counters, and — the one with no equivalent elsewhere — ports that are down but have carried real traffic before, which identifies a run that used to work. An access point that loses its wired uplink does not fail; it meshes, keeps serving clients, and every uptime check stays green while latency quietly goes from sub-millisecond to tens of milliseconds and jittery. `uplink.type` flipping from `wire` to `wireless` is a boolean, so it is asserted exactly rather than thresholded. Read-only: never writes to the controller. Emits a flat Prometheus-ready metric series alongside the verdict, including for healthy devices, so alerts can fire on a series dropping to zero rather than on a document changing shape. Authenticates with an API key over `X-API-KEY`, which sidesteps the HTTP 499 that MFA-enabled SSO accounts return for password logins.