Skip to main content

EXTENSIONS

Built by operatives — models, drivers, vaults, and reports, the parts that plug into Swamp.

Filter by what you need and pull what fits.

Selection
36 results
label:monitoring

Gcp/monitoring

@swamp/gcp/monitoring · v2026.10.06.1

Google Cloud monitoring infrastructure models

upd Oct 620 pullsA100/100

Gcp/vertex Usage

@webframp/gcp/vertex-usage · v2026.10.01.1

GCP Vertex AI and Gemini API usage and cost analysis. Reads token and request

upd Oct 226 pullsA100/100

Systemd Panel

@aaronge/systemd-panel · v2026.10.01.3

Enable, disable, and inspect the status of any systemd unit through swamp — one typed, versioned model instance per unit, with enablement, activation, and last-run outcome (parsed from journalctl for timers, from systemctl show for services) all in one place. Not tied to any particular workload or hardware.

upd Oct 124 pullsA100/100

Credential Expiry

@sntxrr/credential-expiry · v2026.09.30.1

Probe the credentials a fleet actually holds and report how long each has left, distinguishing expiry from an outage in progress

upd Sep 3048 pullsA100/100

Ai Usage

@webframp/ai-usage · v2026.09.25.1

Unified cross-provider AI token usage monitoring — workflow, model, and

upd Sep 2524 pullsA100/100

Aws/kiro Usage

@webframp/aws/kiro-usage · v2026.09.24.1

AWS Kiro per-user usage and spend monitoring — queries the Cost and Usage

upd Sep 2521 pullsA100/100

Aws/metrics

@webframp/aws/metrics · v2026.09.24.2

Query and analyze CloudWatch Metrics for operational visibility and performance monitoring.

upd Sep 2567 pullsA100/100

Aws/bedrock Usage

@webframp/aws/bedrock-usage · v2026.09.24.1

AWS Bedrock token usage monitoring — multi-account fan-out scanning of

upd Sep 2530 pullsA100/100

Backrest

@sntxrr/backrest · v2026.09.23.2

Keep a Backrest server's snapshot index current for restic repositories it does not itself back up, and report how fresh each one is. Backrest indexes snapshots only for repositories it runs backups for; one that is merely configured — the normal shape when restic runs from systemd timers on each host and Backrest is only the console — is never indexed at all, and reports no error while doing so. The repository simply stays empty in the UI, which reads as `no backups` for a host whose backups are in fact current. `sync` reads the operation log and reports each repository's newest indexed snapshot. `reindex` triggers Backrest's own TASK_INDEX_SNAPSHOTS for every configured repository, waits for them to appear, and reports the same shape; it reads repositories and never creates, forgets or prunes a snapshot, so it is safe to schedule. Four server behaviours shape the implementation because each one silently produces a wrong answer if ignored. GetOperations' repoId selector does not filter — a selector matching nothing returns the ENTIRE operation log rather than an empty set, so a per-repository query makes every repository report the whole fleet's totals; operations are therefore fetched once and grouped on each operation's own repoId. A failed index task leaves no trace in the operation log, only successes being recorded, so a repository the server cannot read is indistinguishable through the API from one whose task has not run yet — both are reported as `unindexed` with the reason named as the server log rather than invented. The task queue is serial, so triggering N repositories enqueues N tasks behind each other and one whose credentials were revoked does not fail fast but retries with exponential backoff for six minutes or more while everything behind it waits, which is why the wait is a deadline over the whole set rather than a per-task timeout. And a repository with no indexed snapshot at all is a different failure from one whose snapshots have stopped advancing — the first is a credential this server holds that no longer exists, the second is a backup that has stopped running — so they are counted separately as `unindexed` and `stale` instead of being folded into one unhelpful total. Because it reaches repositories through the server's own stored credentials rather than the ones the backup hosts use, disagreement with a host-side view is informative: it means exactly one of the two credential sets has gone bad.

upd Sep 2338 pullsA100/100

Datadog/synthetics

@webframp/datadog/synthetics · v2026.09.18.1

Datadog Synthetics — synthetic monitoring tests, results, and locations

upd Sep 2015 pullsA100/100

Azure/openai Usage

@webframp/azure/openai-usage · v2026.09.18.1

Azure OpenAI / AI Services token usage monitoring — multi-subscription

upd Sep 2031 pullsA100/100

Swamp Watch

@magistr/swamp-watch · v2026.09.19.2

Per-workflow observability for a swamp repo's own scheduled work.

upd Sep 1917 pullsA100/100

Observability Agent

@magistr/observability-agent · v2026.09.19.2

Install and configure a host-native metrics + logs agent on a remote

upd Sep 1914 pullsA100/100

Cadvisor

@magistr/cadvisor · v2026.09.19.2

cAdvisor container metrics for swamp — deploy a cAdvisor container over SSH,

upd Sep 1915 pullsA100/100

Statuspage

@hmcrum/statuspage · v2026.09.17.1

Manage an Atlassian Statuspage from swamp — component and incident model types wrapping the Statuspage REST API v1. Full CRUD lifecycle plus drift-detecting sync, idempotent create/delete, and rate-limit-aware retries.

upd Sep 1814 pullsB85/100

Tls Cert Expiry

@hmcrum/tls-cert-expiry · v2026.09.17.1

Standalone TLS certificate expiry checker. Give it hostnames, get back days-until-expiry — no external service dependency.

upd Sep 1714 pullsA100/100

Zabbix Retention

@acameron17/zabbix-retention · v2026.09.15.1

Zabbix database-growth diagnostics — find what fills the history tables and

upd Sep 1515 pullsA100/100

Zabbix

@figura/zabbix · v2026.09.10.1

Zabbix Monitoring — read-only integration for troubleshooting and observing monitored infrastructure via the Zabbix JSON-RPC 2.0 API. Retrieves hosts, problems, triggers, items, history, host groups, maintenance windows, events, network maps, and template linkage.

upd Sep 1044 pullsA100/100

Onion

@zocc/onion · v2026.09.10.1

Tor onion-service observability model: bootstrap a persistent Tor daemon, probe onion endpoints, capture sha256-attributed evidence snapshots, crawl dump-inventory listings (filebrowser HTML or enumeration JSON) with PII-bearing filename masking. Network layer of the breach-verification bench.

upd Sep 1015 pullsA92/100

Speedport Plus 2

@dieter/speedport-plus-2 · v2026.09.07.4

Authenticated observation and explicit session management for Arcadyan Speedport Plus 2 routers.

upd Sep 714 pullsA100/100

Openobserve

@sntxrr/openobserve · v2026.08.29.1

Operate a self-hosted OpenObserve instance from swamp, and decide safely whether a newer upstream release should be rolled out to it. OpenObserve publishes release candidates into the same tag namespace as stable releases -- on 2026-08-29 the newest tag was v1.0.0-rc1, published eleven days AFTER the newest stable v0.92.2 -- so anything that reads the tag list and sorts it pins an RC. That is worse here than for a stateless app: OpenObserve is a log store, and an RC that migrates the on-disk schema is not undone by re-pinning the previous tag, because the previous binary can no longer read what the new one wrote. The `check_update` method reads GitHub *releases* rather than tags, excludes drafts and prereleases (by the GitHub flag OR a semver prerelease suffix, since the flag is hand-set and occasionally wrong), applies real semver precedence so a prerelease sorts before its own release, and reports the newer prereleases it skipped so a pending major stays visible without being auto-applied. It then confirms the candidate tag actually resolves in the registry before offering the update -- a GitHub release and a pushed image are separate events, and proposing a bump whose image does not exist yet fails the deploy at `compose pull`, after the running container has already been stopped. The `health` method probes /healthz, treating a refused connection as a health result rather than a model error so a down instance is distinguishable from a broken check.

upd Aug 2932 pullsA100/100

Prometheus

@dieter/prometheus · v2026.08.22.2

Evidence-preserving instant and range PromQL queries against Prometheus, with explicit timestamps and optional SSH transport.

upd Aug 2220 pullsA100/100

Prometheus Pushgateway

@sntxrr/prometheus-pushgateway · v2026.08.20.1

Push metrics from a scheduled swamp method into a Prometheus Pushgateway. Scheduled jobs are exactly what Pushgateway exists for: a batch that runs, computes numbers, and exits long before any scrape could reach it. The `@sntxrr/prometheus/pushgateway` model's `push` method takes a flat `{name, value, labels}` series — the shape several models already emit alongside their verdicts — renders it as text exposition format, and writes it to a grouping key. Defaults to PUT so a series that disappears from the source disappears from the gateway, rather than lingering at its last value forever. Validates metric and label names and rejects non-finite values before sending, because Pushgateway answers 400 without naming the offending series and a single NaN discards the whole batch. Handles the base64 grouping-key escape for values containing a slash or empty values, which would otherwise corrupt or collapse a path segment. Emits an optional heartbeat counter, because pushed metrics go stale silently: the gateway serves the last value forever, so a job that stops running leaves every dashboard green while nothing is being checked.

upd Aug 2014 pullsA100/100

Unifi Fabric

@sntxrr/unifi-fabric · v2026.08.20.1

Structural health monitoring for a UniFi fabric. The `@sntxrr/unifi-fabric/topology` model's `check` method compares a declared topology against live `/stat/device` rows and reports the failures that outcome-based monitoring cannot see: a device expected on the wire that has silently fallen back to a wireless mesh uplink, attachment to the wrong upstream device, links negotiated below their expected speed, ports carrying error counters, and — the one with no equivalent elsewhere — ports that are down but have carried real traffic before, which identifies a run that used to work. An access point that loses its wired uplink does not fail; it meshes, keeps serving clients, and every uptime check stays green while latency quietly goes from sub-millisecond to tens of milliseconds and jittery. `uplink.type` flipping from `wire` to `wireless` is a boolean, so it is asserted exactly rather than thresholded. Read-only: never writes to the controller. Emits a flat Prometheus-ready metric series alongside the verdict, including for healthy devices, so alerts can fire on a series dropping to zero rather than on a document changing shape. Authenticates with an API key over `X-API-KEY`, which sidesteps the HTTP 499 that MFA-enabled SSO accounts return for password logins.

upd Aug 2014 pullsA100/100