Skip to main content

Gcp/vertex Usage

@webframp/gcp/vertex-usagev2026.10.01.1· 5d agoMODELS
01README

GCP Vertex AI and Gemini API usage and cost analysis. Reads token and request volume from Cloud Monitoring and dollar cost from the Cloud Billing export in BigQuery, over the same complete UTC days (standard and FOCUS exports are both supported). Projects are listed explicitly or discovered at runtime through Cloud Resource Manager. Results are one resource instance per project, filterable with CEL.

Authentication

Service account JSON key (serviceAccountJson arg), a GCP_ACCESS_TOKEN env var, or the file named by GOOGLE_APPLICATION_CREDENTIALS (service account or authorized_user). No gcloud CLI dependency.

Required Permissions

  • monitoring.timeSeries.list on each project (Monitoring Viewer)
  • resourcemanager.projects.get for project discovery
  • BigQuery Data Viewer on the export dataset and BigQuery Job User for the billing methods

Usage

swamp model create @webframp/gcp/vertex-usage vertex-usage \
  --global-arg 'serviceAccountJson=<vault:gcp/sa-key>' \
  --global-arg 'billingTable=my-billing.billing.gcp_billing_export_v1_XXXXXX'

# Daily token/request usage for every discovered project
swamp model method run vertex-usage scan_usage

# Confirm which billing services to filter on, then pull cost
swamp model method run vertex-usage discover_billing_services
swamp model method run vertex-usage get_billing_costs

Methods

  • discover_projects — List ACTIVE projects via Cloud Resource Manager
  • scan_usage — Daily rows per project/location/publisher/model/request type, with per-project status so failures are never mistaken for zero usage
  • scan_gemini_api_usage — Gemini API (generativelanguage) daily output tokens and request counts by method and API key
  • discover_billing_services — AI-related billing services and SKUs with cost
  • get_billing_costs — Gross cost, credits and net cost per day, project and SKU
  • scan_projects, get_token_usage — Legacy single-bucket scans
02Release Notes

2026.10.01.1

Fixed: scan_projects set no flag when a project failed to scan, so a partial scan looked complete. A failed project now sets truncated: true. Monitoring requests that hit a 429 or 5xx are retried with backoff instead of failing the project on the first response. billing_summary.complete is now false when no billing rows match the filters, so a wrong service name no longer reads as a valid empty result.

Added: Gemini API coverage and cost analysis alongside Vertex AI.

  • discover_projects, and project discovery in every scan, find ACTIVE projects through Cloud Resource Manager. projects is now optional.
  • scan_usage writes daily token and request rows per project, location, publisher, model and request type, with error and 429 counts, as one usage-<project> instance per project. scan_summary records every project as ok, no_data or error, and complete is false if any project failed.
  • scan_gemini_api_usage does the same for the Gemini API: daily output tokens by model, and API-wide request counts by method, response code and API key. Monitoring has no input-token metric for this API, so input volume comes from billing. gemini_usage.requestsAvailable is false when request counts could not be read, and the scan's complete follows it.
  • BigQuery queries fail with the job's error when the response carries an errors array, instead of returning zero rows.
  • Daily Monitoring points are assigned to a day by the midpoint of their interval, so an inclusive or missing endTime cannot shift a point to the neighbouring day.
  • discover_billing_services and get_billing_costs read the Cloud Billing export in BigQuery. Gross cost, credits and net cost are reported per day, project and SKU, never mixing currencies or cost types. Both the standard and the FOCUS export are supported (billingSchema, auto-detected). The standard export path is covered by unit tests only.
  • Authentication also accepts GCP_ACCESS_TOKEN and authorized_user files named by GOOGLE_APPLICATION_CREDENTIALS.
  • New optional global arguments: billingTable, billingSchema, billingQueryProject.

Changed: scan_projects and get_token_usage keep their output schemas. scan_projects now discovers projects when none are configured. Scans pace Monitoring calls with maxRequestsPerMinute (default 120) to stay under the default quota of 180 requests per minute per user. A GCP_ACCESS_TOKEN environment variable now takes precedence over GOOGLE_APPLICATION_CREDENTIALS when both are set, so unset a stale token if a host exports one for another tool. maxRequestsPerMinute paces first attempts only: a retry no longer takes a limiter slot, but retries still count against Google's quota, so leave headroom below the real limit. A BigQuery response with no job id now fails with its own message instead of reporting a timeout or a page-cap error.

Upgrade note: Existing definitions keep working without changes, and @webframp/ai-usage needs no co-upgrade. The billing methods need billingTable, BigQuery Data Viewer on the export dataset and BigQuery Job User on the project that runs the query. Project discovery needs resourcemanager.projects.get.

03Models1
@webframp/gcp/vertex-usagev2026.10.01.1gcp/vertex_usage.ts
fn discover_projects()
List ACTIVE GCP projects visible to the credential via Cloud Resource Manager. Needs resourcemanager.projects.get; grant it at the organization or folder level.
fn scan_usage(days: number, concurrency: number)
Fan-out scan of Vertex AI token and request usage as flat daily rows per project, location, publisher, model and request type. Writes one usage-<project> instance per project plus a scan_summary with per-project status. Covers the last N complete UTC days; Cloud Monitoring retains these metrics for a limited time, so use get_billing_costs for older history.
ArgumentTypeDescription
daysnumberNumber of complete UTC days to scan, ending at today's UTC midnight
concurrencynumberProjects scanned in parallel
fn scan_gemini_api_usage(days: number, concurrency: number)
Fan-out scan of Gemini API (generativelanguage.googleapis.com, the Gemini Developer API) usage as daily rows. Cloud Monitoring exposes output tokens only: there is no input-token metric, so use get_billing_costs for input volume. Request counts come from the API-wide request_count metric and include credentialId (which API key made the call) and every API method. Writes one gemini-usage-<project> instance per project plus a gemini_scan_summary with per-project status.
ArgumentTypeDescription
daysnumberNumber of complete UTC days to scan, ending at today's UTC midnight
concurrencynumberProjects scanned in parallel
fn discover_billing_services(days: number, pattern: string)
List billing services and SKUs in the Cloud Billing export that match an AI-related pattern, with cost. Run this first to confirm which service names to pass to get_billing_costs, since partner models can bill under a different service than Google's own.
ArgumentTypeDescription
daysnumber
patternstringRE2 regex matched against 'service.description sku.description'
fn get_billing_costs(days: number, services: array, skuPattern?: string)
Query the Cloud Billing export in BigQuery for Vertex AI cost over the last N complete UTC days. Writes one billing-<project> instance per project (gross cost, credits and net cost per day and SKU, never mixing currencies or cost types) plus a billing_summary that flags an export that has not caught up to the window end.
ArgumentTypeDescription
daysnumberNumber of complete UTC days, ending at today's UTC midnight
servicesarrayExact service.description values to include. Confirm with discover_billing_services.
skuPattern?stringOptional RE2 regex on sku.description to narrow results
fn scan_projects(days: number)
Legacy: fan-out scan returning one total per model per project, with no daily detail. Prefer scan_usage. Projects come from the model's projects global argument or runtime discovery.
ArgumentTypeDescription
daysnumberLookback period in days
fn get_token_usage(project: string, days: number)
Legacy: token usage for a single GCP project with a per-model total. Prefer scan_usage with projects set.
ArgumentTypeDescription
projectstringGCP project ID
daysnumberLookback period in days

Resources

scan_results(6h)— Legacy multi-project Vertex AI token usage scan (single bucket)
single_scan(6h)— Legacy single project Vertex AI token usage scan
projects(7d)— GCP projects discovered through Cloud Resource Manager
usage(30d)— Daily Vertex AI token and request usage for one project (instance usage-<project>)
scan_summary(30d)— Per-project status and totals for the latest scan_usage run, including errors
gemini_usage(30d)— Daily Gemini API output tokens and request counts for one project (instance gemini-usage-<project>)
gemini_scan_summary(30d)— Per-project status and totals for the latest scan_gemini_api_usage run, including errors
billing_costs(30d)— Vertex AI billing export cost rows for one project (instance billing-<project>)
billing_summary(30d)— Totals and export-freshness check for the latest get_billing_costs run
billing_services(7d)— Billing services and SKUs matching an AI-related pattern, for choosing the cost filter
04Previous Versions16
2026.09.18.1

2026.09.18.1

Upgrade note: Normalized npm:zod dependency version to 4.6.5 across the repo. No behavioral changes in this extension.

2026.09.15.1

2026.09.15.1

Changed: Bump zod 4.4.3 → 4.6.5

2026.08.28.1

Changed: Normalized the extension license to Apache-2.0 and corrected the copyright holder to "Sean Escriva". Extensions that previously shipped an MIT LICENSE.md are now Apache-2.0, consistent with the repository root and every other extension. No code or behavioral changes.

Upgrade note: License text only. No API, schema, or runtime behavior changed.

2026.08.26.2

Fixed: Restored inline npm:zod@4.4.3 import specifiers so the registry quality scorer can resolve dependencies and score the extension. An earlier release used a bare "zod" import-map specifier, which published but scored as unscored.

Changed: Retained explicit compilerOptions.strict in deno.json. No behavioral or schema changes.

2026.08.28.1

2026.08.28.1

Changed: Normalized the extension license to Apache-2.0 and corrected the copyright holder to "Sean Escriva". Extensions that previously shipped an MIT LICENSE.md are now Apache-2.0, consistent with the repository root and every other extension. No code or behavioral changes.

Upgrade note: License text only. No API, schema, or runtime behavior changed.

2026.08.26.2

Fixed: Restored inline npm:zod@4.4.3 import specifiers so the registry quality scorer can resolve dependencies and score the extension. An earlier release used a bare "zod" import-map specifier, which published but scored as unscored.

Changed: Retained explicit compilerOptions.strict in deno.json. No behavioral or schema changes.

2026.08.26.2

2026.08.26.2

Fixed: Restored inline npm:zod@4.4.3 import specifiers so the registry quality scorer can resolve dependencies and score the extension. An earlier release used a bare "zod" import-map specifier, which published but scored as unscored.

Changed: Retained explicit compilerOptions.strict in deno.json. No behavioral or schema changes.

2026.08.25.1

2026.08.25.1

Changed: Updated labels for improved extension discoverability. Added cross-cutting category labels (security, observability, finops, infrastructure, networking, compliance, devops, ai, incident-response) where applicable.

updated labels

2026.08.24.1

2026.08.24.1

Added: Output metadata attributes for observability.

  • durationMs: Method execution duration in milliseconds.
  • collectedBy: Extension name that produced the data.
  • fetchedAt: ISO 8601 timestamp when data was fetched (added to resources that previously lacked it).
2026.08.23.1

2026.08.23.1

Changed: Documentation only — no code changes. Clarified the behavioral difference between scan_projects (per-project try/catch, warns and continues) and get_token_usage (no catch — a single project's failure fails the whole call), with a new days=90 usage example. Added a ## Troubleshooting section covering the silent per-project skip when a metric query returns "Cannot find metric," the warn-and-drop path for genuine per-project failures, the MAX_PAGES = 50 pagination cap, the four thrown-error cases in resolveServiceAccount, and the GCP token exchange failed error format from getAccessToken.

2026.08.21.2

2026.08.21.2

Changed: The projects global argument now requires at least one non-empty project ID; previously an empty list silently produced a scan of zero projects with no explanation. Service-account credential loading errors — a missing or unreadable GOOGLE_APPLICATION_CREDENTIALS file, or a malformed JSON key — now name the path/field that failed instead of surfacing a bare filesystem or JSON.parse error. Cloud Monitoring API failures now include the response body and the project ID in the error message instead of just an HTTP status code, and a malformed JSON response from either the OAuth token endpoint or the Monitoring API now raises a clear "returned malformed JSON" error naming the request that failed.

2026.08.21.1

Changed: Added .describe(...) documentation to previously undocumented fields in ModelUsageSchema, ProjectUsageSchema, and ScanResultsSchema (model/project identity fields, token counts, period and rate fields, and the truncated flag). Tightened the project argument on get_token_usage to require a non-empty string. No behavioral changes.

2026.07.31.1

Fixed: README incorrectly stated authentication uses gcloud CLI (Application Default Credentials). The extension actually uses a service account JSON key with signed JWT exchange. README now documents the correct auth mechanism, required role (roles/monitoring.viewer), and all global arguments.

2026.08.21.1

2026.08.21.1

Changed: Added .describe(...) documentation to previously undocumented fields in ModelUsageSchema, ProjectUsageSchema, and ScanResultsSchema (model/project identity fields, token counts, period and rate fields, and the truncated flag). Tightened the project argument on get_token_usage to require a non-empty string. No behavioral changes.

2026.07.31.1

Fixed: README incorrectly stated authentication uses gcloud CLI (Application Default Credentials). The extension actually uses a service account JSON key with signed JWT exchange. README now documents the correct auth mechanism, required role (roles/monitoring.viewer), and all global arguments.

2026.07.31.1

2026.07.31.1

Fixed: README incorrectly stated authentication uses gcloud CLI (Application Default Credentials). The extension actually uses a service account JSON key with signed JWT exchange. README now documents the correct auth mechanism, required role (roles/monitoring.viewer), and all global arguments.

2026.07.21.1

2026.07.21.1

Changed: Authentication no longer shells out to gcloud auth print-access-token. Auth now uses a GCP service account JSON key — the extension signs a JWT (RS256) and exchanges it for an access token at Google's token endpoint. This eliminates the gcloud CLI runtime dependency.

Added: serviceAccountJson optional global argument (sensitive). Accepts a stringified service account JSON key. Falls back to reading the file at GOOGLE_APPLICATION_CREDENTIALS if omitted.

Upgrade note: Existing model instances must provide credentials via one of:

  1. --global-arg 'serviceAccountJson=<vault:path/to/sa-key>' (recommended)
  2. Set GOOGLE_APPLICATION_CREDENTIALS env var pointing to a key file

The extension no longer requires gcloud to be installed or authenticated.

2026.07.18.1

2026.07.18.1

Added: An upgrades array entry (no-op) to vertex_usage.ts for proper typeVersion tracking on existing instances. No schema or behavior changes.

2026.07.10.1

Changed: Replaced internal org identifiers in the manifest usage examples with generic placeholders. No functional, API, or schema changes.

2026.07.10.1

2026.07.10.1

Changed: Replaced internal org identifiers in the manifest usage examples with generic placeholders. No functional, API, or schema changes.

2026.06.21.1
2026.06.15.1
2026.05.12.1
05Stats
A
100 / 100
Downloads
19
Archive size
46.1 KB
  • Has README or module doc2/2earned
  • README has a code example1/1earned
  • README is substantive1/1earned
  • Most symbols documented1/1earned
  • No slow types (deprecated)1/1earned
  • Dependencies pass trust audit2/2earned
  • Has description1/1earned
  • Platform support declared (or universal)2/2earned
  • License declared1/1earned
  • Verified public repository2/2earned
06Platforms
07Labels