Cost Projection
GPU inference cost projection across cloud, rental, and capex scenarios.
Three model types store scenario inputs (instance rates, hardware costs, facility expenses) and compute normalized $/GPU-hour projections. A comparison report queries all scenarios and produces a ranked table with crossover analysis.
No live API calls — rates are manually entered from quotes, pricing pages, or account team conversations. All scenarios in a comparison must share a currency.
Model Types
- gpu-cloud — Hyperscaler GPU instances (AWS, Azure, GCP) under various capacity models: on-demand, reserved, FTP, Capacity Blocks, etc.
- gpu-rental — Third-party providers (CoreWeave, Lambda Labs, RunPod) with simpler per-GPU-hour pricing and optional commitment discounts.
- gpu-capex — On-premises/colo hardware. Amortizes a capital purchase into synthetic recurring cost, adds facility/staff/maintenance, and normalizes to $/GPU-hour. Includes sensitivity analysis across utilization and useful-life assumptions.
Usage
swamp extension pull @webframp/cost-projection
# Cloud scenario
swamp model create @webframp/cost-projection/gpu-cloud kimi-k3-hyperpod
swamp model method run kimi-k3-hyperpod record \
--input name=kimi-k3-hyperpod \
--input provider=aws \
--input region=us-east-1 \
--input instanceType=ml.p6-b300.48xlarge \
--input gpuCount=8 \
--input gpuModel="NVIDIA B300" \
--input capacityModel=flexible-training-plan \
--input instanceRatePerHour=98.50
# Rental scenario
swamp model create @webframp/cost-projection/gpu-rental kimi-k3-coreweave
swamp model method run kimi-k3-coreweave record \
--input name=kimi-k3-coreweave \
--input provider=coreweave \
--input gpuModel="NVIDIA H100 SXM" \
--input gpuCount=8 \
--input ratePerGpuHour=2.49
# Capex scenario
swamp model create @webframp/cost-projection/gpu-capex dc-east-b300
swamp model method run dc-east-b300 record \
--input name=dc-east-b300 \
--input gpuModel="NVIDIA B300" \
--input gpuCount=8 \
--input gpuCostPerUnit=35000 \
--input serverCost=45000 \
--input networkingCost=15000 \
--input totalHardwareCost=340000 \
--input usefulLifeMonths=36 \
--input coloCostPerKwMonth=150 \
--input powerDrawKw=10.5
# Run sensitivity analysis
swamp model method run dc-east-b300 sensitivity
# The comparison report runs automatically after any method call on any
# gpu-cloud, gpu-rental, or gpu-capex instance. Retrieve its latest output:
swamp report get @webframp/cost-projection-comparison --model dc-east-b300 --markdownPer-Token Break-Even (Optional)
To compare self-hosted cost against per-token API pricing, provide
apiComparisonRatePerMToken ($/million tokens from the API provider) and
estimatedTokensPerGpuHour (from vLLM benchmarks or model card throughput
figures) when recording a scenario. The projection will compute the monthly
token volume at which self-hosting breaks even.
2026.09.19.1
Changed: Removed the deprecated drivers: manifest key. Swamp core removed
extension driver support (swamp-club/swamp#2333, 2026-09-01); the key now emits
a deprecation warning and is ignored by the manifest parser. The field was an
empty list, so this is a no-op for behavior — the extension ships the same three
models and one report as before.
Upgrade note: No migration required. No model schema, method, or resource changed.
| Argument | Type | Description |
|---|---|---|
| name | string | Human-readable name for this scenario |
| region | string | Cloud region for the instance (e.g. us-east-1) |
| instanceType | string | Cloud provider instance type or SKU |
| gpuCount | number | Number of GPUs per instance |
| gpuModel | string | GPU model (e.g. H100, A100) |
| commitmentTermMonths? | number | Length of the commitment term in months, if the capacity model requires one |
| instanceRatePerHour | number | Quoted hourly rate per instance, in currency units |
| currency | string | Currency of all rate fields in this scenario |
| hoursPerDay | number | Expected hours of instance utilization per day |
| daysPerMonth | number | Expected days of instance utilization per month |
| replicas | number | Number of identical instances running concurrently |
| storageGb | number | Attached storage size per instance, in GB |
| storageRatePerGbMonth | number | Storage cost per GB per month |
| dataTransferGbMonth | number | Expected monthly data transfer volume, in GB |
| dataTransferRatePerGb | number | Data transfer cost per GB |
| managementFeePerMonth | number | Flat monthly management/platform fee, independent of replicas |
| apiComparisonRatePerMToken? | number | Comparable managed-API rate per million tokens, used to compute break-even |
| estimatedTokensPerGpuHour? | number | Estimated throughput in tokens per GPU-hour, from benchmarks or model card figures |
| notes? | string | Free-form notes about this scenario |
| sourceUrl? | string | URL of the quote or pricing page this scenario is based on |
| quotedAt? | string | Date the rate was quoted (YYYY-MM-DD) |
| fetchedAt? | string | ISO 8601 timestamp when data was fetched |
| durationMs? | number | Method execution duration in milliseconds |
| collectedBy? | string | Extension that collected this data |
| Argument | Type | Description |
|---|---|---|
| instanceRatePerHour | number | New quoted hourly rate per instance, in the scenario's currency |
| quotedAt? | string | Date the new rate was quoted (YYYY-MM-DD); defaults to today |
Resources
| Argument | Type | Description |
|---|---|---|
| name | string | Human-readable name for this scenario |
| provider | string | GPU rental provider name (e.g. CoreWeave, Lambda Labs) |
| region? | string | Provider region or datacenter location |
| gpuModel | string | GPU model (e.g. H100, A100) |
| gpuCount | number | Number of GPUs in this scenario |
| gpuMemoryGb? | number | Memory per GPU, in GB |
| ratePerGpuHour | number | Quoted list price per GPU per hour |
| currency | string | Currency of all rate fields in this scenario |
| commitmentDiscountPct | number | Discount off ratePerGpuHour granted for the commitment term, as a percentage |
| hoursPerDay | number | Expected hours of GPU utilization per day |
| daysPerMonth | number | Expected days of GPU utilization per month |
| storageGb | number | Attached storage size, in GB |
| storageRatePerGbMonth | number | Storage cost per GB per month |
| networkEgressGbMonth | number | Expected monthly network egress volume, in GB |
| networkEgressRatePerGb | number | Network egress cost per GB |
| apiComparisonRatePerMToken? | number | Comparable managed-API rate per million tokens, used to compute break-even |
| estimatedTokensPerGpuHour? | number | Estimated throughput in tokens per GPU-hour, from benchmarks or model card figures |
| notes? | string | Free-form notes about this scenario |
| sourceUrl? | string | URL of the quote or pricing page this scenario is based on |
| quotedAt? | string | Date the rate was quoted (YYYY-MM-DD) |
| fetchedAt? | string | ISO 8601 timestamp when data was fetched |
| durationMs? | number | Method execution duration in milliseconds |
| collectedBy? | string | Extension that collected this data |
| Argument | Type | Description |
|---|---|---|
| ratePerGpuHour | number | New quoted list price per GPU per hour |
| quotedAt? | string | Date the new rate was quoted (YYYY-MM-DD); defaults to today |
Resources
| Argument | Type | Description |
|---|---|---|
| name | string | Human-readable name for this scenario |
| site? | string | Physical site or datacenter location |
| gpuModel | string | GPU model (e.g. H100, A100) |
| gpuCount | number | Number of GPUs purchased |
| gpuCostPerUnit | number | Purchase price per GPU |
| serverCost | number | Server/chassis hardware cost, excluding GPUs |
| networkingCost | number | Networking hardware cost (switches, NICs, cabling) |
| totalHardwareCost | number | Total upfront hardware cost; should equal gpuCostPerUnit * gpuCount + serverCost + networkingCost |
| usefulLifeMonths | number | Amortization period for the hardware, in months |
| residualValuePct | number | Expected resale/residual value at end of useful life, as a percentage of totalHardwareCost |
| coloCostPerKwMonth | number | Colocation/power cost per kW per month |
| powerDrawKw | number | IT power draw of the hardware, in kW |
| pue | number | Power usage effectiveness multiplier applied to powerDrawKw to account for cooling/overhead |
| networkBandwidthCostPerMonth | number | Flat monthly network bandwidth cost |
| staffFteAllocation | number | Fraction of staff FTE allocated to operating this hardware |
| staffCostPerFteMonth | number | Fully-loaded staff cost per FTE per month |
| failureRatePctPerYear | number | Expected annual hardware failure rate, as a percentage of totalHardwareCost, used to estimate replacement spend |
| spareBudgetPerMonth | number | Flat monthly budget reserved for spare parts |
| warrantyMonths | number | Manufacturer warranty coverage, in months |
| targetUtilizationPct | number | Target GPU utilization used to compute the effective cost per GPU-hour |
| hoursPerDay | number | Expected hours of hardware availability per day |
| daysPerMonth | number | Expected days of hardware availability per month |
| apiComparisonRatePerMToken? | number | Comparable managed-API rate per million tokens, used to compute break-even |
| estimatedTokensPerGpuHour? | number | Estimated throughput in tokens per GPU-hour, from benchmarks or model card figures |
| currency | string | Currency of all cost fields in this scenario |
| notes? | string | Free-form notes about this scenario |
| quotedAt? | string | Date the hardware cost was quoted (YYYY-MM-DD) |
| fetchedAt? | string | ISO 8601 timestamp when data was fetched |
| durationMs? | number | Method execution duration in milliseconds |
| collectedBy? | string | Extension that collected this data |
| Argument | Type | Description |
|---|---|---|
| totalHardwareCost | number | New total upfront hardware cost |
| gpuCostPerUnit? | number | New purchase price per GPU; leave unset to keep the stored value |
| quotedAt? | string | Date the new cost was quoted (YYYY-MM-DD); defaults to today |
Resources
Cross-scenario GPU inference cost comparison, normalized to $/GPU-hour.
2026.09.18.1
Upgrade note: Normalized npm:zod dependency version to 4.6.5 across the
repo. No behavioral changes in this extension.
2026.09.15.1
Changed: Bump zod 4.4.3 → 4.6.5
2026.08.28.1
Changed: Normalized the extension license to Apache-2.0 and corrected the copyright holder to "Sean Escriva". Extensions that previously shipped an MIT LICENSE.md are now Apache-2.0, consistent with the repository root and every other extension. No code or behavioral changes.
Upgrade note: License text only. No API, schema, or runtime behavior changed.
2026.08.26.3
Fixed: Restored inline npm:zod@4.4.3 import specifiers so the registry
quality scorer can resolve dependencies and score the extension. An earlier
release used a bare "zod" import-map specifier, which published but scored as
unscored.
Changed: Retained explicit compilerOptions.strict in deno.json. No
behavioral or schema changes.
2026.08.28.1
Changed: Normalized the extension license to Apache-2.0 and corrected the copyright holder to "Sean Escriva". Extensions that previously shipped an MIT LICENSE.md are now Apache-2.0, consistent with the repository root and every other extension. No code or behavioral changes.
Upgrade note: License text only. No API, schema, or runtime behavior changed.
2026.08.26.3
Fixed: Restored inline npm:zod@4.4.3 import specifiers so the registry
quality scorer can resolve dependencies and score the extension. An earlier
release used a bare "zod" import-map specifier, which published but scored as
unscored.
Changed: Retained explicit compilerOptions.strict in deno.json. No
behavioral or schema changes.
2026.08.26.3
Fixed: Restored inline npm:zod@4.4.3 import specifiers so the registry
quality scorer can resolve dependencies and score the extension. An earlier
release used a bare "zod" import-map specifier, which published but scored as
unscored.
Changed: Retained explicit compilerOptions.strict in deno.json. No
behavioral or schema changes.
2026.08.26.1
Changed: Normalized deno.json configuration for repo-wide consistency:
added explicit compilerOptions.strict and migrated zod dependency to the
import map (bare "zod" specifier instead of inline npm:zod@4.4.3). No
behavioral changes — runtime resolution is identical.
2026.08.25.1
Changed: Updated labels for improved extension discoverability. Added cross-cutting category labels (security, observability, finops, infrastructure, networking, compliance, devops, ai, incident-response) where applicable.
updated labels
2026.08.24.2
Added: Output metadata attributes for observability.
durationMs: Method execution duration in milliseconds.collectedBy: Extension name that produced the data.fetchedAt: ISO 8601 timestamp when data was fetched (added to resources that previously lacked it).
2026.08.24.1
Fixed source/manifest version mismatch and unquoted manifest version. Added Troubleshooting section documenting record prerequisite, sensitivity array inputs, conditional break-even computation, and independent model state.
2026.08.21.2
Changed: The gpu-capex model's sensitivity method now rejects an empty
usefulLifeMonthsRange or empty utilizationPctRange at validation time,
with a clear "at least 1 element" error, instead of silently returning a
sensitivity matrix with zero rows when either range is passed as [].
No changes to the record, project, or update_hardware_cost methods, and
no changes to gpu-cloud, gpu-rental, or the comparison report.
Release Notes
2026.08.21.1
Changed: Added .describe(...) documentation to previously undocumented
fields across ScenarioSchema, ProjectionSchema, and (for gpu-capex)
SensitivityRowSchema/SensitivitySchema in all three model types
(gpu-cloud, gpu-rental, gpu-capex), plus the update_rate,
update_hardware_cost, and sensitivity method argument schemas. No
behavioral changes.
2026.08.02.1
Fixed: The comparison report declared scope: "workspace", which isn't a
valid ReportScope ("method" | "model" | "workflow"). The extension failed
to load entirely — swamp doctor extensions reported ValidationFailed for
@webframp/cost-projection, and the report could never run.
Changed: The report now runs at model scope and is attached as a
default report on gpu-cloud, gpu-rental, and gpu-capex, so it fires
after any method call on any instance of the three model types (including
sensitivity and update_hardware_cost on gpu-capex, which has no
update_rate method). Internally, execute() no longer relies on the single
instance's
context.dataHandles — it scans every cost-projection instance in the repo
via dataRepository.findAllGlobal(), so the table actually compares sibling
scenarios (e.g. an on-demand quote against a capacity-block quote) instead of
only ever showing the one instance that triggered the run.
Upgrade note: No input/output schema changes. Re-pull the extension to
pick up the fix — running any method on a gpu-cloud/gpu-rental/
gpu-capex instance now produces the cross-scenario table, retrievable with
swamp report get @webframp/cost-projection-comparison --model <instance>.
2026.08.01.1
Fixed: Report name now uses collective prefix (@webframp/cost-projection-comparison).
Previous name (webframp/cost-projection-comparison) failed publish validation.
2026.07.31.1
Initial release.
- Three model types:
gpu-cloud,gpu-rental,gpu-capex - All normalize to $/GPU-hour for cross-scenario comparison
- Capex model includes sensitivity analysis (utilization × useful-life matrix)
- Workspace-scoped comparison report with crossover analysis
- Single-currency assumption (all scenarios must share a currency)
- Optional per-token break-even calculation against API pricing
Release Notes
2026.08.02.1
Fixed: The comparison report declared scope: "workspace", which isn't a
valid ReportScope ("method" | "model" | "workflow"). The extension failed
to load entirely — swamp doctor extensions reported ValidationFailed for
@webframp/cost-projection, and the report could never run.
Changed: The report now runs at model scope and is attached as a
default report on gpu-cloud, gpu-rental, and gpu-capex, so it fires
after any method call on any instance of the three model types (including
sensitivity and update_hardware_cost on gpu-capex, which has no
update_rate method). Internally, execute() no longer relies on the single
instance's
context.dataHandles — it scans every cost-projection instance in the repo
via dataRepository.findAllGlobal(), so the table actually compares sibling
scenarios (e.g. an on-demand quote against a capacity-block quote) instead of
only ever showing the one instance that triggered the run.
Upgrade note: No input/output schema changes. Re-pull the extension to
pick up the fix — running any method on a gpu-cloud/gpu-rental/
gpu-capex instance now produces the cross-scenario table, retrievable with
swamp report get @webframp/cost-projection-comparison --model <instance>.
2026.08.01.1
Fixed: Report name now uses collective prefix (@webframp/cost-projection-comparison).
Previous name (webframp/cost-projection-comparison) failed publish validation.
2026.07.31.1
Initial release.
- Three model types:
gpu-cloud,gpu-rental,gpu-capex - All normalize to $/GPU-hour for cross-scenario comparison
- Capex model includes sensitivity analysis (utilization × useful-life matrix)
- Workspace-scoped comparison report with crossover analysis
- Single-currency assumption (all scenarios must share a currency)
- Optional per-token break-even calculation against API pricing
Modified 1 reports
Release Notes
2026.08.01.1
Fixed: Report name now uses collective prefix (@webframp/cost-projection-comparison).
Previous name (webframp/cost-projection-comparison) failed publish validation.
2026.07.31.1
Initial release.
- Three model types:
gpu-cloud,gpu-rental,gpu-capex - All normalize to $/GPU-hour for cross-scenario comparison
- Capex model includes sensitivity analysis (utilization × useful-life matrix)
- Workspace-scoped comparison report with crossover analysis
- Single-currency assumption (all scenarios must share a currency)
- Optional per-token break-even calculation against API pricing
- Has README or module doc2/2earned
- README has a code example1/1earned
- README is substantive1/1earned
- Most symbols documented1/1earned
- No slow types (deprecated)1/1earned
- Dependencies pass trust audit2/2earned
- Has description1/1earned
- Platform support declared (or universal)2/2earned
- License declared1/1earned
- Verified public repository2/2earned