What will your inference actually cost?
Describe the workload, pick a model, and get the monthly bill three ways: self-hosted on provisioned GPUs, self-hosted as batch jobs, and pay-per-token APIs. First-principles throughput modeling against live GPU pricing, with the break-even shown.
The monthly bill
Self-host options, ranked
Per-GPU-configuration cost for this exact workload. Throughput = min(bandwidth bound at 50% MBU, compute bound at 40% MFU) for decode, compute-bound prefill for input tokens. GPU count is the minimum that fits the model at 8K context. Prices are cheapest tracked on-demand, 2026-09; see GPU pricing.
| Setup | $/mo | $/1k queries | $/M out tok | Cheapest at |
|---|
How to read this. Provisioned mode divides compute cost by your utilization, because a fleet sized for your traffic bills 24/7 whether busy or not; batch mode assumes you spin capacity up, run at ~95% busy, and shut it down, the right model for overnight jobs and evals. API pricing is an indicative market midpoint across major open-model hosts (2026-09), billed only for tokens used, which is why APIs win at low volume and lose at sustained scale. The setup column shows the minimum configuration that fits the model; high volumes imply running several replicas of it, and cost scales linearly with replicas, so the dollar figures hold. Real throughput varies with batch depth and stack; treat results as ranking numbers, and verify with a pilot. Methodology on the about page.