What will your inference actually cost?

Describe the workload, pick a model, and get the monthly bill three ways: self-hosted on provisioned GPUs, self-hosted as batch jobs, and pay-per-token APIs. First-principles throughput modeling against live GPU pricing, with the break-even shown.

The monthly bill

Self-host options, ranked

Per-GPU-configuration cost for this exact workload. Throughput = min(bandwidth bound at 50% MBU, compute bound at 40% MFU) for decode, compute-bound prefill for input tokens. GPU count is the minimum that fits the model at 8K context. Prices are cheapest tracked on-demand, 2026-09; see GPU pricing.

SetupOut tok/s$/mo$/1k queries$/M out tokCheapest at

How to read this. Provisioned mode divides compute cost by your utilization, because a fleet sized for your traffic bills 24/7 whether busy or not; batch mode assumes you spin capacity up, run at ~95% busy, and shut it down, the right model for overnight jobs and evals. API pricing is an indicative market midpoint across major open-model hosts (2026-09), billed only for tokens used, which is why APIs win at low volume and lose at sustained scale. The setup column shows the minimum configuration that fits the model; high volumes imply running several replicas of it, and cost scales linearly with replicas, so the dollar figures hold. Real throughput varies with batch depth and stack; treat results as ranking numbers, and verify with a pilot. Methodology on the about page.