Renting vs buying GPUs: the decision framework
9 min read · updated 2026-09
Every team serving models at scale eventually does this math, usually after the first big cloud invoice. The answer is not a matter of taste: it is one variable, utilization, plus a handful of costs people forget to count. Here is the whole framework, with live numbers.
The one number that decides it
A rented GPU costs the same whether it is busy or idle only if you remember to turn it off. An owned GPU costs the same whether it is busy or idle, period: the capital is spent. So the comparison is really between a rental rate you pay per hour used, and an ownership cost that gets divided across however many useful hours you extract.
That division is the whole game. A $6,800 card amortized over 36 months is $0.26/hr if it works every hour, $0.43/hr at 60% utilization, and $2.59/hr if it is busy one hour in ten, worse than renting an H100. Buying a GPU is not buying compute, it is buying an obligation to keep it busy.
The math, worked at live prices
Effective ownership cost per hour = purchase price ÷ (36 months × 730 hrs × utilization) + power. Power = 70% of TDP × 1.4 PUE × $0.08/kWh. Rentals are the cheapest on-demand rate in our index, 2026-09.
| GPU | Used price | Own, 60% util | Own, 90% util |
|---|---|---|---|
| A100 40GB | $3,400 | $0.25/hr | $0.18/hr |
| A100 80GB | $6,800 | $0.46/hr | $0.32/hr |
| H100 SXM | $23,000 | $1.51/hr | $1.03/hr |
| RTX 4090 | $1,500 | $0.13/hr | $0.10/hr |
| RTX 3090 | $650 | $0.07/hr | $0.05/hr |
Indicative used-market prices, reviewed 2026-09. Run your own numbers, your utilization, your power rate, in the GPU ROI calculator.
Read the table honestly and both columns win somewhere. At 60% utilization, every card here beats its own cheapest rental. But 60% sustained utilization is genuinely hard: it means the card is doing paid work 14.4 hours of every day, weekends included, for three years. Most self-assessments of future utilization are off by 2×, in the optimistic direction.
When renting wins
- Spiky or exploratory workloads. Fine-tuning runs, evals, experiments. If the GPU would sit dark between bursts, on-demand pricing is the discount, not the premium.
- You need current-generation silicon. H100s and newer still command prices where the capital outlay is large and the depreciation curve is steepest. Renting pushes that depreciation risk onto someone else.
- Scale uncertainty. If you might need 4 GPUs or 40 next quarter, owning 4 solves nothing and owning 40 might be a write-off.
- Nobody owns infrastructure at your company. A GPU box needs an owner: drivers, firmware, thermals, a plan for the fan that dies at 2am. That salary fraction belongs in the math.
When owning wins
- Sustained, predictable inference load. A production model serving real traffic around the clock is the textbook case: utilization is high and forecastable, which is exactly what the amortization math rewards.
- The hardware is past its depreciation cliff. This is why the used A100 dominates this conversation: the steep part of its depreciation already happened to its first owner. At second-life prices, you are buying bandwidth per dollar that no new card matches.
- Batch and throughput work. Overnight pipelines, synthetic data, agents nobody is watching: work that can always fill the troughs pushes utilization toward the 90% column above.
- Data or compliance gravity. Some workloads cannot leave your walls. Then the rent column is not available at any price.
The costs owners forget
The purchase price is the visible cost. The recurring ones decide whether the spreadsheet was honest:
- Power and cooling: a 400W datacenter card at realistic draw and PUE is roughly $0.03-0.06 per busy hour at $0.08/kWh, and double that at European rates.
- Space and hosting: colocation for a GPU server runs from tens to hundreds of dollars per month depending on density and market.
- Failure risk: used cards typically carry no warranty. Price in a failure rate; on a multi-GPU box, budget a spare. Our diligence checklist covers what to test before wiring money.
- Resale timing: ownership math improves if you sell the card while it still clears a good price. That optionality is real but it requires actually doing it.
The middle path most teams actually want
Rent-vs-buy is a false binary for a lot of workloads. Between on-demand rental and a pallet of hardware sits contracted dedicated capacity: single-tenant machines, committed terms, priced well below on-demand because the provider gets utilization certainty and you get ownership-like economics without the capital outlay, the hosting problem, or the failure risk.
This is where previous-generation fleets shine. Dedicated capacity on validated, hyperscaler-grade A100 fleets acquired at a recovered cost basis can price below what new-build capacity structurally supports, while the silicon still tops the bandwidth-per-dollar table for throughput inference.
Sustained throughput workload? Charg, from the team behind Virtualized, operates dedicated single-tenant A100 clusters in U.S. datacenters, validated hyperscaler-grade fleets, under contract, deployable now. Reserve capacity →
The decision, compressed
- Estimate honest utilization over 36 months. Then halve it.
- Below ~40%: rent on-demand, and use the price index to rent well.
- Above ~60% with capital and an infra owner: buy used, past the depreciation cliff, with the ROI calculator open.
- Above ~60% without the appetite for hardware: contract dedicated capacity and keep the balance sheet clean.
- In between: rent, raise utilization with batch work, and rerun the math quarterly. Prices move; we chart them.