Cutting inference cost per token

10 min read · updated 2026-09

Inference is the AI cost that compounds: every user, every agent step, every retry is tokens, consumed the moment they're made. Training is a capital project; inference is a utility bill. Here are the levers that actually move it, ranked by how hard they pull.

First, know which buyer you are

Every inference workload optimizes one of two things, and the entire hardware decision follows from which:

Most fleets serve both and should split them: pinning batch work to latency-priced infrastructure is the single most common overspend in AI infra.

The seven levers, ranked

1. Quantize (2-4× cheaper)

Weights at 4-8 bit halve-to-quarter the bytes moved per token, and bytes moved per token is the whole game. Modern 4-bit builds of 70B-class models cost a point or two of benchmark; measure on your task, then take the 2-4×.

2. Batch properly (up to 10×)

Continuous batching (vLLM, SGLang, TensorRT-LLM) turns idle FLOPS into throughput, aggregate tokens/sec scales near-linearly with batch depth until compute-bound. If you're serving batch-1 on a datacenter GPU, this lever alone dwarfs every hardware choice.

3. Right-size the silicon (1.5-3×)

Decode cost tracks memory bandwidth per dollar, and the ranking is not sorted by launch date: an A100 at $0.85-1.29/hr frequently beats an H100 at $2.49 on $/token for INT8/4-bit throughput work. The A100 vs H100 line is one question: FP8 plus latency SLO → new silicon; otherwise → cheapest adequate bandwidth.

4. Buy sustained, not spot-priced convenience (1.4-3×)

On-demand list price is the ceiling. Reserved and contracted capacity runs 30-70% below it, and dedicated single-tenant contracts on previous-generation fleets sit below that, capacity acquired at a recovered cost basis prices at levels new-build capacity structurally can't reach.

5. Raise utilization (up to 3× on your effective rate)

A GPU billing 24/7 but busy 30% of the time costs 3.3× its sticker per useful hour. Fill troughs with batch backlog, evals, and synthetic-data generation; the ROI calculator shows how brutally utilization dominates the own-vs-rent math.

6. Shrink the model (2-10×, workload permitting)

Distill, fine-tune a smaller base, or route: a tuned 8B handling 80% of traffic with a 70B fallback cuts blended cost dramatically. This lever trades eng time for unit cost, it pulls hardest at scale.

7. Reconsider rent vs own

At sustained high utilization, owning, or contracting dedicated capacity on, redeployed enterprise fleets beats every rental rate in our index (the full decision framework). It's the last lever because it's the least reversible; it's on the list because at scale it's the biggest.

Why this matters more every quarter

Inference demand is structural, not speculative: tokens are consumed the instant they're produced, by products with users, there is no token inventory. Every efficiency gain gets spent on more usage (agents got cheap, so now they run for hours). Which means $/token isn't a one-time optimization; it's an operating discipline. Run your stack through the estimator, split your latency and throughput traffic, and re-check the hardware ranking whenever prices move, they move often.

Throughput buyer with sustained volume? Charg, from the team behind Virtualized, operates dedicated single-tenant A100 clusters (validated hyperscaler-grade fleets, U.S. datacenters, InfiniBand) at a recovered cost basis, under contract, deployable now. Reserve capacity →

Related