Serverless GPU platform

Modal

Serverless GPUs as a Python decorator: cold starts in seconds, scale to zero, pay per second.

You write Python; Modal containers it and runs it on A10G/A100/H100/B200 with per-second billing. The unit economics differ from renting boxes, you pay a premium per GPU-second but zero for idle. For bursty inference it's often the cheapest real-world option.

Visit Modal ↗

GPU pricing on Modal

On-demand list per GPU-hour, last reviewed 2026-09. Methodology.

GPU $/GPU-hr vs cheapest Notes
T4 $0.59 cheapest Serverless, per-second billing
L4 $0.80 1.4× Serverless, per-second billing
A10 $1.10 1.9× Serverless A10G
L40S $1.95 3.3× Serverless (L40S class)
A100 80GB $2.50 4.2× Serverless, per-second billing
H100 SXM $3.95 6.7× Serverless, per-second billing
H200 $4.54 7.7× Serverless, per-second billing
B200 $6.25 10.6× Serverless, per-second billing

Strengths

  • True scale-to-zero
  • Developer experience
  • No idle cost

Watch for

  • Premium per-second rates vs raw rental
  • Not for long training runs

Alternatives