Serverless GPU platform
Modal
Serverless GPUs as a Python decorator: cold starts in seconds, scale to zero, pay per second.
You write Python; Modal containers it and runs it on A10G/A100/H100/B200 with per-second billing. The unit economics differ from renting boxes, you pay a premium per GPU-second but zero for idle. For bursty inference it's often the cheapest real-world option.
GPU pricing on Modal
On-demand list per GPU-hour, last reviewed 2026-09. Methodology.
| GPU | $/GPU-hr | Notes |
|---|---|---|
| T4 | $0.59 | Serverless, per-second billing |
| L4 | $0.80 | Serverless, per-second billing |
| A10 | $1.10 | Serverless A10G |
| L40S | $1.95 | Serverless (L40S class) |
| A100 80GB | $2.50 | Serverless, per-second billing |
| H100 SXM | $3.95 | Serverless, per-second billing |
| H200 | $4.54 | Serverless, per-second billing |
| B200 | $6.25 | Serverless, per-second billing |
Strengths
- True scale-to-zero
- Developer experience
- No idle cost
Watch for
- Premium per-second rates vs raw rental
- Not for long training runs