NVIDIA · Ada Lovelace · 2023
NVIDIA L40S
Ada silicon with 48 GB: strong FP8 compute, modest bandwidth. A favorite for image gen and mid-size LLMs.
Essentially an RTX 4090-class die with double the memory and datacenter drivers. GDDR6 bandwidth (0.86 TB/s) is the bottleneck for LLM decoding; compute-heavy diffusion workloads suit it better.
Where to rent a L40S
On-demand prices per GPU-hour; marketplace rates refresh daily, list prices reviewed 2026-09. Reserved and spot run 30-70% lower. Full breakdown with the used market and rent-vs-buy math: L40S pricing.
| Provider | $/GPU-hr | Notes |
|---|---|---|
| Vast.ai | $0.80 | Verified-listing rate, refreshed daily |
| Cudo | $0.87 | marketplace |
| Packet.ai | $0.92 | Dedicated |
| Hyperstack | $0.95 | neocloud |
| RunPod | $1.09 | neocloud |
| CoreWeave | $1.28 | neocloud |
| DigitalOcean | $1.57 | neocloud |
| Modal | $1.95 | Serverless (L40S class) |
| AWS | $1.96 | g6e.xlarge |
Specifications in context
Bars scaled against the best value in the whole catalog (B200 / MI325X era).
| Architecture | Ada Lovelace (2023) |
|---|---|
| Memory | 48 GB GDDR6, 0.864 TB/s |
| Interconnect | PCIe Gen4 · 64 GB/s |
| Form factor | PCIe |
| Partitioning | No MIG (time-slicing / vGPU only) |
What fits on one L40S
Weights + ~2 GB runtime overhead against 43 GB usable VRAM. Longer contexts and bigger batches need more, check the VRAM calculator.
| Model | Params | Highest precision that fits |
|---|---|---|
| Llama 3.2 1B | 1.24B | FP16 |
| Llama 3.2 3B | 3.21B | FP16 |
| Llama 3.1 8B | 8.03B | FP16 |
| Qwen2.5 7B | 7.62B | FP16 |
| Qwen2.5 14B | 14.8B | FP16 |
| Qwen2.5 32B | 32.8B | 8-bit |
| Qwen2.5 Coder 32B | 32.8B | 8-bit |
| QwQ 32B (reasoning) | 32.8B | 8-bit |
| Mistral 7B | 7.25B | FP16 |
| Mixtral 8x7B | 46.7B | 4-bit |
| Gemma 2 9B | 9.24B | FP16 |
| Gemma 2 27B | 27.2B | 8-bit |
| Phi-4 14B | 14.7B | FP16 |
| gpt-oss-20b | 20.9B | 8-bit |
Best for
- Stable Diffusion / Flux image generation
- 13B-34B quantized LLM serving
- Fine-tuning small models
Compare
A100 80GB vs L40S
The GPU that started the LLM era, now a value pick for fine-tuning and mid-size inference.
L40S vs RTX 4090
The people's inference GPU. On marketplaces, the best $/FLOP in the catalog.
L40S vs L4
72 watts. The efficiency play for video, small-model inference, and high-density serving.