NVIDIA · Ada Lovelace · 2023

NVIDIA L40S

Ada silicon with 48 GB: strong FP8 compute, modest bandwidth. A favorite for image gen and mid-size LLMs.

48 GB
GDDR6
0.864 TB/s
Mem bandwidth
362 TF
FP16 dense
733 TF
FP8 dense
350 W
TDP
$0.80 /hr
From · Vast.ai

Essentially an RTX 4090-class die with double the memory and datacenter drivers. GDDR6 bandwidth (0.86 TB/s) is the bottleneck for LLM decoding; compute-heavy diffusion workloads suit it better.

Where to rent a L40S

On-demand prices per GPU-hour; marketplace rates refresh daily, list prices reviewed 2026-09. Reserved and spot run 30-70% lower. Full breakdown with the used market and rent-vs-buy math: L40S pricing.

Provider $/GPU-hr vs cheapest Notes
Vast.ai $0.80 cheapest Verified-listing rate, refreshed daily
Cudo $0.87 1.1× marketplace
Packet.ai $0.92 1.1× Dedicated
Hyperstack $0.95 1.2× neocloud
RunPod $1.09 1.4× neocloud
CoreWeave $1.28 1.6× neocloud
DigitalOcean $1.57 2.0× neocloud
Modal $1.95 2.4× Serverless (L40S class)
AWS $1.96 2.4× g6e.xlarge

Specifications in context

Bars scaled against the best value in the whole catalog (B200 / MI325X era).

Memory 48 GB
Memory bandwidth 0.864 TB/s
FP16 dense 362 TFLOPS
FP8 dense 733 TFLOPS
Power (TDP) 350 W
ArchitectureAda Lovelace (2023)
Memory48 GB GDDR6, 0.864 TB/s
InterconnectPCIe Gen4 · 64 GB/s
Form factorPCIe
PartitioningNo MIG (time-slicing / vGPU only)

What fits on one L40S

Weights + ~2 GB runtime overhead against 43 GB usable VRAM. Longer contexts and bigger batches need more, check the VRAM calculator.

ModelParamsHighest precision that fits
Llama 3.2 1B 1.24B FP16
Llama 3.2 3B 3.21B FP16
Llama 3.1 8B 8.03B FP16
Qwen2.5 7B 7.62B FP16
Qwen2.5 14B 14.8B FP16
Qwen2.5 32B 32.8B 8-bit
Qwen2.5 Coder 32B 32.8B 8-bit
QwQ 32B (reasoning) 32.8B 8-bit
Mistral 7B 7.25B FP16
Mixtral 8x7B 46.7B 4-bit
Gemma 2 9B 9.24B FP16
Gemma 2 27B 27.2B 8-bit
Phi-4 14B 14.7B FP16
gpt-oss-20b 20.9B 8-bit

Best for

Compare

A100 80GB vs L40S

The GPU that started the LLM era, now a value pick for fine-tuning and mid-size inference.

L40S vs RTX 4090

The people's inference GPU. On marketplaces, the best $/FLOP in the catalog.

L40S vs L4

72 watts. The efficiency play for video, small-model inference, and high-density serving.