NVIDIA · Ada Lovelace · 2023

NVIDIA L4

72 watts. The efficiency play for video, small-model inference, and high-density serving.

24 GB
GDDR6
0.3 TB/s
Mem bandwidth
121 TF
FP16 dense
242 TF
FP8 dense
72 W
TDP
$0.32 /hr
From · Vast.ai

Single-slot, no external power connector, clouds deploy it densely and price it low. Think of it as the successor to the T4.

Where to rent a L4

On-demand prices per GPU-hour; marketplace rates refresh daily, list prices reviewed 2026-09. Reserved and spot run 30-70% lower. Full breakdown with the used market and rent-vs-buy math: L4 pricing.

Provider $/GPU-hr vs cheapest Notes
Vast.ai $0.32 cheapest Verified-listing rate, refreshed daily
RunPod $0.49 1.5× neocloud
GCP $0.71 2.2× g2-standard-4
OVHcloud $0.75 2.3× hyperscaler
Modal $0.80 2.5× Serverless, per-second billing
AWS $0.98 3.1× g6.xlarge

Specifications in context

Bars scaled against the best value in the whole catalog (B200 / MI325X era).

Memory 24 GB
Memory bandwidth 0.3 TB/s
FP16 dense 121 TFLOPS
FP8 dense 242 TFLOPS
Power (TDP) 72 W
ArchitectureAda Lovelace (2023)
Memory24 GB GDDR6, 0.3 TB/s
InterconnectPCIe Gen4
Form factorPCIe (low-profile)
PartitioningNo MIG (time-slicing / vGPU only)

What fits on one L4

Weights + ~2 GB runtime overhead against 22 GB usable VRAM. Longer contexts and bigger batches need more, check the VRAM calculator.

ModelParamsHighest precision that fits
Llama 3.2 1B 1.24B FP16
Llama 3.2 3B 3.21B FP16
Llama 3.1 8B 8.03B FP16
Qwen2.5 7B 7.62B FP16
Qwen2.5 14B 14.8B 8-bit
Mistral 7B 7.25B FP16
Gemma 2 9B 9.24B 8-bit
Gemma 2 27B 27.2B 4-bit
Phi-4 14B 14.7B 8-bit
gpt-oss-20b 20.9B 4-bit

Best for

Compare

L40S vs L4

Ada silicon with 48 GB: strong FP8 compute, modest bandwidth. A favorite for image gen and mid-size LLMs.

L4 vs T4

The 2018 inference stalwart. Slow by modern standards, but everywhere and nearly free.

A10 vs L4

The quiet default for 7B-class production inference on AWS and Oracle.