NVIDIA L4
Compare NVIDIA L4 against another machine in the Arena
Reference: Ada Lovelace (microarchitecture) on Wikipedia
- Class
- GPU
- Memory
- 24 GB GDDR6
- Runs (Q4_K_M)
- —
- Bandwidth
- 300 GB/s
- TDP
- 72 W
- Released
- 2023-03-01
- Price (US)
- $3500 used · $2750 new as of 2026-08
Sourcing and disambiguation notes
Single-slot, 72W, no external power connector — designed to drop into dense inference servers. Low memory bandwidth (300 GB/s) relative to price/VRAM makes it a poor decode-speed choice versus an A5000/A6000 despite being newer; its value proposition is power/density, not raw tok/s. No sourced llama.cpp or vLLM decode-speed benchmark found this session (see gaps file). New-old-stock 'new' figure of $2,750 (Newegg) is now available via the same gpudojo.com tracker -- notably BELOW the $3,500 used/eBay figure, which is unusual and likely reflects thin/noisy used listings for this still-in-production card rather than a stable secondary-market premium; treat the used figure with caution.
Specification sources
- https://gpudojo.com/l4 — press, 2026-08-22
- https://cputronic.com/en/gpu/nvidia-l4 — press, 2026-08-22
Measurements
Decode is token generation — the speed you feel while an answer streams. Prefill is prompt processing — the wait before it starts. Why bandwidth predicts decode speed.
No records yet for this hardware. Know of a published benchmark on it? Submit the link.
What verified, single-source and estimated mean, and the same rows with every filter and sort in the benchmarks explorer.
Speed over time
What it can run
Fit is arithmetic, not a measurement — how it is computed. Against 24 GB; models that fit are listed largest first.
No modeled quant fits in 24 GB.