NVIDIA L4

Compare NVIDIA L4 against another machine in the Arena

Class
GPU
Memory
24 GB GDDR6
Runs (Q4_K_M)
Bandwidth
300 GB/s
TDP
72 W
Released
2023-03-01
Price (US)
$3500 used · $2750 new as of 2026-08
Sourcing and disambiguation notes

Single-slot, 72W, no external power connector — designed to drop into dense inference servers. Low memory bandwidth (300 GB/s) relative to price/VRAM makes it a poor decode-speed choice versus an A5000/A6000 despite being newer; its value proposition is power/density, not raw tok/s. No sourced llama.cpp or vLLM decode-speed benchmark found this session (see gaps file). New-old-stock 'new' figure of $2,750 (Newegg) is now available via the same gpudojo.com tracker -- notably BELOW the $3,500 used/eBay figure, which is unusual and likely reflects thin/noisy used listings for this still-in-production card rather than a stable secondary-market premium; treat the used figure with caution.

Specification sources

Measurements

Decode is token generation — the speed you feel while an answer streams. Prefill is prompt processing — the wait before it starts. Why bandwidth predicts decode speed.

What verified, single-source and estimated mean, and the same rows with every filter and sort in the benchmarks explorer.

Speed over time

What it can run

Fit is arithmetic, not a measurement — how it is computed. Against 24 GB; models that fit are listed largest first.