NVIDIA Tesla P40
Compare NVIDIA Tesla P40 against another machine in the Arena
Reference: Nvidia Tesla on Wikipedia
- Class
- GPU
- Memory
- 24 GB GDDR5
- Runs (Q4_K_M)
- —
- Bandwidth
- 346 GB/s
- TDP
- 250 W
- Released
- 2016-09-01
- Price (US)
- $225 used · $329 new as of 2026-08
- Price (Canada)
- CA$339.49 used sourced, as of 2026-08
Sourcing and disambiguation notes
Passive-cooled server card: needs a 3rd-party blower shroud + high-static-pressure fan and an EPS-to-PCIe (CPU 8-pin) power adapter to run outside a server chassis. No video output. Extremely poor native FP16 (~1/64 of FP32 rate on Pascal), so GGUF/INT8-style quantized inference (llama.cpp) is the only practical path, not FP16 frameworks. No flash-attention kernel support. Price source: gpudojo.com price tracker (eBay-sourced, refreshed twice daily), accessed 2026-08-22. New-old-stock (NOS) 'new' figure of $329 is via Amazon, per the same gpudojo.com tracker cited above (used $225 via eBay). Canadian price: eBay.ca active listing, pre-owned 24GB GDDR5, plus ~$29 CAD shipping (price given is pre-tax item price only).
Specification sources
- https://images.nvidia.com/content/pdf/tesla/184427-Tesla-P40-Datasheet-NV-Final-Letter-Web.pdf — vendor, 2016-09-01
- https://gpudojo.com/tesla-p40 — press, 2026-08-22
Measurements
Decode is token generation — the speed you feel while an answer streams. Prefill is prompt processing — the wait before it starts. Why bandwidth predicts decode speed.
No records yet for this hardware. Know of a published benchmark on it? Submit the link.
What verified, single-source and estimated mean, and the same rows with every filter and sort in the benchmarks explorer.
Speed over time
What it can run
Fit is arithmetic, not a measurement — how it is computed. Against 24 GB; models that fit are listed largest first.
No modeled quant fits in 24 GB.