NVIDIA Tesla P100 16GB (PCIe)
Compare NVIDIA Tesla P100 16GB (PCIe) against another machine in the Arena
Reference: Nvidia Tesla on Wikipedia
- Class
- GPU
- Memory
- 16 GB HBM2
- Runs (Q4_K_M)
- —
- Bandwidth
- 732 GB/s
- TDP
- 250 W
- Released
- 2016-06-01
- Price (US)
- $78 used · $94 new as of 2026-08
- Price (Canada)
- CA$184.22 used sourced, as of 2026-08
Sourcing and disambiguation notes
Also sold as an SXM2 module with the same 732 GB/s HBM2, which needs a server carrier board and does not fit a standard desktop PCIe slot; a separate SXM2 llama.cpp result exists but is not merged here, to avoid mixing form factors. No video output, and passive cooling on most pulled units. Real-world decode on this chip is oddly bandwidth-underutilised in llama.cpp, with contributors reporting 41-43% memory-controller utilisation, so it under-performs its bandwidth advantage over the P40. The Canadian figure is for the PCIe card; cheaper SXM2 listings exist around CA$113 but need an adapter board, so the PCIe price is used as representative.
Specification sources
- https://gpudojo.com/tesla-p100 — press, 2026-08-22
- https://cputronic.com/en/gpu/nvidia-tesla-p100-pcie-16-gb — press, 2026-08-22
Measurements
Decode is token generation — the speed you feel while an answer streams. Prefill is prompt processing — the wait before it starts. Why bandwidth predicts decode speed.
No records yet for this hardware. Know of a published benchmark on it? Submit the link.
What verified, single-source and estimated mean, and the same rows with every filter and sort in the benchmarks explorer.
Speed over time
What it can run
Fit is arithmetic, not a measurement — how it is computed. Against 16 GB; models that fit are listed largest first.
No modeled quant fits in 16 GB.