AMD Instinct MI100 32GB
Compare AMD Instinct MI100 32GB against another machine in the Arena
Reference: AMD Instinct on Wikipedia
- Class
- GPU
- Memory
- 32 GB HBM2
- Runs (Q4_K_M)
- —
- Bandwidth
- 1229 GB/s
- TDP
- 300 W
- Released
- 2020-11-16
- Price (US)
- $880 used · $2499 new as of 2026-08
Sourcing and disambiguation notes
No sourced llama.cpp/vLLM/ROCm decode tokens/second figure could be found this session despite searching GitHub discussions and the exllama repo — a turboderp/exllama discussion (#247) only discusses MI100 being roughly "20% more tok/s than a 3090" in general terms with no reproducible number, model, or date, so it was not usable. Recorded as a benchmark gap below. ROCm-only; no Vulkan fallback maturity comparable to MI50/MI60 as of this session's research. 2026-08 pricing (GPUDojo aggregation, updated 2026-08-22): new $2,499, used $880.
Specification sources
- https://gpudojo.com/mi100 — press, 2026-08-22
- https://cputronic.com/gpu/amd-radeon-instinct-mi100 — press, 2026-08-22
Measurements
Decode is token generation — the speed you feel while an answer streams. Prefill is prompt processing — the wait before it starts. Why bandwidth predicts decode speed.
No records yet for this hardware. Know of a published benchmark on it? Submit the link.
What verified, single-source and estimated mean, and the same rows with every filter and sort in the benchmarks explorer.
Speed over time
What it can run
Fit is arithmetic, not a measurement — how it is computed. Against 32 GB; models that fit are listed largest first.
No modeled quant fits in 32 GB.