AMD Instinct MI100 32GB

Compare AMD Instinct MI100 32GB against another machine in the Arena

Class
GPU
Memory
32 GB HBM2
Runs (Q4_K_M)
Bandwidth
1229 GB/s
TDP
300 W
Released
2020-11-16
Price (US)
$880 used · $2499 new as of 2026-08
Sourcing and disambiguation notes

No sourced llama.cpp/vLLM/ROCm decode tokens/second figure could be found this session despite searching GitHub discussions and the exllama repo — a turboderp/exllama discussion (#247) only discusses MI100 being roughly "20% more tok/s than a 3090" in general terms with no reproducible number, model, or date, so it was not usable. Recorded as a benchmark gap below. ROCm-only; no Vulkan fallback maturity comparable to MI50/MI60 as of this session's research. 2026-08 pricing (GPUDojo aggregation, updated 2026-08-22): new $2,499, used $880.

Specification sources

Measurements

Decode is token generation — the speed you feel while an answer streams. Prefill is prompt processing — the wait before it starts. Why bandwidth predicts decode speed.

What verified, single-source and estimated mean, and the same rows with every filter and sort in the benchmarks explorer.

Speed over time

What it can run

Fit is arithmetic, not a measurement — how it is computed. Against 32 GB; models that fit are listed largest first.