AMD EPYC 7532 + 7x RTX 3090 (ik_llama.cpp, partial GPU offload)
Reference: Epyc on Wikipedia
- Class
- CPU
- Memory
- 512 GB DDR4-2933 (8 channel, Rome SP3 standard)
- Runs (Q4_K_M)
- —
- Bandwidth
- 187.7 GB/s
- TDP
- 200 W
- Released
- 2020-02-07
- Price (US)
- $11200 used as of 2026-08
Sourcing and disambiguation notes
Hybrid build: seven RTX 3090s assist a 671B model that mostly lives in CPU RAM under ik_llama.cpp's expert-offload scheme, so it is modelled as a CPU entry per this project's bottleneck convention -- about 2GB of VRAM is left per card after load, meaning the GPUs hold KV cache and attention layers, not expert weights. 32-core EPYC 7532. Bandwidth is computed as 8ch x DDR4-2933 x 8 bytes; the channel count is not stated by the source and standard Rome 8ch is assumed. The $11,200 used rig total breaks down as CPU ~$180, SP3 board ~$510, 8x32GB DDR4-2933 ECC RDIMM ~$1,200, and seven used RTX 3090s ~$8,820 -- the GPUs are about 79% of the cost, unlike the CPU-only builds here where DDR5 RDIMM dominates. Ancillary items (two 1600W PSUs, risers, an open-air frame; ~$466 together) are rough estimates. No new figure: a Rome CPU has no new-retail channel, and used cards are the only realistic way to source seven of them.
Specification sources
- https://huggingface.co/anikifoss/DeepSeek-R1-0528-DQ4_K_R4/discussions/2 — forum, 2025-06-04
Measurements
Decode is token generation — the speed you feel while an answer streams. Prefill is prompt processing — the wait before it starts. Why bandwidth predicts decode speed.
No records yet for this hardware. Know of a published benchmark on it? Submit the link.
What verified, single-source and estimated mean, and the same rows with every filter and sort in the benchmarks explorer.
Speed over time
What it can run
Fit is arithmetic, not a measurement — how it is computed. Against 512 GB; models that fit are listed largest first.
No modeled quant fits in 512 GB.