AMD EPYC 7532 + 7x RTX 3090 (ik_llama.cpp, partial GPU offload)

Compare AMD EPYC 7532 + 7x RTX 3090 (ik_llama.cpp, partial GPU offload) against another machine in the Arena

Class
CPU
Memory
512 GB DDR4-2933 (8 channel, Rome SP3 standard)
Runs (Q4_K_M)
Bandwidth
187.7 GB/s
TDP
200 W
Released
2020-02-07
Price (US)
$11200 used as of 2026-08
Sourcing and disambiguation notes

Hybrid build: seven RTX 3090s assist a 671B model that mostly lives in CPU RAM under ik_llama.cpp's expert-offload scheme, so it is modelled as a CPU entry per this project's bottleneck convention -- about 2GB of VRAM is left per card after load, meaning the GPUs hold KV cache and attention layers, not expert weights. 32-core EPYC 7532. Bandwidth is computed as 8ch x DDR4-2933 x 8 bytes; the channel count is not stated by the source and standard Rome 8ch is assumed. The $11,200 used rig total breaks down as CPU ~$180, SP3 board ~$510, 8x32GB DDR4-2933 ECC RDIMM ~$1,200, and seven used RTX 3090s ~$8,820 -- the GPUs are about 79% of the cost, unlike the CPU-only builds here where DDR5 RDIMM dominates. Ancillary items (two 1600W PSUs, risers, an open-air frame; ~$466 together) are rough estimates. No new figure: a Rome CPU has no new-retail channel, and used cards are the only realistic way to source seven of them.

Specification sources

Measurements

Decode is token generation — the speed you feel while an answer streams. Prefill is prompt processing — the wait before it starts. Why bandwidth predicts decode speed.

What verified, single-source and estimated mean, and the same rows with every filter and sort in the benchmarks explorer.

Speed over time

What it can run

Fit is arithmetic, not a measurement — how it is computed. Against 512 GB; models that fit are listed largest first.