AMD EPYC 7702 + 4x RTX 3090 (partial GPU offload, DeepSeek MoE experts on CPU RAM)
Reference: Epyc on Wikipedia
- Class
- CPU
- Memory
- 512 GB DDR4-2400 (32GB ECC DIMMs; channel count not explicitly stated, 8ch assumed for Rome SP3)
- Runs (Q4_K_M)
- —
- Bandwidth
- 153.6 GB/s
- TDP
- 200 W
- Released
- 2019-08-07
- Price (US)
- $7700 used as of 2026-08
Sourcing and disambiguation notes
Hybrid build: four RTX 3090s (96GB of VRAM) are present, but the 671B model does not fit in VRAM, so decode speed reflects CPU RAM bandwidth and the machine is modelled as a CPU entry per this project's bottleneck convention. TDP and bandwidth describe the CPU and RAM subsystem only, with GPU power not summed in. 64-core EPYC 7702 on a Gigabyte MZ32-AR0. Bandwidth is computed as 8ch x DDR4-2400 x 8 bytes; the channel count is not stated by the source and standard Rome 8ch is assumed. The $7,700 used rig total breaks down as CPU ~$650, SP3 board ~$510, 8x32GB DDR4-2400 ECC RDIMM ~$1,200, and four used RTX 3090s ~$5,040, so the GPUs are about 65% of the cost. Ancillary items (PSU, risers, frame; ~$263 together) are rough estimates. No new figure: a Rome CPU has no new-retail channel, and used cards are the only realistic sourcing route.
Specification sources
- https://digitalspaceport.com/running-deepseek-r1-locally-not-a-distilled-qwen-or-llama/ — press, 2025-01-28
Measurements
Decode is token generation — the speed you feel while an answer streams. Prefill is prompt processing — the wait before it starts. Why bandwidth predicts decode speed.
No records yet for this hardware. Know of a published benchmark on it? Submit the link.
What verified, single-source and estimated mean, and the same rows with every filter and sort in the benchmarks explorer.
Speed over time
What it can run
Fit is arithmetic, not a measurement — how it is computed. Against 512 GB; models that fit are listed largest first.
No modeled quant fits in 512 GB.