AMD EPYC 7702 + 4x RTX 3090 (partial GPU offload, DeepSeek MoE experts on CPU RAM)

Compare AMD EPYC 7702 + 4x RTX 3090 (partial GPU offload, DeepSeek MoE experts on CPU RAM) against another machine in the Arena

Class
CPU
Memory
512 GB DDR4-2400 (32GB ECC DIMMs; channel count not explicitly stated, 8ch assumed for Rome SP3)
Runs (Q4_K_M)
Bandwidth
153.6 GB/s
TDP
200 W
Released
2019-08-07
Price (US)
$7700 used as of 2026-08
Sourcing and disambiguation notes

Hybrid build: four RTX 3090s (96GB of VRAM) are present, but the 671B model does not fit in VRAM, so decode speed reflects CPU RAM bandwidth and the machine is modelled as a CPU entry per this project's bottleneck convention. TDP and bandwidth describe the CPU and RAM subsystem only, with GPU power not summed in. 64-core EPYC 7702 on a Gigabyte MZ32-AR0. Bandwidth is computed as 8ch x DDR4-2400 x 8 bytes; the channel count is not stated by the source and standard Rome 8ch is assumed. The $7,700 used rig total breaks down as CPU ~$650, SP3 board ~$510, 8x32GB DDR4-2400 ECC RDIMM ~$1,200, and four used RTX 3090s ~$5,040, so the GPUs are about 65% of the cost. Ancillary items (PSU, risers, frame; ~$263 together) are rough estimates. No new figure: a Rome CPU has no new-retail channel, and used cards are the only realistic sourcing route.

Specification sources

Measurements

Decode is token generation — the speed you feel while an answer streams. Prefill is prompt processing — the wait before it starts. Why bandwidth predicts decode speed.

What verified, single-source and estimated mean, and the same rows with every filter and sort in the benchmarks explorer.

Speed over time

What it can run

Fit is arithmetic, not a measurement — how it is computed. Against 512 GB; models that fit are listed largest first.