RTX 5070 Ti + RTX 3070 Ti (mixed dual-GPU)

Compare RTX 5070 Ti + RTX 3070 Ti (mixed dual-GPU) against another machine in the Arena

Class
Multi-GPU
Memory
24 GB GDDR7 + GDDR6X
Runs (Q4_K_M)
Bandwidth
608 GB/s
TDP
590 W
Released
2025-02-20
Price (US)
as of 2026-08
Sourcing and disambiguation notes

A mixed-generation pair, not two of the same card: 16GB GDDR7 (RTX 5070 Ti, 896 GB/s) plus 8GB GDDR6X (RTX 3070 Ti, 608 GB/s), split via llama.cpp --tensor-split so layers are sized proportional to each card's free VRAM. memory_gb is the sum (24GB combined, the usable pool for a tensor-split model). memory_bandwidth_gbs is deliberately set to the SLOWER card's figure (608, RTX 3070 Ti) rather than a sum or average: unlike the identical-card dual-4090/dual-3090 entries in this dataset, per-token decode on a heterogeneous split is gated by whichever card is still working on its layer slice, so the slower card is the honest bandwidth ceiling for this pair rather than an additive one. TDP summed at stock (300W + 290W). Released date is the newer card's launch (2025-02-20); this combination did not exist as a shippable product before then. No pair-bundle price exists for a mismatched pair -- left null rather than guessed, unlike the identical-card dual entries which calculate 2x a single-card price.

Specification sources

Measurements

Decode is token generation — the speed you feel while an answer streams. Prefill is prompt processing — the wait before it starts. Why bandwidth predicts decode speed.

What verified, single-source and estimated mean, and the same rows with every filter and sort in the benchmarks explorer.

Speed over time

What it can run

Fit is arithmetic, not a measurement — how it is computed. Against 24 GB; models that fit are listed largest first.