Local AI Hardware Bench
Which machine runs the model you want, how fast, and what it costs. Every figure is a real measurement with a named source and a date.
- Confidence is computed from the sources, never asserted —
verified,single-source,disputedorestimated. - Every measurement is dated and sortable, because the same hardware gets faster as backends improve.
- Anyone can add to it or argue with it — submit a benchmark, correct a number, or say what is wrong with the whole approach.
- 197 measured runs
- 261 citations
- 17 sources
- 115 contributors
- 34 verified
- 152 single-source
- 11 disputed
- 85 hardware
- 61 models
- 1849 estimated
What should I buy?
Uses measured runs where they exist; unmeasured picks are marked as an estimate. Full ranking on the advisor.
Ranked recommendations
Your ranked recommendations will appear here after you press Recommend.
Browse the data How the numbers are judged About this project
Most recent measurements
| Hardware | Model | Quant | Decode tok/s | Measured | Evidence |
|---|---|---|---|---|---|
| AMD Ryzen AI Max+ 395 (128GB unified, Strix Halo) | Qwen3.8 27B | Q4_K_M | 12.3 | 2026-08-22 | single-source: One credible source |
| AMD Ryzen AI Max+ 395 (128GB unified, Strix Halo) | Qwen3.8 27B | ROCmFP4_FAST | 34.8 | 2026-08-22 | single-source: One credible source |
| AMD Ryzen AI Max+ 395 (128GB unified, Strix Halo) | Qwen3.8 27B | ROCmFP4_FAST | 14 | 2026-08-22 | single-source: One credible source |
| NVIDIA GeForce RTX 3060 12GB | Llama 2 7B | Q4_0 | 75.6 | 2026-08-21 | single-source: One credible source |
| NVIDIA GeForce RTX 4090 | Qwen3.8 27B | Q4_K_M | 47.4 | 2026-08-15 | verified: Two or more independent sources agreeing within 25% |