Apple M4 Max (40-core GPU, 128GB)
Compare Apple M4 Max (40-core GPU, 128GB) against another machine in the Arena
Reference: Apple M4 on Wikipedia
- Class
- Mac
- Memory
- 128 GB LPDDR5X
- Runs (Q4_K_M)
- —
- Bandwidth
- 546 GB/s
- Released
- 2024-10-30
Sourcing and disambiguation notes
The machine most MLX benchmarking is done on, which is why it is here: a large share of published Apple Silicon results say only "M4 Max, 128GB" and never state the GPU core count. That omission does not have to block anything, because Apple offers 128GB on ONE bin only -- its Mac Studio specs read "48GB, 64GB, or 128GB unified memory (M4 Max with 16-core CPU and 40-core GPU)". So a 128GB M4 Max is necessarily the 40-core part at 546 GB/s, and the 32-core bin tops out at 36GB. Sold in both the Mac Studio and the 16-inch MacBook Pro; the form factor is recorded as desktop because the Studio is the configuration these runs usually come from, and nothing in this dataset distinguishes the two thermally. No price: 128GB is a build-to-order upgrade rather than one of the standard tiles Apple's buy page renders, so no figure could be tied to this configuration at the source.
Specification sources
- https://support.apple.com/en-us/122211 — vendor, 2026-08-23
- https://www.apple.com/newsroom/2024/10/apple-introduces-m4-pro-and-m4-max/ — vendor, 2024-10-30
Measurements
Decode is token generation — the speed you feel while an answer streams. Prefill is prompt processing — the wait before it starts. Why bandwidth predicts decode speed.
No records yet for this hardware. Know of a published benchmark on it? Submit the link.
What verified, single-source and estimated mean, and the same rows with every filter and sort in the benchmarks explorer.
Speed over time
What it can run
Fit is arithmetic, not a measurement — how it is computed. Against 128 GB; models that fit are listed largest first.
No modeled quant fits in 128 GB.