AI PC · desktop
Apple Mac Studio (M4 Max)
The fastest memory of the unified-memory boxes, in a silent desktop - if you're happy on macOS.
- Memory
- 128 GB
- Bandwidth
- 546 GB/s
- AI compute
- -
- £ / GB
- £27
What it runs
With 128 GB of memory you can load, at 4-bit, a model up to roughly
~252B params
Verdict
The quickest tokens per second of the unified-memory machines, with lovely MLX tooling, on macOS only.
BEST FOR
- + fastest tokens of the three
- + people already on a Mac
NOT FOR
- - anyone who needs CUDA
- - value hunters
The Mac Studio M4 Max is the quiet surprise of local AI: a small, near-silent desktop that happens to have the fastest memory of any machine near this price. Load it with 128GB and it’ll run big models quicker than the PC-based unified-memory boxes, provided you’re happy living in macOS to do it.
The take
Of the three unified-memory machines worth considering, this is the fast one. The M4 Max moves memory at up to 546 GB/s on the full chip, roughly double what the DGX Spark or the Framework Desktop manage, and bandwidth is exactly what decides tokens per second once a model fits. So on any model that lives in both machines’ memory, the Mac answers noticeably quicker. Pair that with Apple’s MLX tooling, which is pleasant to use, and it’s a lovely local-AI box. The trade is plain: it’s macOS only, there’s no Nvidia CUDA, and Apple’s memory upgrades are famously not cheap.
What it’ll actually run
Kit it out with 128GB and the same big models come into range: a 70B at Q4 fits easily, with headroom for a healthy context, and you can push into larger models than any 24GB card allows. Nothing new there against its rivals - the capacity is similar.
What’s different is the pace. Thanks to that 546 GB/s of bandwidth on the full M4 Max, a 70B at Q4 runs appreciably faster here than on the 256-to-273 GB/s PC boxes - think low-to-mid double figures of tokens per second where they’d be stuck in single figures. Drop to a 30B or an 8B and it’s properly snappy. One thing to watch at the till: the binned M4 Max runs at 410 GB/s, the full chip at 546 GB/s, so pick the 40-core GPU option if bandwidth is why you’re here. On software, MLX and llama.cpp both run well on Apple silicon, and the machine idles cool and silent. It will spin its fan up and get warm under a long, sustained prompt - owners report it becoming clearly audible then - but for an always-on assistant ticking over, it’s usually the quietest thing on the desk.
Who should buy it
Buy it if you want the fastest tokens per second of the unified-memory machines and you’re already comfortable on a Mac. For someone who lives in macOS, wants a silent desktop, and fancies playing with MLX, it’s hard to beat - and it doubles as a properly capable workstation for everything else.
Steer clear if you need CUDA, or if value per pound is the deciding factor. Nvidia’s software still rules a lot of the AI world, and none of it runs here. A used 3090 is far cheaper for models that fit in 24GB, and the Framework Desktop loads the same big models for less if raw speed isn’t the point. But if you want big models answering quickly, quietly, and you don’t mind the Apple tax, the Mac Studio is the pick.
Settings people actually run
The configs owners land on, pulled from the community. A sensible starting point, not gospel - tune to your own kit.
Biggest model that fits
104B-class dense (Command R+) Q4 (~62GB)
Runs but slow, around 6.5 tokens/s; the Q6 version (~85GB) drops to about 4.5. Uses 90-113GB with no swap.
Everyday sweet spot
Sub-32B models (Gemma 3 27B, Qwen 32B) Q4/Q8 via MLX
Comfortable speeds; owners report the machine is happiest below 32B parameters.
Fast inference
MLX quants via LM Studio, not GGUF/llama.cpp
MLX is noticeably quicker than llama.cpp on Apple silicon for the same model.
Order tip
40-core GPU (546 GB/s), not the binned 32-core (410 GB/s)
Pick the full M4 Max if memory bandwidth, and therefore speed, is why you are buying.
What owners report
Real first-hand experience gathered from owners and the community.
- “
Owner test on an M4 Max 128GB: a 104B Command R+ at Q6 (~85GB) ran at 4.5 tokens/s, the Q4 version (~62GB) at 6.5 (about reading speed), using 90-113GB with no swap. A large prompt pushed the GPU to 90C and the fan to 1,700 rpm, audible in a quiet room, so it is not silent under sustained load.
- “
Same thread: the hardware struggles above 32B parameters and slows further as context grows; DeepSeek R1 only stays usable because it is a mixture-of-experts model with about 37B active parameters.
- “
Community figures put a dense Llama 70B at Q4 around 10-13 tokens/s on the full M4 Max, and MLX consistently beats GGUF/llama.cpp on Apple silicon.
2 claim(s) we couldn't fully verify
- · Power draw of 160W (power_w) - Apple only publishes a maximum continuous power of 480W for the whole unit, shared with the M3 Ultra config. 160W reflects an approximate real-world M4 Max inference-load draw, not a vendor figure.
- · A 70B at Q4 runs in low-to-mid double figures of tokens per second - Community-reported at roughly 10-13 tokens/s across several sources, but no single authoritative vendor or first-party benchmark.
Hands-on reviews we drew on
We don't just copy the spec sheet. These are the teardowns and hands-on reviews behind this page - worth watching in their own right.
▶Heavy Metal Cloud
Mac Studio vs. Nvidia: Running Large Models Locally!
▶Nomaditsu
DeepSeek R1: $5000 vs $1000 Computer | M4 Max 128GB vs M1 Air 16GB
- Mac Studio (2025) - Tech Specs (vendor spec)Apple · primary source
- M4 Max Studio 128GB - LLM testing (owner thread)JSRinUK (MacRumors) · forum
Common questions
Why is the Mac Studio M4 Max faster than the DGX Spark for local models?+
Bandwidth. The full M4 Max feeds memory at up to 546 GB/s, roughly double the DGX Spark or Framework Desktop, and bandwidth is what sets tokens per second once a model fits. So on a model both machines can load, the Mac answers noticeably quicker.
Does the binned M4 Max run slower?+
A little. The binned M4 Max runs at 410 GB/s, the full chip at 546 GB/s. If bandwidth is the reason you're buying, pick the 40-core GPU option at checkout so you get the full 546 GB/s rather than the cut-down version.
Can I run CUDA models on a Mac Studio?+
No. There's no Nvidia CUDA on Apple silicon. You run models through MLX or llama.cpp, both of which work well here, but anything that specifically requires CUDA won't run. That's the main trade-off to weigh before buying.