← Back to list

Macbook Pro M4 Max vs M5 Max : Quick LLM Speed Test

Let’s make it short and sweet

Laurent-Philippe Albou in GoPenAI · 2026-03-12 23:22 · 2 claps · 2.3 min read
#m4 #m5s #macbook-pro #llm #speed
Open on Medium ↗
Wiki topics: LLM · Large Language Models

Macbook Pro M4 Max vs M5 Max : Quick LLM Speed Test

Let’s make it short and sweet

M4 Max Nano Texture 16C / 40GPU — 128gb (left) vs M5 Max NO Nano Texture 18C / 40GPU — 128gb (right)

M4 Max Nano Texture 16C / 40GPU — 128gb (left) vs M5 Max NO Nano Texture 18C / 40GPU — 128gb (right)

Hi everyone, as I am certain this is a burning question for many, here is a micro test of the M4 Max vs M5 Max in terms of LLM preprocessing and processing speed. I also put a photo side-by-side for those hesitating to take the Nano Texture option : brigthness and colors are about the same so the Nano Texture actually doesn’t degrade the image quality as I initially feared — the difference in the photo above is due to True Tone that I forgot to deactivate on the M5 max.

The LLM Speed Test

LMStudio 0.4.6+1 Model : Qwen3.5 9B Context : 102k

Repeated twice for consistency, with the model unloaded to avoid any cache interference

Both laptops are in High Power modes to avoid biases from energy regulations.

Why Qwen3.5 9B ? First, it’s an excellent model for its size, but more importantly, it is a dense model, not a Mixture of Experts (MoE). And as it turns out, 9B parameters is actually not that fast. That’s also why MoE models with 2B (lfm2 24b A2B), 3B (the various qwen3 and qwen3.5 with A3B) or 5B (gpt-oss-20b) active parameters are so appreciated : they are blazing fast in comparison.

M4 Max 16C / 40GPU — 128gb

On average, the M4 Max preprocessed 102k tokens in 453s (around 225 tk/s) and it generated 16.16 tk/s.

M5 Max 18C / 40GPU — 128gb

On average, the M5 Max preprocessed the same 102k tokens in 227s (around 449 tk/s) and it generated 27.87 tk/s.

In a Nutshell

Yes, the M5 Max is much faster, but unsurprisingly not as fast as the 4x claim by Apple.

On average, expect the M5 max to be 2x faster for preprocessing (which is already great) and about 1.7x faster for generating tokens.

I am long overdue for an Open Source LLM Benchmark in 2026, so stay tuned for the next tests. This specific result has to be taken with a grain of salt : e.g. tested only with lmstudio on a single dense model. But it gives you at least a ballpark idea of what to expect.

Enjoy !


메타데이터
post_id
e678eb18e4d2
slug
macbook-pro-m4-max-vs-m5-max-quick-llm-speed-test-e678eb18e4d2
url
https://blog.gopenai.com/macbook-pro-m4-max-vs-m5-max-quick-llm-speed-test-e678eb18e4d2
canonical_url
https://blog.gopenai.com/macbook-pro-m4-max-vs-m5-max-quick-llm-speed-test-e678eb18e4d2
author_url
https://medium.com/@lpalbou
status
ok
fetched_at
2026-06-12 07:40:50