M5 Ultra 256GB Benchmarks: 27B LLM 51 tok/s, GPU +82% in Cyberpunk 4K RT
Hands-on benchmarks of the M5 Ultra 256GB Mac Studio show 45.7% faster single-stream LLM decoding, 3.6x higher multi-user prefill throughput, 43.6% CPU multi-core gains, 67.8% GPU improvement, and 82.8% higher Cyberpunk 2077 4K ray-tracing frame rates versus M3 Ultra, with purchasing guidance.
The 256GB unified memory M5 Ultra Mac Studio finally has real-device data. This article compiles all raw numbers from a hands-on test video, comparing each item against the M3 Ultra — conclusion first: CPU and GPU across the board, local LLM single-stream decoding speed up nearly 50%.
Test Methodology: Clarifying the Scope
This machine tests Qwen3.8-27B-4bit , with weights of 15.7GB , running on Mac Studio. Two key methodological points must be stated upfront to avoid misinterpretation:
MTP disabled : Inference uses LM-only baseline (speculative decoding MTP turned off), so the reported decoding speeds represent a conservative lower bound , not the maximum achievable in daily use.
ANE alignment enabled : Test items enable ANE alignment, causing pp1024/pp4096 to appear as pp1025/pp4097 in the tables (an extra token to align prefill to 2048-block boundaries).
In short: these are conservative "capped" numbers; real-world daily usage will be faster.
Table 1: Qwen3.8-27B-4bit local inference, M5 Ultra vs M3 Ultra (MTP disabled conservative baseline)
1. AI Local Inference: Single-Stream Decoding +45.7%, Multi-Concurrency Prefill Up 3.6x
First, the most watched metric — how fast can it run large models locally .
Single-stream decoding speed is the most direct daily experience indicator: M5 Ultra reaches 51 tok/s , M3 Ultra 35 tok/s , a +45.7% improvement. Converting to effective bandwidth, the 48.4 tok/s tier corresponds to ~760 GB/s, while 51 tok/s reaches ~800 GB/s — essentially hitting the practical ceiling of this generation's unified memory bandwidth. Single-stream speed has plateaued; further gains rely on MTP and continuous batching.
Multi-concurrency is where this machine truly shines:
Aggregate decoding: 2 concurrent +46.3%, 4 concurrent +30.3%, but at 8 concurrent M5 (248.9) slightly trails M3 (257.4) by −3.3% — the only test where M5 falls behind.
Prefill : 2 concurrent 1243.3 vs 345.5 ( 3.6× ), 8 concurrent 1016.1 vs 282.6 ( 3.6× ) — massive advantage when multiple users query simultaneously.
8-concurrent end-to-end latency (pp1024/tg512): M5 12.176s vs M3 32.971s, 2.8× faster .
Conclusion: running a single large model, M5 is ~50% faster than M3; under multi-user/multi-task concurrency, M5's lead widens further — only in extreme 8-stream high-concurrency decoding do the two essentially tie.
2. CPU: Single-Core +25%, Multi-Core +43.6%
Table 2: CPU, GPU synthetic benchmarks and Cyberpunk 2077 real-world (vs M3 Ultra)
CPU tested with Geekbench 7 and Cinebench 2026:
Geekbench 7 single-core 3,711 (M3 2,963, +25.2% ), multi-core 51,760 (M3 40,271, +28.5% ).
Cinebench 2026 single-core 744 (M3 676, +10.1% ), multi-core 17,348 (M3 12,082, +43.6% ).
Single-core gains are modest, multi-core gains pronounced — consistent with the large-chip strategy of adding more cores.
3. GPU: Overall +34% to +68%, Graphics Performance Nears Desktop Class
GPU shows the largest generational leap:
Geekbench 7 GPU: 352,444 (M3 263,242, +33.9% ).
Cinebench 2026 GPU: 140,734 (M3 83,865, +67.8% ), now surpassing the RTX 5080 (100k-class) in this benchmark.
Blender 5.2.1 LTS (samples/minute, higher better): Junkshop scene 3,795 (M3 2,048, +85.3% ), Lion scene 4,806 (M3 2,861, +68% ) — nearly double the performance.
3DMark Steel Nomad: 8,052 (M3 5,548, +45.1% ); 3DMark Solar Bay Extreme: 26,105 (M3 16,426, +58.9% ).
In absolute terms, M5 Ultra's GPU scores exceed most desktop cards beyond the RTX 5080 (Steel Nomad 8,800) and RTX 5070 Ti (Solar Bay 27,915), representing a tier-breaking lead within the Mac lineup.
4. Real-World Gaming: Cyberpunk 2077 with Ray Tracing +82.8%
Ray tracing has been Apple Silicon's weak spot, but the gap is visibly narrowing:
At 4K resolution · Ultra Ray Tracing preset · MetalFX Quality mode , Cyberpunk 2077 average frame rates:
Ray tracing on: M5 37.28 FPS (M3 20.39, +82.8% ).
Ray tracing off: M5 60.32 FPS (M3 47.78, +26.2% ).
The ray-tracing improvement (+82.8%) is the second-highest single-item gain, indicating major generational GPU improvements in the ray-tracing pipeline. However, 37 FPS remains "playable but not smooth"; 4K ray-traced gaming is still not Apple Silicon's forte.
5. Gain Breakdown: Which Areas Improved Most?
Table 3: M5 Ultra vs M3 Ultra gain ranking (red indicates sole negative item)
Sorting the 14 test items by gain reveals a clear pattern:
First tier (+60% and above) : Blender Junkshop +85.3%, Cyberpunk RT on +82.8%, Blender Lion +68%, Cinebench GPU +67.8%, Solar Bay +58.9% — all graphics/rendering/ray-tracing related.
Second tier (+30% to +46%) : Multi-concurrent decoding, 3DMark, CPU multi-core, GPU overall.
Third tier (~+25%) : Geekbench CPU single/multi-core, Cyberpunk RT off.
Sole negative item : 8-concurrent aggregate decoding −3.3% .
In one sentence: M5 Ultra concentrates its gains on GPU and concurrent throughput , CPU improves steadily, extreme high-concurrency decoding roughly matches previous generation — fully consistent with the "larger die + more cores + stronger GPU" generational positioning.
6. Methodology Notes and Buying Advice
Several methodological caveats:
Inference data is MTP-disabled conservative baseline ; enabling MTP/continuous batching yields higher speeds.
Comparisons must account for identical engine settings — both MTP off, both LM-only baseline — otherwise architectural differences conflate with configuration differences.
For multi-concurrency, examine both "aggregate tok/s" and "per-stream tok/s" to avoid masking per-stream slowdown at 8-way behind aggregate throughput.
Only Qwen3.8-27B-4bit was tested; the community awaits 70B/200B head-to-heads, which are not yet available.
Buying advice :
If you need "single-machine local LLM + occasional video editing/rendering" , the M5 Ultra 256GB's graphics and concurrency gains are compelling , especially for multi-user shared inference and local agent orchestration.
If you only run single chat sessions, code, or edit 4K video — M4 Max (64–128GB) is already ample . The 256GB version costs nearly 7× the entry model; the value of paying for memory and concurrency only exists when you actually saturate them .
Final Thoughts
M5 Ultra 256GB proves with real benchmarks that it is not merely "more memory" but a generational upgrade across CPU, GPU, and local inference — especially in graphics rendering and concurrent throughput where it nearly doubles performance. 256GB unified memory lets you fit 200B+ parameter models, and these benchmarks show it runs them far faster than the previous generation.
Signed-in readers can open the original source through BestHub's protected redirect.
This article has been distilled and summarized from source material, then republished for learning and reference. If you believe it infringes your rights, please contactand we will review it promptly.
Lao Guo's Learning Space
AI learning, discussion, and hands‑on practice with self‑reflection
How this landed with the community
Was this worth your time?
0 Comments
Thoughtful readers leave field notes, pushback, and hard-won operational detail here.
