Tagged articles

M5 Ultra

6 articles · Page 1 of 1
Lao Guo's Learning Space
Lao Guo's Learning Space
Sep 23, 2026 · Artificial Intelligence

M5 Ultra 256GB vs RTX 5090 32GB: Local LLM Inference Compared

This article compares Apple M5 Ultra (256GB unified memory) and NVIDIA RTX 5090 (32GB GDDR7) for local LLM inference across 27B–100B+ models, measuring capacity limits, token throughput, power draw, and CUDA vs MLX ecosystems, concluding M5 Ultra excels beyond 32GB while RTX 5090 leads on smaller models.

CUDALLM inferenceM5 Ultra
0 likes · 10 min read
M5 Ultra 256GB vs RTX 5090 32GB: Local LLM Inference Compared
Lao Guo's Learning Space
Lao Guo's Learning Space
Sep 23, 2026 · Artificial Intelligence

M5 Ultra Mac Studio: The Dream Machine for Local AI Agents?

Federico Viticci's four-day deep test of the M5 Ultra Mac Studio (256GB unified memory) reveals 1.2TB/s bandwidth, 2.5x faster prefill, up to 93.5% faster long-context generation, and superior concurrency for 24/7 local AI agent workflows at near-zero cost, outperforming RTX 5090 in usability.

BenchmarkConcurrencyM5 Ultra
0 likes · 12 min read
M5 Ultra Mac Studio: The Dream Machine for Local AI Agents?
Lao Guo's Learning Space
Lao Guo's Learning Space
Sep 22, 2026 · Industry Insights

M5 Ultra 256GB Benchmarks: 27B LLM 51 tok/s, GPU +82% in Cyberpunk 4K RT

Hands-on benchmarks of the M5 Ultra 256GB Mac Studio show 45.7% faster single-stream LLM decoding, 3.6x higher multi-user prefill throughput, 43.6% CPU multi-core gains, 67.8% GPU improvement, and 82.8% higher Cyberpunk 2077 4K ray-tracing frame rates versus M3 Ultra, with purchasing guidance.

BenchmarkGPU performanceLLM inference
0 likes · 13 min read
M5 Ultra 256GB Benchmarks: 27B LLM 51 tok/s, GPU +82% in Cyberpunk 4K RT
Lao Guo's Learning Space
Lao Guo's Learning Space
Sep 21, 2026 · Artificial Intelligence

M5 Ultra 256GB: Which LLMs Fit? 27B to 200B+ Capacity Analysis

The article analyzes Apple M5 Ultra 256GB unified memory capacity for local LLM inference, showing Q4-quantized models from 27B to 200B+ fit, but KV cache overhead for long contexts reduces headroom; it compares with multi-GPU setups and advises on purchase decisions.

Apple SiliconHardware AnalysisKV Cache
0 likes · 8 min read
M5 Ultra 256GB: Which LLMs Fit? 27B to 200B+ Capacity Analysis