Tagged articles

unified memory

21 articles · Page 1 of 1
Lao Guo's Learning Space
Lao Guo's Learning Space
Sep 29, 2026 · Artificial Intelligence

MiniMax H3 vs Wan 2.2: 33B RAM vs 14B VRAM for Local Deployment

This article compares MiniMax H3 and Wan 2.2 for local deployment, contrasting H3's 33B unified-memory multimodal system with native stereo audio against Wan 2.2's 14B MoE visual family requiring GPU VRAM, covering architecture, quality, licensing, hardware requirements, and a decision framework.

Apache-2.0MiniMax H3VRAM
0 likes · 12 min read
MiniMax H3 vs Wan 2.2: 33B RAM vs 14B VRAM for Local Deployment
Ops Development & AI Practice
Ops Development & AI Practice
Sep 26, 2026 · Artificial Intelligence

Mac mini M6 32GB LLM Inference: Real TPS, Bandwidth Limits & Hybrid Cloud Strategy

This analysis dissects the 32GB Mac mini M6's true LLM inference capabilities, revealing 170 GB/s unified memory bandwidth limits, GPU compute constraints versus AMD Strix Halo, measured tokens-per-second for 8B–32B models, and a cost-driven argument for hybrid local-cloud deployment over expensive local-only workstations.

Apple SiliconLLM inferenceM6 chip
0 likes · 21 min read
Mac mini M6 32GB LLM Inference: Real TPS, Bandwidth Limits & Hybrid Cloud Strategy
Lao Guo's Learning Space
Lao Guo's Learning Space
Sep 23, 2026 · Artificial Intelligence

Mac mini 16GB to Studio 256GB: Which Mac Runs Your LLMs Best?

This article analyzes Mac mini and Mac Studio configurations for local LLM inference, showing how unified memory capacity determines which models fit and memory bandwidth dictates token generation speed, with real-world benchmarks for 8B to 235B models across M6, M5 Pro, M5 Max, and M5 Ultra chips.

Apple SiliconLLM inferenceMac Studio
0 likes · 13 min read
Mac mini 16GB to Studio 256GB: Which Mac Runs Your LLMs Best?
Lao Guo's Learning Space
Lao Guo's Learning Space
Sep 23, 2026 · Artificial Intelligence

M5 Ultra 256GB vs RTX 5090 32GB: Local LLM Inference Compared

This article compares Apple M5 Ultra (256GB unified memory) and NVIDIA RTX 5090 (32GB GDDR7) for local LLM inference across 27B–100B+ models, measuring capacity limits, token throughput, power draw, and CUDA vs MLX ecosystems, concluding M5 Ultra excels beyond 32GB while RTX 5090 leads on smaller models.

CUDALLM inferenceM5 Ultra
0 likes · 10 min read
M5 Ultra 256GB vs RTX 5090 32GB: Local LLM Inference Compared
Lao Guo's Learning Space
Lao Guo's Learning Space
Sep 23, 2026 · Artificial Intelligence

M5 Ultra Mac Studio: The Dream Machine for Local AI Agents?

Federico Viticci's four-day deep test of the M5 Ultra Mac Studio (256GB unified memory) reveals 1.2TB/s bandwidth, 2.5x faster prefill, up to 93.5% faster long-context generation, and superior concurrency for 24/7 local AI agent workflows at near-zero cost, outperforming RTX 5090 in usability.

BenchmarkConcurrencyM5 Ultra
0 likes · 12 min read
M5 Ultra Mac Studio: The Dream Machine for Local AI Agents?
Lao Guo's Learning Space
Lao Guo's Learning Space
Sep 22, 2026 · Industry Insights

M5 Ultra 256GB Benchmarks: 27B LLM 51 tok/s, GPU +82% in Cyberpunk 4K RT

Hands-on benchmarks of the M5 Ultra 256GB Mac Studio show 45.7% faster single-stream LLM decoding, 3.6x higher multi-user prefill throughput, 43.6% CPU multi-core gains, 67.8% GPU improvement, and 82.8% higher Cyberpunk 2077 4K ray-tracing frame rates versus M3 Ultra, with purchasing guidance.

BenchmarkGPU performanceLLM inference
0 likes · 13 min read
M5 Ultra 256GB Benchmarks: 27B LLM 51 tok/s, GPU +82% in Cyberpunk 4K RT
Lao Guo's Learning Space
Lao Guo's Learning Space
Sep 21, 2026 · Artificial Intelligence

M5 Ultra 256GB: Which LLMs Fit? 27B to 200B+ Capacity Analysis

The article analyzes Apple M5 Ultra 256GB unified memory capacity for local LLM inference, showing Q4-quantized models from 27B to 200B+ fit, but KV cache overhead for long contexts reduces headroom; it compares with multi-GPU setups and advises on purchase decisions.

Apple SiliconHardware AnalysisKV Cache
0 likes · 8 min read
M5 Ultra 256GB: Which LLMs Fit? 27B to 200B+ Capacity Analysis
Architects' Tech Alliance
Architects' Tech Alliance
Aug 26, 2026 · Artificial Intelligence

Inside Huawei’s Atlas 850E & 950 Supernodes: Key Technical Innovations

Huawei’s Atlas 850E wind‑cooled supernode delivers 14.27 PFLOPS, 768 GB HBM, 4 TB/s bandwidth and VCE phase‑change cooling for inference in standard data centers, while the Atlas 950 SuperPoD provides a flagship 1 EFLOPS FP8 training platform with 256 TB unified memory, 1.72 PB/s interconnect, sub‑3 µs RTT, full liquid cooling and scalability up to 8 192 NPU cards.

AI hardwareAtlas 850EAtlas 950
0 likes · 8 min read
Inside Huawei’s Atlas 850E & 950 Supernodes: Key Technical Innovations
Architects' Tech Alliance
Architects' Tech Alliance
Jul 20, 2026 · Artificial Intelligence

Supernode Architecture Explained: Definitions, Core Features, and Practical Use Cases

The whitepaper defines supernodes as high‑speed, tightly‑connected compute systems with unified memory addressing, microsecond‑level latency and terabyte‑per‑second bandwidth, outlines their physical, transaction, function and topology layers, demonstrates AI training and inference gains such as 80% communication reduction and 98.4% cluster scaling efficiency, and discusses industry impact, future scaling, standardization and green energy trends.

AI infrastructureHardware‑software co‑designhigh-speed interconnect
0 likes · 8 min read
Supernode Architecture Explained: Definitions, Core Features, and Practical Use Cases
Lao Guo's Learning Space
Lao Guo's Learning Space
Jun 3, 2026 · Industry Insights

Can Apple’s M5 Ultra Still Compete After NVIDIA’s RTX Spark Launch?

The RTX Spark desktop processor delivers 1 PFLOP of AI compute—about 14 times the M5 Ultra—while the M5 Ultra retains a three‑times higher memory bandwidth and twice the memory capacity, making it superior for certain inference workloads; the article breaks down specs, benchmarks, ecosystem differences, pricing and market positioning to show how each platform fits distinct AI use cases.

AI ComputeApple M5 UltraCUDA
0 likes · 12 min read
Can Apple’s M5 Ultra Still Compete After NVIDIA’s RTX Spark Launch?
Architects' Tech Alliance
Architects' Tech Alliance
Mar 7, 2024 · Industry Insights

How Nvidia GH200 and AMD MI300A Are Redefining CPU‑GPU Memory Integration

The article examines Nvidia’s GH200 and AMD’s MI300A processors, highlighting their unified memory domains that eliminate PCIe bottlenecks, detailing benchmark results, power‑measurement challenges, and the broader industry shift toward integrated CPU‑GPU architectures for high‑performance and generative‑AI workloads.

AMD MI300ABenchmarkCPU‑GPU Integration
0 likes · 11 min read
How Nvidia GH200 and AMD MI300A Are Redefining CPU‑GPU Memory Integration
Big Data Technology & Architecture
Big Data Technology & Architecture
Dec 6, 2021 · Big Data

Understanding Spark’s Memory Model: Unified Memory Management, On‑Heap and Off‑Heap Memory, and Configuration

This article explains Spark’s unified memory management model, detailing the division between on‑heap and off‑heap memory, the roles of execution, storage, user, and reserved memory, configuration parameters, dynamic allocation, and how these concepts affect performance and resource utilization.

Execution MemoryOff‑HeapSpark
0 likes · 17 min read
Understanding Spark’s Memory Model: Unified Memory Management, On‑Heap and Off‑Heap Memory, and Configuration