Tagged articles

memory bandwidth

18 articles · Page 1 of 1
Ops Development & AI Practice
Ops Development & AI Practice
Sep 26, 2026 · Artificial Intelligence

Mac mini M6 32GB LLM Inference: Real TPS, Bandwidth Limits & Hybrid Cloud Strategy

This analysis dissects the 32GB Mac mini M6's true LLM inference capabilities, revealing 170 GB/s unified memory bandwidth limits, GPU compute constraints versus AMD Strix Halo, measured tokens-per-second for 8B–32B models, and a cost-driven argument for hybrid local-cloud deployment over expensive local-only workstations.

Apple SiliconLLM inferenceM6 chip
0 likes · 21 min read
Mac mini M6 32GB LLM Inference: Real TPS, Bandwidth Limits & Hybrid Cloud Strategy
Lao Guo's Learning Space
Lao Guo's Learning Space
Sep 23, 2026 · Artificial Intelligence

Mac mini 16GB to Studio 256GB: Which Mac Runs Your LLMs Best?

This article analyzes Mac mini and Mac Studio configurations for local LLM inference, showing how unified memory capacity determines which models fit and memory bandwidth dictates token generation speed, with real-world benchmarks for 8B to 235B models across M6, M5 Pro, M5 Max, and M5 Ultra chips.

Apple SiliconLLM inferenceMac Studio
0 likes · 13 min read
Mac mini 16GB to Studio 256GB: Which Mac Runs Your LLMs Best?
Lao Guo's Learning Space
Lao Guo's Learning Space
Sep 22, 2026 · Industry Insights

M5 Ultra 64 vs 80 Core GPU: What ¥9750 Buys – Not Faster LLM Decoding

The article compares Apple M5 Ultra 64-core and 80-core GPU variants, revealing that the ¥9750 price difference adds 16 GPU cores, 6 CPU cores, and 512GB memory support, but identical 1.2TB/s memory bandwidth means LLM decoding speed remains unchanged; the upgrade only benefits GPU-intensive tasks like 3D rendering and video effects.

3D renderingApple M5 UltraGPU cores
0 likes · 10 min read
M5 Ultra 64 vs 80 Core GPU: What ¥9750 Buys – Not Faster LLM Decoding
Refining Core Development Skills
Refining Core Development Skills
Aug 25, 2026 · Fundamentals

NVIDIA Maxwell Architecture: How SM Partitioning Drove 20x FP32 Growth in a Decade

The article analyzes NVIDIA's Maxwell GM200 architecture, detailing how splitting each SM into four independent processing blocks improved scheduling efficiency and core utilization, boosting FP32 performance to 6.84 TFLOPS on the Tesla M40 — a 20x increase over 10 years — while highlighting limited FP64 capability and memory bandwidth scaling challenges.

FP32 performanceGM200GPU architecture
0 likes · 12 min read
NVIDIA Maxwell Architecture: How SM Partitioning Drove 20x FP32 Growth in a Decade
Machine Heart
Machine Heart
May 16, 2026 · Artificial Intelligence

Why More Compute Can't Fix LLM Inference Lag and Why RL Leads to Overtraining

In a deep interview, former Google TPU architect Reiner Pope explains that low‑concurrency fast‑mode services trade higher fees for faster streaming but are limited by memory‑bandwidth bottlenecks, that optimal concurrency balances compute and memory costs, and that pipeline‑parallel sparse expert models and reinforcement‑learning fine‑tuning introduce new inefficiencies and overtraining risks.

LLMOvertrainingReinforcement Learning
0 likes · 7 min read
Why More Compute Can't Fix LLM Inference Lag and Why RL Leads to Overtraining
AI Frontier Lectures
AI Frontier Lectures
Jan 12, 2026 · Industry Insights

Why LLM Inference Hits a Memory Wall – Four Hardware Research Directions

The article analyses the challenges of large‑language‑model inference, highlighting memory bandwidth and interconnect as the primary bottlenecks, and presents four research opportunities—high‑bandwidth flash, processing‑near‑memory, 3D memory‑logic stacking, and low‑latency interconnect—while evaluating current Nvidia solutions and proposing integrated architectural approaches.

3D stackingAI hardware researchLLM inference
0 likes · 22 min read
Why LLM Inference Hits a Memory Wall – Four Hardware Research Directions
Architects' Tech Alliance
Architects' Tech Alliance
Mar 31, 2025 · Industry Insights

GPGPU vs ASIC: Who Wins the AI Compute Race?

This article analyzes the trade‑offs between GPGPU and ASIC for AI workloads, covering precision, compute density, power efficiency, memory bandwidth, interconnect technologies like NVLink, and the strategic reasons why leading firms are investing in custom AI chips.

AI chipsASICGPGPU
0 likes · 8 min read
GPGPU vs ASIC: Who Wins the AI Compute Race?
Architects' Tech Alliance
Architects' Tech Alliance
Mar 30, 2025 · Industry Insights

Why Memory, Not Compute, Is the Bottleneck for Next‑Gen AI Chips

The article analyzes the rapid growth of AI model memory and compute demands, the slow increase of chip memory capacity, and argues that memory bandwidth and energy consumption, rather than raw compute, will dominate AI chip design, emphasizing multi‑tenancy, DSA flexibility, and data‑flow optimization.

AI chipsDSAEnergy Efficiency
0 likes · 7 min read
Why Memory, Not Compute, Is the Bottleneck for Next‑Gen AI Chips
Architects' Tech Alliance
Architects' Tech Alliance
Mar 13, 2025 · Fundamentals

How Memory Bandwidth and Latency Shape CPU Performance

The article explains how CPU computation latency arises from memory speed, bandwidth, and access delays, detailing the relationships among memory, bandwidth, and latency, and examines key factors such as clock frequency, pipelining, parallelism, cache hit rate, and signal propagation distances that together determine overall system performance.

CPUPerformance Optimizationcomputer architecture
0 likes · 9 min read
How Memory Bandwidth and Latency Shape CPU Performance
Alibaba Cloud Infrastructure
Alibaba Cloud Infrastructure
Mar 24, 2021 · Cloud Computing

LIBRA and CARE: Memory Bandwidth Management and Fault‑Tolerance Innovations Presented at HPCA 2021

The article reviews two HPCA 2021 papers from Alibaba Cloud—LIBRA, a dynamic memory‑bandwidth management framework that boosts data‑center utilization, and CARE, a cache‑based fault‑tolerance architecture that delivers near‑Chipkill reliability with minimal overhead—while also highlighting future research directions in ML systems, quantum computing, and cache computing.

Cloud ComputingHPCA2021Resource Utilization
0 likes · 4 min read
LIBRA and CARE: Memory Bandwidth Management and Fault‑Tolerance Innovations Presented at HPCA 2021