Fundamentals 12 min read

NVIDIA Maxwell Architecture: How SM Partitioning Drove 20x FP32 Growth in a Decade

The article analyzes NVIDIA's Maxwell GM200 architecture, detailing how splitting each SM into four independent processing blocks improved scheduling efficiency and core utilization, boosting FP32 performance to 6.84 TFLOPS on the Tesla M40 — a 20x increase over 10 years — while highlighting limited FP64 capability and memory bandwidth scaling challenges.

Refining Core Development Skills
Refining Core Development Skills
Refining Core Development Skills
NVIDIA Maxwell Architecture: How SM Partitioning Drove 20x FP32 Growth in a Decade

Process Technology

Maxwell GM200 retains the 28 nm process but uses the 28 HP (High Performance) variant. Strained silicon technology raises electron mobility by 15–20%, targeting high-compute GPUs.

Maxwell Architecture Overview

The GM200 die contains 6 GPCs, each with 4 SMs (SMMs). The chip keeps PCIe 3.0, adds six memory controllers connecting to GDDR5, and expands the shared L2 cache to 2–3 MB (2–4× Kepler). Larger on-chip cache reduces DRAM requests and lowers total board power.

SM (SMM) Internal Design

Each SMM is divided into four independent processing blocks. Each block holds 32 CUDA cores, 8 LD/ST units, and 8 SFUs. Total CUDA cores per SMM drop from 192 (Kepler) to 128, but the partitioned scheduler raises maximum active thread blocks per SMM from 16 to 32. NVIDIA states this improves average per-core realized performance by over 40%.

L1 cache is now separate from shared memory (previously shared in Kepler). Shared memory per SMM doubles from 48 KB to 96 KB. (Pascal later merges them again for hardware simplicity.)

Tesla M40 24 GB Compute Figures

FP32 Peak Performance

GM200 has 24 SMMs × 128 cores = 3,072 CUDA cores. Base clock 948 MHz; Boost reaches 1,114 MHz. FP32 TFLOPS = clock (GHz) × cores × 2 (FMA). At boost: 1.114 × 3,072 × 2 = 6,844 GFLOPS ≈ 6.84 TFLOPS.

FP32算力(TFLOPS) = 核心频率(GHz) × CUDA核心数 × 2(FMA双倍率)
FP32算力 = 核心频率(GHz) × CUDA核心数 × 2(FMA双倍率)
         = 1.114 * 3072 * 2
         = 6844.416 GFLOPS
         ≈ 6.84 TFLOPS

Generational FP32 Comparison

Tesla (GeForce 8800 Ultra): 128 cores @ 1,512 MHz → 387 GFLOPS

Fermi (Tesla M2070): 448 cores @ 1,150 MHz → 1.03 TFLOPS

Kepler (Tesla K20X): 2,688 cores @ 732 MHz → 3.94 TFLOPS

Maxwell 2.0 (Tesla M40): 3,072 cores @ 1,114 MHz → 6.84 TFLOPS

Over ~10 years FP32 peak grew ~20×.

CPU Comparison

Xeon E5-2699 v4 (22 cores, 2.2 GHz, AVX2+FMA3 → 16 FLOP/cycle): 22 × 2.2 × 16 = 761.6 GFLOPS. GPU FP32 throughput exceeds this CPU by >10×.

FP32算力  = 单 CPU 核数 × 单核主频 × 单个周期浮点计算值
          = 22 * 2.2 GHz * 16
          = 761.6 GFLOPS。

FP64 Performance

GM200 FP64 ratio is 1/32. Tesla M40 FP64 peak = 6.84 TFLOPS / 32 ≈ 0.213 TFLOPS, matching official specs.

FP64算力 = FP32算力 / 32
         ≈ 0.213 TFLOPS(理论值)

Memory Bandwidth

GDDR5 at 1,502 MHz (6 Gbps effective) on a 384-bit bus: 384 × 6 / 8 = 288 GB/s.

内存带宽 = 内存位宽 * 数据频率 / 8(换算成字节数)
         = 384 bit * 6 Gbps / 8
         = 288 GB/s

Generational Memory Comparison

Tesla (8800 Ultra): GDDR3, 2.2 Gbps, 103.7 GB/s

Fermi (M2070): GDDR5, 3.1 Gbps, 150.3 GB/s

Kepler (K20X): GDDR5, 5.2 Gbps, 249.6 GB/s

Maxwell (M40): GDDR5, 6 Gbps, 288.4 GB/s

Memory capacity jumped from 6 GB (K20X) to 24 GB. Bandwidth grew <3× over the same period while compute grew 20×, exposing a widening gap.

Summary of Maxwell Optimizations

SM control-block partitioning : four independent schedulers per SMM raise utilization and cut wasted power.

Larger L2 cache : 2–3 MB (2–4× Kepler) accelerates core data access and reduces DRAM traffic.

Independent L1 cache : no longer contends with shared memory; shared memory doubled to 96 KB.

Tesla M40 delivers 6.84 TFLOPS FP32 but only 0.213 TFLOPS FP64. The article notes that raising GDDR frequency to chase bandwidth increases power and hits limits; future generations (Pascal onward) adopt HBM for high bandwidth at lower frequency.

Maxwell GM200 package photo
Maxwell GM200 package photo
GM200 die shot
GM200 die shot
GM200 architecture block diagram
GM200 architecture block diagram
Maxwell SMM internal structure
Maxwell SMM internal structure
Official Tesla M40 specification table
Official Tesla M40 specification table
QR code or footer image
QR code or footer image
Original Source

Signed-in readers can open the original source through BestHub's protected redirect.

Sign in to view source
Republication Notice

This article has been distilled and summarized from source material, then republished for learning and reference. If you believe it infringes your rights, please contactadmin@besthub.devand we will review it promptly.

NVIDIAGPU architecturememory bandwidthFP32 performanceGM200Maxwell architectureSM partitioningTesla M40
Refining Core Development Skills
Written by

Refining Core Development Skills

Fei has over 10 years of development experience at Tencent and Sogou. Through this account, he shares his deep insights on performance.

0 followers
Reader feedback

How this landed with the community

Sign in to like

Rate this article

Was this worth your time?

Sign in to rate
Discussion

0 Comments

Thoughtful readers leave field notes, pushback, and hard-won operational detail here.