Tesla P100 Deep Dive: Pascal's FP16, HBM2, and NVLink Break Performance Ceilings
This article dissects NVIDIA's Pascal architecture via the Tesla P100, detailing its 16nm process, HBM2 memory delivering 732 GB/s bandwidth, FP16 support via FP32 units achieving 21.2 TFLOPS, and NVLink 1.0 providing 160 GB/s GPU interconnect bandwidth, with comparative calculations across GPU generations.
Process Technology and GP100 Core Architecture
Pascal moves from Maxwell's 28 nm to a 16 nm process, nearly doubling transistor count to 15.3 billion. The GP100 die contains 6 GPCs (Graphics Processing Clusters), each with 10 SMs (Streaming Multiprocessors), for a total of 60 SMs. However, the Tesla P100 enables only 56 SMs, each with 64 CUDA cores, yielding 3,584 CUDA cores. The chip includes 8 memory controllers connected to 4 HBM2 stacks, a 4 MB L2 cache, and a High Speed Hub driving 4 NVLink interfaces.
HBM2 Memory: Solving the Memory Wall
Maxwell's Tesla M40 used 384-bit GDDR5 at 288.4 GB/s, nearing GDDR5's ~320 GB/s ceiling. Pascal adopts HBM2 (High Bandwidth Memory), stacking DRAM dies vertically with Through-Silicon Vias (TSVs). Each of the 4 stacks contains 8 layers of 128-bit wide DRAM, giving a 4096-bit aggregate bus. At 715 MHz clock (1.43 Gbps effective DDR), peak bandwidth reaches 732.2 GB/s — a 2.5× jump over M40. The trade-off is a slight capacity reduction to 16 GB.
Memory Bandwidth = Bus Width (bits) × Data Rate (Gbps) / 8
= 4096 × 1.43 / 8
≈ 732.2 GB/sSM Internal Architecture and FP16 Support
Each SM is partitioned into two processing blocks. Each block contains 32 CUDA cores (FP32/FP16/INT), 16 DP units (FP64), 8 LD/ST units, 8 SFUs (transcendental functions), 64 KB shared memory, and an L1 cache. Warp scheduling uses one scheduler with two dispatch units, unchanged from Maxwell.
FP16 is implemented by extending FP16 operands to FP32 in the FP32 units, then truncating results back to FP16. Pascal adds dedicated FP16 instructions (HADD, HMUL) to reduce conversion overhead. Each CUDA core can process 2 FP16 values per cycle (2-way SIMD), so FP16 throughput is theoretically 2× FP32.
NVLink 1.0: Breaking the PCIe Bottleneck
PCIe 3.0 x16 offers only 32 GB/s bidirectional bandwidth and microsecond latency, far below GP100's 732 GB/s memory bandwidth. NVLink 1.0 provides 4 links per GPU. Each link has 8 lanes per direction at 20 Gbps, yielding 20 GB/s unidirectional (40 GB/s bidirectional) per link. Four links give 160 GB/s aggregate bidirectional bandwidth — 5× PCIe 3.0 x16.
Key features: point-to-point connections (no intermediate switches), multi-link parallelism (bandwidth aggregates), and unified memory sharing with hardware page-fault support, allowing direct GPU-to-GPU memory access without CPU involvement.
Multi-GPU topologies:
2 GPUs: 4 NVLinks between them → 160 GB/s bidirectional.
3 GPUs: each pair gets 2 links → 80 GB/s per pair.
4 GPUs: some pairs have 2 links (80 GB/s), others only 1 link (40 GB/s).
8 GPUs: each GPU connects to 4 others with 1 link each → 40 GB/s per pair; some pairs require multi-hop routing, increasing latency.
Interconnect Bandwidth = NVLink Count × Per-Link Bandwidth × 2 (bidirectional)
= 4 × 20 GB/s × 2
= 160 GB/sCompute Performance: FP16, FP32, FP64 Calculations
Boost clock: 1.480 GHz. CUDA cores: 3,584. DP units: 56 SMs × 2 blocks/SM × 16 DP/block = 1,792.
FP16: 1.480 GHz × 3584 × 4 ops/cycle = 21,217 GFLOPS ≈ 21.22 TFLOPS
FP32: 1.480 GHz × 3584 × 2 ops/cycle = 10,609 GFLOPS ≈ 10.61 TFLOPS
FP64: 1.480 GHz × 1792 × 2 ops/cycle = 5,304 GFLOPS ≈ 5.30 TFLOPSThese match NVIDIA's published specs. Compared to previous generations:
Architecture Product CUDA Cores Clock (MHz) FP32 (TFLOPS) FP16 (TFLOPS) FP64 (TFLOPS)
Tesla GeForce 8800 Ultra 128 1512 0.387 - -
Fermi Tesla M2070 448 1150 1.030 - 0.515
Kepler Tesla K20X 2688 732 3.935 - 1.311
Maxwell 2.0 Tesla M40 3072 1114 6.84 - 0.214
Pascal Tesla P100 SXM2 3584 1480 10.61 21.22 5.30Memory Bandwidth and Interconnect Bandwidth Comparisons
Architecture Product Memory Type Clock (MHz) Data Rate (Gbps) Bandwidth (GB/s)
Tesla GeForce 8800 Ultra 768 MB GDDR3 1080 2.2 103.7
Fermi Tesla M2070 6 GB GDDR5 783 3.1 150.3
Kepler Tesla K20X 6 GB GDDR5 1300 5.2 249.6
Maxwell 2.0 Tesla M40 24 GB GDDR5 1502 6.0 288.4
Pascal Tesla P100 SXM2 16 GB HBM2 715 1.43 732.2 Architecture Product PCIe Version PCIe BW (GB/s) NVLink Version NVLink BW (GB/s)
Tesla GeForce 8800 Ultra PCIe 1.0 x16 8 - -
Fermi Tesla M2070 PCIe 2.0 x16 16 - -
Kepler Tesla K20X PCIe 3.0 x16 32 - -
Maxwell 2.0 Tesla M40 PCIe 3.0 x16 32 - -
Pascal Tesla P100 SXM2 PCIe 3.0 x16 32 1.0 × 4 160Power Consumption
Despite massive performance gains, TDP rises only from 250 W (M40) to 300 W (P100), and recommended PSU from 600 W to 700 W, demonstrating improved performance-per-watt.
Signed-in readers can open the original source through BestHub's protected redirect.
This article has been distilled and summarized from source material, then republished for learning and reference. If you believe it infringes your rights, please contactand we will review it promptly.
Refining Core Development Skills
Fei has over 10 years of development experience at Tencent and Sogou. Through this account, he shares his deep insights on performance.
How this landed with the community
Was this worth your time?
0 Comments
Thoughtful readers leave field notes, pushback, and hard-won operational detail here.
