Fundamentals 21 min read

Tesla P100 Deep Dive: Pascal's FP16, HBM2, and NVLink Break Performance Ceilings

This article dissects NVIDIA's Pascal architecture via the Tesla P100, detailing its 16nm process, HBM2 memory delivering 732 GB/s bandwidth, FP16 support via FP32 units achieving 21.2 TFLOPS, and NVLink 1.0 providing 160 GB/s GPU interconnect bandwidth, with comparative calculations across GPU generations.

Refining Core Development Skills
Refining Core Development Skills
Refining Core Development Skills
Tesla P100 Deep Dive: Pascal's FP16, HBM2, and NVLink Break Performance Ceilings

Process Technology and GP100 Core Architecture

Pascal moves from Maxwell's 28 nm to a 16 nm process, nearly doubling transistor count to 15.3 billion. The GP100 die contains 6 GPCs (Graphics Processing Clusters), each with 10 SMs (Streaming Multiprocessors), for a total of 60 SMs. However, the Tesla P100 enables only 56 SMs, each with 64 CUDA cores, yielding 3,584 CUDA cores. The chip includes 8 memory controllers connected to 4 HBM2 stacks, a 4 MB L2 cache, and a High Speed Hub driving 4 NVLink interfaces.

HBM2 Memory: Solving the Memory Wall

Maxwell's Tesla M40 used 384-bit GDDR5 at 288.4 GB/s, nearing GDDR5's ~320 GB/s ceiling. Pascal adopts HBM2 (High Bandwidth Memory), stacking DRAM dies vertically with Through-Silicon Vias (TSVs). Each of the 4 stacks contains 8 layers of 128-bit wide DRAM, giving a 4096-bit aggregate bus. At 715 MHz clock (1.43 Gbps effective DDR), peak bandwidth reaches 732.2 GB/s — a 2.5× jump over M40. The trade-off is a slight capacity reduction to 16 GB.

HBM2 stack diagram
HBM2 stack diagram
GP100-HBM2 cross-section
GP100-HBM2 cross-section
GP100 package with four HBM2 stacks highlighted
GP100 package with four HBM2 stacks highlighted
Memory Bandwidth = Bus Width (bits) × Data Rate (Gbps) / 8
                = 4096 × 1.43 / 8
                ≈ 732.2 GB/s

SM Internal Architecture and FP16 Support

Each SM is partitioned into two processing blocks. Each block contains 32 CUDA cores (FP32/FP16/INT), 16 DP units (FP64), 8 LD/ST units, 8 SFUs (transcendental functions), 64 KB shared memory, and an L1 cache. Warp scheduling uses one scheduler with two dispatch units, unchanged from Maxwell.

FP16 is implemented by extending FP16 operands to FP32 in the FP32 units, then truncating results back to FP16. Pascal adds dedicated FP16 instructions (HADD, HMUL) to reduce conversion overhead. Each CUDA core can process 2 FP16 values per cycle (2-way SIMD), so FP16 throughput is theoretically 2× FP32.

GP100 SM block diagram
GP100 SM block diagram

NVLink 1.0: Breaking the PCIe Bottleneck

PCIe 3.0 x16 offers only 32 GB/s bidirectional bandwidth and microsecond latency, far below GP100's 732 GB/s memory bandwidth. NVLink 1.0 provides 4 links per GPU. Each link has 8 lanes per direction at 20 Gbps, yielding 20 GB/s unidirectional (40 GB/s bidirectional) per link. Four links give 160 GB/s aggregate bidirectional bandwidth — 5× PCIe 3.0 x16.

Key features: point-to-point connections (no intermediate switches), multi-link parallelism (bandwidth aggregates), and unified memory sharing with hardware page-fault support, allowing direct GPU-to-GPU memory access without CPU involvement.

Multi-GPU topologies:

2 GPUs: 4 NVLinks between them → 160 GB/s bidirectional.

3 GPUs: each pair gets 2 links → 80 GB/s per pair.

4 GPUs: some pairs have 2 links (80 GB/s), others only 1 link (40 GB/s).

8 GPUs: each GPU connects to 4 others with 1 link each → 40 GB/s per pair; some pairs require multi-hop routing, increasing latency.

NVLink 1.0 link structure
NVLink 1.0 link structure
8-GPU NVLink topology
8-GPU NVLink topology
Interconnect Bandwidth = NVLink Count × Per-Link Bandwidth × 2 (bidirectional)
                     = 4 × 20 GB/s × 2
                     = 160 GB/s

Compute Performance: FP16, FP32, FP64 Calculations

Boost clock: 1.480 GHz. CUDA cores: 3,584. DP units: 56 SMs × 2 blocks/SM × 16 DP/block = 1,792.

FP16: 1.480 GHz × 3584 × 4 ops/cycle = 21,217 GFLOPS ≈ 21.22 TFLOPS
FP32: 1.480 GHz × 3584 × 2 ops/cycle = 10,609 GFLOPS ≈ 10.61 TFLOPS
FP64: 1.480 GHz × 1792 × 2 ops/cycle = 5,304 GFLOPS ≈ 5.30 TFLOPS

These match NVIDIA's published specs. Compared to previous generations:

Architecture   Product            CUDA Cores   Clock (MHz)   FP32 (TFLOPS)   FP16 (TFLOPS)   FP64 (TFLOPS)
Tesla          GeForce 8800 Ultra 128          1512          0.387           -               -
Fermi          Tesla M2070        448          1150          1.030           -               0.515
Kepler         Tesla K20X         2688         732           3.935           -               1.311
Maxwell 2.0    Tesla M40          3072         1114          6.84            -               0.214
Pascal         Tesla P100 SXM2    3584         1480          10.61           21.22           5.30

Memory Bandwidth and Interconnect Bandwidth Comparisons

Architecture   Product            Memory   Type   Clock (MHz)   Data Rate (Gbps)   Bandwidth (GB/s)
Tesla          GeForce 8800 Ultra 768 MB   GDDR3  1080          2.2                103.7
Fermi          Tesla M2070        6 GB     GDDR5  783           3.1                150.3
Kepler         Tesla K20X         6 GB     GDDR5  1300          5.2                249.6
Maxwell 2.0    Tesla M40          24 GB    GDDR5  1502          6.0                288.4
Pascal         Tesla P100 SXM2    16 GB    HBM2   715           1.43               732.2
Architecture   Product            PCIe Version   PCIe BW (GB/s)   NVLink Version   NVLink BW (GB/s)
Tesla          GeForce 8800 Ultra PCIe 1.0 x16   8                -                -
Fermi          Tesla M2070        PCIe 2.0 x16   16               -                -
Kepler         Tesla K20X         PCIe 3.0 x16   32               -                -
Maxwell 2.0    Tesla M40          PCIe 3.0 x16   32               -                -
Pascal         Tesla P100 SXM2    PCIe 3.0 x16   32               1.0 × 4          160

Power Consumption

Despite massive performance gains, TDP rises only from 250 W (M40) to 300 W (P100), and recommended PSU from 600 W to 700 W, demonstrating improved performance-per-watt.

Original Source

Signed-in readers can open the original source through BestHub's protected redirect.

Sign in to view source
Republication Notice

This article has been distilled and summarized from source material, then republished for learning and reference. If you believe it infringes your rights, please contactadmin@besthub.devand we will review it promptly.

GPU architectureGPU interconnectNVLinkmemory bandwidthFP16HBM2FP32Tesla P100FP64GP100Pascal architecture
Refining Core Development Skills
Written by

Refining Core Development Skills

Fei has over 10 years of development experience at Tencent and Sogou. Through this account, he shares his deep insights on performance.

0 followers
Reader feedback

How this landed with the community

Sign in to like

Rate this article

Was this worth your time?

Sign in to rate
Discussion

0 Comments

Thoughtful readers leave field notes, pushback, and hard-won operational detail here.