Tagged articles

Tesla M40

1 articles · Page 1 of 1
Refining Core Development Skills
Refining Core Development Skills
Aug 25, 2026 · Fundamentals

NVIDIA Maxwell Architecture: How SM Partitioning Drove 20x FP32 Growth in a Decade

The article analyzes NVIDIA's Maxwell GM200 architecture, detailing how splitting each SM into four independent processing blocks improved scheduling efficiency and core utilization, boosting FP32 performance to 6.84 TFLOPS on the Tesla M40 — a 20x increase over 10 years — while highlighting limited FP64 capability and memory bandwidth scaling challenges.

FP32 performanceGM200GPU architecture
0 likes · 12 min read
NVIDIA Maxwell Architecture: How SM Partitioning Drove 20x FP32 Growth in a Decade