Go Performance 13-Year Retrospective: 26 Versions Benchmarked, 1.27 Beats 1.23 by 11.5%
Praneos Labs benchmarked 26 Go versions (1.2 to 1.27) on identical hardware, revealing three performance phases: compiler overhaul (1.5–1.7), a seven-year throughput plateau (1.8–1.23) focused on GC latency, and recent runtime optimizations (1.24–1.27) delivering an 11.5% speedup over 1.23, with map-heavy and allocation-heavy code benefiting most while numeric loops stagnated since 1.7.
Praneos Labs conducted a comprehensive benchmark of 26 Go versions — from 1.2.2 to 1.27.1 — using the same 12 programs (10 from the Benchmarks Game plus two pure-Go variants) on a single GCP c3-standard-8 VM (4 physical Xeon Platinum 8481C cores, 32 GB RAM, SMT disabled). Each version ran five iterations with page cache cleared between runs; median run-to-run variation was only 0.29%, and all outputs matched byte-for-byte. The full test suite, raw data, and scripts are open-sourced at github.com/nrnjn42/go-bench.
Methodology
Programs: 10 Benchmarks Game tasks; pidigits and regex-redux had C-library dependencies (GMP, PCRE), so pure-Go versions were added, totaling 12 programs.
Source adjustments: Five programs used APIs absent in Go 1.2 (e.g., strings.Builder, sort.Slice); each was patched with a one-line change that preserved output and performance on 1.27.
Versions: Latest patch of each minor release (26 total), fetched via Go module proxy and verified against the checksum database.
Hardware: GCP c3-standard-8, 4 physical cores, 32 GB RAM, SMT off.
Runs: 5 rounds, each program × each version once per round, page cache dropped before each round.
Quality checks: All 26 versions produced byte-identical output; VM steal time was 0 seconds across 1,560 timed runs; median run-to-run fluctuation 0.29%.
13-Year Evolution in Three Phases
Relative elapsed time normalized to Go 1.27 (= 1.00):
Version: 1.2 1.4 1.5 1.6 1.7 1.8 1.10 1.13
Rel. time: 1.69 1.83 1.42 1.39 1.21 1.16 1.16 1.12
Version: 1.16 1.20 1.23 1.24 1.25 1.26 1.27
Rel. time: 1.14 1.14 1.13 1.06 1.04 1.01 1.00The curve shows three distinct phases:
Phase 1 (Go 1.2 → 1.7): Compiler Overhaul
Go 1.5 rewrote the compiler from C to Go (self-hosting); Go 1.7 introduced the SSA backend. Immediate gains: n-body CPU time halved , fannkuch-redux dropped ~30% . Note a brief regression at 1.4→1.5 (1.83→1.42) reflecting the compiler rewrite “growing pains”; build times also slowed in 1.5 and 1.6.
Phase 2 (Go 1.8 → 1.23): Seven-Year Throughput Plateau
Throughput changed only ~3% across seven years. The runtime team focused on sub-millisecond GC pause reduction — benefits invisible in a throughput-oriented benchmark.
Phase 3 (Go 1.24 → 1.27): Runtime Push
Four versions delivered ~12% cumulative improvement via Swiss Table map (1.24), math/big (1.25), cgo (1.26), and GC/allocator tweaks (1.27).
Go 1.23 vs 1.27: Per-Program Breakdown
k-nucleotide: 6.67 s → 2.83 s (−57.5%) — Swiss Table map (1.24)
binary-trees: 8.77 s → 6.08 s (−30.7%) — GC & allocator; 1.27 alone gave 14%
pidigits (pure Go): 1.74 s → 1.46 s (−15.9%) — math/big (1.25)
regex-redux (PCRE): 2.31 s → 2.15 s (−7.1%) — cgo call speedup (1.26)
regex-redux (pure Go): 18.27 s → 17.92 s (−1.9%) — no significant change
fannkuch-redux: 7.46 s → 7.35 s (−1.5%) — no significant change
n-body, mandelbrot, fasta, spectral-norm, pidigits (GMP): ±1% — essentially unchanged
reverse-complement: 1.66 s → 1.72 s (+3.1%) — peak memory +29%
Geometric mean: −11.5% (CPU time −12.3%)
Programs heavy on maps, allocations, and GC sped up dramatically; pure numeric loops have barely moved since Go 1.7.
Machine code for n-body, mandelbrot, spectral-norm is nearly identical to years ago — the compiler has not significantly advanced on scalar numeric code. A side effect: reverse-complement slowed 3.1% with a 29% peak-memory increase; upgrade validation should include memory-profile monitoring.
Energy: Finishing Fast Beats Using Fewer Cores
Using the model from van Kempen et al. (2024) “It’s Not Easy Being Green” (RAPL data, 3,192 runs, 13 languages, R²=0.97):
energy ≈ 1.99 J × CPU-seconds + 286.6 J × elapsed-secondsCPU time alone explains almost none of the energy variance (R²≈0) because the idle server draws ~287 W. Therefore, wall-clock duration dominates energy consumption , not core count.
Go 1.27.1 vs 1.23.12: 11.5% less energy (model estimate).
Go 1.27.1 vs 1.8: 14% less energy .
Note: These are model projections, not direct power measurements.
Go vs C: The Real Gap Is Hand-Written SIMD
The 2017 Pereira et al. study (“Energy Efficiency across Programming Languages”) reported Go consuming 3.23× the energy of C (rank 14/27). That study likely used Go 1.7 or 1.8 (Ubuntu 16.10, code published days after Go 1.9 release). Re-running the original Go programs on 1.8 and 1.27 and scaling the 2017 results yields:
2017 original (Go ~1.7/1.8): Energy 3.25×C, Time 2.83×C, Rank ~14/27
Same programs, Go 1.27: Energy ~2.25×C, Time 2.07×C, Rank ~8/27
Go 1.27 + pre-allocated binary-trees: Energy ~1.57×C, Time 1.56×C, Rank ~4/27
Why does Go lag C on the official leaderboard? Five of the ten Benchmarks Game tasks have top C entries that use hand-written SIMD intrinsics:
n-body, spectral-norm: AVX ( __m256d)
mandelbrot: SSE2
fannkuch-redux, reverse-complement: _mm_shuffle_epi8 byte shuffling
Comparing Go against the fastest C (with SIMD) vs. the fastest scalar C (no SIMD):
n-body: Go ÷ Fastest C (SIMD) = 3.0×; Go ÷ Fastest Scalar C = 1.28×
spectral-norm: Go ÷ Fastest C (SIMD) = 3.6×; Go ÷ Fastest Scalar C = 1.00×
mandelbrot: Go ÷ Fastest C (SIMD) = 2.9×; Go ÷ Fastest Scalar C = 0.93×
fannkuch-redux: Go ÷ Fastest C (SIMD) = 3.9×; Go ÷ Fastest Scalar C = 1.15×
Once SIMD is removed, the gap collapses to 0–30%, with Go occasionally edging ahead. Remaining differences stem from: fewer compiler optimizations (Go prioritizes compile speed), C’s -march=<cpu> vs. Go’s GOAMD64=v2, bounds checks, GC overhead in allocation-heavy tasks, and some C entries relying on libraries like PCRE2-JIT or khash.
10× Speedup: Removing GC from binary-trees
binary-trees allocates millions of tiny nodes. A custom version was written that:
Pre-allocates all nodes in a slice per worker.
Uses int32 indices instead of pointers.
Eliminates GC scanning of those nodes.
Performs zero allocations after startup.
Results on Go 1.27:
Elapsed: Original (Go #2) 11.28 s → Pre-allocated 1.09 s
CPU time: 42.85 s → 3.41 s
Peak memory: 594 MB → 131 MB
~10× faster, reaching C-level performance. The Benchmarks Game rules forbid custom memory pools, but the C entry also uses APR pools, so the comparison is fair. An early version suffered false sharing : four 24-byte slice headers shared a cache line, making 4 cores slower than 1. Padding each header to 128 bytes cut time from 8.3 s to 1.09 s.
Key takeaway: For allocation-heavy workloads, redesigning memory layout yields far larger gains than upgrading Go versions.
Four Recommendations for Developers
Upgrade to the latest version. No code changes needed; 1.27 is 11.5% faster than 1.23. Map- and allocation-intensive services see even larger gains.
Expect runtime improvements, not compiler magic. Recent wins come from map (Swiss Table), GC, math/big, and cgo. Numeric loop performance has barely budged since Go 1.7.
In high-allocation scenarios, fix the memory layout first. Pointer-free arenas or pre-allocated pools can deliver 10× speedups — dwarfing version upgrades.
Scrutinize “Go vs C” benchmarks. Many reported gaps are actually SIMD gaps, not language gaps.
Caveats & Boundaries
All tests are small synthetic benchmarks measuring throughput ; latency and GC pause times are not captured.
Results come from one Intel Sapphire Rapids machine ; relative trends likely transfer, absolute numbers do not.
Energy figures are model estimates , not direct measurements.
The Pereira comparison and the pre-allocated binary-trees run used a shared dev VM with slightly lower precision than the main GCP suite.
Real-world workloads must be validated with your own load tests.
Original article:
https://praneos.com/blog/go-performance-1-2-to-1-27/Signed-in readers can open the original source through BestHub's protected redirect.
This article has been distilled and summarized from source material, then republished for learning and reference. If you believe it infringes your rights, please contactand we will review it promptly.
TonyBai
Tony Bai's tech world (tonybai.com). Not satisfied with just "knowing how", we strive for mastery. Focused on Go language internals, high-quality engineering practices, and cloud‑native architecture, exploring cutting‑edge intersections of Go and AI. Gophers who pursue technology are welcome—follow me and evolve with Go.
How this landed with the community
Was this worth your time?
0 Comments
Thoughtful readers leave field notes, pushback, and hard-won operational detail here.
