Reading Assembly: The Only Way to Know What Your Compiler Actually Optimized

This article explains how reading compiler-generated assembly helps C++ developers verify optimizations like SIMD vectorization, function inlining, and memory layout efficiency, using tools like perf and Compiler Explorer to replace guesswork with evidence-based performance tuning.

IT Services Circle
IT Services Circle
IT Services Circle
Reading Assembly: The Only Way to Know What Your Compiler Actually Optimized

Many C++ developers assume the compiler will automatically inline functions, vectorize loops, and eliminate abstraction overhead. The only way to verify these assumptions is to examine the generated assembly.

Don't Start by Reading Assembly Line by Line

Effective performance analysis begins with profiling tools such as perf, VTune, or Instruments to locate hotspot functions. Then inspect the corresponding assembly in Compiler Explorer (Godbolt) using typical flags:

-O2
-O3
-march=native
-masm=intel

With GCC, add -fopt-info-vec and -fopt-info-vec-missed to see which loops vectorized and why others failed.

Key Signals in Hotspot Assembly

1. Loop SIMD Vectorization

Consider a simple scaling loop:

void scale(float* data, float factor, int n) {
  for (int i = 0; i < n; ++i) {
    data[i] *= factor;
  }
}

Scalar output shows vmulss xmm0, xmm1, xmm2 ( ss = scalar single precision). Successful vectorization shows vmulps ymm0, ymm1, ymm2 ( ps = packed single precision, ymm = 256-bit register handling 8 floats). However, vmulps xmm0, xmm1, xmm2 is also SIMD (128-bit, 4 floats). The decisive indicator is ss (scalar) vs ps (packed) suffixes. Vectorization does not guarantee 8× speedup; memory bandwidth, cache, and load/store throughput also matter.

2. Function Call Overhead and Inlining

Abstractions like inline float square(float x) { return x * x; } should disappear after inlining. If the hotspot assembly contains call square, inlining failed. This loss also prevents constant propagation, dead-code elimination, and cross-function optimizations. Virtual calls can be devirtualized and inlined if the compiler proves the concrete type. Functions defined in separate .cpp files without LTO also lose inlining opportunities.

3. Memory Access Patterns and Data Layout

An Array-of-Structures (AoS) layout:

struct Particle {
  bool active;
  float x, y, z;
  float vx, vy, vz;
  int id;
};
std::vector<Particle> particles;

If the hotspot loop only updates positions:

for (auto& p : particles) {
  p.x += p.vx;
  p.y += p.vy;
  p.z += p.vz;
}

Unused fields ( active, id, padding) consume cache lines, and consecutive x fields are spaced by sizeof(Particle). Switching to Structure-of-Arrays (SoA):

struct Particles {
  std::vector<float> x, y, z;
  std::vector<float> vx, vy, vz;
};

places all x values contiguously ( x0 x1 x2 x3 ...), enabling sequential loads, hardware prefetch, and SIMD. The optimal layout matches the hotspot access pattern; AoS is not universally slower.

Why the Compiler Refuses to Optimize

Often the compiler cannot prove an optimization is safe. Example:

void add(float* dst, const float* src, int n) {
  for (int i = 0; i < n; ++i) {
    dst[i] += src[i];
  }
}

The compiler must consider pointer aliasing: add(data + 1, data, n - 1) creates overlapping reads/writes, preventing loop unrolling or vectorization. Adding __restrict__ informs the compiler the pointers do not alias:

void add(float* __restrict__ dst, const float* __restrict__ src, int n)

This does not guarantee speedup but removes a constraint, allowing more aggressive optimization.

Similarly, an early-exit branch inside a loop:

float sum(const float* data, int n) {
  float total = 0.0f;
  for (int i = 0; i < n; ++i) {
    total += data[i];
    if (total < 0) break;
  }
  return total;
}

The if (total < 0) break; forces the compiler to assume any iteration may terminate the loop, inhibiting SIMD parallelization of subsequent iterations.

Performance optimization is not about teaching the compiler how to optimize, but providing enough information (via restrict, loop structure, data layout) so the compiler dares to optimize.

Conclusion

The practical workflow is: Profile → Find hotspot → Read assembly → Modify code → Benchmark. Assembly reveals what the compiler did ; benchmarking reveals whether it helped . Assembly's value is not a return to hand-written assembly, but a window into the compiler's decisions. Source code expresses intent; machine code is what the CPU executes. When facing a performance hotspot, look at the assembly instead of guessing.

Original Source

Signed-in readers can open the original source through BestHub's protected redirect.

Sign in to view source
Republication Notice

This article has been distilled and summarized from source material, then republished for learning and reference. If you believe it infringes your rights, please contactadmin@besthub.devand we will review it promptly.

performance profilingC++assemblymemory layoutSIMDcompiler optimizationperfAoS vs SoAfunction inliningGodboltpointer aliasingrestrict keyword
IT Services Circle
Written by

IT Services Circle

Delivering cutting-edge internet insights and practical learning resources. We're a passionate and principled IT media platform.

0 followers
Reader feedback

How this landed with the community

Sign in to like

Rate this article

Was this worth your time?

Sign in to rate
Discussion

0 Comments

Thoughtful readers leave field notes, pushback, and hard-won operational detail here.