---
title: "8. Compilers and codegen"
url: https://cpuperf.com/learn/compilers-and-codegen/
source: https://github.com/usamahz/cpu-performance-engineering/blob/deb5a0bac46760503b6f4a2608bdfed470c8532e/README.md#L376
commit: deb5a0bac46760503b6f4a2608bdfed470c8532e
---
## 8. Compilers and codegen

No -O level changes the target instruction set: without -march or -mcpu, every instruction emitted belongs to the default target ISA, so target flags come before any judgement of codegen.

### Reading emitted code

- [Compiler Explorer](https://godbolt.org/) - Shows how a source change alters the emitted instructions across compilers, versions and flags, with nothing installed.
- [What Every C Programmer Should Know About Undefined Behavior](https://blog.llvm.org/2011/05/what-every-c-programmer-should-know.html) - Explains how the signed-overflow and aliasing rules let a trip count be known and a store loop become memset.
- [llvm-objdump](https://llvm.org/docs/CommandGuide/llvm-objdump.html) - Reads the binary that shipped, with source lines and symbolised branch targets, rather than a recompiled snippet.
- [llvm-mca](https://llvm.org/docs/CommandGuide/llvm-mca.html) - Predicts loop throughput and port pressure from the scheduling model, and states it models neither front end nor caches.
- [llvm-exegesis](https://llvm.org/docs/CommandGuide/llvm-exegesis.html) - Measures instruction latency and throughput with counters, so the model llvm-mca predicts from is checked, not trusted.

### Optimisation levels, inlining and link time

- [Options That Control Optimization (GCC)](https://gcc.gnu.org/onlinedocs/gcc/Optimize-Options.html) - Lists what each -O level turns on, the inlining limits, and that -Ofast admits transforms invalid for conforming code.
- [There Are No Zero-cost Abstractions (CppCon 2019)](https://www.youtube.com/watch?v=rHIkrotSwcc) - Shows with real codegen that an abstraction is free only when inlining and the ABI allow it, and the cost when either refuses.
- [Itanium C++ ABI](https://itanium-cxx-abi.github.io/cxx-abi/abi.html) - Fixes the rule that a non-trivial class goes by reference to a caller-made temporary, the cost a wrapped pointer pays.
- [How To Write Shared Libraries](https://www.akkadia.org/drepper/dsohowto.pdf) - States what PLT calls and interposition cost, and the visibility controls a library needs to inline its own exports.
- [LTO Overview (GCC Internals)](https://gcc.gnu.org/onlinedocs/gccint/LTO-Overview.html) - Defines whole-program LTO against partitioned WHOPR, and the LGEN, WPA and LTRANS stages that run -flto in parallel.
- [ThinLTO](https://clang.llvm.org/docs/ThinLTO.html) - Defines the thin link, summaries analysed whole-program then parallel backends, and the cache for incremental rebuilds.

### Target flags and auto-vectorisation

- [x86 Options (GCC)](https://gcc.gnu.org/onlinedocs/gcc/x86-Options.html) - Defines -march against -mtune, the psABI levels, and -mprefer-vector-width, the switch for full-width AVX-512 code.
- [Function Multiversioning (GCC)](https://gcc.gnu.org/onlinedocs/gcc/Function-Multiversioning.html) - Defines target_clones, one function per ISA behind a resolver the dynamic linker runs, so a generic build ships AVX-512.
- [Controlling Floating Point Behavior (Clang)](https://clang.llvm.org/docs/UsersManual.html#controlling-floating-point-behavior) - Lists what -ffast-math implies, of which -fassociative-math alone frees a float reduction, and -ffp-contract for FMA.
- [Auto-Vectorization in LLVM](https://llvm.org/docs/Vectorizers.html) - States what the vectorisers need, aliasing disproved or checked at run time, and where a float reduction stays in order.
- [Options to Emit Optimization Reports (Clang)](https://clang.llvm.org/docs/UsersManual.html#options-to-emit-optimization-reports) - Defines the remarks that make the compiler say which loop it left scalar and why, so the fix targets the real blocker.

Reproduce it: [misc/benchmarks/08-autovectorization-aliasing](https://cpuperf.com/benchmarks/08-autovectorization-aliasing/), the vectoriser with and without restrict.

### Profile-guided and post-link optimisation

- [Profile Guided Optimization (Clang)](https://clang.llvm.org/docs/UsersManual.html#profile-guided-optimization) - Defines the instrumented and sampled workflows, why their profiles cannot mix, and the cost of a wrong training input.
- [AutoFDO: Automatic Feedback-Directed Optimization for Warehouse-Scale Applications](https://research.google/pubs/autofdo-automatic-feedback-directed-optimization-for-warehouse-scale-applications/) - Defines the address-to-source mapping with discriminators that lets a stale production profile still drive FDO.
- [BOLT: A Practical Binary Optimizer for Data Centers and Beyond](https://arxiv.org/abs/1807.06735) - States why a profile applied to the final binary beats one mapped to source, and the layout passes accuracy enables.
- [BOLT (llvm-project/bolt)](https://github.com/llvm/llvm-project/tree/main/bolt) - States what full effect needs, relocations kept at link time and a branch-stack sample profile, neither on by default.
- [RFC: Propeller: A frame work for Post Link Optimizations](https://lists.llvm.org/pipermail/llvm-dev/2019-September/135393.html) - States the design, not a result: basic-block sections and a relink, no binary rewrite, so layout needs no disassembly.
