Learn / Out
Compilers and codegen
Read first
No -O level changes the target instruction set: without -march or -mcpu, every instruction emitted belongs to the default target ISA, so target flags come before any judgement of codegen.
- Sources
- 21 entries in 4 parts
- Reproduce it
- 08-autovectorization-aliasing
- Related sections
- §7 Single-thread optimisation
- In the MCP server
cpuperf://section/8
Reading emitted code
-
01 Compiler Explorer repository
Shows how a source change alters the emitted instructions across compilers, versions and flags, with nothing installed.
-
02
Explains how the signed-overflow and aliasing rules let a trip count be known and a store loop become memset.
-
03 llvm-objdump manual
Reads the binary that shipped, with source lines and symbolised branch targets, rather than a recompiled snippet.
-
04 llvm-mca manual
Predicts loop throughput and port pressure from the scheduling model, and states it models neither front end nor caches.
-
05 llvm-exegesis manual
Measures instruction latency and throughput with counters, so the model llvm-mca predicts from is checked, not trusted.
Optimisation levels, inlining and link time
-
01
Lists what each -O level turns on, the inlining limits, and that -Ofast admits transforms invalid for conforming code.
-
02
Shows with real codegen that an abstraction is free only when inlining and the ABI allow it, and the cost when either refuses.
-
03 Itanium C++ ABI manual
Fixes the rule that a non-trivial class goes by reference to a caller-made temporary, the cost a wrapped pointer pays.
-
04
States what PLT calls and interposition cost, and the visibility controls a library needs to inline its own exports.
-
05 LTO Overview (GCC Internals) manual
Defines whole-program LTO against partitioned WHOPR, and the LGEN, WPA and LTRANS stages that run -flto in parallel.
-
06 ThinLTO manual
Defines the thin link, summaries analysed whole-program then parallel backends, and the cache for incremental rebuilds.
Target flags and auto-vectorisation
-
01 x86 Options (GCC) manual
Defines -march against -mtune, the psABI levels, and -mprefer-vector-width, the switch for full-width AVX-512 code.
-
02
Defines target_clones, one function per ISA behind a resolver the dynamic linker runs, so a generic build ships AVX-512.
-
03
Lists what -ffast-math implies, of which -fassociative-math alone frees a float reduction, and -ffp-contract for FMA.
-
04 Auto-Vectorization in LLVM manual
States what the vectorisers need, aliasing disproved or checked at run time, and where a float reduction stays in order.
-
05
Defines the remarks that make the compiler say which loop it left scalar and why, so the fix targets the real blocker.
Reproduce it
the vectoriser with and without restrict.
Reproduce it · 08-autovectorization-aliasing Auto-vectorisation and aliasing The vectoriser gives up on possible aliasing; a qualifier fixes itProfile-guided and post-link optimisation
-
01
Defines the instrumented and sampled workflows, why their profiles cannot mix, and the cost of a wrong training input.
-
02
Defines the address-to-source mapping with discriminators that lets a stale production profile still drive FDO.
-
03
States why a profile applied to the final binary beats one mapped to source, and the layout passes accuracy enables.
-
04 BOLT (llvm-project/bolt) repository
States what full effect needs, relocations kept at link time and a branch-stack sample profile, neither on by default.
-
05
States the design, not a result: basic-block sections and a relink, no binary rewrite, so layout needs no disassembly.