Reading path
From one instruction to a model served on CPU.
Start here
§1Ten numbered items, read top to bottom, each assuming only the ones before it.
Down
§2–6The core, the memory hierarchy, measurement, models.
-
02
One instruction, end to end
Vendors name the same structures differently (Intel's decoded ICache is AMD's op cache, and AMD's macro-op is Arm's MOP), so the stage names below are generic.
-
03
Microarchitecture
Where a design paper, the vendor manual and a measurement disagree about a core, the measurement is the one to trust and re-run.
-
04
Memory hierarchy
Line size and page size are machine parameters, not constants, so every padding and alignment rule below is applied against the target's own values.
-
05
Measurement
Cloud instances often virtualise the hardware counters away and macOS runs no Linux perf, so perf stat has to show a non-zero cycles count before any counter entry below is trusted.
-
06
Models
A roofline is a bound built from measured roofs and counted bytes, so a point above a roof means a wrong roof or a wrong byte count, not fast code.
Out
§7–15One thread, the compiler, many threads, NUMA, the kernel boundary, tail latency, CPU inference, the parts themselves, the benchmark suites.
-
07
Single-thread optimisation
A loop the compiler reports as vectorised can still run at scalar speed: a float reduction stays one serial chain until reassociation is permitted.
-
08
Compilers and codegen
No -O level changes the target instruction set: without -march or -mcpu, every instruction emitted belongs to the default target ISA, so target flags come before any judgement of codegen.
-
09
Concurrency
Every cost below is a cache line moving between cores, so the ordering models and the measured line-transfer cost in the memory hierarchy section come first.
-
10
NUMA and multi-socket
A page's node is decided at first touch, not when memory is allocated or a policy is set, and every vendor table below depends on the BIOS node mode of the machine it ran on.
-
11
OS and I/O
A syscall's cost depends on the mitigation state, the governor and the idle state the core was in, three sysfs settings that change after boot, so each is recorded beside any number below.
-
12
Tail latency and production systems
A latency figure means nothing without its percentile, its load model and the way it was recorded.
-
13
Inference on CPU
A decode step at batch one reads every weight once for a few flops, so it runs at the memory system's rate, and a CPU figure compares with a GPU figure only at the same batch and precision.
-
14
Hardware generations
The measurement articles below state no compiler, flags or run count, so each is kept for the structure it exposes and no figure from it is repeated.
-
15
Benchmarks
A score means what its suite's run rules say it means, so the rules come before the number.
Watchlist
§16Dated, for things whose evidence is still moving.