Learn / Start here
Start here
Read first
Each entry assumes only the ones before it. The first and sixth are paid books; the third is a manual to open at the chapters its reason names.
- Sources
- 10
- Reproduce it
- 04-cache-latency
- Related sections
- §3 Microarchitecture, §4 Memory hierarchy, §5 Measurement, §6 Models, §7 Single-thread optimisation
- In the MCP server
cpuperf://section/1
-
01
Its pipelining appendix and memory chapters define the hazard, speculation and cache vocabulary the list assumes.
-
02 Optimizing software in C++ manual
Maps C++ onto pipeline mechanisms and shows why a loop-carried dependency chain, not instruction count, paces a loop.
-
03
Its opening chapters show how a shipping x86 core implements the textbook pipeline, each rule tied to a mechanism.
-
04
Measures the step in cost per access at each cache boundary and the gap a prefetcher hides.
-
05
Explains why a second core makes loads and stores reorder and what a barrier drains.
-
06
Puts the method before the tools: what to measure, in what order, and how benchmarks mislead.
-
07
Places a loop from a byte count and a datasheet bandwidth alone, before any counter is read.
-
08
Defines the split of pipeline slots into front end, bad speculation, back end and retiring, the tree profilers report.
-
09
Walks from a noisy timing to counters to a named bottleneck, applying roofline and top-down to whole programs.
-
10
Shows how to read emitted assembly against its source, so each mechanism is checked in a listing, not assumed.
Work the exercises in Performance Ninja alongside them; reading alone will not build the instinct.
Reproduce it
the cost per dependent load stepping up at each cache boundary, the curve the fourth entry measures.
Reproduce it · 04-cache-latency Dependent-load latency across the memory hierarchy Dependent-load latency steps at each cache level, and page-random access adds TLB cost