---
title: "1. Start here"
url: https://cpuperf.com/learn/start-here/
source: https://github.com/usamahz/cpu-performance-engineering/blob/deb5a0bac46760503b6f4a2608bdfed470c8532e/README.md#L130
commit: deb5a0bac46760503b6f4a2608bdfed470c8532e
---
## 1. Start here

Each entry assumes only the ones before it. The first and sixth are paid books; the third is a manual to open at the chapters its reason names.

1. [Computer Architecture: A Quantitative Approach, 7th Edition](https://shop.elsevier.com/books/computer-architecture/hennessy/978-0-443-15406-5) - Its pipelining appendix and memory chapters define the hazard, speculation and cache vocabulary the list assumes.
2. [Optimizing software in C++](https://www.agner.org/optimize/optimizing_cpp.pdf) - Maps C++ onto pipeline mechanisms and shows why a loop-carried dependency chain, not instruction count, paces a loop.
3. [Intel Optimization Reference Manual](https://www.intel.com/content/www/us/en/content-details/671488/intel-64-and-ia-32-architectures-optimization-reference-manual-volume-1.html) - Its opening chapters show how a shipping x86 core implements the textbook pipeline, each rule tied to a mechanism.
4. [What Every Programmer Should Know About Memory](https://www.akkadia.org/drepper/cpumemory.pdf) - Measures the step in cost per access at each cache boundary and the gap a prefetcher hides.
5. [Memory Barriers: a Hardware View for Software Hackers](http://www.rdrop.com/users/paulmck/scalability/paper/whymb.2010.07.23a.pdf) - Explains why a second core makes loads and stores reorder and what a barrier drains.
6. [Systems Performance: Enterprise and the Cloud, 2nd Edition](https://www.brendangregg.com/systems-performance-2nd-edition-book.html) - Puts the method before the tools: what to measure, in what order, and how benchmarks mislead.
7. [Roofline: An Insightful Visual Performance Model for Multicore Architectures](https://cacm.acm.org/research/roofline-an-insightful-visual-performance-model-for-multicore-architectures/) - Places a loop from a byte count and a datasheet bandwidth alone, before any counter is read.
8. [A Top-Down Method for Performance Analysis and Counters Architecture](https://sites.google.com/site/analysismethods/yasin-pubs) - Defines the split of pipeline slots into front end, bad speculation, back end and retiring, the tree profilers report.
9. [Performance Analysis and Tuning on Modern CPUs](https://github.com/dendibakh/perf-book) - Walks from a noisy timing to counters to a named bottleneck, applying roofline and top-down to whole programs.
10. [What Has My Compiler Done for Me Lately? Unbolting the Compiler's Lid](https://www.youtube.com/watch?v=bSkpMdDe4g4) - Shows how to read emitted assembly against its source, so each mechanism is checked in a listing, not assumed.

Work the exercises in [Performance Ninja](https://github.com/dendibakh/perf-ninja) alongside them; reading alone will not build the instinct.

Reproduce it: [misc/benchmarks/04-cache-latency](https://cpuperf.com/benchmarks/04-cache-latency/), the cost per dependent load stepping up at each cache boundary, the curve the fourth entry measures.
