# CPU Performance Engineering > Making a program fast on a modern CPU means knowing what the core does with each instruction, where the time actually goes, and how to prove a change helped. This is the reading that gets you there, in the order that makes the next piece legible. 16 sections, 304 primary sources, 14 benchmarks with committed results, and the editorial record of 541 candidates left out. Every page has a Markdown twin at the same path with .md. ## Sections - [1. Start here](https://cpuperf.com/learn/start-here.md): Each entry assumes only the ones before it. The first and sixth are paid books; the third is a manual to open at the chapters its reason names. - [2. One instruction, end to end](https://cpuperf.com/learn/one-instruction-end-to-end.md): Vendors name the same structures differently (Intel's decoded ICache is AMD's op cache, and AMD's macro-op is Arm's MOP), so the stage names below are generic. - [3. Microarchitecture](https://cpuperf.com/learn/microarchitecture.md): Where a design paper, the vendor manual and a measurement disagree about a core, the measurement is the one to trust and re-run. - [4. Memory hierarchy](https://cpuperf.com/learn/memory-hierarchy.md): Line size and page size are machine parameters, not constants, so every padding and alignment rule below is applied against the target's own values. - [5. Measurement](https://cpuperf.com/learn/measurement.md): Cloud instances often virtualise the hardware counters away and macOS runs no Linux perf, so perf stat has to show a non-zero cycles count before any counter entry below is trusted. - [6. Models](https://cpuperf.com/learn/models.md): A roofline is a bound built from measured roofs and counted bytes, so a point above a roof means a wrong roof or a wrong byte count, not fast code. - [7. Single-thread optimisation](https://cpuperf.com/learn/single-thread-optimisation.md): A loop the compiler reports as vectorised can still run at scalar speed: a float reduction stays one serial chain until reassociation is permitted. - [8. Compilers and codegen](https://cpuperf.com/learn/compilers-and-codegen.md): No -O level changes the target instruction set: without -march or -mcpu, every instruction emitted belongs to the default target ISA, so target flags come before any judgement of codegen. - [9. Concurrency](https://cpuperf.com/learn/concurrency.md): Every cost below is a cache line moving between cores, so the ordering models and the measured line-transfer cost in the memory hierarchy section come first. - [10. NUMA and multi-socket](https://cpuperf.com/learn/numa-and-multi-socket.md): A page's node is decided at first touch, not when memory is allocated or a policy is set, and every vendor table below depends on the BIOS node mode of the machine it ran on. - [11. OS and I/O](https://cpuperf.com/learn/os-and-io.md): A syscall's cost depends on the mitigation state, the governor and the idle state the core was in, three sysfs settings that change after boot, so each is recorded beside any number below. - [12. Tail latency and production systems](https://cpuperf.com/learn/tail-latency-and-production-systems.md): A latency figure means nothing without its percentile, its load model and the way it was recorded. - [13. Inference on CPU](https://cpuperf.com/learn/inference-on-cpu.md): A decode step at batch one reads every weight once for a few flops, so it runs at the memory system's rate, and a CPU figure compares with a GPU figure only at the same batch and precision. - [14. Hardware generations](https://cpuperf.com/learn/hardware-generations.md): The measurement articles below state no compiler, flags or run count, so each is kept for the structure it exposes and no figure from it is repeated. - [15. Benchmarks](https://cpuperf.com/learn/benchmarks.md): A score means what its suite's run rules say it means, so the rules come before the number. - [16. Watchlist](https://cpuperf.com/watchlist.md): Everything below is real but unproven: no item yet has all three of a written specification, a part you can buy, and a public measurement stating every one of the seven fields. Each line says what would promote it. Vendor multiples never qualify. Last checked 2026-09-15. ## Benchmarks - [Branch misprediction](https://cpuperf.com/benchmarks/02-branch-misprediction.md): A mispredicted branch costs on the order of the pipeline depth; sorted or branchless data removes it - [Latency versus throughput](https://cpuperf.com/benchmarks/03-latency-vs-throughput.md): A dependency chain is bound by latency; independent accumulators are bound by issue throughput - [Dependent-load latency across the memory hierarchy](https://cpuperf.com/benchmarks/04-cache-latency.md): Dependent-load latency steps at each cache level, and page-random access adds TLB cost - [Measurement pitfalls](https://cpuperf.com/benchmarks/05-measurement-pitfalls.md): An unused result measures nothing; a single run is not a measurement - [Roofline](https://cpuperf.com/benchmarks/06-roofline.md): Arithmetic intensity predicts which roof binds a loop - [Array of structs versus structure of arrays](https://cpuperf.com/benchmarks/07-aos-vs-soa-simd.md): Layout decides bytes moved and whether the loop vectorises - [Auto-vectorisation and aliasing](https://cpuperf.com/benchmarks/08-autovectorization-aliasing.md): The vectoriser gives up on possible aliasing; a qualifier fixes it - [False sharing](https://cpuperf.com/benchmarks/09-false-sharing.md): Writers sharing a cache line serialise; padding restores scaling - [First touch](https://cpuperf.com/benchmarks/10-first-touch.md): Allocation is not placement; the first touch pays the fault and picks the home - [Syscall cost](https://cpuperf.com/benchmarks/11-syscall-cost.md): A kernel crossing has a fixed cost that request size does not amortise - [Coordinated omission](https://cpuperf.com/benchmarks/12-coordinated-omission.md): A closed-loop load generator hides stalls that an open-loop one reports - [SGEMM: naive loop, microkernel, vendor BLAS](https://cpuperf.com/benchmarks/13-sgemm-naive-vs-blas.md): A packed, register-blocked microkernel takes a naive loop to the vector unit's ceiling; the vendor BLAS is further ahead only where it owns a matrix unit - [P-core versus E-core](https://cpuperf.com/benchmarks/14-pcore-vs-ecore.md): The same code runs at different speeds by core type; a result without the core type is not comparable - [Memory bandwidth by thread count and working set](https://cpuperf.com/benchmarks/15-stream-bandwidth.md): Vendor bandwidth is a package number; one thread and cache-resident data cannot reveal it ## MCP server - [cpu-perf](https://cpuperf.com/mcp.md): An [MCP](https://modelcontextprotocol.io) server that gives any AI client the whole [CPU Performance Engineering](https://github.com/usamahz/cpu-performance-engineering) list as something it can search and reason over, instead of a page it has to be pasted. Connect it to Claude Code, Codex, Claude Desktop, Cursor or VS Code and use it for your own performance work: ask questions, paste `perf stat` or compiler output, and the client's own model writes the answer from what the server returns, with citations back to the sources. ## Data - [corpus.json](https://cpuperf.com/corpus.json): the parsed list, benchmarks and editorial record as JSON, schema at https://cpuperf.com/schema/corpus-v1.json ## Optional - [The whole list](https://cpuperf.com/learn/all.md): every section on one page - [Evidence rules](https://cpuperf.com/evidence.md): what earns a place, and the contributing rules - [Editorial record](https://cpuperf.com/evidence/record.md): what was left out, and why