---
title: "4. Memory hierarchy"
url: https://cpuperf.com/learn/memory-hierarchy/
source: https://github.com/usamahz/cpu-performance-engineering/blob/deb5a0bac46760503b6f4a2608bdfed470c8532e/README.md#L222
commit: deb5a0bac46760503b6f4a2608bdfed470c8532e
---
## 4. Memory hierarchy

Line size and page size are machine parameters, not constants, so every padding and alignment rule below is applied against the target's own values.

### Cache geometry, replacement and misses in flight

- [What Every Programmer Should Know About Memory](https://www.akkadia.org/drepper/cpumemory.pdf) - One measured account of DRAM timing, cache geometry, TLBs and prefetchers that sets the padding and prefetch rules.
- [Measuring Cache and TLB Performance and Their Effect on Benchmark Run Times](https://www2.eecs.berkeley.edu/Pubs/TechRpts/1993/Archive/CSD-93-767.pdf) - The origin of the strided-loop method that recovers cache and TLB size, line size, associativity and miss cost.
- [Achieving Non-Inclusive Cache Performance with Inclusive Caches](https://www.jaleels.org/ajaleel/publications/micro2010-tla.pdf) - Names inclusion victims as the cost of an inclusive last-level cache, the case for non-inclusive and victim designs.
- [Adaptive Insertion Policies for High Performance Caching](https://dl.acm.org/doi/10.1145/1250662.1250709) - The origin of set duelling and bimodal insertion, the adaptive replacement that survives a streaming pass.
- [Lockup-Free Instruction Fetch/Prefetch Cache Organization](https://dl.acm.org/doi/10.1145/285930.285979) - Origin of the lockup-free cache, whose miss-status registers keep misses in flight, so a stream beats a chase.

Reproduce it: [misc/benchmarks/04-cache-latency](https://cpuperf.com/benchmarks/04-cache-latency/), dependent-load latency from L1 to DRAM, with and without TLB pressure.

### TLBs, page walks and prefetchers

- [Intel SDM Volume 3A: System Programming Guide, Part 1](https://www.intel.com/content/www/us/en/content-details/671190/intel-64-and-ia-32-architectures-software-developer-s-manual-volume-3a-system-programming-guide-part-1.html) - Fixes the paging walk, the page-walk caches, the TLB invalidation rules and what each memory type permits.
- [Translation Caching: Skip, Don't Walk (the Page Table)](https://www.cs.rice.edu/CS/Architecture/docs/barr-isca10.pdf) - Defines the page-walk cache design space and shows why cached partial translations let a walk skip page-table levels.
- [Improving Direct-Mapped Cache Performance](https://ieeexplore.ieee.org/document/134547) - The origin of the stream buffer that every vendor stream prefetcher descends from, and of the victim cache.
- [Intel Optimization Reference Manual Volume 1](https://www.intel.com/content/www/us/en/content-details/671488/intel-64-and-ia-32-architectures-optimization-reference-manual-volume-1.html) - Names each prefetcher and what trains it, and which stop at a page boundary and which cross it.
- [Arm Neoverse V2 Core Technical Reference Manual](https://support.arm.com/documentation/102375/latest/) - Names an Arm server core's load-side and store-side prefetchers, its TLB levels and the bits that disable them.

### Store buffers, ordering and cache-line contention

- [Memory Barriers: a Hardware View for Software Hackers](http://www.rdrop.com/users/paulmck/scalability/paper/whymb.2010.07.23a.pdf) - Derives store buffers and invalidate queues from the cost of coherence, the reason reordering exists at all.
- [x86-TSO: A Rigorous and Usable Programmer's Model for x86 Multiprocessors](https://www.cl.cam.ac.uk/~pes20/weakmemory/cacm.pdf) - The store-buffer model of Intel and AMD ordering, tested on hardware, fixing which reorderings a fence pays for.
- [Simplifying ARM Concurrency: Multicopy-Atomic Axiomatic and Operational Models for ARMv8](https://www.cl.cam.ac.uk/~pes20/armv8-mca/) - The formal model Arm adopted into its architecture, and the record of why the non-multicopy-atomic option was dropped.
- [Everything You Always Wanted to Know About Synchronization but Were Afraid to Ask](https://sigops.org/s/conferences/sosp/2013/papers/p33-david.pdf) - Measures line-transfer cost between cores by coherence state and distance, the figure behind every contention rule.
- [C2C - False Sharing Detection in Linux Perf](https://joemario.github.io/blog/2016/09/01/c2c-blog/) - The implementers' account of perf c2c, whose worked example reads contended lines, offsets and callers off the report.

### Struct layout, software prefetch and page size

- [Cache-Conscious Structure Definition](https://www.microsoft.com/en-us/research/publication/cache-conscious-structure-definition-2/) - Introduces and measures structure splitting and field reordering on real programs, the origin of hot-cold layout rules.
- [CppCon 2014: Data-Oriented Design and C++](https://www.youtube.com/watch?v=rX0ItVEVjHc) - Argues from cache-line utilisation arithmetic that layout must follow the access pattern, the case behind structure of arrays.
- [dwarves (pahole)](https://github.com/acmel/dwarves) - Prints a struct's holes, padding and cache-line boundaries from DWARF, so a layout is seen rather than guessed.
- [When Prefetching Works, When It Doesn't, and Why](https://faculty.cc.gatech.edu/~hyesoon/lee_taco12.pdf) - Sorts software prefetch into the cases where it helps and hurts, and shows how it mistrains hardware prefetchers.
- [Transparent Hugepage Support](https://docs.kernel.org/admin-guide/mm/transhuge.html) - The kernel's statement of THP: the enabled and defrag knobs, khugepaged, and the counters showing what was obtained.
- [Coordinated and Efficient Huge Page Management with Ingens](https://www.usenix.org/conference/osdi16/technical-sessions/presentation/kwon) - Measures the fault latency, memory bloat and unfairness that eager THP promotion causes, and shows what removes them.
