Learn / Down
Memory hierarchy
Read first
Line size and page size are machine parameters, not constants, so every padding and alignment rule below is applied against the target's own values.
- Sources
- 21 entries in 4 parts
- Reproduce it
- 04-cache-latency
- Related sections
- §1 Start here, §7 Single-thread optimisation, §9 Concurrency
- In the MCP server
cpuperf://section/4
Cache geometry, replacement and misses in flight
-
01
One measured account of DRAM timing, cache geometry, TLBs and prefetchers that sets the padding and prefetch rules.
-
02
The origin of the strided-loop method that recovers cache and TLB size, line size, associativity and miss cost.
-
03
Names inclusion victims as the cost of an inclusive last-level cache, the case for non-inclusive and victim designs.
-
04
The origin of set duelling and bimodal insertion, the adaptive replacement that survives a streaming pass.
-
05
Origin of the lockup-free cache, whose miss-status registers keep misses in flight, so a stream beats a chase.
Reproduce it
dependent-load latency from L1 to DRAM, with and without TLB pressure.
Reproduce it · 04-cache-latency Dependent-load latency across the memory hierarchy Dependent-load latency steps at each cache level, and page-random access adds TLB costTLBs, page walks and prefetchers
-
01
Fixes the paging walk, the page-walk caches, the TLB invalidation rules and what each memory type permits.
-
02
Defines the page-walk cache design space and shows why cached partial translations let a walk skip page-table levels.
-
03
The origin of the stream buffer that every vendor stream prefetcher descends from, and of the victim cache.
-
04
Names each prefetcher and what trains it, and which stop at a page boundary and which cross it.
-
05
Names an Arm server core's load-side and store-side prefetchers, its TLB levels and the bits that disable them.
Store buffers, ordering and cache-line contention
-
01
Derives store buffers and invalidate queues from the cost of coherence, the reason reordering exists at all.
-
02
The store-buffer model of Intel and AMD ordering, tested on hardware, fixing which reorderings a fence pays for.
-
03
The formal model Arm adopted into its architecture, and the record of why the non-multicopy-atomic option was dropped.
-
04
Measures line-transfer cost between cores by coherence state and distance, the figure behind every contention rule.
-
05
The implementers' account of perf c2c, whose worked example reads contended lines, offsets and callers off the report.
Struct layout, software prefetch and page size
-
01
Introduces and measures structure splitting and field reordering on real programs, the origin of hot-cold layout rules.
-
02
Argues from cache-line utilisation arithmetic that layout must follow the access pattern, the case behind structure of arrays.
-
03 dwarves (pahole) repository
Prints a struct's holes, padding and cache-line boundaries from DWARF, so a layout is seen rather than guessed.
-
04
Sorts software prefetch into the cases where it helps and hurts, and shows how it mistrains hardware prefetchers.
-
05 Transparent Hugepage Support manual
The kernel's statement of THP: the enabled and defrag knobs, khugepaged, and the counters showing what was obtained.
-
06
Measures the fault latency, memory bloat and unfairness that eager THP promotion causes, and shows what removes them.