Sources
21 entries in 4 parts
Reproduce it
04-cache-latency
In the MCP server
cpuperf://section/4

Cache geometry, replacement and misses in flight

  1. 01
    What Every Programmer Should Know About Memory paper

    One measured account of DRAM timing, cache geometry, TLBs and prefetchers that sets the padding and prefetch rules.

  2. 02
    Measuring Cache and TLB Performance and Their Effect on Benchmark Run Times paper

    The origin of the strided-loop method that recovers cache and TLB size, line size, associativity and miss cost.

  3. 03
    Achieving Non-Inclusive Cache Performance with Inclusive Caches paper

    Names inclusion victims as the cost of an inclusive last-level cache, the case for non-inclusive and victim designs.

  4. 04
    Adaptive Insertion Policies for High Performance Caching paper

    The origin of set duelling and bimodal insertion, the adaptive replacement that survives a streaming pass.

  5. 05
    Lockup-Free Instruction Fetch/Prefetch Cache Organization paper

    Origin of the lockup-free cache, whose miss-status registers keep misses in flight, so a stream beats a chase.

Reproduce it

dependent-load latency from L1 to DRAM, with and without TLB pressure.

Reproduce it · 04-cache-latency Dependent-load latency across the memory hierarchy Dependent-load latency steps at each cache level, and page-random access adds TLB cost

TLBs, page walks and prefetchers

  1. 01
    Intel SDM Volume 3A: System Programming Guide, Part 1 manual

    Fixes the paging walk, the page-walk caches, the TLB invalidation rules and what each memory type permits.

  2. 02
    Translation Caching: Skip, Don't Walk (the Page Table) paper

    Defines the page-walk cache design space and shows why cached partial translations let a walk skip page-table levels.

  3. 03
    Improving Direct-Mapped Cache Performance paper

    The origin of the stream buffer that every vendor stream prefetcher descends from, and of the victim cache.

  4. 04
    Intel Optimization Reference Manual Volume 1 manual

    Names each prefetcher and what trains it, and which stop at a page boundary and which cross it.

  5. 05
    Arm Neoverse V2 Core Technical Reference Manual manual

    Names an Arm server core's load-side and store-side prefetchers, its TLB levels and the bits that disable them.

Store buffers, ordering and cache-line contention

  1. 01
    Memory Barriers: a Hardware View for Software Hackers paper

    Derives store buffers and invalidate queues from the cost of coherence, the reason reordering exists at all.

  2. 02
    x86-TSO: A Rigorous and Usable Programmer's Model for x86 Multiprocessors paper

    The store-buffer model of Intel and AMD ordering, tested on hardware, fixing which reorderings a fence pays for.

  3. 03
    Simplifying ARM Concurrency: Multicopy-Atomic Axiomatic and Operational Models for ARMv8 report

    The formal model Arm adopted into its architecture, and the record of why the non-multicopy-atomic option was dropped.

  4. 04
    Everything You Always Wanted to Know About Synchronization but Were Afraid to Ask paper

    Measures line-transfer cost between cores by coherence state and distance, the figure behind every contention rule.

  5. 05
    C2C - False Sharing Detection in Linux Perf report

    The implementers' account of perf c2c, whose worked example reads contended lines, offsets and callers off the report.

Struct layout, software prefetch and page size

  1. 01
    Cache-Conscious Structure Definition paper

    Introduces and measures structure splitting and field reordering on real programs, the origin of hot-cold layout rules.

  2. 02
    CppCon 2014: Data-Oriented Design and C++ talk

    Argues from cache-line utilisation arithmetic that layout must follow the access pattern, the case behind structure of arrays.

  3. 03
    dwarves (pahole) repository

    Prints a struct's holes, padding and cache-line boundaries from DWARF, so a layout is seen rather than guessed.

  4. 04
    When Prefetching Works, When It Doesn't, and Why paper

    Sorts software prefetch into the cases where it helps and hurts, and shows how it mistrains hardware prefetchers.

  5. 05
    Transparent Hugepage Support manual

    The kernel's statement of THP: the enabled and defrag knobs, khugepaged, and the counters showing what was obtained.

  6. 06
    Coordinated and Efficient Huge Page Management with Ingens paper

    Measures the fault latency, memory bloat and unfairness that eager THP promotion causes, and shows what removes them.