Sources
18 entries in 4 parts
In the MCP server
cpuperf://section/2

Fetch and decode

  1. 01
    Fetch Directed Instruction Prefetching paper

    The origin of the decoupled front end, where the predictor runs ahead of fetch and drives instruction prefetch.

  2. 02
    Micro-Operation Cache: A Power Aware Frontend for Variable Instruction Length ISA paper

    Introduces the decoded micro-op cache and its power case, the structure both x86 vendors later built.

  3. 03
    Software Optimization Guide for the AMD Zen5 Microarchitecture manual

    Where AMD states when the op cache feeds micro-ops, and the fusion and alignment rules for hot loops.

  4. 04
    The microarchitecture of Intel, AMD, and VIA CPUs manual

    Measures rather than quotes each x86 core's misprediction penalty, micro-op cache and loop buffer behaviour.

  5. 05
    Intel Mitigations for Jump Conditional Code Erratum manual

    States what a microcode fix evicts from the decoded cache and which counters show the fall back to legacy decode.

Reproduce it

the cost of a mispredicted branch, sorted against unsorted against branchless.

Reproduce it · 02-branch-misprediction Branch misprediction A mispredicted branch costs on the order of the pipeline depth; sorted or branchless data removes it

Rename and issue

  1. 01
    An Efficient Algorithm for Exploiting Multiple Arithmetic Units paper

    The origin of tag-based renaming and reservation stations, from which every out-of-order issue queue descends.

  2. 02
    Optimizing subroutines in assembly language manual

    Shows which chains renaming cannot break, from partial registers to flags, and the zero-cost idioms that do.

  3. 03
    Measuring Reorder Buffer Capacity report

    Sets the user-space method that measures the window and register files, and which idioms take no physical register.

Execute

  1. 01
    Focusing Processor Policies via Critical-Path Prediction paper

    Defines the dependence graph of an out-of-order core and the critical path through it that sets run time.

  2. 02
    Instruction tables: latencies, throughputs and micro-operation breakdowns manual

    Measures latency, reciprocal throughput and port assignment per instruction, the weights a dependence graph needs.

  3. 03
    Arm Neoverse V2 Core Software Optimization Guide manual

    States every stage of one Arm core from fetch to issue, then per-instruction latency, throughput and pipe assignment.

  4. 04
    Entropy Decoding in Oodle Data: x86-64 3-Stream Huffman Decoders report

    Predicts a real decoder loop's cycle count from its carried chain and issue slots, then measures the prediction.

  5. 05
    Faster zlib/DEFLATE decompression on the Apple M1 (and x86) report

    Predicts a decoder's refill chain from Arm core latencies, fits extra work under it, and measures the gain.

Memory access and retire

  1. 01
    Memory Dependence Prediction using Store Sets paper

    The origin of memory dependence prediction, which lets a load pass older stores and flushes on a wrong guess.

  2. 02
    Store-to-Load Forwarding and Memory Disambiguation in x86 Processors report

    Measures which store and load size and offset pairs forward or stall, and which cores predict memory dependences.

  3. 03
    Microarchitecture Optimizations for Exploiting Memory-Level Parallelism paper

    Defines memory-level parallelism as the misses one window overlaps, and shows which core limits cap it.

  4. 04
    Implementing Precise Interrupts in Pipelined Processors paper

    Introduces the reorder buffer and defines a precise exception as in-order commit of out-of-order results.

  5. 05
    Where Do Interrupts Happen? report

    Measures that an interrupt lands on the oldest unretired instruction, which decides what a sampling profiler blames.