---
title: "2. One instruction, end to end"
url: https://cpuperf.com/learn/one-instruction-end-to-end/
source: https://github.com/usamahz/cpu-performance-engineering/blob/deb5a0bac46760503b6f4a2608bdfed470c8532e/README.md#L149
commit: deb5a0bac46760503b6f4a2608bdfed470c8532e
---
## 2. One instruction, end to end

Vendors name the same structures differently (Intel's decoded ICache is AMD's op cache, and AMD's macro-op is Arm's MOP), so the stage names below are generic.

### Fetch and decode

- [Fetch Directed Instruction Prefetching](https://cseweb.ucsd.edu/~calder/papers/MICRO-99-FDP.pdf) - The origin of the decoupled front end, where the predictor runs ahead of fetch and drives instruction prefetch.
- [Micro-Operation Cache: A Power Aware Frontend for Variable Instruction Length ISA](https://ieeexplore.ieee.org/document/945363) - Introduces the decoded micro-op cache and its power case, the structure both x86 vendors later built.
- [Software Optimization Guide for the AMD Zen5 Microarchitecture](https://docs.amd.com/v/u/en-US/58455_1.00) - Where AMD states when the op cache feeds micro-ops, and the fusion and alignment rules for hot loops.
- [The microarchitecture of Intel, AMD, and VIA CPUs](https://www.agner.org/optimize/microarchitecture.pdf) - Measures rather than quotes each x86 core's misprediction penalty, micro-op cache and loop buffer behaviour.
- [Intel Mitigations for Jump Conditional Code Erratum](https://www.intel.com/content/www/us/en/content-details/841076/intel-mitigations-for-jump-conditional-code-erratum.html) - States what a microcode fix evicts from the decoded cache and which counters show the fall back to legacy decode.

Reproduce it: [misc/benchmarks/02-branch-misprediction](https://cpuperf.com/benchmarks/02-branch-misprediction/), the cost of a mispredicted branch, sorted against unsorted against branchless.

### Rename and issue

- [An Efficient Algorithm for Exploiting Multiple Arithmetic Units](https://ieeexplore.ieee.org/document/5392028) - The origin of tag-based renaming and reservation stations, from which every out-of-order issue queue descends.
- [Optimizing subroutines in assembly language](https://www.agner.org/optimize/optimizing_assembly.pdf) - Shows which chains renaming cannot break, from partial registers to flags, and the zero-cost idioms that do.
- [Measuring Reorder Buffer Capacity](https://blog.stuffedcow.net/2013/05/measuring-rob-capacity/) - Sets the user-space method that measures the window and register files, and which idioms take no physical register.

### Execute

- [Focusing Processor Policies via Critical-Path Prediction](https://dl.acm.org/doi/10.1145/379240.379253) - Defines the dependence graph of an out-of-order core and the critical path through it that sets run time.
- [Instruction tables: latencies, throughputs and micro-operation breakdowns](https://www.agner.org/optimize/instruction_tables.pdf) - Measures latency, reciprocal throughput and port assignment per instruction, the weights a dependence graph needs.
- [Arm Neoverse V2 Core Software Optimization Guide](https://support.arm.com/documentation/109898/latest/) - States every stage of one Arm core from fetch to issue, then per-instruction latency, throughput and pipe assignment.
- [Entropy Decoding in Oodle Data: x86-64 3-Stream Huffman Decoders](https://fgiesen.wordpress.com/2022/09/05/entropy-decoding-in-oodle-data-x86-64-3-stream-huffman-decoders/) - Predicts a real decoder loop's cycle count from its carried chain and issue slots, then measures the prediction.
- [Faster zlib/DEFLATE decompression on the Apple M1 (and x86)](https://dougallj.wordpress.com/2022/08/20/faster-zlib-deflate-decompression-on-the-apple-m1-and-x86/) - Predicts a decoder's refill chain from Arm core latencies, fits extra work under it, and measures the gain.

### Memory access and retire

- [Memory Dependence Prediction using Store Sets](https://people.csail.mit.edu/emer/media/papers/1990s/1998/1998.06.isca.storesets.pdf) - The origin of memory dependence prediction, which lets a load pass older stores and flushes on a wrong guess.
- [Store-to-Load Forwarding and Memory Disambiguation in x86 Processors](https://blog.stuffedcow.net/2014/01/x86-memory-disambiguation/) - Measures which store and load size and offset pairs forward or stall, and which cores predict memory dependences.
- [Microarchitecture Optimizations for Exploiting Memory-Level Parallelism](https://ieeexplore.ieee.org/document/1310765) - Defines memory-level parallelism as the misses one window overlaps, and shows which core limits cap it.
- [Implementing Precise Interrupts in Pipelined Processors](https://ieeexplore.ieee.org/document/4607) - Introduces the reorder buffer and defines a precise exception as in-order commit of out-of-order results.
- [Where Do Interrupts Happen?](https://travisdowns.github.io/blog/2019/08/20/interrupts.html) - Measures that an interrupt lands on the oldest unretired instruction, which decides what a sampling profiler blames.
