---
title: "3. Microarchitecture"
url: https://cpuperf.com/learn/microarchitecture/
source: https://github.com/usamahz/cpu-performance-engineering/blob/deb5a0bac46760503b6f4a2608bdfed470c8532e/README.md#L185
commit: deb5a0bac46760503b6f4a2608bdfed470c8532e
---
## 3. Microarchitecture

Where a design paper, the vendor manual and a measurement disagree about a core, the measurement is the one to trust and re-run.

### Limits of ILP and SMT

- [The MIPS R10000 Superscalar Microprocessor](https://ieeexplore.ieee.org/document/491460) - Shows rename, the active list and precise branch recovery fitted together in one shipped out-of-order core.
- [Limits of Instruction-Level Parallelism (WRL Research Report 93/6)](https://davidwall.info/papers/WRL-TR-93.6.pdf) - Separates perfect from realistic prediction, renaming and aliasing, and shows how little ILP survives the real ones.
- [Complexity-Effective Superscalar Processors](https://ftp.cs.wisc.edu/sohi/papers/1997/isca.complexity.pdf) - Puts circuit delay on wakeup, select and bypass, capping issue width and window size before ILP runs out.
- [Simultaneous Multithreading: Maximizing On-Chip Parallelism](https://cseweb.ucsd.edu/~tullsen/isca95.pdf) - The origin of SMT, where another thread fills one thread's idle issue slots at a cost to both.
- [A Mechanistic Performance Model for Superscalar Out-of-Order Processors](https://users.elis.ugent.be/~leeckhou/papers/tocs09.pdf) - Turns window size, issue width and miss events into cycles lost, making an ILP limit a cycle count.

### Branch prediction and speculation

- [A Case for (Partially) TAgged GEometric History Length Branch Prediction](https://jilp.org/vol8/v8paper1.pdf) - Defines the tagged geometric-history predictor shipped designs converge on, and the limit of what it can learn.
- [A 64-Kbytes ITTAGE indirect branch predictor](https://jilp.org/jwac-2/program/cbp3_07_seznec.pdf) - Carries the tagged geometric scheme to indirect jumps and calls, where interpreters and virtual dispatch stall.
- [Characterizing the Branch Misprediction Penalty](https://users.elis.ugent.be/~leeckhou/papers/ispass06-eyerman.pdf) - Defines the penalty as pipeline refill plus window drain, so it exceeds pipeline depth and is not constant.
- [Spectre Attacks: Exploiting Speculative Execution](https://arxiv.org/abs/1801.01203) - Establishes that predictor state is shared and trainable across contexts, the root of every mitigation and its cost.
- [Speculative Execution Side Channel Mitigations](https://www.intel.com/content/www/us/en/developer/articles/technical/software-security-guidance/technical-documentation/speculative-execution-side-channel-mitigations.html) - Defines the indirect-branch and store-bypass controls (IBRS, STIBP, IBPB, SSBD) and states which carry a large cost.

### Vendor estimates and measured tables

- [Software Optimization Guide for the AMD Zen5 Microarchitecture](https://docs.amd.com/v/u/en-US/58455_1.00) - Calls its own latency spreadsheet an estimate and lists its assumptions, the vendor claim the measured tables check.
- [The microarchitecture of Intel, AMD, and VIA CPUs](https://www.agner.org/optimize/microarchitecture.pdf) - States where measured predictor, op cache and port findings disagree with the vendor's account of an x86 core.
- [Instruction tables: latencies, throughputs and micro-operation breakdowns](https://www.agner.org/optimize/instruction_tables.pdf) - States why its measured figures differ from the vendor's and names which latencies cannot be measured accurately.
- [uops.info](https://uops.info/) - Where each latency, throughput and port entry links to its microbenchmark, so any value can be re-run.
- [applecpu: Firestorm Overview](https://dougallj.github.io/applecpu/firestorm.html) - Measured per-instruction tables for an Apple AArch64 core, each entry linked to the counter experiment behind it.

### What the manuals leave out

- [Performance Speed Limits](https://travisdowns.github.io/blog/2019/06/11/speed-limits.html) - Sets the method for finding which hard bound binds a loop, testing its cycles against each in turn.
- [uiCA: Accurate Throughput Prediction of Basic Blocks on Recent Intel Microarchitectures](https://arxiv.org/abs/2107.14210) - Models the predecoder, decoders, micro-op cache and loop buffer closely enough to predict front-end bound loops.
- [Gathering Intel on Intel AVX-512 Transitions](https://travisdowns.github.io/blog/2020/01/17/avxfreq1.html) - Measures what the vendor leaves untimed in an AVX-512 licence change, a throttled phase then a frequency step.
- [Hardware Store Elimination](https://travisdowns.github.io/blog/2020/05/13/intel-zero-opt.html) - Finds an undocumented optimisation by its counter signature and bandwidth, and sets the method for what manuals omit.

Reproduce it: [misc/benchmarks/03-latency-vs-throughput](https://cpuperf.com/benchmarks/03-latency-vs-throughput/), one dependency chain against eight independent accumulators.
