Learn / Down
Microarchitecture
Read first
Where a design paper, the vendor manual and a measurement disagree about a core, the measurement is the one to trust and re-run.
- Sources
- 19 entries in 4 parts
- Reproduce it
- 03-latency-vs-throughput
- Related sections
- §1 Start here, §2 One instruction, end to end, §14 Hardware generations
- In the MCP server
cpuperf://section/3
Limits of ILP and SMT
-
01
Shows rename, the active list and precise branch recovery fitted together in one shipped out-of-order core.
-
02
Separates perfect from realistic prediction, renaming and aliasing, and shows how little ILP survives the real ones.
-
03
Puts circuit delay on wakeup, select and bypass, capping issue width and window size before ILP runs out.
-
04
The origin of SMT, where another thread fills one thread's idle issue slots at a cost to both.
-
05
Turns window size, issue width and miss events into cycles lost, making an ILP limit a cycle count.
Branch prediction and speculation
-
01
Defines the tagged geometric-history predictor shipped designs converge on, and the limit of what it can learn.
-
02
Carries the tagged geometric scheme to indirect jumps and calls, where interpreters and virtual dispatch stall.
-
03
Defines the penalty as pipeline refill plus window drain, so it exceeds pipeline depth and is not constant.
-
04
Establishes that predictor state is shared and trainable across contexts, the root of every mitigation and its cost.
-
05
Defines the indirect-branch and store-bypass controls (IBRS, STIBP, IBPB, SSBD) and states which carry a large cost.
Vendor estimates and measured tables
-
01
Calls its own latency spreadsheet an estimate and lists its assumptions, the vendor claim the measured tables check.
-
02
States where measured predictor, op cache and port findings disagree with the vendor's account of an x86 core.
-
03
States why its measured figures differ from the vendor's and names which latencies cannot be measured accurately.
-
04 uops.info report
Where each latency, throughput and port entry links to its microbenchmark, so any value can be re-run.
-
05 applecpu: Firestorm Overview report
Measured per-instruction tables for an Apple AArch64 core, each entry linked to the counter experiment behind it.
What the manuals leave out
-
01 Performance Speed Limits report
Sets the method for finding which hard bound binds a loop, testing its cycles against each in turn.
-
02
Models the predecoder, decoders, micro-op cache and loop buffer closely enough to predict front-end bound loops.
-
03
Measures what the vendor leaves untimed in an AVX-512 licence change, a throttled phase then a frequency step.
-
04 Hardware Store Elimination report
Finds an undocumented optimisation by its counter signature and bandwidth, and sets the method for what manuals omit.
Reproduce it
one dependency chain against eight independent accumulators.
Reproduce it · 03-latency-vs-throughput Latency versus throughput A dependency chain is bound by latency; independent accumulators are bound by issue throughput