Sources
19 entries in 4 parts
In the MCP server
cpuperf://section/3

Limits of ILP and SMT

  1. 01
    The MIPS R10000 Superscalar Microprocessor paper

    Shows rename, the active list and precise branch recovery fitted together in one shipped out-of-order core.

  2. 02
    Limits of Instruction-Level Parallelism (WRL Research Report 93/6) paper

    Separates perfect from realistic prediction, renaming and aliasing, and shows how little ILP survives the real ones.

  3. 03
    Complexity-Effective Superscalar Processors paper

    Puts circuit delay on wakeup, select and bypass, capping issue width and window size before ILP runs out.

  4. 04
    Simultaneous Multithreading: Maximizing On-Chip Parallelism paper

    The origin of SMT, where another thread fills one thread's idle issue slots at a cost to both.

  5. 05
    A Mechanistic Performance Model for Superscalar Out-of-Order Processors paper

    Turns window size, issue width and miss events into cycles lost, making an ILP limit a cycle count.

Branch prediction and speculation

  1. 01
    A Case for (Partially) TAgged GEometric History Length Branch Prediction paper

    Defines the tagged geometric-history predictor shipped designs converge on, and the limit of what it can learn.

  2. 02
    A 64-Kbytes ITTAGE indirect branch predictor paper

    Carries the tagged geometric scheme to indirect jumps and calls, where interpreters and virtual dispatch stall.

  3. 03
    Characterizing the Branch Misprediction Penalty paper

    Defines the penalty as pipeline refill plus window drain, so it exceeds pipeline depth and is not constant.

  4. 04
    Spectre Attacks: Exploiting Speculative Execution paper

    Establishes that predictor state is shared and trainable across contexts, the root of every mitigation and its cost.

  5. 05
    Speculative Execution Side Channel Mitigations manual

    Defines the indirect-branch and store-bypass controls (IBRS, STIBP, IBPB, SSBD) and states which carry a large cost.

Vendor estimates and measured tables

  1. 01
    Software Optimization Guide for the AMD Zen5 Microarchitecture manual

    Calls its own latency spreadsheet an estimate and lists its assumptions, the vendor claim the measured tables check.

  2. 02
    The microarchitecture of Intel, AMD, and VIA CPUs manual

    States where measured predictor, op cache and port findings disagree with the vendor's account of an x86 core.

  3. 03
    Instruction tables: latencies, throughputs and micro-operation breakdowns manual

    States why its measured figures differ from the vendor's and names which latencies cannot be measured accurately.

  4. 04
    uops.info report

    Where each latency, throughput and port entry links to its microbenchmark, so any value can be re-run.

  5. 05
    applecpu: Firestorm Overview report

    Measured per-instruction tables for an Apple AArch64 core, each entry linked to the counter experiment behind it.

What the manuals leave out

  1. 01
    Performance Speed Limits report

    Sets the method for finding which hard bound binds a loop, testing its cycles against each in turn.

  2. 02
    uiCA: Accurate Throughput Prediction of Basic Blocks on Recent Intel Microarchitectures paper

    Models the predecoder, decoders, micro-op cache and loop buffer closely enough to predict front-end bound loops.

  3. 03
    Gathering Intel on Intel AVX-512 Transitions report

    Measures what the vendor leaves untimed in an AVX-512 licence change, a throttled phase then a frequency step.

  4. 04
    Hardware Store Elimination report

    Finds an undocumented optimisation by its counter signature and bandwidth, and sets the method for what manuals omit.

Reproduce it

one dependency chain against eight independent accumulators.

Reproduce it · 03-latency-vs-throughput Latency versus throughput A dependency chain is bound by latency; independent accumulators are bound by issue throughput