Sources
20 entries in 4 parts
Reproduce it
06-roofline
Related sections
§1 Start here
In the MCP server
cpuperf://section/6

Roofline and the execution-cache-memory model

  1. 01
    Roofline: An Insightful Visual Performance Model for Multicore Architectures paper

    Defines operational intensity as traffic past the caches, and the ceilings that say which optimisation can pay.

  2. 02
    Applying the Roofline Model paper

    States the counter set and method that put a measured point on the plot in place of a hand-counted intensity.

  3. 03
    LIKWID repository

    Measures the roofs with likwid-bench and the point from likwid-perfctr groups on Intel, AMD and Arm server cores.

  4. 04
    Cache-aware Roofline model: Upgrading the loft paper

    Adds a roof per cache level against traffic at the core, so a kernel that hits in cache is not plotted as DRAM bound.

  5. 05
    Performance bottlenecks of stencil computations using the Execution-Cache-Memory model paper

    Times each level's transfer with overlap rules, predicting a core's rate and the core count where bandwidth saturates.

Reproduce it

measured roofs and three kernels of rising arithmetic intensity.

Reproduce it · 06-roofline Roofline Arithmetic intensity predicts which roof binds a loop

Top-down analysis

  1. 01
    A Top-Down Method for Performance Analysis and Counters Architecture paper

    Defines the slot accounting that turns raw counters into a weighted tree of bottlenecks, and why the unit is a slot.

  2. 02
    TMA_Metrics-full.xlsx (intel/perfmon) repository

    The official home of every TMA formula, event, threshold and level per microarchitecture, from which toplev derives.

  3. 03
    pmu-tools repository

    Runs the TMA tree on Linux from the spreadsheet and states why multiplexed levels mislead on short or varied workloads.

  4. 04
    pipeline.json (perf pmu-events, amdzen4) repository

    States the AMD top-down formulas in events for both levels, the metric groups perf stat runs on Zen with no vendor tool.

  5. 05
    Arm CPU Telemetry Solution Topdown Methodology Specification manual

    Defines Arm's top-down as staged stall accounting, first locating the bottleneck and then measuring its resource.

Scaling laws

  1. 01
    Validity of the single processor approach to achieving large scale computing capabilities paper

    States the serial-fraction bound on speedup, the argument every later scaling law is written against.

  2. 02
    Reevaluating Amdahl's Law paper

    Defines scaled speedup, the bound that holds when the problem grows with the processor count instead of staying fixed.

  3. 03
    Amdahl's Law in the Multicore Era paper

    Extends the bound to chips of unequal cores under a fixed area budget, the arithmetic behind big and little cores.

  4. 04
    A Simple Capacity Model of Massively Parallel Transaction Systems paper

    The origin of the coherency term that makes throughput fall, not merely flatten, as processors are added.

  5. 05
    Guerrilla Capacity Planning book

    Derives the universal scalability law and states the procedure that fits its coefficients to measured throughput.

Queueing

  1. 01
    A Proof for the Queuing Formula: L = λW paper

    Proves occupancy equals arrival rate times time in system with no assumption on arrivals or service.

  2. 02
    Quantitative System Performance report

    Origin of the bound-and-bottleneck analysis roofline names as its ancestor, and of the asymptotic bounds on throughput.

  3. 03
    Performance Modeling and Design of Computer Systems book

    Proves why open and closed systems answer a load question differently, and when scheduling, not capacity, sets latency.

  4. 04
    Stochastic Processes Occurring in the Theory of Queues paper

    Origin of the A/S/c notation, and of the embedded chain that solves a queue with non-memoryless arrivals or service.

  5. 05
    The single server queue in heavy traffic paper

    Derives the wait near saturation from utilisation and arrival and service variance, the formula behind the latency knee.