---
title: "6. Models"
url: https://cpuperf.com/learn/models/
source: https://github.com/usamahz/cpu-performance-engineering/blob/deb5a0bac46760503b6f4a2608bdfed470c8532e/README.md#L299
commit: deb5a0bac46760503b6f4a2608bdfed470c8532e
---
## 6. Models

A roofline is a bound built from measured roofs and counted bytes, so a point above a roof means a wrong roof or a wrong byte count, not fast code.

### Roofline and the execution-cache-memory model

- [Roofline: An Insightful Visual Performance Model for Multicore Architectures](https://cacm.acm.org/research/roofline-an-insightful-visual-performance-model-for-multicore-architectures/) - Defines operational intensity as traffic past the caches, and the ceilings that say which optimisation can pay.
- [Applying the Roofline Model](https://spiral.ece.cmu.edu/pub-spiral/pubfile/ispass-2013_177.pdf) - States the counter set and method that put a measured point on the plot in place of a hand-counted intensity.
- [LIKWID](https://github.com/RRZE-HPC/likwid) - Measures the roofs with likwid-bench and the point from likwid-perfctr groups on Intel, AMD and Arm server cores.
- [Cache-aware Roofline model: Upgrading the loft](https://ieeexplore.ieee.org/document/6506838) - Adds a roof per cache level against traffic at the core, so a kernel that hits in cache is not plotted as DRAM bound.
- [Performance bottlenecks of stencil computations using the Execution-Cache-Memory model](https://arxiv.org/abs/1410.5010) - Times each level's transfer with overlap rules, predicting a core's rate and the core count where bandwidth saturates.

Reproduce it: [misc/benchmarks/06-roofline](https://cpuperf.com/benchmarks/06-roofline/), measured roofs and three kernels of rising arithmetic intensity.

### Top-down analysis

- [A Top-Down Method for Performance Analysis and Counters Architecture](https://sites.google.com/site/analysismethods/yasin-pubs) - Defines the slot accounting that turns raw counters into a weighted tree of bottlenecks, and why the unit is a slot.
- [TMA_Metrics-full.xlsx (intel/perfmon)](https://github.com/intel/perfmon/blob/main/TMA_Metrics-full.xlsx) - The official home of every TMA formula, event, threshold and level per microarchitecture, from which toplev derives.
- [pmu-tools](https://github.com/andikleen/pmu-tools) - Runs the TMA tree on Linux from the spreadsheet and states why multiplexed levels mislead on short or varied workloads.
- [pipeline.json (perf pmu-events, amdzen4)](https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/tree/tools/perf/pmu-events/arch/x86/amdzen4/pipeline.json) - States the AMD top-down formulas in events for both levels, the metric groups perf stat runs on Zen with no vendor tool.
- [Arm CPU Telemetry Solution Topdown Methodology Specification](https://support.arm.com/documentation/109542/latest/) - Defines Arm's top-down as staged stall accounting, first locating the bottleneck and then measuring its resource.

### Scaling laws

- [Validity of the single processor approach to achieving large scale computing capabilities](https://dl.acm.org/doi/10.1145/1465482.1465560) - States the serial-fraction bound on speedup, the argument every later scaling law is written against.
- [Reevaluating Amdahl's Law](http://www.johngustafson.net/pubs/pub13/amdahl.htm) - Defines scaled speedup, the bound that holds when the problem grows with the processor count instead of staying fixed.
- [Amdahl's Law in the Multicore Era](https://research.cs.wisc.edu/multifacet/papers/ieeecomputer08_amdahl_multicore.pdf) - Extends the bound to chips of unequal cores under a fixed area budget, the arithmetic behind big and little cores.
- [A Simple Capacity Model of Massively Parallel Transaction Systems](https://www.perfdynamics.com/Papers/njgCMG93.pdf) - The origin of the coherency term that makes throughput fall, not merely flatten, as processors are added.
- [Guerrilla Capacity Planning](https://www.perfdynamics.com/iBook/gcap.html) - Derives the universal scalability law and states the procedure that fits its coefficients to measured throughput.

### Queueing

- [A Proof for the Queuing Formula: L = λW](https://pubsonline.informs.org/doi/10.1287/opre.9.3.383) - Proves occupancy equals arrival rate times time in system with no assumption on arrivals or service.
- [Quantitative System Performance](https://homes.cs.washington.edu/~lazowska/qsp/) - Origin of the bound-and-bottleneck analysis roofline names as its ancestor, and of the asymptotic bounds on throughput.
- [Performance Modeling and Design of Computer Systems](https://www.cs.cmu.edu/~harchol/PerformanceModeling/book.html) - Proves why open and closed systems answer a load question differently, and when scheduling, not capacity, sets latency.
- [Stochastic Processes Occurring in the Theory of Queues](https://projecteuclid.org/journals/annals-of-mathematical-statistics/volume-24/issue-3/Stochastic-Processes-Occurring-in-the-Theory-of-Queues-and-their-Analysis-by-the/10.1214/aoms/1177728975.full) - Origin of the A/S/c notation, and of the embedded chain that solves a queue with non-memoryless arrivals or service.
- [The single server queue in heavy traffic](https://www.cambridge.org/core/journals/mathematical-proceedings-of-the-cambridge-philosophical-society/article/abs/single-server-queue-in-heavy-traffic/81C55BC00A68FE6D5385638AA0B0AF37) - Derives the wait near saturation from utilisation and arrival and service variance, the formula behind the latency knee.
