---
title: "5. Measurement"
url: https://cpuperf.com/learn/measurement/
source: https://github.com/usamahz/cpu-performance-engineering/blob/deb5a0bac46760503b6f4a2608bdfed470c8532e/README.md#L261
commit: deb5a0bac46760503b6f4a2608bdfed470c8532e
---
## 5. Measurement

Cloud instances often virtualise the hardware counters away and macOS runs no Linux perf, so perf stat has to show a non-zero cycles count before any counter entry below is trusted.

### Method and the whole-system view

- [The USE Method](https://www.brendangregg.com/usemethod.html) - Sets the checklist that finds the saturated resource before any profiler is opened.
- [Performance Analysis and Tuning on Modern CPUs](https://github.com/dendibakh/perf-book) - Draws the line between counting, sampling, instrumentation and tracing, so a question is matched to its tool.
- [BPF Performance Tools](https://www.brendangregg.com/bpf-performance-tools-book.html) - The reference for time a CPU sampler cannot see, off-CPU, scheduler and I/O waits, traced at bounded cost.
- [bpftrace](https://github.com/bpftrace/bpftrace) - Makes a tracing hypothesis a one-line experiment over kprobes, uprobes, tracepoints and PMU events.
- [Google-Wide Profiling: A Continuous Profiling Infrastructure for Data Centers](https://research.google/pubs/google-wide-profiling-a-continuous-profiling-infrastructure-for-data-centers/) - The design continuous profilers descend from, always-on sampling across a fleet, cheap enough to leave running.

### Counters, events and precise sampling

- [perf_event_open(2)](https://man7.org/linux/man-pages/man2/perf_event_open.2.html) - Defines the counting and sampling modes every Linux profiler uses, and the sample record fields, branch stack included.
- [Instruction-Based Sampling: A New Performance Analysis Technique](https://www.amd.com/content/dam/amd/en/documents/archived-tech-docs/white-papers/AMD_IBS_paper_EN.pdf) - Defines skid, why a sample lands after the culprit, and how tagging one op through the pipeline removes it.
- [Intel Software Developer Manuals](https://www.intel.com/content/www/us/en/developer/articles/technical/intel-sdm.html) - Defines the architectural counters, what a PEBS sample captures, and why rdtsc counts time, not cycles, under DVFS.
- [Processor Programming Reference for AMD Family 1Ah Model 02h](https://docs.amd.com/v/u/en-US/57238) - Defines the event encodings and IBS registers for one Zen core, the tables perf's AMD events are derived from.
- [perf-arm-spe(1)](https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/tree/tools/perf/Documentation/perf-arm-spe.txt) - Defines Arm SPE as perf drives it, one sampled op in flight, the filters, and what a record holds.

### CPU profilers and flame graphs

- [Linux perf wiki: Tutorial](https://perfwiki.github.io/main/tutorial/) - The maintainers' walk from perf stat to perf record to perf annotate, with the sample fields each flag sets.
- [Intel VTune Profiler Documentation](https://www.intel.com/content/www/us/en/developer/tools/oneapi/vtune-profiler-documentation.html) - Home of the user guide and cookbook, where each hardware analysis is defined by the events behind it.
- [AMD uProf User Guide](https://docs.amd.com/r/en-US/57368-uProf-user-guide) - The vendor's reference for IBS-driven profiling on Zen, with metric presets defined per core generation.
- [The Flame Graph](https://queue.acm.org/detail.cfm?id=2927301) - Records the design decisions, width as sample share and alphabetical rather than time order, behind its reading rules.
- [Coz: Finding Code that Counts with Causal Profiling](https://arxiv.org/abs/1608.03676) - Proves a hot function need not be worth optimising, by measuring what speeding up a line does to end-to-end time.

### Microbenchmarks that lie

- [Producing Wrong Data Without Doing Anything Obviously Wrong!](https://sape.inf.usi.ch/publications/asplos09) - The origin of measurement bias as a term, link order and environment size alone flipping a compiler flag comparison.
- [Non-Determinism and Overcount on Modern Hardware Performance Counter Implementations](https://web.eece.maine.edu/~vweaver/projects/deterministic/ispass2013_deterministic.pdf) - Traces run-to-run variation in x86 retired-instruction counts to one extra count per interrupt and per fault.
- [clock_gettime(2)](https://man7.org/linux/man-pages/man2/clock_gettime.2.html) - Defines what each clock counts, NTP-slewed or raw monotonic time, or CPU time, and the resolution call.
- [Benchmarking tips (LLVM)](https://llvm.org/docs/Benchmarking.html) - The compiler project's recipe for a quiet Linux host, governor, boost, SMT siblings, ASLR, a cpuset and tmpfs.
- [Google Benchmark User Guide](https://github.com/google/benchmark/blob/main/docs/user_guide.md) - Documents the barriers that keep the optimiser from deleting the work under test, and repetition statistics.

Reproduce it: [misc/benchmarks/05-measurement-pitfalls](https://cpuperf.com/benchmarks/05-measurement-pitfalls/), dead-code elimination, run-to-run spread, cold against warm.
