Learn / Down
Measurement
Read first
Cloud instances often virtualise the hardware counters away and macOS runs no Linux perf, so perf stat has to show a non-zero cycles count before any counter entry below is trusted.
- Sources
- 20 entries in 4 parts
- Reproduce it
- 05-measurement-pitfalls
- Related sections
- §1 Start here, §7 Single-thread optimisation, §13 Inference on CPU, §15 Benchmarks
- In the MCP server
cpuperf://section/5
Method and the whole-system view
-
01 The USE Method report
Sets the checklist that finds the saturated resource before any profiler is opened.
-
02
Draws the line between counting, sampling, instrumentation and tracing, so a question is matched to its tool.
-
03
The reference for time a CPU sampler cannot see, off-CPU, scheduler and I/O waits, traced at bounded cost.
-
04 bpftrace repository
Makes a tracing hypothesis a one-line experiment over kprobes, uprobes, tracepoints and PMU events.
-
05
The design continuous profilers descend from, always-on sampling across a fleet, cheap enough to leave running.
Counters, events and precise sampling
-
01 perf_event_open(2) manual
Defines the counting and sampling modes every Linux profiler uses, and the sample record fields, branch stack included.
-
02
Defines skid, why a sample lands after the culprit, and how tagging one op through the pipeline removes it.
-
03
Defines the architectural counters, what a PEBS sample captures, and why rdtsc counts time, not cycles, under DVFS.
-
04
Defines the event encodings and IBS registers for one Zen core, the tables perf's AMD events are derived from.
-
05 perf-arm-spe(1) repository
Defines Arm SPE as perf drives it, one sampled op in flight, the filters, and what a record holds.
CPU profilers and flame graphs
-
01 Linux perf wiki: Tutorial manual
The maintainers' walk from perf stat to perf record to perf annotate, with the sample fields each flag sets.
-
02
Home of the user guide and cookbook, where each hardware analysis is defined by the events behind it.
-
03 AMD uProf User Guide manual
The vendor's reference for IBS-driven profiling on Zen, with metric presets defined per core generation.
-
04 The Flame Graph paper
Records the design decisions, width as sample share and alphabetical rather than time order, behind its reading rules.
-
05
Proves a hot function need not be worth optimising, by measuring what speeding up a line does to end-to-end time.
Microbenchmarks that lie
-
01
The origin of measurement bias as a term, link order and environment size alone flipping a compiler flag comparison.
-
02
Traces run-to-run variation in x86 retired-instruction counts to one extra count per interrupt and per fault.
-
03 clock_gettime(2) manual
Defines what each clock counts, NTP-slewed or raw monotonic time, or CPU time, and the resolution call.
-
04 Benchmarking tips (LLVM) manual
The compiler project's recipe for a quiet Linux host, governor, boost, SMT siblings, ASLR, a cpuset and tmpfs.
-
05 Google Benchmark User Guide repository
Documents the barriers that keep the optimiser from deleting the work under test, and repetition statistics.
Reproduce it
dead-code elimination, run-to-run spread, cold against warm.
Reproduce it · 05-measurement-pitfalls Measurement pitfalls An unused result measures nothing; a single run is not a measurement