Sources
20 entries in 4 parts
In the MCP server
cpuperf://section/5

Method and the whole-system view

  1. 01
    The USE Method report

    Sets the checklist that finds the saturated resource before any profiler is opened.

  2. 02
    Performance Analysis and Tuning on Modern CPUs book

    Draws the line between counting, sampling, instrumentation and tracing, so a question is matched to its tool.

  3. 03
    BPF Performance Tools book

    The reference for time a CPU sampler cannot see, off-CPU, scheduler and I/O waits, traced at bounded cost.

  4. 04
    bpftrace repository

    Makes a tracing hypothesis a one-line experiment over kprobes, uprobes, tracepoints and PMU events.

  5. 05
    Google-Wide Profiling: A Continuous Profiling Infrastructure for Data Centers paper

    The design continuous profilers descend from, always-on sampling across a fleet, cheap enough to leave running.

Counters, events and precise sampling

  1. 01
    perf_event_open(2) manual

    Defines the counting and sampling modes every Linux profiler uses, and the sample record fields, branch stack included.

  2. 02
    Instruction-Based Sampling: A New Performance Analysis Technique paper

    Defines skid, why a sample lands after the culprit, and how tagging one op through the pipeline removes it.

  3. 03
    Intel Software Developer Manuals manual

    Defines the architectural counters, what a PEBS sample captures, and why rdtsc counts time, not cycles, under DVFS.

  4. 04
    Processor Programming Reference for AMD Family 1Ah Model 02h manual

    Defines the event encodings and IBS registers for one Zen core, the tables perf's AMD events are derived from.

  5. 05
    perf-arm-spe(1) repository

    Defines Arm SPE as perf drives it, one sampled op in flight, the filters, and what a record holds.

CPU profilers and flame graphs

  1. 01
    Linux perf wiki: Tutorial manual

    The maintainers' walk from perf stat to perf record to perf annotate, with the sample fields each flag sets.

  2. 02
    Intel VTune Profiler Documentation manual

    Home of the user guide and cookbook, where each hardware analysis is defined by the events behind it.

  3. 03
    AMD uProf User Guide manual

    The vendor's reference for IBS-driven profiling on Zen, with metric presets defined per core generation.

  4. 04
    The Flame Graph paper

    Records the design decisions, width as sample share and alphabetical rather than time order, behind its reading rules.

  5. 05
    Coz: Finding Code that Counts with Causal Profiling paper

    Proves a hot function need not be worth optimising, by measuring what speeding up a line does to end-to-end time.

Microbenchmarks that lie

  1. 01
    Producing Wrong Data Without Doing Anything Obviously Wrong! paper

    The origin of measurement bias as a term, link order and environment size alone flipping a compiler flag comparison.

  2. 02
    Non-Determinism and Overcount on Modern Hardware Performance Counter Implementations paper

    Traces run-to-run variation in x86 retired-instruction counts to one extra count per interrupt and per fault.

  3. 03
    clock_gettime(2) manual

    Defines what each clock counts, NTP-slewed or raw monotonic time, or CPU time, and the resolution call.

  4. 04
    Benchmarking tips (LLVM) manual

    The compiler project's recipe for a quiet Linux host, governor, boost, SMT siblings, ASLR, a cpuset and tmpfs.

  5. 05
    Google Benchmark User Guide repository

    Documents the barriers that keep the optimiser from deleting the work under test, and repetition statistics.

Reproduce it

dead-code elimination, run-to-run spread, cold against warm.

Reproduce it · 05-measurement-pitfalls Measurement pitfalls An unused result measures nothing; a single run is not a measurement