---
title: "Benchmark lab"
url: https://cpuperf.com/benchmarks/
source: https://github.com/usamahz/cpu-performance-engineering/blob/deb5a0bac46760503b6f4a2608bdfed470c8532e/misc/benchmarks/README.md
commit: deb5a0bac46760503b6f4a2608bdfed470c8532e
---
# Benchmarks

One reproducible measurement for each numbered section of the README from
2 to 15. Each directory holds the source, the exact build line, a run
script, the machine description, the raw numbers from the last run, and
the analysis. Nothing here is quoted from a vendor; every number was
produced by the code in the directory on the machine named in its README.

| Directory | Section | Claim it reproduces |
|---|---|---|
| [02-branch-misprediction](https://cpuperf.com/benchmarks/02-branch-misprediction/) | 2. One instruction, end to end | A mispredicted branch costs on the order of the pipeline depth; sorted or branchless data removes it |
| [03-latency-vs-throughput](https://cpuperf.com/benchmarks/03-latency-vs-throughput/) | 3. Microarchitecture | A dependency chain is bound by latency; independent accumulators are bound by issue throughput |
| [04-cache-latency](https://cpuperf.com/benchmarks/04-cache-latency/) | 4. Memory hierarchy | Dependent-load latency steps at each cache level, and page-random access adds TLB cost |
| [05-measurement-pitfalls](https://cpuperf.com/benchmarks/05-measurement-pitfalls/) | 5. Measurement | An unused result measures nothing; a single run is not a measurement |
| [06-roofline](https://cpuperf.com/benchmarks/06-roofline/) | 6. Models | Arithmetic intensity predicts which roof binds a loop |
| [07-aos-vs-soa-simd](https://cpuperf.com/benchmarks/07-aos-vs-soa-simd/) | 7. Single-thread optimisation | Layout decides bytes moved and whether the loop vectorises |
| [08-autovectorization-aliasing](https://cpuperf.com/benchmarks/08-autovectorization-aliasing/) | 8. Compilers and codegen | The vectoriser gives up on possible aliasing; a qualifier fixes it |
| [09-false-sharing](https://cpuperf.com/benchmarks/09-false-sharing/) | 9. Concurrency | Writers sharing a cache line serialise; padding restores scaling |
| [10-first-touch](https://cpuperf.com/benchmarks/10-first-touch/) | 10. NUMA and multi-socket | Allocation is not placement; the first touch pays the fault and picks the home |
| [11-syscall-cost](https://cpuperf.com/benchmarks/11-syscall-cost/) | 11. OS and I/O | A kernel crossing has a fixed cost that request size does not amortise |
| [12-coordinated-omission](https://cpuperf.com/benchmarks/12-coordinated-omission/) | 12. Tail latency and production systems | A closed-loop load generator hides stalls that an open-loop one reports |
| [13-sgemm-naive-vs-blas](https://cpuperf.com/benchmarks/13-sgemm-naive-vs-blas/) | 13. Inference on CPU | A packed, register-blocked microkernel takes a naive loop to the vector unit's ceiling; the vendor BLAS is further ahead only where it owns a matrix unit |
| [14-pcore-vs-ecore](https://cpuperf.com/benchmarks/14-pcore-vs-ecore/) | 14. Hardware generations | The same code runs at different speeds by core type; a result without the core type is not comparable |
| [15-stream-bandwidth](https://cpuperf.com/benchmarks/15-stream-bandwidth/) | 15. Benchmarks | Vendor bandwidth is a package number; one thread and cache-resident data cannot reveal it |

## The machine

Apple M4 Pro, macOS 26.2, Apple clang 17.0.0 (clang-1700.4.4.1), arm64.
10 performance cores and 4 efficiency cores, no SMT, no NUMA, 128-byte
cache lines, 16 KiB pages, 24 GiB unified memory. P-core: 128 KiB L1d and
a 16 MiB L2 shared by a cluster of 5. E-core: 64 KiB L1d and a 4 MiB L2
shared by 4. ISA: NEON, FP16, BF16, I8MM, DotProd, LSE2, SME, SME2. Apple
publishes no clock, so every run records an estimate from a dependent
one-cycle add chain (`common/clock_estimate.c`); it reads about 4.4 GHz on
a performance core and about 1.5 GHz on an efficiency core under
background QoS. `common/machine.sh` prints the description that heads
every `results/raw.txt`.

## The seven fields

Every README states, for its numbers: the CPU model and microarchitecture;
the cores used; the frequency (estimated as above) with the DVFS and SMT
state; the compiler and flags (from `build.sh`); the workload; the
baseline; and the method (timer, repetitions, warmup, statistic). A number
without all seven does not appear anywhere in this repository.

## Running

    ./run_all.sh            # every benchmark, serially, full sizes
    QUICK=1 ./run_all.sh    # smoke test, small sizes, under ten seconds each
    ./compile_all.sh        # compile only, for a platform you cannot run on

Each directory also runs on its own with `./run.sh`. A run rewrites
`results/raw.txt`, `results/summary.md` and the results table in the
directory's README; commit them together if you re-measure on a different
machine, and change the machine section of the README to match.

Timing uses `clock_gettime(CLOCK_MONOTONIC_RAW)`, at least ten repetitions
after a discarded warmup, the minimum for latency-like quantities and the
median for throughput-like ones, and the coefficient of variation is
printed with every row. Every result is consumed with a compiler barrier
or checked against a reference so the optimiser cannot delete the work;
each README quotes the relevant assembly.

## Portability

The sources are C11 with pthreads. NEON, Accelerate and macOS QoS calls
are behind `#if` guards so every file compiles on Linux x86-64 and arm64
(`compile_all.sh` runs in CI on Linux); kernels that need a feature the
host lacks skip with a message rather than fail. Results on another
machine will differ and are only comparable once the seven fields are
restated for it.
