Benchmark lab
§ Benchmark and claim Section Type
02 Branch misprediction

A mispredicted branch costs on the order of the pipeline depth; sorted or branchless data removes it

One instruction, end to end single-threaded
03 Latency versus throughput

A dependency chain is bound by latency; independent accumulators are bound by issue throughput

Microarchitecture single-threaded
04 Dependent-load latency across the memory hierarchy

Dependent-load latency steps at each cache level, and page-random access adds TLB cost

Memory hierarchy single-threaded
05 Measurement pitfalls

An unused result measures nothing; a single run is not a measurement

Measurement single-threaded
06 Roofline

Arithmetic intensity predicts which roof binds a loop

Models multi-threaded
NEON intrinsics
07 Array of structs versus structure of arrays

Layout decides bytes moved and whether the loop vectorises

Single-thread optimisation single-threaded
NEON intrinsics
08 Auto-vectorisation and aliasing

The vectoriser gives up on possible aliasing; a qualifier fixes it

Compilers and codegen single-threaded
09 False sharing

Writers sharing a cache line serialise; padding restores scaling

Concurrency multi-threaded
10 First touch

Allocation is not placement; the first touch pays the fault and picks the home

NUMA and multi-socket single-threaded
11 Syscall cost

A kernel crossing has a fixed cost that request size does not amortise

OS and I/O single-threaded
12 Coordinated omission

A closed-loop load generator hides stalls that an open-loop one reports

Tail latency and production systems single-threaded
13 SGEMM: naive loop, microkernel, vendor BLAS

A packed, register-blocked microkernel takes a naive loop to the vector unit's ceiling; the vendor BLAS is further ahead only where it owns a matrix unit

Inference on CPU single-threaded
NEON intrinsics
14 P-core versus E-core

The same code runs at different speeds by core type; a result without the core type is not comparable

Hardware generations multi-threaded
NEON intrinsics
15 Memory bandwidth by thread count and working set

Vendor bandwidth is a package number; one thread and cache-resident data cannot reveal it

Benchmarks multi-threaded

The machine

Apple M4 Pro, macOS 26.2, Apple clang 17.0.0 (clang-1700.4.4.1), arm64. 10 performance cores and 4 efficiency cores, no SMT, no NUMA, 128-byte cache lines, 16 KiB pages, 24 GiB unified memory. P-core: 128 KiB L1d and a 16 MiB L2 shared by a cluster of 5. E-core: 64 KiB L1d and a 4 MiB L2 shared by 4. ISA: NEON, FP16, BF16, I8MM, DotProd, LSE2, SME, SME2. Apple publishes no clock, so every run records an estimate from a dependent one-cycle add chain (common/clock_estimate.c); it reads about 4.4 GHz on a performance core and about 1.5 GHz on an efficiency core under background QoS. common/machine.sh prints the description that heads every results/raw.txt.

The seven fields

Every README states, for its numbers: the CPU model and microarchitecture; the cores used; the frequency (estimated as above) with the DVFS and SMT state; the compiler and flags (from build.sh); the workload; the baseline; and the method (timer, repetitions, warmup, statistic). A number without all seven does not appear anywhere in this repository.

Running

./run_all.sh            # every benchmark, serially, full sizes
QUICK=1 ./run_all.sh    # smoke test, small sizes, under ten seconds each
./compile_all.sh        # compile only, for a platform you cannot run on

Each directory also runs on its own with ./run.sh. A run rewrites results/raw.txt, results/summary.md and the results table in the directory's README; commit them together if you re-measure on a different machine, and change the machine section of the README to match.

Timing uses clock_gettime(CLOCK_MONOTONIC_RAW), at least ten repetitions after a discarded warmup, the minimum for latency-like quantities and the median for throughput-like ones, and the coefficient of variation is printed with every row. Every result is consumed with a compiler barrier or checked against a reference so the optimiser cannot delete the work; each README quotes the relevant assembly.

Portability

The sources are C11 with pthreads. NEON, Accelerate and macOS QoS calls are behind #if guards so every file compiles on Linux x86-64 and arm64 (compile_all.sh runs in CI on Linux); kernels that need a feature the host lacks skip with a message rather than fail. Results on another machine will differ and are only comparable once the seven fields are restated for it.