---
title: "15. Benchmarks"
url: https://cpuperf.com/learn/benchmarks/
source: https://github.com/usamahz/cpu-performance-engineering/blob/deb5a0bac46760503b6f4a2608bdfed470c8532e/README.md#L651
commit: deb5a0bac46760503b6f4a2608bdfed470c8532e
---
## 15. Benchmarks

A score means what its suite's run rules say it means, so the rules come before the number.

### Standard suites

- [SPEC CPU 2026 Run and Reporting Rules](https://www.spec.org/cpu2026/Docs/runrules.html) - Defines base against peak, rate against speed, the threading models a speed run may use and an Arm reference machine.
- [SPEC CPU: The Next Generation](https://arxiv.org/abs/2605.01575) - Where the committee states how workloads were chosen and hardened, and defines the rolling round-robin rate.
- [SPEC CPU2026: Characterization, Representativeness, and Cross-Suite Comparison](https://arxiv.org/abs/2605.03713) - Measures with counters what each workload stresses on x86 and Arm server parts, beside data-centre and inference suites.
- [MLPerf Inference Rules](https://github.com/mlcommons/inference_policies/blob/master/inference_rules.adoc) - Fixes model, accuracy floor and query pattern per scenario, so a Server score is throughput under a latency bound.
- [DCPerf: An Open-Source, Battle-Tested Performance Benchmark Suite for Datacenter Workloads](https://aisystemcodesign.github.io/papers/DCPerf-ISCA25.pdf) - Shows standard suites misproject data-centre servers, and states the fleet-matching method the suite is built by.

### Microbenchmark suites

- [Memory Bandwidth and Machine Balance in Current High Performance Computers](https://www.cs.virginia.edu/~mccalpin/papers/balance/) - Defines sustainable bandwidth as what unit-stride loops get, not bus peak, and machine balance as flops per access.
- [STREAM Benchmark Reference Information](https://www.cs.virginia.edu/stream/ref.html) - Sets the array size rule, timing over repeated trials, and counting bytes a loop asks for, not what the cache moved.
- [lmbench: Portable Tools for Performance Analysis](https://www.usenix.org/legacy/publications/library/proceedings/sd96/mcvoy.html) - Origin of the one-mechanism-per-test method for memory, system call, pipe and socket latency, and what each leaves out.
- [uarch-bench](https://github.com/travisdowns/uarch-bench) - Isolates memory-level parallelism from load latency as separate tests, with DVFS held off before timing, x86 Linux only.
- [nanoBench: A Low-Overhead Tool for Running Microbenchmarks on x86 Systems](https://arxiv.org/abs/1911.03282) - Shows why kernel mode with interrupts off matters, removes harness overhead, then recovers cache replacement policies.

Reproduce it: [misc/benchmarks/15-stream-bandwidth](https://cpuperf.com/benchmarks/15-stream-bandwidth/), triad bandwidth by thread count against the vendor figure.

### Methodology and what suites miss

- [How Not to Lie with Statistics: The Correct Way to Summarize Benchmark Results](https://dl.acm.org/doi/10.1145/5666.5673) - Origin of the rule that normalised results take the geometric mean, which a SPEC ratio and the crimes list rest on.
- [Systems Benchmarking Crimes](https://gernot-heiser.org/benchmarking-crimes.html) - Checklist of evaluation faults from sub-setting and improper baselines to arithmetic means of ratios, each with a fix.
- [Scientific Benchmarking of Parallel Computing Systems](https://htor.inf.ethz.ch/publications/img/hoefler-scientific-benchmarking.pdf) - Sets which mean fits costs, rates and ratios, when confidence intervals are owed, and the absolute base a speedup needs.
- [Rigorous Benchmarking in Reasonable Time](https://kar.kent.ac.uk/33611/) - Decides how many builds, runs and iterations an experiment needs by measuring at which level the variation arises.
- [Profiling a warehouse-scale computer](https://static.googleusercontent.com/media/research.google.com/en//pubs/archive/44271.pdf) - Fleet counter profile showing services stall on instruction fetch and burn cycles in shared routines, which SPEC lacks.
