Learn / Out
Benchmarks
Read first
A score means what its suite's run rules say it means, so the rules come before the number.
- Sources
- 15 entries in 3 parts
- Reproduce it
- 15-stream-bandwidth
- Related sections
- §2 One instruction, end to end, §5 Measurement, §12 Tail latency and production systems, §13 Inference on CPU
- In the MCP server
cpuperf://section/15
Standard suites
-
01
Defines base against peak, rate against speed, the threading models a speed run may use and an Arm reference machine.
-
02
Where the committee states how workloads were chosen and hardened, and defines the rolling round-robin rate.
-
03
Measures with counters what each workload stresses on x86 and Arm server parts, beside data-centre and inference suites.
-
04 MLPerf Inference Rules repository
Fixes model, accuracy floor and query pattern per scenario, so a Server score is throughput under a latency bound.
-
05
Shows standard suites misproject data-centre servers, and states the fleet-matching method the suite is built by.
Microbenchmark suites
-
01
Defines sustainable bandwidth as what unit-stride loops get, not bus peak, and machine balance as flops per access.
-
02
Sets the array size rule, timing over repeated trials, and counting bytes a loop asks for, not what the cache moved.
-
03
Origin of the one-mechanism-per-test method for memory, system call, pipe and socket latency, and what each leaves out.
-
04 uarch-bench repository
Isolates memory-level parallelism from load latency as separate tests, with DVFS held off before timing, x86 Linux only.
-
05
Shows why kernel mode with interrupts off matters, removes harness overhead, then recovers cache replacement policies.
Reproduce it
triad bandwidth by thread count against the vendor figure.
Reproduce it · 15-stream-bandwidth Memory bandwidth by thread count and working set Vendor bandwidth is a package number; one thread and cache-resident data cannot reveal itMethodology and what suites miss
-
01
Origin of the rule that normalised results take the geometric mean, which a SPEC ratio and the crimes list rest on.
-
02 Systems Benchmarking Crimes report
Checklist of evaluation faults from sub-setting and improper baselines to arithmetic means of ratios, each with a fix.
-
03
Sets which mean fits costs, rates and ratios, when confidence intervals are owed, and the absolute base a speedup needs.
-
04
Decides how many builds, runs and iterations an experiment needs by measuring at which level the variation arises.
-
05
Fleet counter profile showing services stall on instruction fetch and burn cycles in shared routines, which SPEC lacks.