01Start here

Each entry assumes only the ones before it. The first and sixth are paid books; the third is a manual to open at the chapters its reason names.

  1. 01
    Computer Architecture: A Quantitative Approach, 7th Edition book

    Its pipelining appendix and memory chapters define the hazard, speculation and cache vocabulary the list assumes.

  2. 02
    Optimizing software in C++ manual

    Maps C++ onto pipeline mechanisms and shows why a loop-carried dependency chain, not instruction count, paces a loop.

  3. 03
    Intel Optimization Reference Manual manual

    Its opening chapters show how a shipping x86 core implements the textbook pipeline, each rule tied to a mechanism.

  4. 04
    What Every Programmer Should Know About Memory paper

    Measures the step in cost per access at each cache boundary and the gap a prefetcher hides.

  5. 05
    Memory Barriers: a Hardware View for Software Hackers paper

    Explains why a second core makes loads and stores reorder and what a barrier drains.

  6. 06
    Systems Performance: Enterprise and the Cloud, 2nd Edition book

    Puts the method before the tools: what to measure, in what order, and how benchmarks mislead.

  7. 07
    Roofline: An Insightful Visual Performance Model for Multicore Architectures paper

    Places a loop from a byte count and a datasheet bandwidth alone, before any counter is read.

  8. 08
    A Top-Down Method for Performance Analysis and Counters Architecture paper

    Defines the split of pipeline slots into front end, bad speculation, back end and retiring, the tree profilers report.

  9. 09
    Performance Analysis and Tuning on Modern CPUs book

    Walks from a noisy timing to counters to a named bottleneck, applying roofline and top-down to whole programs.

  10. 10
    What Has My Compiler Done for Me Lately? Unbolting the Compiler's Lid talk

    Shows how to read emitted assembly against its source, so each mechanism is checked in a listing, not assumed.

Work the exercises in Performance Ninja alongside them; reading alone will not build the instinct.

Reproduce it: 04-cache-latency, the cost per dependent load stepping up at each cache boundary, the curve the fourth entry measures.

02One instruction, end to end

Vendors name the same structures differently (Intel's decoded ICache is AMD's op cache, and AMD's macro-op is Arm's MOP), so the stage names below are generic.

Fetch and decode

  1. 01
    Fetch Directed Instruction Prefetching paper

    The origin of the decoupled front end, where the predictor runs ahead of fetch and drives instruction prefetch.

  2. 02
    Micro-Operation Cache: A Power Aware Frontend for Variable Instruction Length ISA paper

    Introduces the decoded micro-op cache and its power case, the structure both x86 vendors later built.

  3. 03
    Software Optimization Guide for the AMD Zen5 Microarchitecture manual

    Where AMD states when the op cache feeds micro-ops, and the fusion and alignment rules for hot loops.

  4. 04
    The microarchitecture of Intel, AMD, and VIA CPUs manual

    Measures rather than quotes each x86 core's misprediction penalty, micro-op cache and loop buffer behaviour.

  5. 05
    Intel Mitigations for Jump Conditional Code Erratum manual

    States what a microcode fix evicts from the decoded cache and which counters show the fall back to legacy decode.

Reproduce it: 02-branch-misprediction, the cost of a mispredicted branch, sorted against unsorted against branchless.

Rename and issue

  1. 01
    An Efficient Algorithm for Exploiting Multiple Arithmetic Units paper

    The origin of tag-based renaming and reservation stations, from which every out-of-order issue queue descends.

  2. 02
    Optimizing subroutines in assembly language manual

    Shows which chains renaming cannot break, from partial registers to flags, and the zero-cost idioms that do.

  3. 03
    Measuring Reorder Buffer Capacity report

    Sets the user-space method that measures the window and register files, and which idioms take no physical register.

Execute

  1. 01
    Focusing Processor Policies via Critical-Path Prediction paper

    Defines the dependence graph of an out-of-order core and the critical path through it that sets run time.

  2. 02
    Instruction tables: latencies, throughputs and micro-operation breakdowns manual

    Measures latency, reciprocal throughput and port assignment per instruction, the weights a dependence graph needs.

  3. 03
    Arm Neoverse V2 Core Software Optimization Guide manual

    States every stage of one Arm core from fetch to issue, then per-instruction latency, throughput and pipe assignment.

  4. 04
    Entropy Decoding in Oodle Data: x86-64 3-Stream Huffman Decoders report

    Predicts a real decoder loop's cycle count from its carried chain and issue slots, then measures the prediction.

  5. 05
    Faster zlib/DEFLATE decompression on the Apple M1 (and x86) report

    Predicts a decoder's refill chain from Arm core latencies, fits extra work under it, and measures the gain.

Memory access and retire

  1. 01
    Memory Dependence Prediction using Store Sets paper

    The origin of memory dependence prediction, which lets a load pass older stores and flushes on a wrong guess.

  2. 02
    Store-to-Load Forwarding and Memory Disambiguation in x86 Processors report

    Measures which store and load size and offset pairs forward or stall, and which cores predict memory dependences.

  3. 03
    Microarchitecture Optimizations for Exploiting Memory-Level Parallelism paper

    Defines memory-level parallelism as the misses one window overlaps, and shows which core limits cap it.

  4. 04
    Implementing Precise Interrupts in Pipelined Processors paper

    Introduces the reorder buffer and defines a precise exception as in-order commit of out-of-order results.

  5. 05
    Where Do Interrupts Happen? report

    Measures that an interrupt lands on the oldest unretired instruction, which decides what a sampling profiler blames.

03Microarchitecture

Where a design paper, the vendor manual and a measurement disagree about a core, the measurement is the one to trust and re-run.

Limits of ILP and SMT

  1. 01
    The MIPS R10000 Superscalar Microprocessor paper

    Shows rename, the active list and precise branch recovery fitted together in one shipped out-of-order core.

  2. 02
    Limits of Instruction-Level Parallelism (WRL Research Report 93/6) paper

    Separates perfect from realistic prediction, renaming and aliasing, and shows how little ILP survives the real ones.

  3. 03
    Complexity-Effective Superscalar Processors paper

    Puts circuit delay on wakeup, select and bypass, capping issue width and window size before ILP runs out.

  4. 04
    Simultaneous Multithreading: Maximizing On-Chip Parallelism paper

    The origin of SMT, where another thread fills one thread's idle issue slots at a cost to both.

  5. 05
    A Mechanistic Performance Model for Superscalar Out-of-Order Processors paper

    Turns window size, issue width and miss events into cycles lost, making an ILP limit a cycle count.

Branch prediction and speculation

  1. 01
    A Case for (Partially) TAgged GEometric History Length Branch Prediction paper

    Defines the tagged geometric-history predictor shipped designs converge on, and the limit of what it can learn.

  2. 02
    A 64-Kbytes ITTAGE indirect branch predictor paper

    Carries the tagged geometric scheme to indirect jumps and calls, where interpreters and virtual dispatch stall.

  3. 03
    Characterizing the Branch Misprediction Penalty paper

    Defines the penalty as pipeline refill plus window drain, so it exceeds pipeline depth and is not constant.

  4. 04
    Spectre Attacks: Exploiting Speculative Execution paper

    Establishes that predictor state is shared and trainable across contexts, the root of every mitigation and its cost.

  5. 05
    Speculative Execution Side Channel Mitigations manual

    Defines the indirect-branch and store-bypass controls (IBRS, STIBP, IBPB, SSBD) and states which carry a large cost.

Vendor estimates and measured tables

  1. 01
    Software Optimization Guide for the AMD Zen5 Microarchitecture manual

    Calls its own latency spreadsheet an estimate and lists its assumptions, the vendor claim the measured tables check.

  2. 02
    The microarchitecture of Intel, AMD, and VIA CPUs manual

    States where measured predictor, op cache and port findings disagree with the vendor's account of an x86 core.

  3. 03
    Instruction tables: latencies, throughputs and micro-operation breakdowns manual

    States why its measured figures differ from the vendor's and names which latencies cannot be measured accurately.

  4. 04
    uops.info report

    Where each latency, throughput and port entry links to its microbenchmark, so any value can be re-run.

  5. 05
    applecpu: Firestorm Overview report

    Measured per-instruction tables for an Apple AArch64 core, each entry linked to the counter experiment behind it.

What the manuals leave out

  1. 01
    Performance Speed Limits report

    Sets the method for finding which hard bound binds a loop, testing its cycles against each in turn.

  2. 02
    uiCA: Accurate Throughput Prediction of Basic Blocks on Recent Intel Microarchitectures paper

    Models the predecoder, decoders, micro-op cache and loop buffer closely enough to predict front-end bound loops.

  3. 03
    Gathering Intel on Intel AVX-512 Transitions report

    Measures what the vendor leaves untimed in an AVX-512 licence change, a throttled phase then a frequency step.

  4. 04
    Hardware Store Elimination report

    Finds an undocumented optimisation by its counter signature and bandwidth, and sets the method for what manuals omit.

Reproduce it: 03-latency-vs-throughput, one dependency chain against eight independent accumulators.

04Memory hierarchy

Line size and page size are machine parameters, not constants, so every padding and alignment rule below is applied against the target's own values.

Cache geometry, replacement and misses in flight

  1. 01
    What Every Programmer Should Know About Memory paper

    One measured account of DRAM timing, cache geometry, TLBs and prefetchers that sets the padding and prefetch rules.

  2. 02
    Measuring Cache and TLB Performance and Their Effect on Benchmark Run Times paper

    The origin of the strided-loop method that recovers cache and TLB size, line size, associativity and miss cost.

  3. 03
    Achieving Non-Inclusive Cache Performance with Inclusive Caches paper

    Names inclusion victims as the cost of an inclusive last-level cache, the case for non-inclusive and victim designs.

  4. 04
    Adaptive Insertion Policies for High Performance Caching paper

    The origin of set duelling and bimodal insertion, the adaptive replacement that survives a streaming pass.

  5. 05
    Lockup-Free Instruction Fetch/Prefetch Cache Organization paper

    Origin of the lockup-free cache, whose miss-status registers keep misses in flight, so a stream beats a chase.

Reproduce it: 04-cache-latency, dependent-load latency from L1 to DRAM, with and without TLB pressure.

TLBs, page walks and prefetchers

  1. 01
    Intel SDM Volume 3A: System Programming Guide, Part 1 manual

    Fixes the paging walk, the page-walk caches, the TLB invalidation rules and what each memory type permits.

  2. 02
    Translation Caching: Skip, Don't Walk (the Page Table) paper

    Defines the page-walk cache design space and shows why cached partial translations let a walk skip page-table levels.

  3. 03
    Improving Direct-Mapped Cache Performance paper

    The origin of the stream buffer that every vendor stream prefetcher descends from, and of the victim cache.

  4. 04
    Intel Optimization Reference Manual Volume 1 manual

    Names each prefetcher and what trains it, and which stop at a page boundary and which cross it.

  5. 05
    Arm Neoverse V2 Core Technical Reference Manual manual

    Names an Arm server core's load-side and store-side prefetchers, its TLB levels and the bits that disable them.

Store buffers, ordering and cache-line contention

  1. 01
    Memory Barriers: a Hardware View for Software Hackers paper

    Derives store buffers and invalidate queues from the cost of coherence, the reason reordering exists at all.

  2. 02
    x86-TSO: A Rigorous and Usable Programmer's Model for x86 Multiprocessors paper

    The store-buffer model of Intel and AMD ordering, tested on hardware, fixing which reorderings a fence pays for.

  3. 03
    Simplifying ARM Concurrency: Multicopy-Atomic Axiomatic and Operational Models for ARMv8 report

    The formal model Arm adopted into its architecture, and the record of why the non-multicopy-atomic option was dropped.

  4. 04
    Everything You Always Wanted to Know About Synchronization but Were Afraid to Ask paper

    Measures line-transfer cost between cores by coherence state and distance, the figure behind every contention rule.

  5. 05
    C2C - False Sharing Detection in Linux Perf report

    The implementers' account of perf c2c, whose worked example reads contended lines, offsets and callers off the report.

Struct layout, software prefetch and page size

  1. 01
    Cache-Conscious Structure Definition paper

    Introduces and measures structure splitting and field reordering on real programs, the origin of hot-cold layout rules.

  2. 02
    CppCon 2014: Data-Oriented Design and C++ talk

    Argues from cache-line utilisation arithmetic that layout must follow the access pattern, the case behind structure of arrays.

  3. 03
    dwarves (pahole) repository

    Prints a struct's holes, padding and cache-line boundaries from DWARF, so a layout is seen rather than guessed.

  4. 04
    When Prefetching Works, When It Doesn't, and Why paper

    Sorts software prefetch into the cases where it helps and hurts, and shows how it mistrains hardware prefetchers.

  5. 05
    Transparent Hugepage Support manual

    The kernel's statement of THP: the enabled and defrag knobs, khugepaged, and the counters showing what was obtained.

  6. 06
    Coordinated and Efficient Huge Page Management with Ingens paper

    Measures the fault latency, memory bloat and unfairness that eager THP promotion causes, and shows what removes them.

05Measurement

Cloud instances often virtualise the hardware counters away and macOS runs no Linux perf, so perf stat has to show a non-zero cycles count before any counter entry below is trusted.

Method and the whole-system view

  1. 01
    The USE Method report

    Sets the checklist that finds the saturated resource before any profiler is opened.

  2. 02
    Performance Analysis and Tuning on Modern CPUs book

    Draws the line between counting, sampling, instrumentation and tracing, so a question is matched to its tool.

  3. 03
    BPF Performance Tools book

    The reference for time a CPU sampler cannot see, off-CPU, scheduler and I/O waits, traced at bounded cost.

  4. 04
    bpftrace repository

    Makes a tracing hypothesis a one-line experiment over kprobes, uprobes, tracepoints and PMU events.

  5. 05
    Google-Wide Profiling: A Continuous Profiling Infrastructure for Data Centers paper

    The design continuous profilers descend from, always-on sampling across a fleet, cheap enough to leave running.

Counters, events and precise sampling

  1. 01
    perf_event_open(2) manual

    Defines the counting and sampling modes every Linux profiler uses, and the sample record fields, branch stack included.

  2. 02
    Instruction-Based Sampling: A New Performance Analysis Technique paper

    Defines skid, why a sample lands after the culprit, and how tagging one op through the pipeline removes it.

  3. 03
    Intel Software Developer Manuals manual

    Defines the architectural counters, what a PEBS sample captures, and why rdtsc counts time, not cycles, under DVFS.

  4. 04
    Processor Programming Reference for AMD Family 1Ah Model 02h manual

    Defines the event encodings and IBS registers for one Zen core, the tables perf's AMD events are derived from.

  5. 05
    perf-arm-spe(1) repository

    Defines Arm SPE as perf drives it, one sampled op in flight, the filters, and what a record holds.

CPU profilers and flame graphs

  1. 01
    Linux perf wiki: Tutorial manual

    The maintainers' walk from perf stat to perf record to perf annotate, with the sample fields each flag sets.

  2. 02
    Intel VTune Profiler Documentation manual

    Home of the user guide and cookbook, where each hardware analysis is defined by the events behind it.

  3. 03
    AMD uProf User Guide manual

    The vendor's reference for IBS-driven profiling on Zen, with metric presets defined per core generation.

  4. 04
    The Flame Graph paper

    Records the design decisions, width as sample share and alphabetical rather than time order, behind its reading rules.

  5. 05
    Coz: Finding Code that Counts with Causal Profiling paper

    Proves a hot function need not be worth optimising, by measuring what speeding up a line does to end-to-end time.

Microbenchmarks that lie

  1. 01
    Producing Wrong Data Without Doing Anything Obviously Wrong! paper

    The origin of measurement bias as a term, link order and environment size alone flipping a compiler flag comparison.

  2. 02
    Non-Determinism and Overcount on Modern Hardware Performance Counter Implementations paper

    Traces run-to-run variation in x86 retired-instruction counts to one extra count per interrupt and per fault.

  3. 03
    clock_gettime(2) manual

    Defines what each clock counts, NTP-slewed or raw monotonic time, or CPU time, and the resolution call.

  4. 04
    Benchmarking tips (LLVM) manual

    The compiler project's recipe for a quiet Linux host, governor, boost, SMT siblings, ASLR, a cpuset and tmpfs.

  5. 05
    Google Benchmark User Guide repository

    Documents the barriers that keep the optimiser from deleting the work under test, and repetition statistics.

Reproduce it: 05-measurement-pitfalls, dead-code elimination, run-to-run spread, cold against warm.

06Models

A roofline is a bound built from measured roofs and counted bytes, so a point above a roof means a wrong roof or a wrong byte count, not fast code.

Roofline and the execution-cache-memory model

  1. 01
    Roofline: An Insightful Visual Performance Model for Multicore Architectures paper

    Defines operational intensity as traffic past the caches, and the ceilings that say which optimisation can pay.

  2. 02
    Applying the Roofline Model paper

    States the counter set and method that put a measured point on the plot in place of a hand-counted intensity.

  3. 03
    LIKWID repository

    Measures the roofs with likwid-bench and the point from likwid-perfctr groups on Intel, AMD and Arm server cores.

  4. 04
    Cache-aware Roofline model: Upgrading the loft paper

    Adds a roof per cache level against traffic at the core, so a kernel that hits in cache is not plotted as DRAM bound.

  5. 05
    Performance bottlenecks of stencil computations using the Execution-Cache-Memory model paper

    Times each level's transfer with overlap rules, predicting a core's rate and the core count where bandwidth saturates.

Reproduce it: 06-roofline, measured roofs and three kernels of rising arithmetic intensity.

Top-down analysis

  1. 01
    A Top-Down Method for Performance Analysis and Counters Architecture paper

    Defines the slot accounting that turns raw counters into a weighted tree of bottlenecks, and why the unit is a slot.

  2. 02
    TMA_Metrics-full.xlsx (intel/perfmon) repository

    The official home of every TMA formula, event, threshold and level per microarchitecture, from which toplev derives.

  3. 03
    pmu-tools repository

    Runs the TMA tree on Linux from the spreadsheet and states why multiplexed levels mislead on short or varied workloads.

  4. 04
    pipeline.json (perf pmu-events, amdzen4) repository

    States the AMD top-down formulas in events for both levels, the metric groups perf stat runs on Zen with no vendor tool.

  5. 05
    Arm CPU Telemetry Solution Topdown Methodology Specification manual

    Defines Arm's top-down as staged stall accounting, first locating the bottleneck and then measuring its resource.

Scaling laws

  1. 01
    Validity of the single processor approach to achieving large scale computing capabilities paper

    States the serial-fraction bound on speedup, the argument every later scaling law is written against.

  2. 02
    Reevaluating Amdahl's Law paper

    Defines scaled speedup, the bound that holds when the problem grows with the processor count instead of staying fixed.

  3. 03
    Amdahl's Law in the Multicore Era paper

    Extends the bound to chips of unequal cores under a fixed area budget, the arithmetic behind big and little cores.

  4. 04
    A Simple Capacity Model of Massively Parallel Transaction Systems paper

    The origin of the coherency term that makes throughput fall, not merely flatten, as processors are added.

  5. 05
    Guerrilla Capacity Planning book

    Derives the universal scalability law and states the procedure that fits its coefficients to measured throughput.

Queueing

  1. 01
    A Proof for the Queuing Formula: L = λW paper

    Proves occupancy equals arrival rate times time in system with no assumption on arrivals or service.

  2. 02
    Quantitative System Performance report

    Origin of the bound-and-bottleneck analysis roofline names as its ancestor, and of the asymptotic bounds on throughput.

  3. 03
    Performance Modeling and Design of Computer Systems book

    Proves why open and closed systems answer a load question differently, and when scheduling, not capacity, sets latency.

  4. 04
    Stochastic Processes Occurring in the Theory of Queues paper

    Origin of the A/S/c notation, and of the embedded chain that solves a queue with non-memoryless arrivals or service.

  5. 05
    The single server queue in heavy traffic paper

    Derives the wait near saturation from utilisation and arrival and service variance, the formula behind the latency knee.

07Single-thread optimisation

A loop the compiler reports as vectorised can still run at scalar speed: a float reduction stays one serial chain until reassociation is permitted.

Data layout and loop transforms

  1. 01
    Intel Optimization Reference Manual manual

    Where Intel states when a structure of arrays beats an array of structures, strided and hybrid cases included.

  2. 02
    ispc: A SPMD Compiler for High-Performance CPU Programming paper

    Measures one kernel in both layouts and traces the gain to the gathers the array-of-structures form forces.

  3. 03
    A Data Locality Optimizing Algorithm paper

    Defines interchange, skewing, reversal and tiling as one family and proves when each keeps a loop nest legal.

  4. 04
    The Cache Performance and Optimizations of Blocked Algorithms paper

    Traces the drops in a blocked loop's speed curve to self-interference misses, and shows when copying a tile pays.

  5. 05
    Auto-Vectorization in LLVM manual

    Where LLVM lists what its loop and SLP vectorisers accept: interleaving, reductions, if-conversion and runtime checks.

Reproduce it: 07-aos-vs-soa-simd, array of structs against structure of arrays, scalar against NEON.

SIMD instruction sets

  1. 01
    Intel Intrinsics Guide manual

    Maps each intrinsic to its instruction and CPUID flag, with the vendor's latency and throughput per microarchitecture.

  2. 02
    Intel Software Developer Manuals manual

    The normative semantics of every Intel vector instruction, with the masking, rounding and fault rules intrinsics hide.

  3. 03
    Introduction to SVE manual

    Where Arm explains vector-length-agnostic loops and predication, the model that removes remainder loops entirely.

  4. 04
    Arm C Language Extensions manual

    The specification the NEON and SVE intrinsics come from, so it settles what a compiler must accept and what is a bug.

  5. 05
    Arm Architecture Reference Manual for A-profile architecture manual

    The normative definition of NEON, SVE and SVE2 instructions and of the scalable vector and predicate register model.

SIMD libraries and measured kernels

  1. 01
    Highway repository

    Defines sizeless vector types with run-time dispatch, so one source serves SVE and every fixed-width ISA.

  2. 02
    xsimd repository

    Fixes a batch type per ISA and width, the compile-time vector length model, and still reaches NEON and SVE.

  3. 03
    Faster Base64 Encoding and Decoding Using AVX2 Instructions paper

    Where shuffle-based lookup replacing a byte loop is worked through step by step, with the measurement method spelt out.

  4. 04
    Parsing Gigabytes of JSON per Second paper

    Shows branch-free structural indexing with carry-less multiply and shuffles, measured against conventional parsers.

  5. 05
    simdjson repository

    Where that technique ships, with a kernel per ISA and the harness that keeps its published comparisons reproducible.

  6. 06
    Hyperscan: A Fast Multi-pattern Regex Matcher for Modern CPUs paper

    Decomposes regexes into string and automaton pieces so both run on SIMD, the design inside the matcher Snort embeds.

Branchless code and bit manipulation

  1. 01
    Branch Prediction and the Performance of Interpreters paper

    Counter evidence that current predictors absorb interpreter dispatch, so a branch removal must be measured, not assumed.

  2. 02
    Hacker's Delight, 2nd Edition book

    Derives the branch-free integer tricks, division by a constant among them, with proofs rather than as a catalogue.

  3. 03
    Faster sorted array unions by reducing branches report

    A worked branchless merge with code whose gain vanishes when the compiler emits no conditional move.

  4. 04
    Array Layouts for Comparison-Based Searching paper

    Measures branch-free search over sorted, Eytzinger and B-tree layouts, the winner changing with size and prefetch.

  5. 05
    Faster Population Counts Using AVX2 Instructions paper

    Shows a carry-save adder tree in vector registers beating the dedicated instruction, timed as the minimum of many runs.

08Compilers and codegen

No -O level changes the target instruction set: without -march or -mcpu, every instruction emitted belongs to the default target ISA, so target flags come before any judgement of codegen.

Reading emitted code

  1. 01
    Compiler Explorer repository

    Shows how a source change alters the emitted instructions across compilers, versions and flags, with nothing installed.

  2. 02
    What Every C Programmer Should Know About Undefined Behavior report

    Explains how the signed-overflow and aliasing rules let a trip count be known and a store loop become memset.

  3. 03
    llvm-objdump manual

    Reads the binary that shipped, with source lines and symbolised branch targets, rather than a recompiled snippet.

  4. 04
    llvm-mca manual

    Predicts loop throughput and port pressure from the scheduling model, and states it models neither front end nor caches.

  5. 05
    llvm-exegesis manual

    Measures instruction latency and throughput with counters, so the model llvm-mca predicts from is checked, not trusted.

  1. 01
    Options That Control Optimization (GCC) manual

    Lists what each -O level turns on, the inlining limits, and that -Ofast admits transforms invalid for conforming code.

  2. 02
    There Are No Zero-cost Abstractions (CppCon 2019) talk

    Shows with real codegen that an abstraction is free only when inlining and the ABI allow it, and the cost when either refuses.

  3. 03
    Itanium C++ ABI manual

    Fixes the rule that a non-trivial class goes by reference to a caller-made temporary, the cost a wrapped pointer pays.

  4. 04
    How To Write Shared Libraries paper

    States what PLT calls and interposition cost, and the visibility controls a library needs to inline its own exports.

  5. 05
    LTO Overview (GCC Internals) manual

    Defines whole-program LTO against partitioned WHOPR, and the LGEN, WPA and LTRANS stages that run -flto in parallel.

  6. 06
    ThinLTO manual

    Defines the thin link, summaries analysed whole-program then parallel backends, and the cache for incremental rebuilds.

Target flags and auto-vectorisation

  1. 01
    x86 Options (GCC) manual

    Defines -march against -mtune, the psABI levels, and -mprefer-vector-width, the switch for full-width AVX-512 code.

  2. 02
    Function Multiversioning (GCC) manual

    Defines target_clones, one function per ISA behind a resolver the dynamic linker runs, so a generic build ships AVX-512.

  3. 03
    Controlling Floating Point Behavior (Clang) manual

    Lists what -ffast-math implies, of which -fassociative-math alone frees a float reduction, and -ffp-contract for FMA.

  4. 04
    Auto-Vectorization in LLVM manual

    States what the vectorisers need, aliasing disproved or checked at run time, and where a float reduction stays in order.

  5. 05
    Options to Emit Optimization Reports (Clang) manual

    Defines the remarks that make the compiler say which loop it left scalar and why, so the fix targets the real blocker.

Reproduce it: 08-autovectorization-aliasing, the vectoriser with and without restrict.

Profile-guided and post-link optimisation

  1. 01
    Profile Guided Optimization (Clang) manual

    Defines the instrumented and sampled workflows, why their profiles cannot mix, and the cost of a wrong training input.

  2. 02
    AutoFDO: Automatic Feedback-Directed Optimization for Warehouse-Scale Applications paper

    Defines the address-to-source mapping with discriminators that lets a stale production profile still drive FDO.

  3. 03
    BOLT: A Practical Binary Optimizer for Data Centers and Beyond paper

    States why a profile applied to the final binary beats one mapped to source, and the layout passes accuracy enables.

  4. 04
    BOLT (llvm-project/bolt) repository

    States what full effect needs, relocations kept at link time and a branch-stack sample profile, neither on by default.

  5. 05
    RFC: Propeller: A frame work for Post Link Optimizations report

    States the design, not a result: basic-block sections and a relink, no binary rewrite, so layout needs no disassembly.

09Concurrency

Every cost below is a cache line moving between cores, so the ordering models and the measured line-transfer cost in the memory hierarchy section come first.

Memory models and atomics

  1. 01
    Foundations of the C++ Concurrency Memory Model paper

    Defines the data-race-free contract, sequential consistency for race-free programs and no meaning for a race.

  2. 02
    Atomic operations, C++ working draft manual

    The normative wording for every memory order, fence and read-modify-write, the text a compiler is checked against.

  3. 03
    C/C++11 mappings to processors report

    The table that turns each memory order into x86 and Arm instructions, so what an order costs is read off the page.

  4. 04
    Linux kernel memory-barriers.txt manual

    States what the kernel assumes any CPU may reorder and what each barrier and access primitive guarantees.

  5. 05
    herdtools7 repository

    Where herd7, litmus7 and klitmus7 live, the tools that run a litmus test against the x86, Arm and kernel models.

Locks, contention and allocators

  1. 01
    Is Parallel Programming Hard, And, If So, What Can You Do About It? book

    Derives counting, partitioning, locking and deferral with code that runs, the textbook the section assumes.

  2. 02
    Algorithms for Scalable Synchronization on Shared-Memory Multiprocessors paper

    The origin of the queue lock, each waiter spinning on its own line, measured against ticket and test-and-set locks.

  3. 03
    Futexes Are Tricky paper

    Derives a correct user-space mutex from futex and shows the lost wakeups and extra kernel entries naive versions pay.

  4. 04
    Hoard: A Scalable Memory Allocator for Multithreaded Applications paper

    Defines blowup and allocator-induced false sharing, which per-processor heaps under a bounded global heap avoid.

  5. 05
    TCMalloc: Thread-Caching Malloc manual

    The design statement for per-CPU caches built on restartable sequences and a hugepage-aware back end for TLB reach.

Reproduce it: 09-false-sharing, adjacent counters against padded counters across threads.

Lock-free structures and RCU

  1. 01
    The Art of Multiprocessor Programming book

    States linearizability and the consensus hierarchy and builds both into working stacks, queues, lists and hash tables.

  2. 02
    Simple, Fast, and Practical Non-Blocking and Blocking Concurrent Queue Algorithms paper

    The lock-free queue later libraries copy, with the counted pointer against ABA and a two-lock queue beside it.

  3. 03
    Hazard Pointers for C++26 paper

    The standard-track form of safe reclamation, fixing when a retired node may be freed while a reader still holds it.

  4. 04
    What is RCU? manual

    The kernel's own statement of RCU as publish, wait for readers and keep old versions, with a free read side.

  5. 05
    User-Level Implementations of Read-Copy Update paper

    Defines liburcu's quiescent-state, signal-based and general RCU flavours and measures each read side against locks.

Thread pools and work stealing

  1. 01
    Cilk: An Efficient Multithreaded Runtime System paper

    Defines work and critical path, proves the work-stealing bound, and shows they alone predict a runtime's speedup.

  2. 02
    The Implementation of the Cilk-5 Multithreaded Language paper

    States the work-first principle, that overhead belongs on the rare steal path and not on every spawn.

  3. 03
    Correct and Efficient Work-Stealing for Weak Memory Models paper

    Gives the work-stealing deque a proven atomics form and derives which fence push, take and steal need on Arm and x86.

  4. 04
    OpenMP Specifications manual

    Fixes fork-join and tasking semantics, and the wait and binding controls deciding if idle workers spin, sleep or move.

  5. 05
    oneTBB repository

    The shipping work-stealing runtime, arenas and task groups over a deque, the home of grain size and spin-before-sleep.

10NUMA and multi-socket

A page's node is decided at first touch, not when memory is allocated or a policy is set, and every vendor table below depends on the BIOS node mode of the machine it ran on.

NUMA and Linux memory placement

  1. 01
    NUMA (Non-Uniform Memory Access): An Overview paper

    The one account tying first touch, policy scope, zone reclaim and page movement together from the implementer's side.

  2. 02
    What is NUMA? manual

    Defines nodes, zonelists and the distance-ordered fallback that places an allocation once local memory runs out.

  3. 03
    NUMA Memory Policy manual

    The normative statement of policy scopes, every mode including weighted interleave, and the cpuset intersection rule.

  4. 04
    Numa policy hit/miss statistics manual

    Defines numa_hit, numa_miss and numa_foreign, the counters that show whether a policy put pages where it said.

  5. 05
    numactl repository

    Reference implementation of the policy API, prints the distance table and binds a binary that cannot be rebuilt.

Reproduce it: 10-first-touch, first touch of fresh pages against the second pass.

Topology and interconnects

  1. 01
    NUMA Memory Performance manual

    Explains the firmware-rated latency and bandwidth per initiator and target, and memory-side caches, that rank nodes.

  2. 02
    Intel Xeon Processor Scalable Family Technical Overview manual

    Where Intel names the mesh, UPI socket links, the directory-running home agent, and how SNC splits the cache.

  3. 03
    Intel Xeon 6 with P-cores Configuration and Tuning Guide for HPC Applications manual

    Defines SNC on current parts as one node per compute die, and fixes the numactl and numastat checks of placement.

  4. 04
    BIOS and Workload Tuning Guide for AMD EPYC 9004 Series Processors manual

    Discloses the I/O die, GMI and xGMI links, the NPS modes with their interleave widths, and the cache-as-NUMA override.

  5. 05
    Arm Neoverse CMN-700 Coherent Mesh Network Technical Reference Manual manual

    Defines the mesh, the home nodes holding the system cache and snoop filter, and the gateways joining sockets or CXL.

Migration, balancing and measured effects

  1. 01
    move_pages(2) manual

    Defines per-page migration of a running process, and a query reporting each page's node, the direct test of first touch.

  2. 02
    sysctl kernel numa_balancing manual

    Defines the hinting-fault sampling behind automatic balancing and tiering, and warns the overhead may not pay off.

  3. 03
    Traffic Management: A Holistic Approach to Memory Placement on NUMA Systems paper

    Proves against the kernel balancer that controller and link congestion, not remote latency, is what placement manages.

  4. 04
    Intel Memory Latency Checker manual

    Measures the node-to-node latency and bandwidth matrix and loaded latency on the x86 at hand, which no datasheet states.

  5. 05
    High Performance Computing Tuning Guide for AMD EPYC 9004 Series Processors manual

    Tabulates measured bandwidth by NPS mode, cores per die, boost and SMT, so the NPS trade-off is shown, not asserted.

11OS and I/O

A syscall's cost depends on the mitigation state, the governor and the idle state the core was in, three sysfs settings that change after boot, so each is recorded beside any number below.

Syscalls and asynchronous I/O

  1. 01
    vdso(7) manual

    Defines the calls the kernel answers without a mode switch, and the clocksource condition for skipping the trap.

  2. 02
    FlexSC: Flexible System Call Scheduling with Exception-Less System Calls paper

    Separates a syscall's trap cost from its cache and TLB pollution, and shows the pollution can dominate.

  3. 03
    An Analysis of Performance Evolution of Linux's Core Operations paper

    Measures syscall and context switch cost across kernel releases and traces each slowdown to a named mitigation.

  4. 04
    Efficient IO with io_uring paper

    States the goals aio failed, and which io_uring features remove a syscall and which remove a copy.

  5. 05
    Understanding Modern Storage APIs: A systematic study of libaio, SPDK, and io_uring paper

    Measures io_uring's polling modes against libaio and SPDK, and shows the kernel poller needs its own core.

Reproduce it: 11-syscall-cost, the fixed cost of a kernel crossing across request sizes.

Scheduling, affinity and isolation

  1. 01
    EEVDF Scheduler manual

    Defines lag and virtual deadline, which the default class schedules by, and the slice request in sched_setattr.

  2. 02
    The Linux Scheduler: a Decade of Wasted Cores paper

    Proves cores sit idle while runnable threads queue, and gives the invariant checker that found the load-balancer bugs.

  3. 03
    Control Group v2 manual

    Defines cpu.max throttling, cpu.weight and the cpusets that bound affinity, the controls behind every container limit.

  4. 04
    CPU Performance Scaling manual

    Defines the governors, driver and boost switch that set a core's frequency, the sysfs state a measurement records.

  5. 05
    CPU Isolation manual

    Ties isolcpus, nohz_full, IRQ affinity, RCU offload and cpusets into one recipe, and lists the jitter it leaves.

Interrupts and kernel bypass

  1. 01
    NAPI manual

    Defines the polling, software coalescing, busy polling and IRQ suspension knobs that trade interrupts against latency.

  2. 02
    DPDK Programmer's Guide manual

    Defines the full bypass model, pinned poll-mode cores with no interrupts, that every kernel path is measured against.

  3. 03
    The eXpress Data Path repository

    Measures an in-kernel programmable path against DPDK and the stack per core, with the full configuration published.

  4. 04
    Kernel vs. User-Level Networking: Don't Throw Out the Stack with the Interrupts paper

    Separates direct and indirect NIC interrupt cost, measures the stack against bypass, and is where IRQ suspension began.

  5. 05
    AF_XDP manual

    Defines the socket and UMEM rings handing XDP frames to user space, and the zero-copy and need-wakeup modes.

Cache and bandwidth partitioning

  1. 01
    Intel Resource Director Technology Architecture Specification manual

    Defines classes of service, cache masks, bandwidth allocation and monitoring IDs, the model resctrl exposes.

  2. 02
    User Interface for Resource Control feature (resctrl) manual

    Defines the filesystem through which Linux exposes Intel, AMD and Arm partitioning, and the schemata format.

  3. 03
    MPAM manual

    Maps Arm's cache portion and bandwidth controls onto resctrl's schemata, and states which platform limits apply.

  4. 04
    CPI2: CPU performance isolation for shared compute clusters paper

    Shows at fleet scale that cycles per instruction alone finds an interfering neighbour and the one to throttle.

  5. 05
    Heracles: Improving Resource Efficiency at Scale paper

    Shows cache ways, cores, bandwidth and power must be partitioned together, or batch work reaches the tail.

12Tail latency and production systems

A latency figure means nothing without its percentile, its load model and the way it was recorded.

Measuring the tail

  1. 01
    The Tail at Scale paper

    Shows why fan-out makes a rare slow server a common slow request, and names the techniques that tolerate variance.

  2. 02
    Attack of the Killer Microseconds paper

    Defines the stall band that out-of-order hardware cannot hide and a context switch cannot amortise.

  3. 03
    How NOT to Measure Latency talk

    Shows that a summary without a max discards the samples that define the tail, and closed-loop load never records them.

  4. 04
    Coordinated Omission report

    The original definition of the recording error, with arithmetic for how far a reported percentile sits from the truth.

  5. 05
    HdrHistogram repository

    Keeps the whole distribution at fixed relative precision in constant time, so the far percentiles and max survive.

Reproduce it: 12-coordinated-omission, closed-loop against open-loop p99 under the same stalls.

Where jitter comes from

  1. 01
    rt-tests repository

    The reference wakeup-latency measurement for Linux, whose README states that an unloaded run proves nothing.

  2. 02
    osnoise tracer manual

    Counts the noise a spinning thread suffers and attributes each event to NMI, IRQ, softirq, thread or hardware.

  3. 03
    Tales of the Tail paper

    Derives the queueing-ideal tail and attributes the excess to scheduling, interrupt placement, power saving and NUMA.

  4. 04
    Latency Implications of Virtual Memory report

    Measures with code the page-fault, TLB-shootdown and writeback stalls that memory mapping hides from the caller.

  5. 05
    The KVM halt polling system manual

    Defines the host-side polling after a vCPU halt that trades idle host CPU for guest wakeup time, unseen by the guest.

Load generation and production workloads

  1. 01
    Open Versus Closed: A Cautionary Tale paper

    Shows that open and closed load models disagree on response time and scheduling gains, with rules for choosing one.

  2. 02
    wrk2 repository

    Issues requests on a fixed schedule and times each from when it was due, so server stalls reach the percentiles.

  3. 03
    Reconciling High Server Utilization and Sub-millisecond Quality-of-Service paper

    Shows the tail, not throughput, caps a latency-critical server's utilisation, and how far co-located work lowers it.

  4. 04
    TailBench report

    Pairs latency-critical services with an open-loop harness that records sojourn against service time per request.

  5. 05
    Workload Analysis of a Large-Scale Key-Value Store paper

    Measures the key, value and inter-arrival distributions of live key-value traffic, the shape load generators imitate.

Mechanical sympathy

  1. 01
    Inter Thread Latency report

    Measures with code the floor for handing a cache line between cores, which every queue and lock is built on.

  2. 02
    Single Writer Principle report

    States the design rule that removes write contention outright, using a contended increment's cost as the argument.

  3. 03
    Optimizing a Ring Buffer for Throughput report

    Adds cached indices to a single-producer single-consumer ring and shows with counters the coherence traffic removed.

  4. 04
    LMAX Disruptor manual

    Applies the single writer rule and cache-line padding to a ring buffer, with the queue comparison that motivated it.

  5. 05
    Aeron repository

    Carries the single writer and batching rules through a whole transport, the reference beyond one in-process queue.

13Inference on CPU

A decode step at batch one reads every weight once for a few flops, so it runs at the memory system's rate, and a CPU figure compares with a GPU figure only at the same batch and precision.

GEMM and BLAS

  1. 01
    Anatomy of High-Performance Matrix Multiplication paper

    Derives the cache blocking every fast CPU GEMM still uses by refining a memory model until the packed kernel falls out.

  2. 02
    BLIS: A Framework for Rapidly Instantiating BLAS Functionality paper

    Reduces a BLAS to one register-tile micro-kernel per architecture and states which loop owns which level of cache.

  3. 03
    LLaMA Now Goes Faster on CPUs report

    Builds a register-tile kernel step by step and shows where outer-loop unrolling wins and where a vendor BLAS still does.

  4. 04
    OpenBLAS repository

    Sets run-time dispatch for a BLAS, one kernel table per microarchitecture chosen as a DYNAMIC_ARCH build loads.

  5. 05
    oneDNN Matrix Multiplication Primitive manual

    Fixes the data-type table, packed weight format and fused post-ops a CPU GEMM must expose to an inference graph.

  6. 06
    Scaled Dot-Product Attention (oneDNN Graph) manual

    Fixes the fused attention pattern, f32 accumulation under bf16 inputs and the shapes the fast CPU path accepts.

Reproduce it: 13-sgemm-naive-vs-blas, naive GEMM, a hand microkernel and the vendor BLAS, plus int8 against float32 dot products.

Runtimes

  1. 01
    oneDNN repository

    Generates VNNI and AMX kernels at run time under PyTorch, TensorFlow and OpenVINO, and calls Compute Library on Arm.

  2. 02
    ggml repository

    Defines the quantized block formats and the per-ISA dot-product kernels over them that every llama.cpp type rests on.

  3. 03
    llama.cpp repository

    Where new quantization types, kernels and thread pools land first, each with the perplexity and speed table behind it.

  4. 04
    ONNX Runtime MLAS repository

    Holds the CPU provider's GEMM, int8 and int4 MatMul kernels, dispatched per ISA at run time from SSE to AMX and SME.

  5. 05
    OpenVINO CPU Device manual

    States precision defaults per ISA, the int8 path through oneDNN and the streams model that turns cores into throughput.

Quantization

  1. 01
    Quantization and Training of Neural Networks for Efficient Integer-Arithmetic-Only Inference paper

    Defines the affine scale and zero-point scheme with integer accumulation that every int8 CPU runtime still implements.

  2. 02
    Nuances of int8 Computations manual

    States where int8 saturates on pre-VNNI x86, the compensation that makes signed by signed work, and what VNNI removes.

  3. 03
    Quantize ONNX models manual

    Defines the QDQ form against QOperator and dynamic against static, and which sign choices are safe on which CPUs.

  4. 04
    k-quants repository

    Defines super-block formats with quantized scales and the perplexity against size curve behind mixing types per tensor.

  5. 05
    T-MAC: CPU Renaissance via Table Lookup for Low-Bit LLM Deployment on Edge paper

    Replaces unpacking and multiplying low-bit weights with bit-sliced lookups, so a narrower weight costs less to apply.

Matrix extensions

  1. 01
    Intel Software Developer Manuals manual

    Defines AMX tiles, tile configuration and TMUL, the VNNI and bf16 dot products, and the XSAVE state a tile load needs.

  2. 02
    Using XSTATE features in user space applications manual

    Defines the arch_prctl permission Linux requires for AMX tile data and the trap a tile instruction takes until granted.

  3. 03
    Add Intel Advanced Matrix Extensions (AMX) support to ggml repository

    Holds the design report for the ggml AMX path, whose double buffering hides applying scales between tile products.

  4. 04
    Arm Architecture Reference Manual for A-profile architecture manual

    Defines streaming SVE mode, the ZA tile array and SME2 multi-vector instructions, with pseudocode a kernel must match.

  5. 05
    SME Programmer's Guide manual

    Walks through SME2 int8 and f32 matmul and gemv kernels and a lookup-table path, on a unit so far only in client parts.

Threading for inference

  1. 01
    Thread management manual

    Sets the physical-core default, the affinity it implies and the spin-wait controls that trade idle CPU for latency.

  2. 02
    Performance Hints and Thread Scheduling manual

    States the vendor defaults, one thread per core, SMT siblings off, core type by precision and one socket for latency.

  3. 03
    Threadpool: take 2 repository

    Defines the explicit thread pool with CPU masks, strict placement, priority and polling that ggml runs without OpenMP.

  4. 04
    llama-bench repository

    Defines the prompt and generation tests, repetitions and mean with deviation behind any comparable llama.cpp number.

  5. 05
    Dual Epyc Genoa/Turin token generation performance bottlenecks repository

    Traces poor decode scaling across sockets to remote NUMA access from weight placement, with numatop counts as evidence.

When CPU beats GPU

  1. 01
    gpt-j example README (ggml) repository

    Origin of ggml's bandwidth argument, NEON threads saturate a laptop's memory bus, so a GPU on that bus gains nothing.

  2. 02
    llama.cpp Performance Testing report

    Holds the memory clock sweep with all else fixed, the token rate following it and flattening after a few threads.

  3. 03
    SparAMX: Accelerating Compressed LLMs Token Generation on AMX-powered CPUs paper

    Measures decode kernels on a Xeon as DRAM-bound, so AMX pays at batch one only once bytes per token shrink, with code.

  4. 04
    MLPerf Inference v6.0 Results repository

    Holds CPU-only entries whose logs show the batched Offline case lost to GPU entries at the same accuracy floor.

14Hardware generations

The measurement articles below state no compiler, flags or run count, so each is kept for the structure it exposes and no figure from it is repeated.

Intel Xeon

  1. 01
    Technical Overview of the 4th Gen Intel Xeon Scalable Processor Family manual

    Where Intel states what Sapphire Rapids added: larger L2 and L3, DDR5, CXL, AMX and on-die accelerators.

  2. 02
    Sapphire Rapids: Golden Cove Hits Servers report

    Measures L3 and memory latency across the tiled mesh and the slow clock ramp the vendor overview omits.

  3. 03
    Emerald Rapids: 5th-Generation Intel Xeon Scalable Processors paper

    The designers' statement of what changed: fewer, larger dies, a bigger shared L3, faster DDR5 and socket links.

  4. 04
    A Look into Intel Xeon 6's Memory Subsystem report

    Measures per-die L3 under sub-NUMA clustering and the die-crossing cost on Granite Rapids beside Turin.

  5. 05
    Benchmarking the Evolution of Performance and Energy Efficiency Across Recent Generations paper

    Bandwidth-bound codes on Sapphire, Emerald and Granite Rapids and Sierra Forest with clocks, SMT and compiler stated.

AMD EPYC

  1. 01
    AMD Next-Generation Zen 4 Core and 4th Gen AMD EPYC Server CPUs paper

    The designers' account of the Zen 4 core and how it yields Genoa, Genoa-X, Bergamo and Siena.

  2. 02
    Testing AMD's Bergamo: Zen 4c Spam report

    Tests the same-core claim for Zen 4c: cache latency, clock ceiling and core-to-core paths beside clock-matched Zen 4.

  3. 03
    Software Optimization Guide for the AMD Zen5 Microarchitecture manual

    Where AMD states what Zen 5 changed in front end, vector datapath and caches, the core Turin carries.

  4. 04
    AMD EPYC 9005 Processor Architecture Overview manual

    Defines Turin: Zen 5 or Zen 5c dies, which parts double die-to-IO links, NUMA modes and full-width AVX-512.

  5. 05
    AMD's Turin: 5th Gen EPYC Launched report

    Measures what wider die-to-IO links and faster DDR5 do for Turin bandwidth, and where latency rose over Genoa.

Arm Neoverse server parts

  1. 01
    AWS Graviton Getting Started repository

    Where AWS states which Neoverse core, ISA revision, mesh, caches and compiler flag each Graviton generation carries.

  2. 02
    Arm Neoverse V2 Core Software Optimization Guide manual

    Sets the pipeline widths and instruction timings of the core Graviton 4, Grace and Axion share.

  3. 03
    NVIDIA Grace Performance Tuning Guide manual

    States the coherency fabric, LPDDR5X fit and MPAM cache and memory partitioning on Grace.

  4. 04
    Arm Neoverse N2 Core Software Optimization Guide manual

    States the timings and fusion rules of the narrower N line core in Cobalt 100 and Yitian.

  5. 05
    Ampere Altra Rev A1 64-Bit Multi-Core Processor Datasheet manual

    Where Ampere states the N1 part: private L2 per core, shared system cache, mesh and DDR4 fit.

Independent measurement across vendors

  1. 01
    Microarchitectural Comparison and In-core Modeling of State-of-the-art CPUs paper

    Measures Neoverse V2, Golden Cove and Zen 4 in-core at fixed clock, and each socket's clock under vector load.

  2. 02
    On the Performance of Cloud-based ARM SVE for Zero-Knowledge Proving Systems paper

    One SVE workload run on Graviton 3, Graviton 4, Yitian and Axion with compiler, runs and spread stated.

  3. 03
    Arm's Neoverse V2, in AWS's Graviton 4 report

    Measures a sustained rename width below the stated one, cache latencies, mesh behaviour and cross-socket cost on V2.

  4. 04
    ARM's Neoverse N2: Cortex A710 for Servers report

    Measures structure sizes, cache latencies and mesh behaviour of N2 on Yitian, the core Cobalt 100 carries.

  5. 05
    AmpereOne at Hot Chips 2024: Maximizing Density report

    Puts vendor slides beside measurements of the predictor, small instruction cache, private L2 and long memory latency.

Reproduce it: 14-pcore-vs-ecore, the same three kernels on a performance core and an efficiency core.

15Benchmarks

A score means what its suite's run rules say it means, so the rules come before the number.

Standard suites

  1. 01
    SPEC CPU 2026 Run and Reporting Rules manual

    Defines base against peak, rate against speed, the threading models a speed run may use and an Arm reference machine.

  2. 02
    SPEC CPU: The Next Generation paper

    Where the committee states how workloads were chosen and hardened, and defines the rolling round-robin rate.

  3. 03
    SPEC CPU2026: Characterization, Representativeness, and Cross-Suite Comparison paper

    Measures with counters what each workload stresses on x86 and Arm server parts, beside data-centre and inference suites.

  4. 04
    MLPerf Inference Rules repository

    Fixes model, accuracy floor and query pattern per scenario, so a Server score is throughput under a latency bound.

  5. 05
    DCPerf: An Open-Source, Battle-Tested Performance Benchmark Suite for Datacenter Workloads paper

    Shows standard suites misproject data-centre servers, and states the fleet-matching method the suite is built by.

Microbenchmark suites

  1. 01
    Memory Bandwidth and Machine Balance in Current High Performance Computers paper

    Defines sustainable bandwidth as what unit-stride loops get, not bus peak, and machine balance as flops per access.

  2. 02
    STREAM Benchmark Reference Information report

    Sets the array size rule, timing over repeated trials, and counting bytes a loop asks for, not what the cache moved.

  3. 03
    lmbench: Portable Tools for Performance Analysis paper

    Origin of the one-mechanism-per-test method for memory, system call, pipe and socket latency, and what each leaves out.

  4. 04
    uarch-bench repository

    Isolates memory-level parallelism from load latency as separate tests, with DVFS held off before timing, x86 Linux only.

  5. 05
    nanoBench: A Low-Overhead Tool for Running Microbenchmarks on x86 Systems paper

    Shows why kernel mode with interrupts off matters, removes harness overhead, then recovers cache replacement policies.

Reproduce it: 15-stream-bandwidth, triad bandwidth by thread count against the vendor figure.

Methodology and what suites miss

  1. 01
    How Not to Lie with Statistics: The Correct Way to Summarize Benchmark Results paper

    Origin of the rule that normalised results take the geometric mean, which a SPEC ratio and the crimes list rest on.

  2. 02
    Systems Benchmarking Crimes report

    Checklist of evaluation faults from sub-setting and improper baselines to arithmetic means of ratios, each with a fix.

  3. 03
    Scientific Benchmarking of Parallel Computing Systems paper

    Sets which mean fits costs, rates and ratios, when confidence intervals are owed, and the absolute base a speedup needs.

  4. 04
    Rigorous Benchmarking in Reasonable Time paper

    Decides how many builds, runs and iterations an experiment needs by measuring at which level the variation arises.

  5. 05
    Profiling a warehouse-scale computer paper

    Fleet counter profile showing services stall on instruction fetch and burn cycles in shared routines, which SPEC lacks.

16Watchlist

Everything below is real but unproven: no item yet has all three of a written specification, a part you can buy, and a public measurement stating every one of the seven fields. Each line says what would promote it. Vendor multiples never qualify. Last checked 2026-09-15.

ISA extensions without a shipped server part

  1. 01
    Intel AVX10.2 Architecture Specification manual

    Folds the AVX-512 subsets into one versioned level and adds BF16 and FP8 forms, pending a shipped part and a run.

    promotes when → a shipped part and a run

  2. 02
    Intel Architecture Instruction Set Extensions Programming Reference manual

    Ties APX, with doubled x86 registers, AVX10.2 and AMX FP8 tiles to Diamond Rapids, pending silicon and a public run.

    promotes when → silicon and a public run

  3. 03
    SME in a Neoverse core no link

    SME in a Neoverse core, so far shipped only in client parts, pending a server core that carries it and a public run.

    promotes when → a server core that carries it and a public run

  4. 04
    RISC-V Vector Extension, Version 1.0 manual

    The ratified vector ISA, shipped only as IP and chiplets, pending a socketed server part and a run against Arm or x86.

    promotes when → a socketed server part and a run against Arm or x86

Parts without a public measurement

  1. 01
    6th Gen AMD EPYC Server CPUs report

    The Zen 6 server family, so far a press release with no shipped part, pending shipment and a public run against Zen 5.

    promotes when → shipment and a public run against Zen 5

  2. 02
    Intel Xeon 6+ Processors manual

    The E-core-only sockets after Sierra Forest, shipped with vendor multiples footnoted off the page, pending a public run.

    promotes when → a public run

  3. 03
    NVIDIA Vera CPU report

    Custom Arm cores with statically partitioned SMT and no architecture document, pending a specification and a public run.

    promotes when → a specification and a public run

  4. 04
    Arm Neoverse V3 Core Software Optimization Guide manual

    Vendor timing tables for the core shipped in Graviton 5 and previewed in Cobalt 200, pending a public run on the core.

    promotes when → a public run on the core

Memory and interconnect

  1. 01
    CXL Specification manual

    Defines memory pooled across hosts on a coherent link, with latency so far estimated, pending a run on a shipped pool.

    promotes when → a run on a shipped pool

  2. 02
    Demystifying CXL Memory with Genuine CXL-Ready Systems and Devices paper

    Measures expansion devices with frequency and SMT fixed, the nearest run to every field, pending compiler and flags.

    promotes when → compiler and flags

  3. 03
    DDR5MDB02 Multiplexed Rank Data Buffer, JESD82-552 manual

    Defines the data buffer behind MRDIMMs, whose vendor bandwidth claims name no method, pending a run against RDIMMs.

    promotes when → a run against RDIMMs

Kernel paths and generated code

  1. 01
    Extensible Scheduler Class manual

    Lets a BPF program schedule at run time with safe fallback, once a run against the default scheduler states every field.

    promotes when → a run against the default scheduler states every field

  2. 02
    io_uring zero copy Rx manual

    Lands payloads straight in user memory on header-splitting NICs, pending the implementer's epoll run naming every field.

    promotes when → the implementer's epoll run naming every field

  3. 03
    T-MAC repository

    Table-lookup kernels for low-bit weights, with a baseline stated but no frequency, compiler or flags, pending those.

    promotes when → those

  4. 04
    Faster sorting algorithms discovered using deep reinforcement learning paper

    Generated small sorts shipped in libc++, timed by CPU family with no model, compiler or flags stated, pending those.

    promotes when → those

What earns a place

Seven fields, and a number without all of them does not appear here:

Miss one and the number is dropped; if the entry rests on that number, it moves to the watchlist or goes.

An entry itself has to be the thing, not writing about the thing: the paper that first described a mechanism, the specification or manual that defines it, the repository the implementation lives in, or a report from whoever did the work with code and reproducible measurements. Summaries, tutorials, surveys, marketing pages, mirrors and repackagings do not qualify. Every URL points at the live canonical copy, and misc/scripts/check_links.py and misc/scripts/check_format.py prove it on every push and again weekly.

CONTRIBUTING.md has the rules in full.

License

MIT. Maintained by @usamahz.

Format inspired by the GPU-side list at wafer-ai/gpu-perf-engineering-resources.