Learn / The whole list
The whole list
01Start here
Each entry assumes only the ones before it. The first and sixth are paid books; the third is a manual to open at the chapters its reason names.
-
01
Its pipelining appendix and memory chapters define the hazard, speculation and cache vocabulary the list assumes.
-
02 Optimizing software in C++ manual
Maps C++ onto pipeline mechanisms and shows why a loop-carried dependency chain, not instruction count, paces a loop.
-
03
Its opening chapters show how a shipping x86 core implements the textbook pipeline, each rule tied to a mechanism.
-
04
Measures the step in cost per access at each cache boundary and the gap a prefetcher hides.
-
05
Explains why a second core makes loads and stores reorder and what a barrier drains.
-
06
Puts the method before the tools: what to measure, in what order, and how benchmarks mislead.
-
07
Places a loop from a byte count and a datasheet bandwidth alone, before any counter is read.
-
08
Defines the split of pipeline slots into front end, bad speculation, back end and retiring, the tree profilers report.
-
09
Walks from a noisy timing to counters to a named bottleneck, applying roofline and top-down to whole programs.
-
10
Shows how to read emitted assembly against its source, so each mechanism is checked in a listing, not assumed.
Work the exercises in Performance Ninja alongside them; reading alone will not build the instinct.
Reproduce it: 04-cache-latency, the cost per dependent load stepping up at each cache boundary, the curve the fourth entry measures.
02One instruction, end to end
Vendors name the same structures differently (Intel's decoded ICache is AMD's op cache, and AMD's macro-op is Arm's MOP), so the stage names below are generic.
Fetch and decode
-
01
The origin of the decoupled front end, where the predictor runs ahead of fetch and drives instruction prefetch.
-
02
Introduces the decoded micro-op cache and its power case, the structure both x86 vendors later built.
-
03
Where AMD states when the op cache feeds micro-ops, and the fusion and alignment rules for hot loops.
-
04
Measures rather than quotes each x86 core's misprediction penalty, micro-op cache and loop buffer behaviour.
-
05
States what a microcode fix evicts from the decoded cache and which counters show the fall back to legacy decode.
Reproduce it: 02-branch-misprediction, the cost of a mispredicted branch, sorted against unsorted against branchless.
Rename and issue
-
01
The origin of tag-based renaming and reservation stations, from which every out-of-order issue queue descends.
-
02
Shows which chains renaming cannot break, from partial registers to flags, and the zero-cost idioms that do.
-
03
Sets the user-space method that measures the window and register files, and which idioms take no physical register.
Execute
-
01
Defines the dependence graph of an out-of-order core and the critical path through it that sets run time.
-
02
Measures latency, reciprocal throughput and port assignment per instruction, the weights a dependence graph needs.
-
03
States every stage of one Arm core from fetch to issue, then per-instruction latency, throughput and pipe assignment.
-
04
Predicts a real decoder loop's cycle count from its carried chain and issue slots, then measures the prediction.
-
05
Predicts a decoder's refill chain from Arm core latencies, fits extra work under it, and measures the gain.
Memory access and retire
-
01
The origin of memory dependence prediction, which lets a load pass older stores and flushes on a wrong guess.
-
02
Measures which store and load size and offset pairs forward or stall, and which cores predict memory dependences.
-
03
Defines memory-level parallelism as the misses one window overlaps, and shows which core limits cap it.
-
04
Introduces the reorder buffer and defines a precise exception as in-order commit of out-of-order results.
-
05 Where Do Interrupts Happen? report
Measures that an interrupt lands on the oldest unretired instruction, which decides what a sampling profiler blames.
03Microarchitecture
Where a design paper, the vendor manual and a measurement disagree about a core, the measurement is the one to trust and re-run.
Limits of ILP and SMT
-
01
Shows rename, the active list and precise branch recovery fitted together in one shipped out-of-order core.
-
02
Separates perfect from realistic prediction, renaming and aliasing, and shows how little ILP survives the real ones.
-
03
Puts circuit delay on wakeup, select and bypass, capping issue width and window size before ILP runs out.
-
04
The origin of SMT, where another thread fills one thread's idle issue slots at a cost to both.
-
05
Turns window size, issue width and miss events into cycles lost, making an ILP limit a cycle count.
Branch prediction and speculation
-
01
Defines the tagged geometric-history predictor shipped designs converge on, and the limit of what it can learn.
-
02
Carries the tagged geometric scheme to indirect jumps and calls, where interpreters and virtual dispatch stall.
-
03
Defines the penalty as pipeline refill plus window drain, so it exceeds pipeline depth and is not constant.
-
04
Establishes that predictor state is shared and trainable across contexts, the root of every mitigation and its cost.
-
05
Defines the indirect-branch and store-bypass controls (IBRS, STIBP, IBPB, SSBD) and states which carry a large cost.
Vendor estimates and measured tables
-
01
Calls its own latency spreadsheet an estimate and lists its assumptions, the vendor claim the measured tables check.
-
02
States where measured predictor, op cache and port findings disagree with the vendor's account of an x86 core.
-
03
States why its measured figures differ from the vendor's and names which latencies cannot be measured accurately.
-
04 uops.info report
Where each latency, throughput and port entry links to its microbenchmark, so any value can be re-run.
-
05 applecpu: Firestorm Overview report
Measured per-instruction tables for an Apple AArch64 core, each entry linked to the counter experiment behind it.
What the manuals leave out
-
01 Performance Speed Limits report
Sets the method for finding which hard bound binds a loop, testing its cycles against each in turn.
-
02
Models the predecoder, decoders, micro-op cache and loop buffer closely enough to predict front-end bound loops.
-
03
Measures what the vendor leaves untimed in an AVX-512 licence change, a throttled phase then a frequency step.
-
04 Hardware Store Elimination report
Finds an undocumented optimisation by its counter signature and bandwidth, and sets the method for what manuals omit.
Reproduce it: 03-latency-vs-throughput, one dependency chain against eight independent accumulators.
04Memory hierarchy
Line size and page size are machine parameters, not constants, so every padding and alignment rule below is applied against the target's own values.
Cache geometry, replacement and misses in flight
-
01
One measured account of DRAM timing, cache geometry, TLBs and prefetchers that sets the padding and prefetch rules.
-
02
The origin of the strided-loop method that recovers cache and TLB size, line size, associativity and miss cost.
-
03
Names inclusion victims as the cost of an inclusive last-level cache, the case for non-inclusive and victim designs.
-
04
The origin of set duelling and bimodal insertion, the adaptive replacement that survives a streaming pass.
-
05
Origin of the lockup-free cache, whose miss-status registers keep misses in flight, so a stream beats a chase.
Reproduce it: 04-cache-latency, dependent-load latency from L1 to DRAM, with and without TLB pressure.
TLBs, page walks and prefetchers
-
01
Fixes the paging walk, the page-walk caches, the TLB invalidation rules and what each memory type permits.
-
02
Defines the page-walk cache design space and shows why cached partial translations let a walk skip page-table levels.
-
03
The origin of the stream buffer that every vendor stream prefetcher descends from, and of the victim cache.
-
04
Names each prefetcher and what trains it, and which stop at a page boundary and which cross it.
-
05
Names an Arm server core's load-side and store-side prefetchers, its TLB levels and the bits that disable them.
Store buffers, ordering and cache-line contention
-
01
Derives store buffers and invalidate queues from the cost of coherence, the reason reordering exists at all.
-
02
The store-buffer model of Intel and AMD ordering, tested on hardware, fixing which reorderings a fence pays for.
-
03
The formal model Arm adopted into its architecture, and the record of why the non-multicopy-atomic option was dropped.
-
04
Measures line-transfer cost between cores by coherence state and distance, the figure behind every contention rule.
-
05
The implementers' account of perf c2c, whose worked example reads contended lines, offsets and callers off the report.
Struct layout, software prefetch and page size
-
01
Introduces and measures structure splitting and field reordering on real programs, the origin of hot-cold layout rules.
-
02
Argues from cache-line utilisation arithmetic that layout must follow the access pattern, the case behind structure of arrays.
-
03 dwarves (pahole) repository
Prints a struct's holes, padding and cache-line boundaries from DWARF, so a layout is seen rather than guessed.
-
04
Sorts software prefetch into the cases where it helps and hurts, and shows how it mistrains hardware prefetchers.
-
05 Transparent Hugepage Support manual
The kernel's statement of THP: the enabled and defrag knobs, khugepaged, and the counters showing what was obtained.
-
06
Measures the fault latency, memory bloat and unfairness that eager THP promotion causes, and shows what removes them.
05Measurement
Cloud instances often virtualise the hardware counters away and macOS runs no Linux perf, so perf stat has to show a non-zero cycles count before any counter entry below is trusted.
Method and the whole-system view
-
01 The USE Method report
Sets the checklist that finds the saturated resource before any profiler is opened.
-
02
Draws the line between counting, sampling, instrumentation and tracing, so a question is matched to its tool.
-
03
The reference for time a CPU sampler cannot see, off-CPU, scheduler and I/O waits, traced at bounded cost.
-
04 bpftrace repository
Makes a tracing hypothesis a one-line experiment over kprobes, uprobes, tracepoints and PMU events.
-
05
The design continuous profilers descend from, always-on sampling across a fleet, cheap enough to leave running.
Counters, events and precise sampling
-
01 perf_event_open(2) manual
Defines the counting and sampling modes every Linux profiler uses, and the sample record fields, branch stack included.
-
02
Defines skid, why a sample lands after the culprit, and how tagging one op through the pipeline removes it.
-
03
Defines the architectural counters, what a PEBS sample captures, and why rdtsc counts time, not cycles, under DVFS.
-
04
Defines the event encodings and IBS registers for one Zen core, the tables perf's AMD events are derived from.
-
05 perf-arm-spe(1) repository
Defines Arm SPE as perf drives it, one sampled op in flight, the filters, and what a record holds.
CPU profilers and flame graphs
-
01 Linux perf wiki: Tutorial manual
The maintainers' walk from perf stat to perf record to perf annotate, with the sample fields each flag sets.
-
02
Home of the user guide and cookbook, where each hardware analysis is defined by the events behind it.
-
03 AMD uProf User Guide manual
The vendor's reference for IBS-driven profiling on Zen, with metric presets defined per core generation.
-
04 The Flame Graph paper
Records the design decisions, width as sample share and alphabetical rather than time order, behind its reading rules.
-
05
Proves a hot function need not be worth optimising, by measuring what speeding up a line does to end-to-end time.
Microbenchmarks that lie
-
01
The origin of measurement bias as a term, link order and environment size alone flipping a compiler flag comparison.
-
02
Traces run-to-run variation in x86 retired-instruction counts to one extra count per interrupt and per fault.
-
03 clock_gettime(2) manual
Defines what each clock counts, NTP-slewed or raw monotonic time, or CPU time, and the resolution call.
-
04 Benchmarking tips (LLVM) manual
The compiler project's recipe for a quiet Linux host, governor, boost, SMT siblings, ASLR, a cpuset and tmpfs.
-
05 Google Benchmark User Guide repository
Documents the barriers that keep the optimiser from deleting the work under test, and repetition statistics.
Reproduce it: 05-measurement-pitfalls, dead-code elimination, run-to-run spread, cold against warm.
06Models
A roofline is a bound built from measured roofs and counted bytes, so a point above a roof means a wrong roof or a wrong byte count, not fast code.
Roofline and the execution-cache-memory model
-
01
Defines operational intensity as traffic past the caches, and the ceilings that say which optimisation can pay.
-
02
States the counter set and method that put a measured point on the plot in place of a hand-counted intensity.
-
03 LIKWID repository
Measures the roofs with likwid-bench and the point from likwid-perfctr groups on Intel, AMD and Arm server cores.
-
04
Adds a roof per cache level against traffic at the core, so a kernel that hits in cache is not plotted as DRAM bound.
-
05
Times each level's transfer with overlap rules, predicting a core's rate and the core count where bandwidth saturates.
Reproduce it: 06-roofline, measured roofs and three kernels of rising arithmetic intensity.
Top-down analysis
-
01
Defines the slot accounting that turns raw counters into a weighted tree of bottlenecks, and why the unit is a slot.
-
02 TMA_Metrics-full.xlsx (intel/perfmon) repository
The official home of every TMA formula, event, threshold and level per microarchitecture, from which toplev derives.
-
03 pmu-tools repository
Runs the TMA tree on Linux from the spreadsheet and states why multiplexed levels mislead on short or varied workloads.
-
04 pipeline.json (perf pmu-events, amdzen4) repository
States the AMD top-down formulas in events for both levels, the metric groups perf stat runs on Zen with no vendor tool.
-
05
Defines Arm's top-down as staged stall accounting, first locating the bottleneck and then measuring its resource.
Scaling laws
-
01
States the serial-fraction bound on speedup, the argument every later scaling law is written against.
-
02
Defines scaled speedup, the bound that holds when the problem grows with the processor count instead of staying fixed.
-
03
Extends the bound to chips of unequal cores under a fixed area budget, the arithmetic behind big and little cores.
-
04
The origin of the coherency term that makes throughput fall, not merely flatten, as processors are added.
-
05
Derives the universal scalability law and states the procedure that fits its coefficients to measured throughput.
Queueing
-
01
Proves occupancy equals arrival rate times time in system with no assumption on arrivals or service.
-
02
Origin of the bound-and-bottleneck analysis roofline names as its ancestor, and of the asymptotic bounds on throughput.
-
03
Proves why open and closed systems answer a load question differently, and when scheduling, not capacity, sets latency.
-
04
Origin of the A/S/c notation, and of the embedded chain that solves a queue with non-memoryless arrivals or service.
-
05
Derives the wait near saturation from utilisation and arrival and service variance, the formula behind the latency knee.
07Single-thread optimisation
A loop the compiler reports as vectorised can still run at scalar speed: a float reduction stays one serial chain until reassociation is permitted.
Data layout and loop transforms
-
01
Where Intel states when a structure of arrays beats an array of structures, strided and hybrid cases included.
-
02
Measures one kernel in both layouts and traces the gain to the gathers the array-of-structures form forces.
-
03
Defines interchange, skewing, reversal and tiling as one family and proves when each keeps a loop nest legal.
-
04
Traces the drops in a blocked loop's speed curve to self-interference misses, and shows when copying a tile pays.
-
05 Auto-Vectorization in LLVM manual
Where LLVM lists what its loop and SLP vectorisers accept: interleaving, reductions, if-conversion and runtime checks.
Reproduce it: 07-aos-vs-soa-simd, array of structs against structure of arrays, scalar against NEON.
SIMD instruction sets
-
01 Intel Intrinsics Guide manual
Maps each intrinsic to its instruction and CPUID flag, with the vendor's latency and throughput per microarchitecture.
-
02
The normative semantics of every Intel vector instruction, with the masking, rounding and fault rules intrinsics hide.
-
03 Introduction to SVE manual
Where Arm explains vector-length-agnostic loops and predication, the model that removes remainder loops entirely.
-
04 Arm C Language Extensions manual
The specification the NEON and SVE intrinsics come from, so it settles what a compiler must accept and what is a bug.
-
05
The normative definition of NEON, SVE and SVE2 instructions and of the scalable vector and predicate register model.
SIMD libraries and measured kernels
-
01 Highway repository
Defines sizeless vector types with run-time dispatch, so one source serves SVE and every fixed-width ISA.
-
02 xsimd repository
Fixes a batch type per ISA and width, the compile-time vector length model, and still reaches NEON and SVE.
-
03
Where shuffle-based lookup replacing a byte loop is worked through step by step, with the measurement method spelt out.
-
04
Shows branch-free structural indexing with carry-less multiply and shuffles, measured against conventional parsers.
-
05 simdjson repository
Where that technique ships, with a kernel per ISA and the harness that keeps its published comparisons reproducible.
-
06
Decomposes regexes into string and automaton pieces so both run on SIMD, the design inside the matcher Snort embeds.
Branchless code and bit manipulation
-
01
Counter evidence that current predictors absorb interpreter dispatch, so a branch removal must be measured, not assumed.
-
02
Derives the branch-free integer tricks, division by a constant among them, with proofs rather than as a catalogue.
-
03
A worked branchless merge with code whose gain vanishes when the compiler emits no conditional move.
-
04
Measures branch-free search over sorted, Eytzinger and B-tree layouts, the winner changing with size and prefetch.
-
05
Shows a carry-save adder tree in vector registers beating the dedicated instruction, timed as the minimum of many runs.
08Compilers and codegen
No -O level changes the target instruction set: without -march or -mcpu, every instruction emitted belongs to the default target ISA, so target flags come before any judgement of codegen.
Reading emitted code
-
01 Compiler Explorer repository
Shows how a source change alters the emitted instructions across compilers, versions and flags, with nothing installed.
-
02
Explains how the signed-overflow and aliasing rules let a trip count be known and a store loop become memset.
-
03 llvm-objdump manual
Reads the binary that shipped, with source lines and symbolised branch targets, rather than a recompiled snippet.
-
04 llvm-mca manual
Predicts loop throughput and port pressure from the scheduling model, and states it models neither front end nor caches.
-
05 llvm-exegesis manual
Measures instruction latency and throughput with counters, so the model llvm-mca predicts from is checked, not trusted.
Optimisation levels, inlining and link time
-
01
Lists what each -O level turns on, the inlining limits, and that -Ofast admits transforms invalid for conforming code.
-
02
Shows with real codegen that an abstraction is free only when inlining and the ABI allow it, and the cost when either refuses.
-
03 Itanium C++ ABI manual
Fixes the rule that a non-trivial class goes by reference to a caller-made temporary, the cost a wrapped pointer pays.
-
04
States what PLT calls and interposition cost, and the visibility controls a library needs to inline its own exports.
-
05 LTO Overview (GCC Internals) manual
Defines whole-program LTO against partitioned WHOPR, and the LGEN, WPA and LTRANS stages that run -flto in parallel.
-
06 ThinLTO manual
Defines the thin link, summaries analysed whole-program then parallel backends, and the cache for incremental rebuilds.
Target flags and auto-vectorisation
-
01 x86 Options (GCC) manual
Defines -march against -mtune, the psABI levels, and -mprefer-vector-width, the switch for full-width AVX-512 code.
-
02
Defines target_clones, one function per ISA behind a resolver the dynamic linker runs, so a generic build ships AVX-512.
-
03
Lists what -ffast-math implies, of which -fassociative-math alone frees a float reduction, and -ffp-contract for FMA.
-
04 Auto-Vectorization in LLVM manual
States what the vectorisers need, aliasing disproved or checked at run time, and where a float reduction stays in order.
-
05
Defines the remarks that make the compiler say which loop it left scalar and why, so the fix targets the real blocker.
Reproduce it: 08-autovectorization-aliasing, the vectoriser with and without restrict.
Profile-guided and post-link optimisation
-
01
Defines the instrumented and sampled workflows, why their profiles cannot mix, and the cost of a wrong training input.
-
02
Defines the address-to-source mapping with discriminators that lets a stale production profile still drive FDO.
-
03
States why a profile applied to the final binary beats one mapped to source, and the layout passes accuracy enables.
-
04 BOLT (llvm-project/bolt) repository
States what full effect needs, relocations kept at link time and a branch-stack sample profile, neither on by default.
-
05
States the design, not a result: basic-block sections and a relink, no binary rewrite, so layout needs no disassembly.
09Concurrency
Every cost below is a cache line moving between cores, so the ordering models and the measured line-transfer cost in the memory hierarchy section come first.
Memory models and atomics
-
01
Defines the data-race-free contract, sequential consistency for race-free programs and no meaning for a race.
-
02
The normative wording for every memory order, fence and read-modify-write, the text a compiler is checked against.
-
03
The table that turns each memory order into x86 and Arm instructions, so what an order costs is read off the page.
-
04
States what the kernel assumes any CPU may reorder and what each barrier and access primitive guarantees.
-
05 herdtools7 repository
Where herd7, litmus7 and klitmus7 live, the tools that run a litmus test against the x86, Arm and kernel models.
Locks, contention and allocators
-
01
Derives counting, partitioning, locking and deferral with code that runs, the textbook the section assumes.
-
02
The origin of the queue lock, each waiter spinning on its own line, measured against ticket and test-and-set locks.
-
03 Futexes Are Tricky paper
Derives a correct user-space mutex from futex and shows the lost wakeups and extra kernel entries naive versions pay.
-
04
Defines blowup and allocator-induced false sharing, which per-processor heaps under a bounded global heap avoid.
-
05
The design statement for per-CPU caches built on restartable sequences and a hugepage-aware back end for TLB reach.
Reproduce it: 09-false-sharing, adjacent counters against padded counters across threads.
Lock-free structures and RCU
-
01
States linearizability and the consensus hierarchy and builds both into working stacks, queues, lists and hash tables.
-
02
The lock-free queue later libraries copy, with the counted pointer against ABA and a two-lock queue beside it.
-
03
The standard-track form of safe reclamation, fixing when a retired node may be freed while a reader still holds it.
-
04 What is RCU? manual
The kernel's own statement of RCU as publish, wait for readers and keep old versions, with a free read side.
-
05
Defines liburcu's quiescent-state, signal-based and general RCU flavours and measures each read side against locks.
Thread pools and work stealing
-
01
Defines work and critical path, proves the work-stealing bound, and shows they alone predict a runtime's speedup.
-
02
States the work-first principle, that overhead belongs on the rare steal path and not on every spawn.
-
03
Gives the work-stealing deque a proven atomics form and derives which fence push, take and steal need on Arm and x86.
-
04 OpenMP Specifications manual
Fixes fork-join and tasking semantics, and the wait and binding controls deciding if idle workers spin, sleep or move.
-
05 oneTBB repository
The shipping work-stealing runtime, arenas and task groups over a deque, the home of grain size and spin-before-sleep.
10NUMA and multi-socket
A page's node is decided at first touch, not when memory is allocated or a policy is set, and every vendor table below depends on the BIOS node mode of the machine it ran on.
NUMA and Linux memory placement
-
01
The one account tying first touch, policy scope, zone reclaim and page movement together from the implementer's side.
-
02 What is NUMA? manual
Defines nodes, zonelists and the distance-ordered fallback that places an allocation once local memory runs out.
-
03 NUMA Memory Policy manual
The normative statement of policy scopes, every mode including weighted interleave, and the cpuset intersection rule.
-
04
Defines numa_hit, numa_miss and numa_foreign, the counters that show whether a policy put pages where it said.
-
05 numactl repository
Reference implementation of the policy API, prints the distance table and binds a binary that cannot be rebuilt.
Reproduce it: 10-first-touch, first touch of fresh pages against the second pass.
Topology and interconnects
-
01 NUMA Memory Performance manual
Explains the firmware-rated latency and bandwidth per initiator and target, and memory-side caches, that rank nodes.
-
02
Where Intel names the mesh, UPI socket links, the directory-running home agent, and how SNC splits the cache.
-
03
Defines SNC on current parts as one node per compute die, and fixes the numactl and numastat checks of placement.
-
04
Discloses the I/O die, GMI and xGMI links, the NPS modes with their interleave widths, and the cache-as-NUMA override.
-
05
Defines the mesh, the home nodes holding the system cache and snoop filter, and the gateways joining sockets or CXL.
Migration, balancing and measured effects
-
01 move_pages(2) manual
Defines per-page migration of a running process, and a query reporting each page's node, the direct test of first touch.
-
02 sysctl kernel numa_balancing manual
Defines the hinting-fault sampling behind automatic balancing and tiering, and warns the overhead may not pay off.
-
03
Proves against the kernel balancer that controller and link congestion, not remote latency, is what placement manages.
-
04 Intel Memory Latency Checker manual
Measures the node-to-node latency and bandwidth matrix and loaded latency on the x86 at hand, which no datasheet states.
-
05
Tabulates measured bandwidth by NPS mode, cores per die, boost and SMT, so the NPS trade-off is shown, not asserted.
11OS and I/O
A syscall's cost depends on the mitigation state, the governor and the idle state the core was in, three sysfs settings that change after boot, so each is recorded beside any number below.
Syscalls and asynchronous I/O
-
01 vdso(7) manual
Defines the calls the kernel answers without a mode switch, and the clocksource condition for skipping the trap.
-
02
Separates a syscall's trap cost from its cache and TLB pollution, and shows the pollution can dominate.
-
03
Measures syscall and context switch cost across kernel releases and traces each slowdown to a named mitigation.
-
04
States the goals aio failed, and which io_uring features remove a syscall and which remove a copy.
-
05
Measures io_uring's polling modes against libaio and SPDK, and shows the kernel poller needs its own core.
Reproduce it: 11-syscall-cost, the fixed cost of a kernel crossing across request sizes.
Scheduling, affinity and isolation
-
01 EEVDF Scheduler manual
Defines lag and virtual deadline, which the default class schedules by, and the slice request in sched_setattr.
-
02
Proves cores sit idle while runnable threads queue, and gives the invariant checker that found the load-balancer bugs.
-
03 Control Group v2 manual
Defines cpu.max throttling, cpu.weight and the cpusets that bound affinity, the controls behind every container limit.
-
04 CPU Performance Scaling manual
Defines the governors, driver and boost switch that set a core's frequency, the sysfs state a measurement records.
-
05 CPU Isolation manual
Ties isolcpus, nohz_full, IRQ affinity, RCU offload and cpusets into one recipe, and lists the jitter it leaves.
Interrupts and kernel bypass
-
01 NAPI manual
Defines the polling, software coalescing, busy polling and IRQ suspension knobs that trade interrupts against latency.
-
02 DPDK Programmer's Guide manual
Defines the full bypass model, pinned poll-mode cores with no interrupts, that every kernel path is measured against.
-
03 The eXpress Data Path repository
Measures an in-kernel programmable path against DPDK and the stack per core, with the full configuration published.
-
04
Separates direct and indirect NIC interrupt cost, measures the stack against bypass, and is where IRQ suspension began.
-
05 AF_XDP manual
Defines the socket and UMEM rings handing XDP frames to user space, and the zero-copy and need-wakeup modes.
Cache and bandwidth partitioning
-
01
Defines classes of service, cache masks, bandwidth allocation and monitoring IDs, the model resctrl exposes.
-
02
Defines the filesystem through which Linux exposes Intel, AMD and Arm partitioning, and the schemata format.
-
03 MPAM manual
Maps Arm's cache portion and bandwidth controls onto resctrl's schemata, and states which platform limits apply.
-
04
Shows at fleet scale that cycles per instruction alone finds an interfering neighbour and the one to throttle.
-
05
Shows cache ways, cores, bandwidth and power must be partitioned together, or batch work reaches the tail.
12Tail latency and production systems
A latency figure means nothing without its percentile, its load model and the way it was recorded.
Measuring the tail
-
01 The Tail at Scale paper
Shows why fan-out makes a rare slow server a common slow request, and names the techniques that tolerate variance.
-
02
Defines the stall band that out-of-order hardware cannot hide and a context switch cannot amortise.
-
03
Shows that a summary without a max discards the samples that define the tail, and closed-loop load never records them.
-
04 Coordinated Omission report
The original definition of the recording error, with arithmetic for how far a reported percentile sits from the truth.
-
05 HdrHistogram repository
Keeps the whole distribution at fixed relative precision in constant time, so the far percentiles and max survive.
Reproduce it: 12-coordinated-omission, closed-loop against open-loop p99 under the same stalls.
Where jitter comes from
-
01 rt-tests repository
The reference wakeup-latency measurement for Linux, whose README states that an unloaded run proves nothing.
-
02 osnoise tracer manual
Counts the noise a spinning thread suffers and attributes each event to NMI, IRQ, softirq, thread or hardware.
-
03 Tales of the Tail paper
Derives the queueing-ideal tail and attributes the excess to scheduling, interrupt placement, power saving and NUMA.
-
04
Measures with code the page-fault, TLB-shootdown and writeback stalls that memory mapping hides from the caller.
-
05 The KVM halt polling system manual
Defines the host-side polling after a vCPU halt that trades idle host CPU for guest wakeup time, unseen by the guest.
Load generation and production workloads
-
01
Shows that open and closed load models disagree on response time and scheduling gains, with rules for choosing one.
-
02 wrk2 repository
Issues requests on a fixed schedule and times each from when it was due, so server stalls reach the percentiles.
-
03
Shows the tail, not throughput, caps a latency-critical server's utilisation, and how far co-located work lowers it.
-
04 TailBench report
Pairs latency-critical services with an open-loop harness that records sojourn against service time per request.
-
05
Measures the key, value and inter-arrival distributions of live key-value traffic, the shape load generators imitate.
Mechanical sympathy
-
01 Inter Thread Latency report
Measures with code the floor for handing a cache line between cores, which every queue and lock is built on.
-
02 Single Writer Principle report
States the design rule that removes write contention outright, using a contended increment's cost as the argument.
-
03
Adds cached indices to a single-producer single-consumer ring and shows with counters the coherence traffic removed.
-
04 LMAX Disruptor manual
Applies the single writer rule and cache-line padding to a ring buffer, with the queue comparison that motivated it.
-
05 Aeron repository
Carries the single writer and batching rules through a whole transport, the reference beyond one in-process queue.
13Inference on CPU
A decode step at batch one reads every weight once for a few flops, so it runs at the memory system's rate, and a CPU figure compares with a GPU figure only at the same batch and precision.
GEMM and BLAS
-
01
Derives the cache blocking every fast CPU GEMM still uses by refining a memory model until the packed kernel falls out.
-
02
Reduces a BLAS to one register-tile micro-kernel per architecture and states which loop owns which level of cache.
-
03
Builds a register-tile kernel step by step and shows where outer-loop unrolling wins and where a vendor BLAS still does.
-
04 OpenBLAS repository
Sets run-time dispatch for a BLAS, one kernel table per microarchitecture chosen as a DYNAMIC_ARCH build loads.
-
05
Fixes the data-type table, packed weight format and fused post-ops a CPU GEMM must expose to an inference graph.
-
06
Fixes the fused attention pattern, f32 accumulation under bf16 inputs and the shapes the fast CPU path accepts.
Reproduce it: 13-sgemm-naive-vs-blas, naive GEMM, a hand microkernel and the vendor BLAS, plus int8 against float32 dot products.
Runtimes
-
01 oneDNN repository
Generates VNNI and AMX kernels at run time under PyTorch, TensorFlow and OpenVINO, and calls Compute Library on Arm.
-
02 ggml repository
Defines the quantized block formats and the per-ISA dot-product kernels over them that every llama.cpp type rests on.
-
03 llama.cpp repository
Where new quantization types, kernels and thread pools land first, each with the perplexity and speed table behind it.
-
04 ONNX Runtime MLAS repository
Holds the CPU provider's GEMM, int8 and int4 MatMul kernels, dispatched per ISA at run time from SSE to AMX and SME.
-
05 OpenVINO CPU Device manual
States precision defaults per ISA, the int8 path through oneDNN and the streams model that turns cores into throughput.
Quantization
-
01
Defines the affine scale and zero-point scheme with integer accumulation that every int8 CPU runtime still implements.
-
02 Nuances of int8 Computations manual
States where int8 saturates on pre-VNNI x86, the compensation that makes signed by signed work, and what VNNI removes.
-
03 Quantize ONNX models manual
Defines the QDQ form against QOperator and dynamic against static, and which sign choices are safe on which CPUs.
-
04 k-quants repository
Defines super-block formats with quantized scales and the perplexity against size curve behind mixing types per tensor.
-
05
Replaces unpacking and multiplying low-bit weights with bit-sliced lookups, so a narrower weight costs less to apply.
Matrix extensions
-
01
Defines AMX tiles, tile configuration and TMUL, the VNNI and bf16 dot products, and the XSAVE state a tile load needs.
-
02
Defines the arch_prctl permission Linux requires for AMX tile data and the trap a tile instruction takes until granted.
-
03
Holds the design report for the ggml AMX path, whose double buffering hides applying scales between tile products.
-
04
Defines streaming SVE mode, the ZA tile array and SME2 multi-vector instructions, with pseudocode a kernel must match.
-
05 SME Programmer's Guide manual
Walks through SME2 int8 and f32 matmul and gemv kernels and a lookup-table path, on a unit so far only in client parts.
Threading for inference
-
01 Thread management manual
Sets the physical-core default, the affinity it implies and the spin-wait controls that trade idle CPU for latency.
-
02
States the vendor defaults, one thread per core, SMT siblings off, core type by precision and one socket for latency.
-
03 Threadpool: take 2 repository
Defines the explicit thread pool with CPU masks, strict placement, priority and polling that ggml runs without OpenMP.
-
04 llama-bench repository
Defines the prompt and generation tests, repetitions and mean with deviation behind any comparable llama.cpp number.
-
05
Traces poor decode scaling across sockets to remote NUMA access from weight placement, with numatop counts as evidence.
When CPU beats GPU
-
01 gpt-j example README (ggml) repository
Origin of ggml's bandwidth argument, NEON threads saturate a laptop's memory bus, so a GPU on that bus gains nothing.
-
02
Holds the memory clock sweep with all else fixed, the token rate following it and flattening after a few threads.
-
03
Measures decode kernels on a Xeon as DRAM-bound, so AMX pays at batch one only once bytes per token shrink, with code.
-
04 MLPerf Inference v6.0 Results repository
Holds CPU-only entries whose logs show the batched Offline case lost to GPU entries at the same accuracy floor.
14Hardware generations
The measurement articles below state no compiler, flags or run count, so each is kept for the structure it exposes and no figure from it is repeated.
Intel Xeon
-
01
Where Intel states what Sapphire Rapids added: larger L2 and L3, DDR5, CXL, AMX and on-die accelerators.
-
02
Measures L3 and memory latency across the tiled mesh and the slow clock ramp the vendor overview omits.
-
03
The designers' statement of what changed: fewer, larger dies, a bigger shared L3, faster DDR5 and socket links.
-
04
Measures per-die L3 under sub-NUMA clustering and the die-crossing cost on Granite Rapids beside Turin.
-
05
Bandwidth-bound codes on Sapphire, Emerald and Granite Rapids and Sierra Forest with clocks, SMT and compiler stated.
AMD EPYC
-
01
The designers' account of the Zen 4 core and how it yields Genoa, Genoa-X, Bergamo and Siena.
-
02
Tests the same-core claim for Zen 4c: cache latency, clock ceiling and core-to-core paths beside clock-matched Zen 4.
-
03
Where AMD states what Zen 5 changed in front end, vector datapath and caches, the core Turin carries.
-
04
Defines Turin: Zen 5 or Zen 5c dies, which parts double die-to-IO links, NUMA modes and full-width AVX-512.
-
05
Measures what wider die-to-IO links and faster DDR5 do for Turin bandwidth, and where latency rose over Genoa.
Arm Neoverse server parts
-
01 AWS Graviton Getting Started repository
Where AWS states which Neoverse core, ISA revision, mesh, caches and compiler flag each Graviton generation carries.
-
02
Sets the pipeline widths and instruction timings of the core Graviton 4, Grace and Axion share.
-
03
States the coherency fabric, LPDDR5X fit and MPAM cache and memory partitioning on Grace.
-
04
States the timings and fusion rules of the narrower N line core in Cobalt 100 and Yitian.
-
05
Where Ampere states the N1 part: private L2 per core, shared system cache, mesh and DDR4 fit.
Independent measurement across vendors
-
01
Measures Neoverse V2, Golden Cove and Zen 4 in-core at fixed clock, and each socket's clock under vector load.
-
02
One SVE workload run on Graviton 3, Graviton 4, Yitian and Axion with compiler, runs and spread stated.
-
03
Measures a sustained rename width below the stated one, cache latencies, mesh behaviour and cross-socket cost on V2.
-
04
Measures structure sizes, cache latencies and mesh behaviour of N2 on Yitian, the core Cobalt 100 carries.
-
05
Puts vendor slides beside measurements of the predictor, small instruction cache, private L2 and long memory latency.
Reproduce it: 14-pcore-vs-ecore, the same three kernels on a performance core and an efficiency core.
15Benchmarks
A score means what its suite's run rules say it means, so the rules come before the number.
Standard suites
-
01
Defines base against peak, rate against speed, the threading models a speed run may use and an Arm reference machine.
-
02
Where the committee states how workloads were chosen and hardened, and defines the rolling round-robin rate.
-
03
Measures with counters what each workload stresses on x86 and Arm server parts, beside data-centre and inference suites.
-
04 MLPerf Inference Rules repository
Fixes model, accuracy floor and query pattern per scenario, so a Server score is throughput under a latency bound.
-
05
Shows standard suites misproject data-centre servers, and states the fleet-matching method the suite is built by.
Microbenchmark suites
-
01
Defines sustainable bandwidth as what unit-stride loops get, not bus peak, and machine balance as flops per access.
-
02
Sets the array size rule, timing over repeated trials, and counting bytes a loop asks for, not what the cache moved.
-
03
Origin of the one-mechanism-per-test method for memory, system call, pipe and socket latency, and what each leaves out.
-
04 uarch-bench repository
Isolates memory-level parallelism from load latency as separate tests, with DVFS held off before timing, x86 Linux only.
-
05
Shows why kernel mode with interrupts off matters, removes harness overhead, then recovers cache replacement policies.
Reproduce it: 15-stream-bandwidth, triad bandwidth by thread count against the vendor figure.
Methodology and what suites miss
-
01
Origin of the rule that normalised results take the geometric mean, which a SPEC ratio and the crimes list rest on.
-
02 Systems Benchmarking Crimes report
Checklist of evaluation faults from sub-setting and improper baselines to arithmetic means of ratios, each with a fix.
-
03
Sets which mean fits costs, rates and ratios, when confidence intervals are owed, and the absolute base a speedup needs.
-
04
Decides how many builds, runs and iterations an experiment needs by measuring at which level the variation arises.
-
05
Fleet counter profile showing services stall on instruction fetch and burn cycles in shared routines, which SPEC lacks.
16Watchlist
Everything below is real but unproven: no item yet has all three of a written specification, a part you can buy, and a public measurement stating every one of the seven fields. Each line says what would promote it. Vendor multiples never qualify. Last checked 2026-09-15.
ISA extensions without a shipped server part
-
01
Folds the AVX-512 subsets into one versioned level and adds BF16 and FP8 forms, pending a shipped part and a run.
promotes when → a shipped part and a run
-
02
Ties APX, with doubled x86 registers, AVX10.2 and AMX FP8 tiles to Diamond Rapids, pending silicon and a public run.
promotes when → silicon and a public run
-
03 SME in a Neoverse core no link
SME in a Neoverse core, so far shipped only in client parts, pending a server core that carries it and a public run.
promotes when → a server core that carries it and a public run
-
04
The ratified vector ISA, shipped only as IP and chiplets, pending a socketed server part and a run against Arm or x86.
promotes when → a socketed server part and a run against Arm or x86
Parts without a public measurement
-
01 6th Gen AMD EPYC Server CPUs report
The Zen 6 server family, so far a press release with no shipped part, pending shipment and a public run against Zen 5.
promotes when → shipment and a public run against Zen 5
-
02 Intel Xeon 6+ Processors manual
The E-core-only sockets after Sierra Forest, shipped with vendor multiples footnoted off the page, pending a public run.
promotes when → a public run
-
03 NVIDIA Vera CPU report
Custom Arm cores with statically partitioned SMT and no architecture document, pending a specification and a public run.
promotes when → a specification and a public run
-
04
Vendor timing tables for the core shipped in Graviton 5 and previewed in Cobalt 200, pending a public run on the core.
promotes when → a public run on the core
Memory and interconnect
-
01 CXL Specification manual
Defines memory pooled across hosts on a coherent link, with latency so far estimated, pending a run on a shipped pool.
promotes when → a run on a shipped pool
-
02
Measures expansion devices with frequency and SMT fixed, the nearest run to every field, pending compiler and flags.
promotes when → compiler and flags
-
03
Defines the data buffer behind MRDIMMs, whose vendor bandwidth claims name no method, pending a run against RDIMMs.
promotes when → a run against RDIMMs
Kernel paths and generated code
-
01 Extensible Scheduler Class manual
Lets a BPF program schedule at run time with safe fallback, once a run against the default scheduler states every field.
promotes when → a run against the default scheduler states every field
-
02 io_uring zero copy Rx manual
Lands payloads straight in user memory on header-splitting NICs, pending the implementer's epoll run naming every field.
promotes when → the implementer's epoll run naming every field
-
03 T-MAC repository
Table-lookup kernels for low-bit weights, with a baseline stated but no frequency, compiler or flags, pending those.
promotes when → those
-
04
Generated small sorts shipped in libc++, timed by CPU family with no model, compiler or flags stated, pending those.
promotes when → those
What earns a place
Seven fields, and a number without all of them does not appear here:
Miss one and the number is dropped; if the entry rests on that number, it moves to the watchlist or goes.
An entry itself has to be the thing, not writing about the thing: the paper
that first described a mechanism, the specification or manual that defines
it, the repository the implementation lives in, or a report from whoever did
the work with code and reproducible measurements. Summaries, tutorials,
surveys, marketing pages, mirrors and repackagings do not qualify. Every URL
points at the live canonical copy, and misc/scripts/check_links.py and
misc/scripts/check_format.py prove it on every push and again weekly.
CONTRIBUTING.md has the rules in full.
License
MIT. Maintained by @usamahz.
Format inspired by the GPU-side list at wafer-ai/gpu-perf-engineering-resources.