Start here

§1

Ten numbered items, read top to bottom, each assuming only the ones before it.

  1. 01

    Start here

    Each entry assumes only the ones before it. The first and sixth are paid books; the third is a manual to open at the chapters its reason names.

    10 sources
    1 benchmark

Down

§2–6

The core, the memory hierarchy, measurement, models.

  1. 02

    One instruction, end to end

    Vendors name the same structures differently (Intel's decoded ICache is AMD's op cache, and AMD's macro-op is Arm's MOP), so the stage names below are generic.

    18 sources · 4 parts
    1 benchmark
  2. 03

    Microarchitecture

    Where a design paper, the vendor manual and a measurement disagree about a core, the measurement is the one to trust and re-run.

    19 sources · 4 parts
    1 benchmark
  3. 04

    Memory hierarchy

    Line size and page size are machine parameters, not constants, so every padding and alignment rule below is applied against the target's own values.

    21 sources · 4 parts
    1 benchmark
  4. 05

    Measurement

    Cloud instances often virtualise the hardware counters away and macOS runs no Linux perf, so perf stat has to show a non-zero cycles count before any counter entry below is trusted.

    20 sources · 4 parts
    1 benchmark
  5. 06

    Models

    A roofline is a bound built from measured roofs and counted bytes, so a point above a roof means a wrong roof or a wrong byte count, not fast code.

    20 sources · 4 parts
    1 benchmark

Out

§7–15

One thread, the compiler, many threads, NUMA, the kernel boundary, tail latency, CPU inference, the parts themselves, the benchmark suites.

  1. 07

    Single-thread optimisation

    A loop the compiler reports as vectorised can still run at scalar speed: a float reduction stays one serial chain until reassociation is permitted.

    21 sources · 4 parts
    1 benchmark
  2. 08

    Compilers and codegen

    No -O level changes the target instruction set: without -march or -mcpu, every instruction emitted belongs to the default target ISA, so target flags come before any judgement of codegen.

    21 sources · 4 parts
    1 benchmark
  3. 09

    Concurrency

    Every cost below is a cache line moving between cores, so the ordering models and the measured line-transfer cost in the memory hierarchy section come first.

    20 sources · 4 parts
    1 benchmark
  4. 10

    NUMA and multi-socket

    A page's node is decided at first touch, not when memory is allocated or a policy is set, and every vendor table below depends on the BIOS node mode of the machine it ran on.

    15 sources · 3 parts
    1 benchmark
  5. 11

    OS and I/O

    A syscall's cost depends on the mitigation state, the governor and the idle state the core was in, three sysfs settings that change after boot, so each is recorded beside any number below.

    20 sources · 4 parts
    1 benchmark
  6. 12

    Tail latency and production systems

    A latency figure means nothing without its percentile, its load model and the way it was recorded.

    20 sources · 4 parts
    1 benchmark
  7. 13

    Inference on CPU

    A decode step at batch one reads every weight once for a few flops, so it runs at the memory system's rate, and a CPU figure compares with a GPU figure only at the same batch and precision.

    30 sources · 6 parts
    1 benchmark
  8. 14

    Hardware generations

    The measurement articles below state no compiler, flags or run count, so each is kept for the structure it exposes and no figure from it is repeated.

    20 sources · 4 parts
    1 benchmark
  9. 15

    Benchmarks

    A score means what its suite's run rules say it means, so the rules come before the number.

    15 sources · 3 parts
    1 benchmark

Watchlist

§16

Dated, for things whose evidence is still moving.

  1. 16

    Watchlist

    Everything below is real but unproven: no item yet has all three of a written specification, a part you can buy, and a public measurement stating every one of the seven fields. Each line says what would promote it. Vendor multiples never qualify. Last checked 2026-09-15.

    15 sources · 4 parts

The whole list on one page →