---
title: "From one instruction to a model served on CPU."
url: https://cpuperf.com/learn/
commit: deb5a0bac46760503b6f4a2608bdfed470c8532e
---

# From one instruction to a model served on CPU.

## Start here

Ten numbered items, read top to bottom, each assuming only the ones before it.

- [1. Start here](https://cpuperf.com/learn/start-here.md): Each entry assumes only the ones before it. The first and sixth are paid books; the third is a manual to open at the chapters its reason names.

## Down

The core, the memory hierarchy, measurement, models.

- [2. One instruction, end to end](https://cpuperf.com/learn/one-instruction-end-to-end.md): Vendors name the same structures differently (Intel's decoded ICache is AMD's op cache, and AMD's macro-op is Arm's MOP), so the stage names below are generic.
- [3. Microarchitecture](https://cpuperf.com/learn/microarchitecture.md): Where a design paper, the vendor manual and a measurement disagree about a core, the measurement is the one to trust and re-run.
- [4. Memory hierarchy](https://cpuperf.com/learn/memory-hierarchy.md): Line size and page size are machine parameters, not constants, so every padding and alignment rule below is applied against the target's own values.
- [5. Measurement](https://cpuperf.com/learn/measurement.md): Cloud instances often virtualise the hardware counters away and macOS runs no Linux perf, so perf stat has to show a non-zero cycles count before any counter entry below is trusted.
- [6. Models](https://cpuperf.com/learn/models.md): A roofline is a bound built from measured roofs and counted bytes, so a point above a roof means a wrong roof or a wrong byte count, not fast code.

## Out

One thread, the compiler, many threads, NUMA, the kernel boundary, tail latency, CPU inference, the parts themselves, the benchmark suites.

- [7. Single-thread optimisation](https://cpuperf.com/learn/single-thread-optimisation.md): A loop the compiler reports as vectorised can still run at scalar speed: a float reduction stays one serial chain until reassociation is permitted.
- [8. Compilers and codegen](https://cpuperf.com/learn/compilers-and-codegen.md): No -O level changes the target instruction set: without -march or -mcpu, every instruction emitted belongs to the default target ISA, so target flags come before any judgement of codegen.
- [9. Concurrency](https://cpuperf.com/learn/concurrency.md): Every cost below is a cache line moving between cores, so the ordering models and the measured line-transfer cost in the memory hierarchy section come first.
- [10. NUMA and multi-socket](https://cpuperf.com/learn/numa-and-multi-socket.md): A page's node is decided at first touch, not when memory is allocated or a policy is set, and every vendor table below depends on the BIOS node mode of the machine it ran on.
- [11. OS and I/O](https://cpuperf.com/learn/os-and-io.md): A syscall's cost depends on the mitigation state, the governor and the idle state the core was in, three sysfs settings that change after boot, so each is recorded beside any number below.
- [12. Tail latency and production systems](https://cpuperf.com/learn/tail-latency-and-production-systems.md): A latency figure means nothing without its percentile, its load model and the way it was recorded.
- [13. Inference on CPU](https://cpuperf.com/learn/inference-on-cpu.md): A decode step at batch one reads every weight once for a few flops, so it runs at the memory system's rate, and a CPU figure compares with a GPU figure only at the same batch and precision.
- [14. Hardware generations](https://cpuperf.com/learn/hardware-generations.md): The measurement articles below state no compiler, flags or run count, so each is kept for the structure it exposes and no figure from it is repeated.
- [15. Benchmarks](https://cpuperf.com/learn/benchmarks.md): A score means what its suite's run rules say it means, so the rules come before the number.

## Watchlist

Dated, for things whose evidence is still moving.

- [16. Watchlist](https://cpuperf.com/watchlist.md): Everything below is real but unproven: no item yet has all three of a written specification, a part you can buy, and a public measurement stating every one of the seven fields. Each line says what would promote it. Vendor multiples never qualify. Last checked 2026-09-15.
