Learn / Out
Inference on CPU
Read first
A decode step at batch one reads every weight once for a few flops, so it runs at the memory system's rate, and a CPU figure compares with a GPU figure only at the same batch and precision.
- Sources
- 30 entries in 6 parts
- Reproduce it
- 13-sgemm-naive-vs-blas
- Related sections
- §5 Measurement, §7 Single-thread optimisation, §9 Concurrency, §15 Benchmarks, §16 Watchlist
- In the MCP server
cpuperf://section/13
GEMM and BLAS
-
01
Derives the cache blocking every fast CPU GEMM still uses by refining a memory model until the packed kernel falls out.
-
02
Reduces a BLAS to one register-tile micro-kernel per architecture and states which loop owns which level of cache.
-
03
Builds a register-tile kernel step by step and shows where outer-loop unrolling wins and where a vendor BLAS still does.
-
04 OpenBLAS repository
Sets run-time dispatch for a BLAS, one kernel table per microarchitecture chosen as a DYNAMIC_ARCH build loads.
-
05
Fixes the data-type table, packed weight format and fused post-ops a CPU GEMM must expose to an inference graph.
-
06
Fixes the fused attention pattern, f32 accumulation under bf16 inputs and the shapes the fast CPU path accepts.
Reproduce it
naive GEMM, a hand microkernel and the vendor BLAS, plus int8 against float32 dot products.
Reproduce it · 13-sgemm-naive-vs-blas SGEMM: naive loop, microkernel, vendor BLAS A packed, register-blocked microkernel takes a naive loop to the vector unit's ceiling; the vendor BLAS is further ahead only where it owns a matrix unitRuntimes
-
01 oneDNN repository
Generates VNNI and AMX kernels at run time under PyTorch, TensorFlow and OpenVINO, and calls Compute Library on Arm.
-
02 ggml repository
Defines the quantized block formats and the per-ISA dot-product kernels over them that every llama.cpp type rests on.
-
03 llama.cpp repository
Where new quantization types, kernels and thread pools land first, each with the perplexity and speed table behind it.
-
04 ONNX Runtime MLAS repository
Holds the CPU provider's GEMM, int8 and int4 MatMul kernels, dispatched per ISA at run time from SSE to AMX and SME.
-
05 OpenVINO CPU Device manual
States precision defaults per ISA, the int8 path through oneDNN and the streams model that turns cores into throughput.
Quantization
-
01
Defines the affine scale and zero-point scheme with integer accumulation that every int8 CPU runtime still implements.
-
02 Nuances of int8 Computations manual
States where int8 saturates on pre-VNNI x86, the compensation that makes signed by signed work, and what VNNI removes.
-
03 Quantize ONNX models manual
Defines the QDQ form against QOperator and dynamic against static, and which sign choices are safe on which CPUs.
-
04 k-quants repository
Defines super-block formats with quantized scales and the perplexity against size curve behind mixing types per tensor.
-
05
Replaces unpacking and multiplying low-bit weights with bit-sliced lookups, so a narrower weight costs less to apply.
Matrix extensions
-
01
Defines AMX tiles, tile configuration and TMUL, the VNNI and bf16 dot products, and the XSAVE state a tile load needs.
-
02
Defines the arch_prctl permission Linux requires for AMX tile data and the trap a tile instruction takes until granted.
-
03
Holds the design report for the ggml AMX path, whose double buffering hides applying scales between tile products.
-
04
Defines streaming SVE mode, the ZA tile array and SME2 multi-vector instructions, with pseudocode a kernel must match.
-
05 SME Programmer's Guide manual
Walks through SME2 int8 and f32 matmul and gemv kernels and a lookup-table path, on a unit so far only in client parts.
Threading for inference
-
01 Thread management manual
Sets the physical-core default, the affinity it implies and the spin-wait controls that trade idle CPU for latency.
-
02
States the vendor defaults, one thread per core, SMT siblings off, core type by precision and one socket for latency.
-
03 Threadpool: take 2 repository
Defines the explicit thread pool with CPU masks, strict placement, priority and polling that ggml runs without OpenMP.
-
04 llama-bench repository
Defines the prompt and generation tests, repetitions and mean with deviation behind any comparable llama.cpp number.
-
05
Traces poor decode scaling across sockets to remote NUMA access from weight placement, with numatop counts as evidence.
When CPU beats GPU
-
01 gpt-j example README (ggml) repository
Origin of ggml's bandwidth argument, NEON threads saturate a laptop's memory bus, so a GPU on that bus gains nothing.
-
02
Holds the memory clock sweep with all else fixed, the token rate following it and flattening after a few threads.
-
03
Measures decode kernels on a Xeon as DRAM-bound, so AMX pays at batch one only once bytes per token shrink, with code.
-
04 MLPerf Inference v6.0 Results repository
Holds CPU-only entries whose logs show the batched Offline case lost to GPU entries at the same accuracy floor.