---
title: "13. Inference on CPU"
url: https://cpuperf.com/learn/inference-on-cpu/
source: https://github.com/usamahz/cpu-performance-engineering/blob/deb5a0bac46760503b6f4a2608bdfed470c8532e/README.md#L559
commit: deb5a0bac46760503b6f4a2608bdfed470c8532e
---
## 13. Inference on CPU

A decode step at batch one reads every weight once for a few flops, so it runs at the memory system's rate, and a CPU figure compares with a GPU figure only at the same batch and precision.

### GEMM and BLAS

- [Anatomy of High-Performance Matrix Multiplication](https://www.cs.utexas.edu/~flame/pubs/GotoTOMS_revision.pdf) - Derives the cache blocking every fast CPU GEMM still uses by refining a memory model until the packed kernel falls out.
- [BLIS: A Framework for Rapidly Instantiating BLAS Functionality](https://www.cs.utexas.edu/~flame/pubs/blis1_toms_rev3.pdf) - Reduces a BLAS to one register-tile micro-kernel per architecture and states which loop owns which level of cache.
- [LLaMA Now Goes Faster on CPUs](http://justine.lol/matmul/) - Builds a register-tile kernel step by step and shows where outer-loop unrolling wins and where a vendor BLAS still does.
- [OpenBLAS](https://github.com/OpenMathLib/OpenBLAS) - Sets run-time dispatch for a BLAS, one kernel table per microarchitecture chosen as a DYNAMIC_ARCH build loads.
- [oneDNN Matrix Multiplication Primitive](https://uxlfoundation.github.io/oneDNN/dev_guide_matmul.html) - Fixes the data-type table, packed weight format and fused post-ops a CPU GEMM must expose to an inference graph.
- [Scaled Dot-Product Attention (oneDNN Graph)](https://uxlfoundation.github.io/oneDNN/dev_guide_graph_sdpa.html) - Fixes the fused attention pattern, f32 accumulation under bf16 inputs and the shapes the fast CPU path accepts.

Reproduce it: [misc/benchmarks/13-sgemm-naive-vs-blas](https://cpuperf.com/benchmarks/13-sgemm-naive-vs-blas/), naive GEMM, a hand microkernel and the vendor BLAS, plus int8 against float32 dot products.

### Runtimes

- [oneDNN](https://github.com/uxlfoundation/oneDNN) - Generates VNNI and AMX kernels at run time under PyTorch, TensorFlow and OpenVINO, and calls Compute Library on Arm.
- [ggml](https://github.com/ggml-org/ggml) - Defines the quantized block formats and the per-ISA dot-product kernels over them that every llama.cpp type rests on.
- [llama.cpp](https://github.com/ggml-org/llama.cpp) - Where new quantization types, kernels and thread pools land first, each with the perplexity and speed table behind it.
- [ONNX Runtime MLAS](https://github.com/microsoft/onnxruntime/tree/main/onnxruntime/core/mlas) - Holds the CPU provider's GEMM, int8 and int4 MatMul kernels, dispatched per ISA at run time from SSE to AMX and SME.
- [OpenVINO CPU Device](https://docs.openvino.ai/2026/openvino-workflow/running-inference/inference-devices-and-modes/cpu-device.html) - States precision defaults per ISA, the int8 path through oneDNN and the streams model that turns cores into throughput.

### Quantization

- [Quantization and Training of Neural Networks for Efficient Integer-Arithmetic-Only Inference](https://arxiv.org/abs/1712.05877) - Defines the affine scale and zero-point scheme with integer accumulation that every int8 CPU runtime still implements.
- [Nuances of int8 Computations](https://uxlfoundation.github.io/oneDNN/dev_guide_int8_computations.html) - States where int8 saturates on pre-VNNI x86, the compensation that makes signed by signed work, and what VNNI removes.
- [Quantize ONNX models](https://onnxruntime.ai/docs/performance/model-optimizations/quantization.html) - Defines the QDQ form against QOperator and dynamic against static, and which sign choices are safe on which CPUs.
- [k-quants](https://github.com/ggml-org/llama.cpp/pull/1684) - Defines super-block formats with quantized scales and the perplexity against size curve behind mixing types per tensor.
- [T-MAC: CPU Renaissance via Table Lookup for Low-Bit LLM Deployment on Edge](https://arxiv.org/abs/2407.00088) - Replaces unpacking and multiplying low-bit weights with bit-sliced lookups, so a narrower weight costs less to apply.

### Matrix extensions

- [Intel Software Developer Manuals](https://www.intel.com/content/www/us/en/developer/articles/technical/intel-sdm.html) - Defines AMX tiles, tile configuration and TMUL, the VNNI and bf16 dot products, and the XSAVE state a tile load needs.
- [Using XSTATE features in user space applications](https://docs.kernel.org/arch/x86/xstate.html) - Defines the arch_prctl permission Linux requires for AMX tile data and the trap a tile instruction takes until granted.
- [Add Intel Advanced Matrix Extensions (AMX) support to ggml](https://github.com/ggml-org/llama.cpp/pull/7707) - Holds the design report for the ggml AMX path, whose double buffering hides applying scales between tile products.
- [Arm Architecture Reference Manual for A-profile architecture](https://support.arm.com/documentation/ddi0487/latest/) - Defines streaming SVE mode, the ZA tile array and SME2 multi-vector instructions, with pseudocode a kernel must match.
- [SME Programmer's Guide](https://support.arm.com/documentation/109246/latest/) - Walks through SME2 int8 and f32 matmul and gemv kernels and a lookup-table path, on a unit so far only in client parts.

### Threading for inference

- [Thread management](https://onnxruntime.ai/docs/performance/tune-performance/threading.html) - Sets the physical-core default, the affinity it implies and the spin-wait controls that trade idle CPU for latency.
- [Performance Hints and Thread Scheduling](https://docs.openvino.ai/2026/openvino-workflow/running-inference/inference-devices-and-modes/cpu-device/performance-hint-and-thread-scheduling.html) - States the vendor defaults, one thread per core, SMT siblings off, core type by precision and one socket for latency.
- [Threadpool: take 2](https://github.com/ggml-org/llama.cpp/pull/8672) - Defines the explicit thread pool with CPU masks, strict placement, priority and polling that ggml runs without OpenMP.
- [llama-bench](https://github.com/ggml-org/llama.cpp/blob/master/tools/llama-bench/README.md) - Defines the prompt and generation tests, repetitions and mean with deviation behind any comparable llama.cpp number.
- [Dual Epyc Genoa/Turin token generation performance bottlenecks](https://github.com/ggml-org/llama.cpp/discussions/11733) - Traces poor decode scaling across sockets to remote NUMA access from weight placement, with numatop counts as evidence.

### When CPU beats GPU

- [gpt-j example README (ggml)](https://github.com/ggml-org/ggml/blob/master/examples/gpt-j/README.md) - Origin of ggml's bandwidth argument, NEON threads saturate a laptop's memory bus, so a GPU on that bus gains nothing.
- [llama.cpp Performance Testing](https://johannesgaessler.github.io/llamacpp_performance) - Holds the memory clock sweep with all else fixed, the token rate following it and flattening after a few threads.
- [SparAMX: Accelerating Compressed LLMs Token Generation on AMX-powered CPUs](https://arxiv.org/abs/2502.12444) - Measures decode kernels on a Xeon as DRAM-bound, so AMX pays at batch one only once bytes per token shrink, with code.
- [MLPerf Inference v6.0 Results](https://github.com/mlcommons/inference_results_v6.0) - Holds CPU-only entries whose logs show the batched Offline case lost to GPU entries at the same accuracy floor.
