Sources
30 entries in 6 parts
In the MCP server
cpuperf://section/13

GEMM and BLAS

  1. 01
    Anatomy of High-Performance Matrix Multiplication paper

    Derives the cache blocking every fast CPU GEMM still uses by refining a memory model until the packed kernel falls out.

  2. 02
    BLIS: A Framework for Rapidly Instantiating BLAS Functionality paper

    Reduces a BLAS to one register-tile micro-kernel per architecture and states which loop owns which level of cache.

  3. 03
    LLaMA Now Goes Faster on CPUs report

    Builds a register-tile kernel step by step and shows where outer-loop unrolling wins and where a vendor BLAS still does.

  4. 04
    OpenBLAS repository

    Sets run-time dispatch for a BLAS, one kernel table per microarchitecture chosen as a DYNAMIC_ARCH build loads.

  5. 05
    oneDNN Matrix Multiplication Primitive manual

    Fixes the data-type table, packed weight format and fused post-ops a CPU GEMM must expose to an inference graph.

  6. 06
    Scaled Dot-Product Attention (oneDNN Graph) manual

    Fixes the fused attention pattern, f32 accumulation under bf16 inputs and the shapes the fast CPU path accepts.

Reproduce it

naive GEMM, a hand microkernel and the vendor BLAS, plus int8 against float32 dot products.

Reproduce it · 13-sgemm-naive-vs-blas SGEMM: naive loop, microkernel, vendor BLAS A packed, register-blocked microkernel takes a naive loop to the vector unit's ceiling; the vendor BLAS is further ahead only where it owns a matrix unit

Runtimes

  1. 01
    oneDNN repository

    Generates VNNI and AMX kernels at run time under PyTorch, TensorFlow and OpenVINO, and calls Compute Library on Arm.

  2. 02
    ggml repository

    Defines the quantized block formats and the per-ISA dot-product kernels over them that every llama.cpp type rests on.

  3. 03
    llama.cpp repository

    Where new quantization types, kernels and thread pools land first, each with the perplexity and speed table behind it.

  4. 04
    ONNX Runtime MLAS repository

    Holds the CPU provider's GEMM, int8 and int4 MatMul kernels, dispatched per ISA at run time from SSE to AMX and SME.

  5. 05
    OpenVINO CPU Device manual

    States precision defaults per ISA, the int8 path through oneDNN and the streams model that turns cores into throughput.

Quantization

  1. 01
    Quantization and Training of Neural Networks for Efficient Integer-Arithmetic-Only Inference paper

    Defines the affine scale and zero-point scheme with integer accumulation that every int8 CPU runtime still implements.

  2. 02
    Nuances of int8 Computations manual

    States where int8 saturates on pre-VNNI x86, the compensation that makes signed by signed work, and what VNNI removes.

  3. 03
    Quantize ONNX models manual

    Defines the QDQ form against QOperator and dynamic against static, and which sign choices are safe on which CPUs.

  4. 04
    k-quants repository

    Defines super-block formats with quantized scales and the perplexity against size curve behind mixing types per tensor.

  5. 05
    T-MAC: CPU Renaissance via Table Lookup for Low-Bit LLM Deployment on Edge paper

    Replaces unpacking and multiplying low-bit weights with bit-sliced lookups, so a narrower weight costs less to apply.

Matrix extensions

  1. 01
    Intel Software Developer Manuals manual

    Defines AMX tiles, tile configuration and TMUL, the VNNI and bf16 dot products, and the XSAVE state a tile load needs.

  2. 02
    Using XSTATE features in user space applications manual

    Defines the arch_prctl permission Linux requires for AMX tile data and the trap a tile instruction takes until granted.

  3. 03
    Add Intel Advanced Matrix Extensions (AMX) support to ggml repository

    Holds the design report for the ggml AMX path, whose double buffering hides applying scales between tile products.

  4. 04
    Arm Architecture Reference Manual for A-profile architecture manual

    Defines streaming SVE mode, the ZA tile array and SME2 multi-vector instructions, with pseudocode a kernel must match.

  5. 05
    SME Programmer's Guide manual

    Walks through SME2 int8 and f32 matmul and gemv kernels and a lookup-table path, on a unit so far only in client parts.

Threading for inference

  1. 01
    Thread management manual

    Sets the physical-core default, the affinity it implies and the spin-wait controls that trade idle CPU for latency.

  2. 02
    Performance Hints and Thread Scheduling manual

    States the vendor defaults, one thread per core, SMT siblings off, core type by precision and one socket for latency.

  3. 03
    Threadpool: take 2 repository

    Defines the explicit thread pool with CPU masks, strict placement, priority and polling that ggml runs without OpenMP.

  4. 04
    llama-bench repository

    Defines the prompt and generation tests, repetitions and mean with deviation behind any comparable llama.cpp number.

  5. 05
    Dual Epyc Genoa/Turin token generation performance bottlenecks repository

    Traces poor decode scaling across sockets to remote NUMA access from weight placement, with numatop counts as evidence.

When CPU beats GPU

  1. 01
    gpt-j example README (ggml) repository

    Origin of ggml's bandwidth argument, NEON threads saturate a laptop's memory bus, so a GPU on that bus gains nothing.

  2. 02
    llama.cpp Performance Testing report

    Holds the memory clock sweep with all else fixed, the token rate following it and flattening after a few threads.

  3. 03
    SparAMX: Accelerating Compressed LLMs Token Generation on AMX-powered CPUs paper

    Measures decode kernels on a Xeon as DRAM-bound, so AMX pays at batch one only once bytes per token shrink, with code.

  4. 04
    MLPerf Inference v6.0 Results repository

    Holds CPU-only entries whose logs show the batched Offline case lost to GPU entries at the same accuracy floor.