Sources
21 entries in 4 parts
Reproduce it
07-aos-vs-soa-simd
In the MCP server
cpuperf://section/7

Data layout and loop transforms

  1. 01
    Intel Optimization Reference Manual manual

    Where Intel states when a structure of arrays beats an array of structures, strided and hybrid cases included.

  2. 02
    ispc: A SPMD Compiler for High-Performance CPU Programming paper

    Measures one kernel in both layouts and traces the gain to the gathers the array-of-structures form forces.

  3. 03
    A Data Locality Optimizing Algorithm paper

    Defines interchange, skewing, reversal and tiling as one family and proves when each keeps a loop nest legal.

  4. 04
    The Cache Performance and Optimizations of Blocked Algorithms paper

    Traces the drops in a blocked loop's speed curve to self-interference misses, and shows when copying a tile pays.

  5. 05
    Auto-Vectorization in LLVM manual

    Where LLVM lists what its loop and SLP vectorisers accept: interleaving, reductions, if-conversion and runtime checks.

Reproduce it

array of structs against structure of arrays, scalar against NEON.

Reproduce it · 07-aos-vs-soa-simd Array of structs versus structure of arrays Layout decides bytes moved and whether the loop vectorises

SIMD instruction sets

  1. 01
    Intel Intrinsics Guide manual

    Maps each intrinsic to its instruction and CPUID flag, with the vendor's latency and throughput per microarchitecture.

  2. 02
    Intel Software Developer Manuals manual

    The normative semantics of every Intel vector instruction, with the masking, rounding and fault rules intrinsics hide.

  3. 03
    Introduction to SVE manual

    Where Arm explains vector-length-agnostic loops and predication, the model that removes remainder loops entirely.

  4. 04
    Arm C Language Extensions manual

    The specification the NEON and SVE intrinsics come from, so it settles what a compiler must accept and what is a bug.

  5. 05
    Arm Architecture Reference Manual for A-profile architecture manual

    The normative definition of NEON, SVE and SVE2 instructions and of the scalable vector and predicate register model.

SIMD libraries and measured kernels

  1. 01
    Highway repository

    Defines sizeless vector types with run-time dispatch, so one source serves SVE and every fixed-width ISA.

  2. 02
    xsimd repository

    Fixes a batch type per ISA and width, the compile-time vector length model, and still reaches NEON and SVE.

  3. 03
    Faster Base64 Encoding and Decoding Using AVX2 Instructions paper

    Where shuffle-based lookup replacing a byte loop is worked through step by step, with the measurement method spelt out.

  4. 04
    Parsing Gigabytes of JSON per Second paper

    Shows branch-free structural indexing with carry-less multiply and shuffles, measured against conventional parsers.

  5. 05
    simdjson repository

    Where that technique ships, with a kernel per ISA and the harness that keeps its published comparisons reproducible.

  6. 06
    Hyperscan: A Fast Multi-pattern Regex Matcher for Modern CPUs paper

    Decomposes regexes into string and automaton pieces so both run on SIMD, the design inside the matcher Snort embeds.

Branchless code and bit manipulation

  1. 01
    Branch Prediction and the Performance of Interpreters paper

    Counter evidence that current predictors absorb interpreter dispatch, so a branch removal must be measured, not assumed.

  2. 02
    Hacker's Delight, 2nd Edition book

    Derives the branch-free integer tricks, division by a constant among them, with proofs rather than as a catalogue.

  3. 03
    Faster sorted array unions by reducing branches report

    A worked branchless merge with code whose gain vanishes when the compiler emits no conditional move.

  4. 04
    Array Layouts for Comparison-Based Searching paper

    Measures branch-free search over sorted, Eytzinger and B-tree layouts, the winner changing with size and prefetch.

  5. 05
    Faster Population Counts Using AVX2 Instructions paper

    Shows a carry-save adder tree in vector registers beating the dedicated instruction, timed as the minimum of many runs.