Learn / Out
Single-thread optimisation
Read first
A loop the compiler reports as vectorised can still run at scalar speed: a float reduction stays one serial chain until reassociation is permitted.
- Sources
- 21 entries in 4 parts
- Reproduce it
- 07-aos-vs-soa-simd
- Related sections
- §1 Start here, §4 Memory hierarchy, §5 Measurement, §8 Compilers and codegen, §13 Inference on CPU
- In the MCP server
cpuperf://section/7
Data layout and loop transforms
-
01
Where Intel states when a structure of arrays beats an array of structures, strided and hybrid cases included.
-
02
Measures one kernel in both layouts and traces the gain to the gathers the array-of-structures form forces.
-
03
Defines interchange, skewing, reversal and tiling as one family and proves when each keeps a loop nest legal.
-
04
Traces the drops in a blocked loop's speed curve to self-interference misses, and shows when copying a tile pays.
-
05 Auto-Vectorization in LLVM manual
Where LLVM lists what its loop and SLP vectorisers accept: interleaving, reductions, if-conversion and runtime checks.
Reproduce it
array of structs against structure of arrays, scalar against NEON.
Reproduce it · 07-aos-vs-soa-simd Array of structs versus structure of arrays Layout decides bytes moved and whether the loop vectorisesSIMD instruction sets
-
01 Intel Intrinsics Guide manual
Maps each intrinsic to its instruction and CPUID flag, with the vendor's latency and throughput per microarchitecture.
-
02
The normative semantics of every Intel vector instruction, with the masking, rounding and fault rules intrinsics hide.
-
03 Introduction to SVE manual
Where Arm explains vector-length-agnostic loops and predication, the model that removes remainder loops entirely.
-
04 Arm C Language Extensions manual
The specification the NEON and SVE intrinsics come from, so it settles what a compiler must accept and what is a bug.
-
05
The normative definition of NEON, SVE and SVE2 instructions and of the scalable vector and predicate register model.
SIMD libraries and measured kernels
-
01 Highway repository
Defines sizeless vector types with run-time dispatch, so one source serves SVE and every fixed-width ISA.
-
02 xsimd repository
Fixes a batch type per ISA and width, the compile-time vector length model, and still reaches NEON and SVE.
-
03
Where shuffle-based lookup replacing a byte loop is worked through step by step, with the measurement method spelt out.
-
04
Shows branch-free structural indexing with carry-less multiply and shuffles, measured against conventional parsers.
-
05 simdjson repository
Where that technique ships, with a kernel per ISA and the harness that keeps its published comparisons reproducible.
-
06
Decomposes regexes into string and automaton pieces so both run on SIMD, the design inside the matcher Snort embeds.
Branchless code and bit manipulation
-
01
Counter evidence that current predictors absorb interpreter dispatch, so a branch removal must be measured, not assumed.
-
02
Derives the branch-free integer tricks, division by a constant among them, with proofs rather than as a catalogue.
-
03
A worked branchless merge with code whose gain vanishes when the compiler emits no conditional move.
-
04
Measures branch-free search over sorted, Eytzinger and B-tree layouts, the winner changing with size and prefetch.
-
05
Shows a carry-save adder tree in vector registers beating the dedicated instruction, timed as the minimum of many runs.