Sources
20 entries in 4 parts
Reproduce it
09-false-sharing
In the MCP server
cpuperf://section/9

Memory models and atomics

  1. 01
    Foundations of the C++ Concurrency Memory Model paper

    Defines the data-race-free contract, sequential consistency for race-free programs and no meaning for a race.

  2. 02
    Atomic operations, C++ working draft manual

    The normative wording for every memory order, fence and read-modify-write, the text a compiler is checked against.

  3. 03
    C/C++11 mappings to processors report

    The table that turns each memory order into x86 and Arm instructions, so what an order costs is read off the page.

  4. 04
    Linux kernel memory-barriers.txt manual

    States what the kernel assumes any CPU may reorder and what each barrier and access primitive guarantees.

  5. 05
    herdtools7 repository

    Where herd7, litmus7 and klitmus7 live, the tools that run a litmus test against the x86, Arm and kernel models.

Locks, contention and allocators

  1. 01
    Is Parallel Programming Hard, And, If So, What Can You Do About It? book

    Derives counting, partitioning, locking and deferral with code that runs, the textbook the section assumes.

  2. 02
    Algorithms for Scalable Synchronization on Shared-Memory Multiprocessors paper

    The origin of the queue lock, each waiter spinning on its own line, measured against ticket and test-and-set locks.

  3. 03
    Futexes Are Tricky paper

    Derives a correct user-space mutex from futex and shows the lost wakeups and extra kernel entries naive versions pay.

  4. 04
    Hoard: A Scalable Memory Allocator for Multithreaded Applications paper

    Defines blowup and allocator-induced false sharing, which per-processor heaps under a bounded global heap avoid.

  5. 05
    TCMalloc: Thread-Caching Malloc manual

    The design statement for per-CPU caches built on restartable sequences and a hugepage-aware back end for TLB reach.

Reproduce it

adjacent counters against padded counters across threads.

Reproduce it · 09-false-sharing False sharing Writers sharing a cache line serialise; padding restores scaling

Lock-free structures and RCU

  1. 01
    The Art of Multiprocessor Programming book

    States linearizability and the consensus hierarchy and builds both into working stacks, queues, lists and hash tables.

  2. 02
    Simple, Fast, and Practical Non-Blocking and Blocking Concurrent Queue Algorithms paper

    The lock-free queue later libraries copy, with the counted pointer against ABA and a two-lock queue beside it.

  3. 03
    Hazard Pointers for C++26 paper

    The standard-track form of safe reclamation, fixing when a retired node may be freed while a reader still holds it.

  4. 04
    What is RCU? manual

    The kernel's own statement of RCU as publish, wait for readers and keep old versions, with a free read side.

  5. 05
    User-Level Implementations of Read-Copy Update paper

    Defines liburcu's quiescent-state, signal-based and general RCU flavours and measures each read side against locks.

Thread pools and work stealing

  1. 01
    Cilk: An Efficient Multithreaded Runtime System paper

    Defines work and critical path, proves the work-stealing bound, and shows they alone predict a runtime's speedup.

  2. 02
    The Implementation of the Cilk-5 Multithreaded Language paper

    States the work-first principle, that overhead belongs on the rare steal path and not on every spawn.

  3. 03
    Correct and Efficient Work-Stealing for Weak Memory Models paper

    Gives the work-stealing deque a proven atomics form and derives which fence push, take and steal need on Arm and x86.

  4. 04
    OpenMP Specifications manual

    Fixes fork-join and tasking semantics, and the wait and binding controls deciding if idle workers spin, sleep or move.

  5. 05
    oneTBB repository

    The shipping work-stealing runtime, arenas and task groups over a deque, the home of grain size and spin-before-sleep.