Sources
15 entries in 3 parts
Reproduce it
10-first-touch
Related sections
None in this section
In the MCP server
cpuperf://section/10

NUMA and Linux memory placement

  1. 01
    NUMA (Non-Uniform Memory Access): An Overview paper

    The one account tying first touch, policy scope, zone reclaim and page movement together from the implementer's side.

  2. 02
    What is NUMA? manual

    Defines nodes, zonelists and the distance-ordered fallback that places an allocation once local memory runs out.

  3. 03
    NUMA Memory Policy manual

    The normative statement of policy scopes, every mode including weighted interleave, and the cpuset intersection rule.

  4. 04
    Numa policy hit/miss statistics manual

    Defines numa_hit, numa_miss and numa_foreign, the counters that show whether a policy put pages where it said.

  5. 05
    numactl repository

    Reference implementation of the policy API, prints the distance table and binds a binary that cannot be rebuilt.

Reproduce it

first touch of fresh pages against the second pass.

Reproduce it · 10-first-touch First touch Allocation is not placement; the first touch pays the fault and picks the home

Topology and interconnects

  1. 01
    NUMA Memory Performance manual

    Explains the firmware-rated latency and bandwidth per initiator and target, and memory-side caches, that rank nodes.

  2. 02
    Intel Xeon Processor Scalable Family Technical Overview manual

    Where Intel names the mesh, UPI socket links, the directory-running home agent, and how SNC splits the cache.

  3. 03
    Intel Xeon 6 with P-cores Configuration and Tuning Guide for HPC Applications manual

    Defines SNC on current parts as one node per compute die, and fixes the numactl and numastat checks of placement.

  4. 04
    BIOS and Workload Tuning Guide for AMD EPYC 9004 Series Processors manual

    Discloses the I/O die, GMI and xGMI links, the NPS modes with their interleave widths, and the cache-as-NUMA override.

  5. 05
    Arm Neoverse CMN-700 Coherent Mesh Network Technical Reference Manual manual

    Defines the mesh, the home nodes holding the system cache and snoop filter, and the gateways joining sockets or CXL.

Migration, balancing and measured effects

  1. 01
    move_pages(2) manual

    Defines per-page migration of a running process, and a query reporting each page's node, the direct test of first touch.

  2. 02
    sysctl kernel numa_balancing manual

    Defines the hinting-fault sampling behind automatic balancing and tiering, and warns the overhead may not pay off.

  3. 03
    Traffic Management: A Holistic Approach to Memory Placement on NUMA Systems paper

    Proves against the kernel balancer that controller and link congestion, not remote latency, is what placement manages.

  4. 04
    Intel Memory Latency Checker manual

    Measures the node-to-node latency and bandwidth matrix and loaded latency on the x86 at hand, which no datasheet states.

  5. 05
    High Performance Computing Tuning Guide for AMD EPYC 9004 Series Processors manual

    Tabulates measured bandwidth by NPS mode, cores per die, boost and SMT, so the NPS trade-off is shown, not asserted.