---
title: "10. NUMA and multi-socket"
url: https://cpuperf.com/learn/numa-and-multi-socket/
source: https://github.com/usamahz/cpu-performance-engineering/blob/deb5a0bac46760503b6f4a2608bdfed470c8532e/README.md#L453
commit: deb5a0bac46760503b6f4a2608bdfed470c8532e
---
## 10. NUMA and multi-socket

A page's node is decided at first touch, not when memory is allocated or a policy is set, and every vendor table below depends on the BIOS node mode of the machine it ran on.

### NUMA and Linux memory placement

- [NUMA (Non-Uniform Memory Access): An Overview](https://queue.acm.org/doi/10.1145/2508834.2513149) - The one account tying first touch, policy scope, zone reclaim and page movement together from the implementer's side.
- [What is NUMA?](https://docs.kernel.org/mm/numa.html) - Defines nodes, zonelists and the distance-ordered fallback that places an allocation once local memory runs out.
- [NUMA Memory Policy](https://docs.kernel.org/admin-guide/mm/numa_memory_policy.html) - The normative statement of policy scopes, every mode including weighted interleave, and the cpuset intersection rule.
- [Numa policy hit/miss statistics](https://docs.kernel.org/admin-guide/numastat.html) - Defines numa_hit, numa_miss and numa_foreign, the counters that show whether a policy put pages where it said.
- [numactl](https://github.com/numactl/numactl) - Reference implementation of the policy API, prints the distance table and binds a binary that cannot be rebuilt.

Reproduce it: [misc/benchmarks/10-first-touch](https://cpuperf.com/benchmarks/10-first-touch/), first touch of fresh pages against the second pass.

### Topology and interconnects

- [NUMA Memory Performance](https://docs.kernel.org/admin-guide/mm/numaperf.html) - Explains the firmware-rated latency and bandwidth per initiator and target, and memory-side caches, that rank nodes.
- [Intel Xeon Processor Scalable Family Technical Overview](https://www.intel.com/content/www/us/en/developer/articles/technical/xeon-processor-scalable-family-technical-overview.html) - Where Intel names the mesh, UPI socket links, the directory-running home agent, and how SNC splits the cache.
- [Intel Xeon 6 with P-cores Configuration and Tuning Guide for HPC Applications](https://www.intel.com/content/www/us/en/content-details/858491/intel-xeon-6-with-p-cores-configuration-and-tuning-guide-for-hpc-applications.html) - Defines SNC on current parts as one node per compute die, and fixes the numactl and numastat checks of placement.
- [BIOS and Workload Tuning Guide for AMD EPYC 9004 Series Processors](https://docs.amd.com/v/u/en-US/58011-epyc-9004-tg-bios-and-workload) - Discloses the I/O die, GMI and xGMI links, the NPS modes with their interleave widths, and the cache-as-NUMA override.
- [Arm Neoverse CMN-700 Coherent Mesh Network Technical Reference Manual](https://support.arm.com/documentation/102308/latest/) - Defines the mesh, the home nodes holding the system cache and snoop filter, and the gateways joining sockets or CXL.

### Migration, balancing and measured effects

- [move_pages(2)](https://man7.org/linux/man-pages/man2/move_pages.2.html) - Defines per-page migration of a running process, and a query reporting each page's node, the direct test of first touch.
- [sysctl kernel numa_balancing](https://docs.kernel.org/admin-guide/sysctl/kernel.html#numa-balancing) - Defines the hinting-fault sampling behind automatic balancing and tiering, and warns the overhead may not pay off.
- [Traffic Management: A Holistic Approach to Memory Placement on NUMA Systems](https://people.ece.ubc.ca/sasha/papers/asplos284-dashti.pdf) - Proves against the kernel balancer that controller and link congestion, not remote latency, is what placement manages.
- [Intel Memory Latency Checker](https://www.intel.com/content/www/us/en/developer/articles/tool/intelr-memory-latency-checker.html) - Measures the node-to-node latency and bandwidth matrix and loaded latency on the x86 at hand, which no datasheet states.
- [High Performance Computing Tuning Guide for AMD EPYC 9004 Series Processors](https://docs.amd.com/v/u/en-US/58002_amd-epyc-9004-tg-hpc) - Tabulates measured bandwidth by NPS mode, cores per die, boost and SMT, so the NPS trade-off is shown, not asserted.
