---
title: "14. Hardware generations"
url: https://cpuperf.com/learn/hardware-generations/
source: https://github.com/usamahz/cpu-performance-engineering/blob/deb5a0bac46760503b6f4a2608bdfed470c8532e/README.md#L613
commit: deb5a0bac46760503b6f4a2608bdfed470c8532e
---
## 14. Hardware generations

The measurement articles below state no compiler, flags or run count, so each is kept for the structure it exposes and no figure from it is repeated.

### Intel Xeon

- [Technical Overview of the 4th Gen Intel Xeon Scalable Processor Family](https://www.intel.com/content/www/us/en/developer/articles/technical/fourth-generation-xeon-scalable-family-overview.html) - Where Intel states what Sapphire Rapids added: larger L2 and L3, DDR5, CXL, AMX and on-die accelerators.
- [Sapphire Rapids: Golden Cove Hits Servers](https://chipsandcheese.com/p/a-peek-at-sapphire-rapids) - Measures L3 and memory latency across the tiled mesh and the slow clock ramp the vendor overview omits.
- [Emerald Rapids: 5th-Generation Intel Xeon Scalable Processors](https://ieeexplore.ieee.org/document/10454434) - The designers' statement of what changed: fewer, larger dies, a bigger shared L3, faster DDR5 and socket links.
- [A Look into Intel Xeon 6's Memory Subsystem](https://chipsandcheese.com/p/a-look-into-intel-xeon-6s-memory) - Measures per-die L3 under sub-NUMA clustering and the die-crossing cost on Granite Rapids beside Turin.
- [Benchmarking the Evolution of Performance and Energy Efficiency Across Recent Generations](https://real.mtak.hu/215594/) - Bandwidth-bound codes on Sapphire, Emerald and Granite Rapids and Sierra Forest with clocks, SMT and compiler stated.

### AMD EPYC

- [AMD Next-Generation Zen 4 Core and 4th Gen AMD EPYC Server CPUs](https://ieeexplore.ieee.org/document/10466769) - The designers' account of the Zen 4 core and how it yields Genoa, Genoa-X, Bergamo and Siena.
- [Testing AMD's Bergamo: Zen 4c Spam](https://chipsandcheese.com/p/testing-amds-bergamo-zen-4c-spam) - Tests the same-core claim for Zen 4c: cache latency, clock ceiling and core-to-core paths beside clock-matched Zen 4.
- [Software Optimization Guide for the AMD Zen5 Microarchitecture](https://docs.amd.com/v/u/en-US/58455_1.00) - Where AMD states what Zen 5 changed in front end, vector datapath and caches, the core Turin carries.
- [AMD EPYC 9005 Processor Architecture Overview](https://docs.amd.com/v/u/en-US/58462_amd-epyc-9005-tg-architecture-overview) - Defines Turin: Zen 5 or Zen 5c dies, which parts double die-to-IO links, NUMA modes and full-width AVX-512.
- [AMD's Turin: 5th Gen EPYC Launched](https://chipsandcheese.com/p/amds-turin-5th-gen-epyc-launched) - Measures what wider die-to-IO links and faster DDR5 do for Turin bandwidth, and where latency rose over Genoa.

### Arm Neoverse server parts

- [AWS Graviton Getting Started](https://github.com/aws/aws-graviton-getting-started) - Where AWS states which Neoverse core, ISA revision, mesh, caches and compiler flag each Graviton generation carries.
- [Arm Neoverse V2 Core Software Optimization Guide](https://support.arm.com/documentation/109898/latest/) - Sets the pipeline widths and instruction timings of the core Graviton 4, Grace and Axion share.
- [NVIDIA Grace Performance Tuning Guide](https://docs.nvidia.com/dccpu/grace-perf-tuning-guide/index.html) - States the coherency fabric, LPDDR5X fit and MPAM cache and memory partitioning on Grace.
- [Arm Neoverse N2 Core Software Optimization Guide](https://support.arm.com/documentation/109914/latest/) - States the timings and fusion rules of the narrower N line core in Cobalt 100 and Yitian.
- [Ampere Altra Rev A1 64-Bit Multi-Core Processor Datasheet](https://amperecomputing.com/assets/Altra_Rev_A1_DS_v1_50_20240130_3375c3dec5_1c5d4604fa.pdf) - Where Ampere states the N1 part: private L2 per core, shared system cache, mesh and DDR4 fit.

### Independent measurement across vendors

- [Microarchitectural Comparison and In-core Modeling of State-of-the-art CPUs](https://arxiv.org/abs/2409.08108) - Measures Neoverse V2, Golden Cove and Zen 4 in-core at fixed clock, and each socket's clock under vector load.
- [On the Performance of Cloud-based ARM SVE for Zero-Knowledge Proving Systems](https://arxiv.org/abs/2506.09505) - One SVE workload run on Graviton 3, Graviton 4, Yitian and Axion with compiler, runs and spread stated.
- [Arm's Neoverse V2, in AWS's Graviton 4](https://chipsandcheese.com/p/arms-neoverse-v2-in-awss-graviton-4) - Measures a sustained rename width below the stated one, cache latencies, mesh behaviour and cross-socket cost on V2.
- [ARM's Neoverse N2: Cortex A710 for Servers](https://chipsandcheese.com/p/arms-neoverse-n2-cortex-a710-for-servers) - Measures structure sizes, cache latencies and mesh behaviour of N2 on Yitian, the core Cobalt 100 carries.
- [AmpereOne at Hot Chips 2024: Maximizing Density](https://chipsandcheese.com/p/ampereone-at-hot-chips-2024-maximizing-density) - Puts vendor slides beside measurements of the predictor, small instruction cache, private L2 and long memory latency.

Reproduce it: [misc/benchmarks/14-pcore-vs-ecore](https://cpuperf.com/benchmarks/14-pcore-vs-ecore/), the same three kernels on a performance core and an efficiency core.
