Learn / Out
Hardware generations
Read first
The measurement articles below state no compiler, flags or run count, so each is kept for the structure it exposes and no figure from it is repeated.
- Sources
- 20 entries in 4 parts
- Reproduce it
- 14-pcore-vs-ecore
- Related sections
- §2 One instruction, end to end, §3 Microarchitecture
- In the MCP server
cpuperf://section/14
Intel Xeon
-
01
Where Intel states what Sapphire Rapids added: larger L2 and L3, DDR5, CXL, AMX and on-die accelerators.
-
02
Measures L3 and memory latency across the tiled mesh and the slow clock ramp the vendor overview omits.
-
03
The designers' statement of what changed: fewer, larger dies, a bigger shared L3, faster DDR5 and socket links.
-
04
Measures per-die L3 under sub-NUMA clustering and the die-crossing cost on Granite Rapids beside Turin.
-
05
Bandwidth-bound codes on Sapphire, Emerald and Granite Rapids and Sierra Forest with clocks, SMT and compiler stated.
AMD EPYC
-
01
The designers' account of the Zen 4 core and how it yields Genoa, Genoa-X, Bergamo and Siena.
-
02
Tests the same-core claim for Zen 4c: cache latency, clock ceiling and core-to-core paths beside clock-matched Zen 4.
-
03
Where AMD states what Zen 5 changed in front end, vector datapath and caches, the core Turin carries.
-
04
Defines Turin: Zen 5 or Zen 5c dies, which parts double die-to-IO links, NUMA modes and full-width AVX-512.
-
05
Measures what wider die-to-IO links and faster DDR5 do for Turin bandwidth, and where latency rose over Genoa.
Arm Neoverse server parts
-
01 AWS Graviton Getting Started repository
Where AWS states which Neoverse core, ISA revision, mesh, caches and compiler flag each Graviton generation carries.
-
02
Sets the pipeline widths and instruction timings of the core Graviton 4, Grace and Axion share.
-
03
States the coherency fabric, LPDDR5X fit and MPAM cache and memory partitioning on Grace.
-
04
States the timings and fusion rules of the narrower N line core in Cobalt 100 and Yitian.
-
05
Where Ampere states the N1 part: private L2 per core, shared system cache, mesh and DDR4 fit.
Independent measurement across vendors
-
01
Measures Neoverse V2, Golden Cove and Zen 4 in-core at fixed clock, and each socket's clock under vector load.
-
02
One SVE workload run on Graviton 3, Graviton 4, Yitian and Axion with compiler, runs and spread stated.
-
03
Measures a sustained rename width below the stated one, cache latencies, mesh behaviour and cross-socket cost on V2.
-
04
Measures structure sizes, cache latencies and mesh behaviour of N2 on Yitian, the core Cobalt 100 carries.
-
05
Puts vendor slides beside measurements of the predictor, small instruction cache, private L2 and long memory latency.
Reproduce it
the same three kernels on a performance core and an efficiency core.
Reproduce it · 14-pcore-vs-ecore P-core versus E-core The same code runs at different speeds by core type; a result without the core type is not comparable