Sources
20 entries in 4 parts
Reproduce it
14-pcore-vs-ecore
In the MCP server
cpuperf://section/14

Intel Xeon

  1. 01
    Technical Overview of the 4th Gen Intel Xeon Scalable Processor Family manual

    Where Intel states what Sapphire Rapids added: larger L2 and L3, DDR5, CXL, AMX and on-die accelerators.

  2. 02
    Sapphire Rapids: Golden Cove Hits Servers report

    Measures L3 and memory latency across the tiled mesh and the slow clock ramp the vendor overview omits.

  3. 03
    Emerald Rapids: 5th-Generation Intel Xeon Scalable Processors paper

    The designers' statement of what changed: fewer, larger dies, a bigger shared L3, faster DDR5 and socket links.

  4. 04
    A Look into Intel Xeon 6's Memory Subsystem report

    Measures per-die L3 under sub-NUMA clustering and the die-crossing cost on Granite Rapids beside Turin.

  5. 05
    Benchmarking the Evolution of Performance and Energy Efficiency Across Recent Generations paper

    Bandwidth-bound codes on Sapphire, Emerald and Granite Rapids and Sierra Forest with clocks, SMT and compiler stated.

AMD EPYC

  1. 01
    AMD Next-Generation Zen 4 Core and 4th Gen AMD EPYC Server CPUs paper

    The designers' account of the Zen 4 core and how it yields Genoa, Genoa-X, Bergamo and Siena.

  2. 02
    Testing AMD's Bergamo: Zen 4c Spam report

    Tests the same-core claim for Zen 4c: cache latency, clock ceiling and core-to-core paths beside clock-matched Zen 4.

  3. 03
    Software Optimization Guide for the AMD Zen5 Microarchitecture manual

    Where AMD states what Zen 5 changed in front end, vector datapath and caches, the core Turin carries.

  4. 04
    AMD EPYC 9005 Processor Architecture Overview manual

    Defines Turin: Zen 5 or Zen 5c dies, which parts double die-to-IO links, NUMA modes and full-width AVX-512.

  5. 05
    AMD's Turin: 5th Gen EPYC Launched report

    Measures what wider die-to-IO links and faster DDR5 do for Turin bandwidth, and where latency rose over Genoa.

Arm Neoverse server parts

  1. 01
    AWS Graviton Getting Started repository

    Where AWS states which Neoverse core, ISA revision, mesh, caches and compiler flag each Graviton generation carries.

  2. 02
    Arm Neoverse V2 Core Software Optimization Guide manual

    Sets the pipeline widths and instruction timings of the core Graviton 4, Grace and Axion share.

  3. 03
    NVIDIA Grace Performance Tuning Guide manual

    States the coherency fabric, LPDDR5X fit and MPAM cache and memory partitioning on Grace.

  4. 04
    Arm Neoverse N2 Core Software Optimization Guide manual

    States the timings and fusion rules of the narrower N line core in Cobalt 100 and Yitian.

  5. 05
    Ampere Altra Rev A1 64-Bit Multi-Core Processor Datasheet manual

    Where Ampere states the N1 part: private L2 per core, shared system cache, mesh and DDR4 fit.

Independent measurement across vendors

  1. 01
    Microarchitectural Comparison and In-core Modeling of State-of-the-art CPUs paper

    Measures Neoverse V2, Golden Cove and Zen 4 in-core at fixed clock, and each socket's clock under vector load.

  2. 02
    On the Performance of Cloud-based ARM SVE for Zero-Knowledge Proving Systems paper

    One SVE workload run on Graviton 3, Graviton 4, Yitian and Axion with compiler, runs and spread stated.

  3. 03
    Arm's Neoverse V2, in AWS's Graviton 4 report

    Measures a sustained rename width below the stated one, cache latencies, mesh behaviour and cross-socket cost on V2.

  4. 04
    ARM's Neoverse N2: Cortex A710 for Servers report

    Measures structure sizes, cache latencies and mesh behaviour of N2 on Yitian, the core Cobalt 100 carries.

  5. 05
    AmpereOne at Hot Chips 2024: Maximizing Density report

    Puts vendor slides beside measurements of the predictor, small instruction cache, private L2 and long memory latency.

Reproduce it

the same three kernels on a performance core and an efficiency core.

Reproduce it · 14-pcore-vs-ecore P-core versus E-core The same code runs at different speeds by core type; a result without the core type is not comparable