---
title: "Editorial record"
url: https://cpuperf.com/evidence/record/
commit: deb5a0bac46760503b6f4a2608bdfed470c8532e
---

# What was left out, and why.

Every candidate the section owners considered and left out, with the rule it failed, and every performance number examined against the seven fields, with its verdict. Parsed from the section drafts in misc/notes/sections/.

## 1. Start here

### Rejected candidates

- r1.1 [Arm Neoverse V2 Core Software Optimization Guide](https://support.arm.com/documentation/109898/latest/): Not a list-rule failure; fails this section's test that a newcomer can read it, being a short pipeline overview followed by per-instruction tables that only the last entry teaches how to use. Listed in four later sections, so it is not lost, and its slot went to the many-threads primer that the concurrency, NUMA, isolation and tail sections all assume.
- r1.2 [Arm's Neoverse V2, in AWS's Graviton 4](https://chipsandcheese.com/p/arms-neoverse-v2-in-awss-graviton-4): Not a list-rule failure; considered as the measured companion the Arm guide needed at this step, and left with it when the guide moved out. Listed in the hardware section.
- r1.3 [The Tail at Scale](https://www.barroso.org/publications/TheTailAtScale.pdf): Not a list-rule failure; the path is at its cap of ten and the freed slot went to the many-threads primer, on which more of the later sections depend. First entry of the tail-latency section.
- r1.4 [Is Parallel Programming Hard, And, If So, What Can You Do About It?](https://mirrors.edge.kernel.org/pub/linux/kernel/people/paulmck/perfbook/perfbook.html): Not a list-rule failure; carries the fifth entry's derivation as one chapter of a long book, and the paper is the shorter reading for a path. Listed in the concurrency section.
- r1.5 [Performance Ninja](https://github.com/dendibakh/perf-ninja): Rule 2: was carried as a trailing companion line outside the entry form and in the reader-addressing voice the list bans, and the path is at its cap of ten. Listed in the measurement section, so it is not lost.
- r1.6 [The microarchitecture of Intel, AMD, and VIA CPUs](https://www.agner.org/optimize/microarchitecture.pdf): Not a list-rule failure; fails this section's test that a newcomer can read it, by the author's own statement that the manual is not for beginners. Belongs in the single-thread throughput section.
- r1.7 [Instruction tables](https://www.agner.org/optimize/instruction_tables.pdf): Not a list-rule failure; reference tables rather than a reading, and they depend on the microarchitecture manual above. Belongs with it.
- r1.8 [What Every Programmer Should Know About Memory, LWN serialisation](https://lwn.net/Articles/250967/): Rule 4: a mirror in parts of a document whose author hosts the canonical copy.
- r1.9 [What Every Programmer Should Know About Memory, FreeBSD mirror](https://people.freebsd.org/~lstewart/articles/cpumemory.pdf): Rule 4: third-party mirror of the same document.
- r1.10 [Roofline: An Insightful Visual Performance Model, Berkeley technical report EECS-2008-134](https://www2.eecs.berkeley.edu/Pubs/TechRpts/2008/EECS-2008-134.html): Rule 4: the earlier long draft of the CACM article; one canonical copy per source, and the published article is free at the publisher.
- r1.11 [A Top-Down Method for Performance Analysis and Counters Architecture, IEEE Xplore](https://ieeexplore.ieee.org/document/6844459): Rule 4: paywalled copy where the author's own page carries the PDF.
- r1.12 [Top-down Microarchitecture Analysis Method, VTune cookbook](https://www.intel.com/content/www/us/en/docs/vtune-profiler/cookbook/2025-0/top-down-microarchitecture-analysis-method.html): Rule 1 and rule 4: a tool page that restates a method whose canonical definitions are the paper and the performance-monitoring chapter of the optimisation manual, both already listed.
- r1.13 [Optimizations in C++ Compilers: A practical journey](https://queue.acm.org/detail.cfm?id=3372264): Not a list-rule failure; the written form of the last entry's material by the same author, and the path has one slot for that step. Candidate for the compilers section.
- r1.14 [Compiler Explorer](https://github.com/compiler-explorer/compiler-explorer): Not a list-rule failure; software rather than a reading, and the last entry teaches its use. Candidate for the compilers section.
- r1.15 [Arm Neoverse N2 Core Software Optimization Guide](https://support.arm.com/documentation/109914/latest/): Not a list-rule failure; the same form as the V2 guide, which is also out of the path for the same reason. Candidate for the hardware section.
- r1.16 [Computer Architecture: A Quantitative Approach, sixth edition](https://shop.elsevier.com/books/computer-architecture/hennessy/978-0-12-811905-1): Rule 4: superseded; the publisher marks the seventh edition as the latest.
- r1.17 [Computer Architecture, sixth edition companion site](https://shop.elsevier.com/books/book-companion/9780128119051): Considered as a free home for the pipelining appendix; it hosts only appendices D to M, so it does not deliver the chapter the first entry is for.
- r1.18 [perf tutorial](https://perfwiki.github.io/main/tutorial/): Not a list-rule failure; kernel documentation for one Linux tool, which the path does not depend on. Belongs in the measurement section.
- r1.19 [perf Examples](https://www.brendangregg.com/perf.html): Not a list-rule failure; a tool cookbook by the author of the sixth entry. Belongs in the measurement section.
- r1.20 [Latency Numbers Every Programmer Should Know](https://gist.github.com/jboner/2841832): Rule 1: an aggregation with no measurement of its own; rule 3: none of the seven fields for any number in it.
- r1.21 [Out-of-order execution, Wikipedia](https://en.wikipedia.org/wiki/Out-of-order_execution): Rule 1: Wikipedia.
- r1.22 [A whirlwind introduction to dataflow graphs](https://fgiesen.wordpress.com/2018/03/05/a-whirlwind-introduction-to-dataflow-graphs/): Rule 1: a personal-blog explainer of latency versus throughput, not an original report of a mechanism or a measurement; the ground is covered by the first and second entries.

### Numbers examined

- c1.1 No reason in this section rests on a number. The numbers below are the first ones a reader meets in the listed sources; they were checked so that the reader knows how far each source's own evidence goes.
- c1.2 "L1d about 3 cycles, L2 about 14, main memory about 240" from https://www.akkadia.org/drepper/cpumemory.pdf (section 3.2): quoted from the vendor for a Pentium M. Fields present: 1 CPU: Pentium M, model not named; 2 cores: one, implicit; 3 freq/turbo/SMT: not stated; 4 compiler/flags: not applicable, vendor figure; 5 workload: not stated; 6 baseline: a register access; 7 method: not stated. Missing: 1 in part, 3, 5, 7. Verdict: watchlist; the entry rests on the mechanism, not on these values. Verdict: watchlist.
- c1.3 "around 4 cycles" per element with the working set in L1d and "only about 9 cycles" per element in the L2 region from https://www.akkadia.org/drepper/cpumemory.pdf (section 3.3.2): fields present: 1 CPU: a Pentium 4 in 64-bit mode, NetBurst, model not named, with the L1d and L2 sizes stated; 2 cores: a single thread; 3 freq/turbo/SMT: not stated; 4 compiler/flags: not stated; 5 workload: a circular linked-list walk with no padding (NPAD=0) in sequential order, Figure 3.10, working set doubled from about a kilobyte upwards; the random-order curves are in Figures 3.15 to 3.17; 6 baseline: the working set resident in L1d; 7 method: cycles per list element, run count and statistic not stated. Missing: 3, 4, 7 in part. Verdict: watchlist; the entry rests on the shape of the curve, which the benchmark proposal below reproduces, not on the values. Verdict: watchlist.
- c1.4 "ten instructions per nanosecond" against "many tens of nanoseconds to fetch a data item from main memory", "more than two orders of magnitude", from http://www.rdrop.com/users/paulmck/scalability/paper/whymb.2010.07.23a.pdf (section 1): an illustration with no machine named. Fields present: none of the seven. Verdict: not quoted; the entry rests on the store-buffer and invalidate-queue mechanism, not on the ratio. Verdict: not_quoted.
- c1.5 "the Retiring category is close to 50%" for single-copy SPEC CPU2006 from the paper linked at https://sites.google.com/site/analysismethods/yasin-pubs (section 5.1, setup in Table 3): fields present: 1 CPU: Intel Core i7-3940XM, Ivy Bridge; 2 cores: one core used of four; 3 freq: fixed at 3 GHz, so turbo off, SMT state not stated; 4 compiler: Intel Compiler 14 targeting SSE4.2, remaining flags not stated; 5 workload: SPEC CPU2006 v1.2 base; 6 baseline: none, a characterisation rather than a speed-up; 7 method: Top-Down metrics computed from PMU events, run count not stated. Missing: 3 in part, 4 in part, 7 in part. Verdict: watchlist; the entry rests on the definition of the hierarchy, not on the characterisation. Verdict: watchlist.
- c1.6 The Roofline article's kernel results are for four machines of its own date and are not carried into this list; the entry rests on the definitions of operational intensity and the ceilings.

## 2. One instruction, end to end

### Rejected candidates

- r2.1 [Arm Neoverse V2 Core Software Optimization Guide](https://support.arm.com/documentation/109898/latest/): Trimmed for length: one home per URL, kept under Execute, whose reason now walks the core from fetch to issue before the per-instruction tables, so the front-end account is not lost.
- r2.2 [Intel Optimization Reference Manual](https://www.intel.com/content/www/us/en/content-details/671488/intel-64-and-ia-32-architectures-optimization-reference-manual-volume-1.html): Trimmed for length: held at Start here, where a reader arriving in order already has its front-end chapters, and under Fetch and decode the decoded cache paths and their cost are carried by the AMD guide, The microarchitecture of Intel, AMD, and VIA CPUs and the erratum note.
- r2.3 [Speculatively speaking](https://fgiesen.wordpress.com/2013/03/04/speculatively-speaking/): Trimmed for length: a worked case of the store-forwarding stall that Store-to-Load Forwarding and Memory Disambiguation in x86 Processors establishes by measurement, and the width-matching fix it applies is the rule that map states.
- r2.4 [What Has Your Microcode Done for You Lately?](https://travisdowns.github.io/blog/2019/03/19/random-writes-and-microcode-oh-my.html): Trimmed for length: a second measurement of the store path after the forwarding map, with the overlap of misses it drains defined by the memory-level parallelism paper, so what leaves the README with it is only the warning that a microcode revision can change how the store buffer drains.
- r2.5 [Computer Architecture: A Quantitative Approach, 7th Edition](https://shop.elsevier.com/books/computer-architecture/hennessy/978-0-443-15406-5): Rule 1 in spirit: a textbook is a secondary source under the four source classes, and it is already the first Start here entry, so a reader arriving in order holds its ILP chapter.
- r2.6 [Performance Speed Limits](https://travisdowns.github.io/blog/2019/06/11/speed-limits.html): Left to section 3: the checklist of hard bounds is listed there with its harness and the benchmark that reproduces it, and the reason here only restated it.
- r2.7 [uarch-bench](https://github.com/travisdowns/uarch-bench): Left to section 15: a suite belongs with the benchmarks, and the reason here described the tool rather than what the Execute stage takes from it.
- r2.8 [uops.info: Characterizing Latency, Throughput, and Port Usage of Instructions on Intel Microarchitectures](https://arxiv.org/abs/1810.04610): Left to section 3, where it sits beside the site whose method it defines; a second listing here carried the same reason.
- r2.9 [uops.info](https://uops.info/): Left to section 3, where the site and the paper that defines its measurements sit together; the Execute stage takes its weights from the tables above.
- r2.10 [Arm Neoverse N2 Core Software Optimization Guide](https://support.arm.com/documentation/109914/latest/): No rule failed; replaced by the V2 guide so that one Arm core is walked from fetch to execute, since both guides carry the same front-end account.
- r2.11 [AsmDB: Understanding and Mitigating Front-End Stalls in Warehouse-Scale Computers](https://research.google/pubs/asmdb-understanding-and-mitigating-front-end-stalls-in-warehouse-scale-computers/): Size: the fleet-scale measurement of where instruction-cache misses come from, held at the cap of seven; it belongs beside the post-link optimiser in the compiler section.
- r2.12 [Entropy Decoding in Oodle Data: x86-64 6-Stream Huffman Decoders](https://fgiesen.wordpress.com/2023/10/29/entropy-decoding-in-oodle-data-x86-64-6-stream-huffman-decoders/): No rule failed; the three-stream report already carries the chain and issue-slot method, and this one adds the wider variant with no Arm measurement.
- r2.13 [A whirlwind introduction to dataflow graphs](https://fgiesen.wordpress.com/2018/03/05/a-whirlwind-introduction-to-dataflow-graphs/): Rule 1: a tutorial explainer, not an original report of a mechanism or a measurement; the critical-path paper that defines the dependence graph model is listed instead.
- r2.14 [Reading bits in far too many ways (part 3)](https://fgiesen.wordpress.com/2018/09/27/reading-bits-in-far-too-many-ways-part-3/): Rule 1 at the margin: dependence graphs on an idealised machine with no measured timings; the same author's measured decoder report is listed under Execute instead.
- r2.15 [Computer Architecture: A Quantitative Approach, 6th Edition](https://shop.elsevier.com/books/computer-architecture/hennessy/978-0-12-811905-1): Rule 4: superseded; the page itself points to the newer edition, which is now held to Start here for the reason above.
- r2.16 [Software Optimization Guide for the AMD Zen4 Microarchitecture](https://docs.amd.com/v/u/en-US/57647_zen4_sog_1.01): Rule 4: the only official home returns 401 inside AMD's portal and the portal search no longer lists document 57647; readmit once AMD restores it.
- r2.17 [Zen 4 guide mirror on numberworld.org](https://www.numberworld.org/blogs/2024_8_7_zen5_avx512_teardown/57647_zen4_sog.pdf): Rule 4: a third-party mirror of a vendor document; not admissible while the vendor copy is missing, since the list links vendor libraries only.
- r2.18 [The microarchitecture of superscalar processors](https://ieeexplore.ieee.org/document/476078): Rule 1: a tutorial survey of the stages, secondary by definition, even though it is well written; the original papers are listed instead.
- r2.19 [Implementing Precise Interrupts in Pipelined Processors on Semantic Scholar](https://www.semanticscholar.org/paper/Implementing-Precise-Interrupts-in-Pipelined-Smith-Pleszkun/341f2c5a2b449df7f0761a866a3ec77aaf13b4d6): Rule 4: an aggregator copy; no author copy exists, so the publisher page is used.
- r2.20 [Microarchitecture Optimizations for Exploiting Memory-Level Parallelism in the ACM Digital Library](https://dl.acm.org/doi/10.1145/1028176.1006708): Rule 4: the same paper as the proceedings page listed above, served as the SIGARCH newsletter copy; one publisher page is enough, and no author copy exists.
- r2.21 [robsize](https://github.com/travisdowns/robsize): Same mechanism and method as the reorder buffer post already listed; a second link for one measurement adds nothing for a newcomer.
- r2.22 [Hardware Store Elimination](https://travisdowns.github.io/blog/2020/05/13/intel-zero-opt.html): Out of scope for this section: the mechanism lives in the cache hierarchy, not the store path of the core; belongs with section 3.
- r2.23 [Clang-format Tanks Performance](https://travisdowns.github.io/blog/2019/11/19/toupper.html): Out of scope: the measured slowdown is an inlining failure caused by header order, not a pipeline stage; belongs with the compiler section.
- r2.24 [Intel 64 and IA-32 Architectures Software Developer Manuals](https://www.intel.com/content/www/us/en/developer/articles/technical/intel-sdm.html): Left to section 9: the fault, trap and abort classes are the kernel boundary's business; retire is covered here by the precise interrupts paper.

### Numbers examined

- c2.1 "168" reorder buffer entries on Ivy Bridge and Sandy Bridge from https://blog.stuffedcow.net/2013/05/measuring-rob-capacity/: fields present: 1 CPU named by microarchitecture only (Ivy Bridge, Sandy Bridge, Lynnfield, Northwood, Yorkfield, Palermo, Coppermine), 2 one thread (a second context is only used to show partitioning), 3 frequency, turbo and SMT state not stated, 4 machine code generated at run time by a C program, no compiler flags stated, 5 two cache-missing loads separated by a varying number of filler instructions, 6 the serialised two-load case, 7 loop time in cycles against filler count, knee read off the curve, single value per point. Missing: 1 (model numbers), 3, 4. Verdict: watchlist for the number; the entry stays for its method and quotes no number. Verdict: watchlist.
- c2.2 "5.07" cycles store-to-load forwarding latency on Haswell from https://blog.stuffedcow.net/2014/01/x86-memory-disambiguation/: fields present: 1 Pentium G3420 Haswell (plus Core 2 Quad Q9550, Core i7-860, Core i5-2500K, Core i5-3570K, Phenom 9550, Phenom II 1090T, FX-8120, FX-8320), 2 one thread, 3 frequency, turbo and SMT state not stated, 4 hand-built instruction sequences, no compiler, 5 a dependent store then load with every combination of size and offset, 6 the aligned same-size case, 7 latency in cycles per pair, all offset combinations, no run count or spread stated. Missing: 3, 7 (runs and statistic). Verdict: watchlist for the number; the entry stays for the map of which cases forward. Verdict: watchlist.
- c2.3 "the throughput is less than half of the old microcode" from https://travisdowns.github.io/blog/2019/03/19/random-writes-and-microcode-oh-my.html: fields present: 1 Core i7-6700HQ Skylake, 2 one thread, 3 not stated, 4 not stated (uarch-bench build), 5 random byte stores into two arrays of varying sizes so that some stores hit and some miss L1, 6 the same binary under microcode revision 0xc2 against revision 0xc6, 7 cycles per store from performance counters, L2 and RFO counters shown, run count not stated. Missing: 3, 4, 7 (runs and statistic). Verdict: watchlist for the number; the entry is in Rejected, trimmed for length, and the store buffer mechanism it isolates goes with it. Verdict: watchlist.
- c2.4 "3.243ms" to "3.110ms" median cull time from https://fgiesen.wordpress.com/2013/03/04/speculatively-speaking/: fields present: 1 not stated beyond an Intel x86 core, 2 not stated, 3 not stated, 4 not stated, 5 the triangle binning pass of a software occlusion culling demo, 6 the previous revision of the same code, 7 minimum, quartiles, median, maximum, mean and standard deviation of repeated runs, VTune counter for loads blocked by store forwarding. Missing: 1, 2, 3, 4. Verdict: watchlist for the number; the entry is in Rejected, trimmed for length, since the forwarding map above establishes the stall it diagnoses. Verdict: watchlist.
- c2.5 "0-4% on many industry-standard benchmarks" from https://www.intel.com/content/www/us/en/content-details/841076/intel-mitigations-for-jump-conditional-code-erratum.html: fields present: 4 Intel Compiler Version 19 update 4 for the SPECrate2017 runs only, no flags, 5 a named list of server and client benchmarks (SPECrate2017 int and fp, Linpack, Stream Triad, FIO, HammerDB PostgreSQL, SPECjbb2015, SPECvirt, SYSmark 2018, PCmark 10, 3DMark, WebXPRT, Cinebench R20), 6 the same platform before the microcode update. Missing: 1, 2, 3 (an unnamed internal reference platform), 7. Verdict: cut for the number; the entry stays for the mechanism and the counter recipe. Verdict: cut.
- c2.6 "15 - 20 clock cycles" misprediction penalty for Haswell, Broadwell and Skylake and "at least 17 clock cycles" for Nehalem, with a passage of the same kind per core, from https://www.agner.org/optimize/microarchitecture.pdf: fields present: 1 named by microarchitecture per chapter, model numbers not stated, 2 one thread, 3 not stated per measurement, 4 the author's own test programs and scripts, published on the same site, no compiler flags since the tests are assembly, 5 branch loops built to mispredict, 6 the predicted case, 7 cycle counts from performance counters, run count and spread not stated. Missing: 3, 7 (runs and statistic). Verdict: watchlist for the numbers; the entry stays because the penalty is measured per core and the benchmark for this section reproduces the measurement, not the figure. Verdict: watchlist.
- c2.7 "benchmarks hover around 41.8 cycles per 15 bytes (i.e. around 2.79 cycles/byte)" against a predicted "38 cycles" critical path from https://fgiesen.wordpress.com/2022/09/05/entropy-decoding-in-oodle-data-x86-64-3-stream-huffman-decoders/: fields present: 1 Skylake, model not stated, 2 one thread, 3 not stated, 4 hand-written assembly, no compiler, 5 the three-stream decoder on an idealised rig and on real data, 6 the critical-path and issue-slot model of the same loop, 7 cycles per iteration, run count and statistic not stated. Missing: 1 (model), 3, 7 (runs and statistic). Verdict: watchlist for the figure; the entry stays for the prediction-then-measurement method, and quotes no number. Verdict: watchlist.
- c2.8 "1.51x speed compared to zlib-cloudflare, and 2.1x speed compared to Apple's system zlib" from https://dougallj.wordpress.com/2022/08/20/faster-zlib-deflate-decompression-on-the-apple-m1-and-x86/: fields present: 1 Apple M1, core type not stated, 2 one thread, 3 not stated (no SMT on the part), 4 Apple clang 13.1.6, flags not stated, 5 the Silesia corpus compressed at zlib level 6, 6 zlib-cloudflare, zlib-ng and the system zlib, 7 compressed megabytes per second per file, run count and statistic not stated. Missing: 1 (core type), 3, 4 (flags), 7 (runs and statistic). Verdict: watchlist for the ratios; the entry stays for the latency-bound refill chain it reasons from on an Arm core, and quotes no number. Verdict: watchlist.
- c2.9 "improves MLP over an in-order issue processor by 12-30%" from https://ieeexplore.ieee.org/document/1310765: fields present: 5 commercial and SPEC workloads on a simulated processor, 6 an in-order issue model of the same width, 7 cycle-accurate simulation. Missing: 1, 2, 3, 4 (a simulated machine, no hardware, no compiler stated in the abstract). Verdict: cut for the number; the entry stays for the definition of memory-level parallelism and the ranking of what caps it. Verdict: cut.

## 3. Microarchitecture

### Rejected candidates

- r3.1 [Computer Architecture: A Quantitative Approach, 7th Edition](https://shop.elsevier.com/books/computer-architecture/hennessy/978-0-443-15406-5): Left to Start here and section 2: the vocabulary it defines is the reason given in both, and a third listing added nothing specific to the section.
- r3.2 [An Efficient Algorithm for Exploiting Multiple Arithmetic Units](https://ieeexplore.ieee.org/document/5392028): Left to section 2: listed there as the origin of renaming and reservation stations, and the reason a second listing would give is the same one.
- r3.3 [Micro-Operation Cache: A Power Aware Frontend for Variable Instruction Length ISA](https://ieeexplore.ieee.org/document/945363): Left to section 2: a front-end design paper, the origin of the op cache that the fetch and decode manuals describe, and misfiled under the limits of ILP.
- r3.4 [KPTI/KAISER Meltdown Initial Performance Regressions](https://www.brendangregg.com/blog/2018-02-09/kpti-kaiser-meltdown-performance.html): Left to section 11: page-table isolation cost against syscall rate is a kernel boundary measurement, not branch prediction, and every figure in it is watchlist-grade.
- r3.5 [uops.info: Characterizing Latency, Throughput, and Port Usage of Instructions](https://arxiv.org/abs/1810.04610): Left to section 2: the per-operand latency definition and the automated method are the reason it is listed there, and the site carries what the tables need.
- r3.6 [Arm Neoverse V2 Core Software Optimization Guide](https://support.arm.com/documentation/109898/latest/): Left to Start here and sections 2, 4 and 14: the per-pipe tables are the reason given in each of them, and provenance is not a reason.
- r3.7 [Intel Optimization Reference Manual](https://www.intel.com/content/www/us/en/content-details/671488/intel-64-and-ia-32-architectures-optimization-reference-manual-volume-1.html): Left to Start here and sections 2, 4, 7, 13 and 14: each lists the chapter it needs, and the terms its counters use belong with the measurement section.
- r3.8 [Intel Optimization Reference Manual Volume 2: Throughput and Latency](https://www.intel.com/content/www/us/en/content-details/787036/intel-64-and-ia-32-architectures-optimization-reference-manual-documentation-volume-2-earlier-generations-of-intel-64-and-ia-32-processor-architectures-throughput-and-latency.html): No rule failed; considered as the vendor's own x86 latency tables, but they stop at Skylake, and the current vendor figures live in the Intrinsics Guide in section 7.
- r3.9 [uarch-bench](https://github.com/travisdowns/uarch-bench): Left to sections 2 and 15: listed there for what it measures and for its ready-made tests, and a third listing as the harness behind the speed-limit post named no mechanism.
- r3.10 [A Whirlwind Introduction to Dataflow Graphs](https://fgiesen.wordpress.com/2018/03/05/a-whirlwind-introduction-to-dataflow-graphs/): Rule 1: an exposition of critical-path reasoning with no measurement and no new mechanism; the decoder report, whose home is section 2, carries the method with numbers.
- r3.11 [Apple Silicon CPU Optimization Guide](https://developer.apple.com/documentation/apple-silicon/cpu-optimization-guide): Rule 4 in spirit: the URL is live but the document needs a registered developer account and an extra agreement, so neither a reader nor the link check can reach the content.
- r3.12 Intel Security Issue Update: Initial Performance Data Results for Client Systems: Rule 4: the vendor's numbered disclosure now lands on a 404 redirector; rule 3 as well, since the SYSmark percentages carried no compiler, flags or run statistics.
- r3.13 Combining Branch Predictors, WRL TN-36: Rule 4: the HP Labs technical report host no longer resolves and no author copy exists; only third-party mirrors remain.
- r3.14 [The Microarchitecture of Superscalar Processors (Proc. IEEE 1995)](https://ieeexplore.ieee.org/document/476078): Rule 1: a tutorial survey of the field rather than a design paper; its content is carried by the R10000 paper in the section and the reservation-station paper in section 2.
- r3.15 [Disabling Zen 5's Op Cache and Exploring its Clustered Decoder](https://chipsandcheese.com/p/disabling-zen-5s-op-cache-and-exploring): Rule 3: throughput figures given without compiler, flags, frequency state or run statistics; useful pointers, not admissible evidence.
- r3.16 [InstLatX64 instruction latency dumps](http://instlatx64.atw.hu/): Rule 3: raw tool output with no stated method, frequency state or run count, and rule 1: no accompanying report of how the numbers were obtained.
- r3.17 [A First-Order Superscalar Processor Model](https://dl.acm.org/doi/10.1145/1028176.1006729): Rule 4 in spirit: the origin of interval analysis, but the ACM page is paywalled and no author copy was found, so the journal form with an author copy is listed instead.
- r3.18 [Two-Level Adaptive Training Branch Prediction](https://hps.ece.utexas.edu/pub/yeh_micro24.pdf): No rule failed; cut at the cap because the TAGE paper restates the two-level scheme it introduced.
- r3.19 [TAGE-SC-L Branch Predictors Again (CBP-5)](https://jilp.org/cbp2016/paper/AndreSeznecLimited.pdf): No rule failed; cut at the cap because the original TAGE paper already defines the structure and this adds the corrector and loop predictor.
- r3.20 [Half&Half: Demystifying Intel's Directional Branch Predictors](https://halfandhalf.cpusec.org/): No rule failed; cut at the cap because the measured predictor behaviour a tuner needs is already in the measured x86 microarchitecture guide the section lists.
- r3.21 [Meltdown](https://arxiv.org/abs/1801.01207): No rule failed; cut at the cap because the Intel mitigations document carries the variant and its controls, and the page-table cost belongs to the OS section.
- r3.22 [Hardware vulnerabilities (Linux admin guide)](https://docs.kernel.org/admin-guide/hw-vuln/index.html): No rule failed; cut here because it belongs to the kernel boundary section, which owns the mitigation switches.
- r3.23 [Affected Processors: Transient Execution Attacks by CPU](https://www.intel.com/content/www/us/en/developer/topic-technology/software-security-guidance/processors-affected-consolidated-product-cpu-model.html): No rule failed; cut at the cap because it lists applicability per model and says nothing about cost.
- r3.24 [nanoBench](https://github.com/andreas-abel/nanoBench): No rule failed; cut here because section 15 lists the harness and its paper, and the uops.info site in the section carries its output.
- r3.25 [Limits of Control Flow on Parallelism](https://dl.acm.org/doi/10.1145/139669.139702): Trimmed for length: the Limits of Instruction-Level Parallelism report beside it already carries the realistic-prediction limit on ILP, with speculation past unresolved branches as one of its models.
- r3.26 [Hyper-Threading Technology Architecture and Microarchitecture](https://www.intel.com/content/dam/www/public/us/en/documents/research/2002-vol06-iss-1-intel-technology-journal.pdf): Trimmed for length: the SMT origin paper beside it carries the shared-issue mechanism and its per-thread cost, and the duplicated, partitioned and shared structures of the shipped x86 core are in the Intel Optimization Reference Manual under Start here.
- r3.27 [Dynamic Branch Prediction with Perceptrons](https://www.cs.utexas.edu/~lin/papers/hpca01.pdf): Trimmed for length: the TAGE paper beside it carries the long global-history predictor the section needs, and shipped designs converge on the tagged geometric scheme rather than the perceptron.
- r3.28 [Software Techniques for Managing Speculation on AMD Processors](https://docs.amd.com/v/u/en-US/software-techniques-for-managing-speculation): Trimmed for length: the Intel mitigations document beside it defines the same speculation controls and their cost, and the per-vendor switches, including the AMD LFENCE case, belong to the kernel boundary section.
- r3.29 [Entropy Decoding in Oodle Data: x86-64 3-Stream Huffman Decoders](https://fgiesen.wordpress.com/2022/09/05/entropy-decoding-in-oodle-data-x86-64-3-stream-huffman-decoders/): Trimmed for length: its home is section 2, Execute, which lists it for the predicted against measured cycle count on a real decoder loop.
- r3.30 [CPU Performance Scaling](https://docs.kernel.org/admin-guide/pm/cpufreq.html): Trimmed for length: its home is sections 5 and 11, which own the P-states, governors and boost that set the clock a core runs at during a measurement.

### Numbers examined

- c3.1 "3.51 cycles per iteration", "4.01 cycles", "2.98 cycles" and "1.00 cycles per iteration" from https://travisdowns.github.io/blog/2019/06/11/speed-limits.html: fields present: 1 CPU "my Intel laptop", model and microarchitecture not stated, 2 cores one thread, 3 freq/turbo/SMT not stated, 4 compiler/flags not stated, 5 workload the cpp/sum-halves, cpp/mul-4 and cpp/mul-chain tests and the split-product loop in uarch-bench, 6 baseline the stated bound for each loop, 7 method uarch-bench with perf counters, run count and statistic not stated. Missing: 1 model, 3, 4, 7 run count and statistic. Verdict: watchlist for the figures; the entry rests on the bounds, which are checked against the vendor tables rather than on these runs. Verdict: watchlist.
- c3.2 "predictions are usually within 1% of measurement results" from https://arxiv.org/abs/2107.14210: fields present: 1 CPU one named part per microarchitecture from Sandy Bridge to Rocket Lake (Table 2), 2 cores all but one core disabled, 3 freq/turbo/SMT core-cycle counts from counters, turbo and SMT state not stated, 4 compiler/flags not applicable (BHive basic blocks), 5 workload BHive basic-block suite in the paper's corrected and original forms, 6 baseline IACA, llvm-mca, OSACA, CQA, Ithemal, DiffTune and an analytical bound, 7 method nanoBench, 100 runs, top and bottom trimmed, median. Missing: turbo and SMT state. Verdict: core; the figure stays out of the annotation. Verdict: core.
- c3.3 "we consistently got parallelism between 4 and 10 for most of the programs in our test suite" from https://davidwall.info/papers/WRL-TR-93.6.pdf: a trace-driven simulation study, so fields 1 to 3 do not apply; 4 compiler/flags stated per benchmark in the report, 5 workload 18 programs, 6 baseline the perfect-everything model against 375 constrained models, 7 method instruction-trace simulation with the parallelism metric defined in the text. Missing: hardware fields, by nature of the study. Verdict: core as a design-limit result; no hardware number is claimed from it. The "rarely exceeds 7" sentence is the abstract of the 1991 technical note (ASPLOS, DOI 10.1145/106972.106991), not of this report. Verdict: core.
- c3.4 "an 8-thread, 8-issue simultaneous multithreaded processor sustains over 5 instructions per cycle, while a single-threaded processor can sustain fewer than 1.5 instructions per cycle with similar resources" from https://cseweb.ucsd.edu/~tullsen/isca95.pdf: a simulation study, so fields 1 to 3 do not apply; 4 compiler the Multiflow trace-scheduling compiler generating Alpha code, flags not stated, 5 workload SPEC92 programs run as independent threads, 6 baseline a single-threaded superscalar and a fine-grain multithreaded processor with the same issue width and units, 7 method simulation of a machine derived from the Alpha 21164, one run per configuration. Missing: hardware fields, by nature of the study. Verdict: core as a design result; the entry rests on the mechanism and quotes no figure. Verdict: core.
- c3.5 "performance gains of up to 30% on common server application benchmarks" and "21% in the cases of the single and dual-processor systems" (OLTP) from https://www.intel.com/content/dam/www/public/us/en/documents/research/2002-vol06-iss-1-intel-technology-journal.pdf: fields present: 1 CPU Intel Xeon processor MP, model not stated, 2 cores one, two and four processors, 3 freq not stated, SMT enabled against disabled, 4 compiler/flags not stated, 5 workload an OLTP run and unnamed server benchmarks, 6 baseline the same system with the technology disabled, 7 method not stated. Missing: 1 model, 3 freq, 4, 5 benchmark names, 7. Verdict: cut for every figure; the entry rests on the duplicated, partitioned and shared resource description alone. Verdict: cut.
- c3.6 "differs from detailed simulation of a 4-wide out-of-order processor by an average of 7%" from https://users.elis.ugent.be/~leeckhou/papers/tocs09.pdf: a model validated against simulation, so fields 1 to 3 do not apply; 4 compiler the SimpleScalar Alpha binaries, flags not stated, 5 workload the SPEC CPU2000 integer benchmarks at SimPoint single simulation points, 6 baseline the SimpleScalar out-of-order simulator at the configurations in Table III, 7 method IPC from the model against IPC from simulation, per benchmark and averaged. Missing: hardware fields, by nature of the study. Verdict: core as a model; no hardware number is claimed from it. Verdict: core.
- c3.7 "reduces the misprediction number by nearly 16 % compared with using a single (taken/not taken) bit for each branch" from https://jilp.org/jwac-2/program/cbp3_07_seznec.pdf: trace simulation, so fields 1 to 3 do not apply; 4 not stated for the traces, 5 workload the forty CBP-3 distributed traces, 6 baseline the same predictor with a one-bit path history per branch, 7 method the CBP-3 simulation framework at the fixed storage budget. Missing: hardware fields and compiler for the traces. Verdict: core as a design result; the figure is not quoted. Verdict: core.
- c3.8 "the branch misprediction penalty (measured in clock cycles) is always larger than the time it takes to traverse the frontend pipeline length" from https://users.elis.ugent.be/~leeckhou/papers/ispass06-eyerman.pdf: simulation, so fields 1 to 3 do not apply; 4 compiler the SimpleScalar website binaries, flags not stated, 5 workload the SPEC CPU2000 integer benchmarks at 100M-instruction SimPoints, 6 baseline the five-stage front end of the modelled processor, 7 method a modified SimpleScalar/Alpha v3.0 with the penalty decomposed by interval analysis. Missing: hardware fields, by nature of the study. Verdict: core as a design result; no cycle count is quoted. Verdict: core.
- c3.9 "improves misprediction rates for the SPEC 2000 benchmarks by 10.1% over the gshare predictor" from https://www.cs.utexas.edu/~lin/papers/hpca01.pdf: simulation, so fields 1 to 3 do not apply; 4 not stated for the traces, 5 workload SPEC 2000 traces, 6 baseline gshare and bi-mode at a 4 KB budget, 7 method trace simulation with hardware budget matched. Missing: hardware fields and compiler for the traces. Verdict: core as a design result; the figure is not quoted. Verdict: core.
- c3.10 Every latency, throughput and uop count in the tables (for example "ADC 1 0.333 1") from https://dougallj.github.io/applecpu/firestorm.html: fields present: 1 CPU Apple M1 and A14 P-core (Firestorm), 2 cores one, 3 freq not stated, cycle counts from the core cycle counter, no SMT on the part, 4 compiler none, generated instruction sequences, 5 workload each instruction in isolation, in a dependency chain and unrolled, 6 baseline none, absolute per-instruction values, 7 method 64 runs of a thousand-way unrolled sequence per test with every non-zero counter published, the statistic not stated. Missing: 3 frequency, 7 statistic. Verdict: core for the method and the per-entry data; no value is quoted in the annotation. Verdict: core.
- c3.11 "benchmarks hover around 41.8 cycles per 15 bytes (i.e. around 2.79 cycles/byte)" from https://fgiesen.wordpress.com/2022/09/05/entropy-decoding-in-oodle-data-x86-64-3-stream-huffman-decoders/: fields present: 1 CPU Skylake, model not stated, 2 cores one thread, 3 freq/turbo/SMT not stated, 4 compiler hand-written assembly, 5 workload the decoder on an idealised rig and on real data, 6 baseline the issue-slot model of 38 cycles per iteration, 7 method not stated (runs, statistic). Missing: 1 model, 3, 7. Verdict: core for the analysis the entry rests on; the figure itself is watchlist-grade and is not quoted. Verdict: core.
- c3.12 "8 to 20 μs" voltage transition and "~10 μs" halted frequency transition from the post's summary table, with "~9 μs", "~11 μs" and "~650 μs" for the individual regions in the body, from https://travisdowns.github.io/blog/2020/01/17/avxfreq1.html: fields present: 1 CPU Intel Xeon W-2104 (Skylake-SP) and i7-6700HQ (Skylake) for the voltage test, 2 cores one, 3 freq licence levels 3.2, 2.8 and 2.4 GHz stated for the W-2104, which has no turbo, 2.6 GHz nominal and 3.5 GHz turbo stated for the i7-6700HQ, SMT state not stated, 4 compiler the post1 branch of the freq-bench repository, version and flags not in the post, 5 workload duty-cycle payload loops of 256-bit and 512-bit instructions, 6 baseline the non-AVX licence, 7 method rdtsc-timed sampling of counters and MSR_PERF_STATUS with plots from the repository. Missing: SMT state, compiler version and flags. Verdict: core for the mechanism and its order of magnitude, reproducible from the repository; watchlist for the exact microsecond figures. Verdict: core.
- c3.13 "a reduction in the number of writes of ~63% (to L3) and ~50% (to memory)" and "about 17% to 18% faster than storing one" from https://travisdowns.github.io/blog/2020/05/13/intel-zero-opt.html: fields present: 1 CPU Skylake client, with a desktop i7-6700 as the second chip, the main chip's model not stated, 2 cores one, 3 freq/turbo/SMT not stated, 4 compiler gcc 9.2.1 with -march=native -O3 -funroll-loops, 5 workload std::fill of zeros versus ones over buffers from 100 bytes to 100 MB, 6 baseline the non-zero fill, 7 method zero-fill-bench with l2_lines_out.silent, l2_lines_out.non_silent and uncore imc data_writes, 27 samples per size with the first 10 discarded as warmup and the median single sample plotted. Missing: 1 model of the main chip, 3. Verdict: core for the counter signature; watchlist for the percentages. Verdict: core.

## 4. Memory hierarchy

### Rejected candidates

- r4.1 [sysfs-devices-system-cpu ABI](https://www.kernel.org/doc/Documentation/ABI/testing/sysfs-devices-system-cpu): Trimmed for length: the geometry it exposes is carried by What Every Programmer Should Know About Memory, which names the same cache files under /sys and reads size, level and shared_cpu_map from them, and by the Berkeley report, which recovers the same values by measurement.
- r4.2 [Intel Ivy Bridge Cache Replacement Policy](https://blog.stuffedcow.net/2013/01/ivb-cache-replacement/): Trimmed for length: the adaptive replacement it confirms on a shipping core is the mechanism the Adaptive Insertion Policies entry establishes, so the measurement was a second source for one mechanism.
- r4.3 [Intel SDM Volume 4: Model-Specific Registers](https://www.intel.com/content/www/us/en/content-details/671098/intel-64-and-ia-32-architectures-software-developer-s-manual-volume-4-model-specific-registers.html): Trimmed for length: the prefetchers its register bits switch off are named, with what trains them, in the Intel Optimization Reference Manual entry, and the Arm control bits stay in the Neoverse V2 entry, so the Intel control register is the one fact that moved out of the list.
- r4.4 [Software Optimization Guide for the AMD Zen5 Microarchitecture](https://docs.amd.com/v/u/en-US/58455_1.00): Trimmed for length: its homes are sections 2, 3 and 14, and the victim last-level cache it documents is the design the Achieving Non-Inclusive Cache Performance entry explains, with TLB levels and prefetchers carried by the Intel and Arm manuals listed.
- r4.5 [Arm Architecture Reference Manual for A-profile architecture](https://support.arm.com/documentation/ddi0487/latest/): Trimmed for length: the memory model it states in prose is the formal model of the Simplifying ARM Concurrency entry beside it, and the manual is listed in sections 5, 7 and 13 for its PMU, vector and SME2 definitions.
- r4.6 [herdtools7](https://github.com/herd/herdtools7): Trimmed for length: its home is section 9, where litmus7 and herd7 are listed with the language-level models they check.
- r4.7 [__builtin_prefetch (GCC Other Builtins)](https://gcc.gnu.org/onlinedocs/gcc/Other-Builtins.html#index-_005f_005fbuiltin_005fprefetch): Trimmed for length: the prefetch intrinsic, its hints and the rule that an invalid pointer does not fault are in the software prefetch chapter of What Every Programmer Should Know About Memory, and the When Prefetching Works entry says when to issue one.
- r4.8 [Efficient Virtual Memory for Big Memory Servers](https://research.cs.wisc.edu/multifacet/papers/isca13_direct_segment.pdf): Trimmed for length: TLB reach as a first-order cost is carried by the TLB measurements in What Every Programmer Should Know About Memory and the miss cost the Berkeley report recovers, and the page-size remedy and its cost stay in the Transparent Hugepage Support and Ingens entries.
- r4.9 [Gallery of Processor Cache Effects](https://igoro.com/archive/gallery-of-processor-cache-effects/): Rule 1 and rule 3: an explainer that restates textbook mechanisms with C# programs on an unnamed quad-core, and three of the seven fields are missing for its figures; the strided-loop method is the Berkeley report above and the false-sharing case is the benchmark proposal below.
- r4.10 [Arm Neoverse V2 Core Software Optimization Guide](https://support.arm.com/documentation/109898/latest/): No rule failed; its store-to-load forwarding, alignment and memory-copy rules are the mechanism it is listed for in the pipeline section, and the Arm geometry and prefetcher facts are in the Technical Reference Manual entry above, so a fifth listing added nothing to the memory hierarchy.
- r4.11 [HugeTLB Pages](https://docs.kernel.org/admin-guide/mm/hugetlbpage.html): No rule failed; the reserved-pool route to huge pages, cut at the subsection cap when the layout and prefetch entries were added, and linked from the THP document listed.
- r4.12 [Reverse Engineering the Stream Prefetcher for Profit](https://ieeexplore.ieee.org/document/9229804): Size cap: a counter-measured account of one shipping L2 stream prefetcher, which would be the only measured hardware-prefetcher source in the section, held until a slot opens and its seven fields are checked.
- r4.13 [Page Tables (Linux kernel documentation)](https://docs.kernel.org/mm/page_tables.html): Size cap: the vendor-neutral statement of the page-table hierarchy and the walk on a TLB miss; the SDM chapter listed fixes the same walk for x86 and the subsection is at the cap.
- r4.14 [A Low-Overhead Coherence Solution for Multiprocessors with Private Cache Memories](https://dl.acm.org/doi/10.1145/800015.808204): Size cap: the origin of the MESI protocol that the barriers paper derives and the SOSP paper measures state by state; first candidate if a slot opens in the ordering subsection.
- r4.15 [Cache-Oblivious Algorithms](https://ieeexplore.ieee.org/document/814600): Size cap: the algorithm family that reaches the cache bound without knowing line size, cache size or page size, the direct answer to the preamble; the single-thread section's tiling entry assumes those parameters are known.
- r4.16 [Microarchitecture Optimizations for Exploiting Memory-Level Parallelism](https://ieeexplore.ieee.org/document/1310765): Size cap: defines memory-level parallelism as a quantity and measures what limits it; the lockup-free cache paper listed is the origin of the mechanism and the benchmark section's uarch-bench measures it.
- r4.17 [Disclosure of H/W prefetcher control on some Intel processors](https://www.intel.com/content/www/us/en/developer/articles/technical/disclosure-of-hw-prefetcher-control-on-some-intel-processors.html): Rule 4: the URL now goes through Intel's not-found redirector to the Developer Zone home page and only third-party mirrors survive; the same register is defined in SDM Volume 4, itself trimmed for length above.
- r4.18 [What every programmer should know about memory, Part 1 (LWN serialisation)](https://lwn.net/Articles/250967/): Rule 4: a serialised copy split across articles; the author's single canonical PDF is listed instead.
- r4.19 [cpumemory.pdf mirror on people.freebsd.org](https://people.freebsd.org/~lstewart/articles/cpumemory.pdf): Rule 4: a third-party mirror of a document with an official home.
- r4.20 [AMD Zen5 Software Optimization Guide on Scribd](https://www.scribd.com/document/830344700/58455-1-00): Rule 4: a mirror; the vendor's own document portal serves the same publication and is listed instead.
- r4.21 [x86-TSO in the ACM Digital Library](https://dl.acm.org/doi/10.1145/1785414.1785443): Rule 4: paywalled and returns 403 to automated clients; the authors' own copy is listed instead.
- r4.22 [Simplifying ARM Concurrency in the ACM Digital Library](https://dl.acm.org/doi/10.1145/3158107): Rule 4: returns 403 to automated clients; the authors' project page, which carries the paper, its errata and the executable model, is listed instead.
- r4.23 [Weak vs. Strong Memory Models](https://preshing.com/20120930/weak-vs-strong-memory-models/): Rule 1: an explanatory blog post, not an original report of a mechanism or a measurement by the person who did the work.
- r4.24 [Data-Oriented Design (online book)](https://www.dataorienteddesign.com/dodbook/): Rule 1: exposition without measurements; the PLDI paper listed above made the same case with measurements and is the original source.
- r4.25 [Linux kernel memory-barriers.txt](https://www.kernel.org/doc/Documentation/memory-barriers.txt): No rule failed; it documents the kernel's barrier API rather than the hardware, so it belongs to the concurrency section.
- r4.26 [The microarchitecture of Intel, AMD, and VIA CPUs](https://www.agner.org/optimize/microarchitecture.pdf): No rule failed; its measured store-forwarding stalls were considered here, but the document is listed in the microarchitecture section and the vendor rules above cover the mechanism.
- r4.27 [Instruction tables](https://www.agner.org/optimize/instruction_tables.pdf): No rule failed; its measured latency and throughput of MFENCE, LFENCE and locked read-modify-write per microarchitecture is the cost side of x86-TSO, but the tables are listed in the instruction and microarchitecture sections and are not repeated here.
- r4.28 [Practical, Transparent Operating System Support for Superpages](https://www.usenix.org/legacy/events/osdi02/tech/navarro.html): No rule failed; the original OS superpage design, cut for size because the kernel documentation and the THP measurement paper cover what a user can act on.

### Numbers examined

- c4.1 "4 cycles" for a first-level hit and "14 cycles" for a second-level hit from https://www.akkadia.org/drepper/cpumemory.pdf (Figure 3.10): fields present: 1 CPU Pentium 4, NetBurst, 64-bit mode; 5 workload pointer-chase list walk over a working set of the stated size; 7 method cycles per element from repeated list traversal. Missing: 2 cores used, 3 frequency and SMT state, 4 compiler and flags, 6 baseline. Verdict: watchlist; the entry stays for its method, the figures are not quoted. Verdict: watchlist.
- c4.2 "80 and 78 ms" for full versus strided loops and "4.3 seconds" versus "0.28 seconds" for false sharing from https://igoro.com/archive/gallery-of-processor-cache-effects/: fields present: 2 quad-core; 4 C# release build, JIT output inspected; 5 array walk and per-thread counter increments; 6 the unpadded variant. Missing: 1 CPU model and microarchitecture, 3 frequency, turbo and SMT state, 7 runs and statistic. Verdict: cut; the entry is under Rejected and the false-sharing case is reproduced by the benchmark proposal below. Verdict: cut.
- c4.3 "as much as 10% of execution cycles on TLB misses, even using large pages" from https://research.cs.wisc.edu/multifacet/papers/isca13_direct_segment.pdf: fields present: 1 dual-socket Intel Xeon E5-2430, Sandy Bridge; 2 six cores per socket, two threads per core; 3 2.2 GHz; 5 graph500, memcached, MySQL, NPB and GUPS; 6 the paged baseline on the same machine; 7 hardware performance counters over the run. Missing: 4 compiler and flags, and whether turbo and SMT were enabled during measurement. Verdict: cut for length; the entry is under Rejected and the percentage is not quoted. Verdict: cut.
- c4.4 Coherence latencies in Table 2 (for example "81" cycles for a load of a Modified line on the same die) from https://sigops.org/s/conferences/sosp/2013/papers/p33-david.pdf: fields present: 1 AMD Opteron 6172, Intel Xeon E7-8867L, UltraSPARC T2, Tilera TILE-Gx36; 2 48, 80, 8 and 36 cores; 3 2.1, 2.13, 1.2 and 1.2 GHz, SMT off on the Xeon (the paper's words: "no hyper-threading"); 5 ccbench load, store and atomic on a line in a given MESI state and distance; 6 same-die access on the same machine; 7 average of 10000 repetitions with under 3% standard deviation. Missing: 4 compiler and flags, turbo state. Verdict: watchlist; the entry stays for the method and the state-by-distance shape, the cycle counts are not quoted. Verdict: watchlist.
- c4.5 "Ivy Bridge L3 uses an adaptive replacement policy" from https://blog.stuffedcow.net/2013/01/ivb-cache-replacement/: fields present: 1 third-generation Core i5 (Ivy Bridge) against Core i5-2500K (Sandy Bridge); 2 one core; 4 -O2 as stated in the comments; 5 random cyclic pointer chase on 2 MB pages; 6 the Sandy Bridge curve; 7 latency per access across working-set sizes. Missing: exact Ivy Bridge model, 3 frequency and turbo state for it, run count. Verdict: cut for length; the qualitative finding rests on the shape of the curve and the code is published, the entry is under Rejected and no cycle count is quoted. Verdict: cut.

## 5. Measurement

### Rejected candidates

- r5.1 [Arm Neoverse V2 PMU Guide](https://support.arm.com/documentation/109709/latest/): Size cap: the architectural events are specified in the architecture manual, listed in sections 4, 7 and 13, and the per-core event definitions for every Neoverse core are carried by the telemetry repository in the models section, so its slot went to the perf SPE document, which describes what a reader runs.
- r5.2 [Flame Graphs](https://www.brendangregg.com/flamegraphs.html): No rule failed; the paper records the design decisions and names the repository that defines the folded format, so the author's catalogue page was a third line on one visualisation, and its slot went to Coz.
- r5.3 [bcc](https://github.com/iovisor/bcc): No rule failed; bpftrace carries the one-line experiment and the BPF book documents what each bcc tool attaches to, so the tool source was a third line on one method and is cut for depth.
- r5.4 [nanoBench](https://github.com/andreas-abel/nanoBench): Duplicate: listed in the benchmarks section beside its paper, whose reasons cover the kernel-mode, interrupts-off method, and no reason specific to measurement pitfalls remained, since it is a harness that does not lie.
- r5.5 [Performance Ninja](https://github.com/dendibakh/perf-ninja): No rule failed; a lab course teaches practice rather than establishing, defining or measuring a mechanism, so no reason of the entry form justified it, and its slot went to the clock and machine-state sources.
- r5.6 [How to Benchmark Code Execution Times on Intel IA-32 and IA-64 Instruction Set Architectures](https://www.intel.com/content/dam/www/public/us/en/documents/white-papers/ia-32-ia-64-benchmark-code-execution-paper.pdf): Rule 4: the intel.com URL now returns a permanent redirect to Intel's not-found handler and only third-party mirrors (a GitHub pdfs collection) remain; mirrors are not admissible, so the rdtsc method has no canonical home to link, and what a timer counts is carried by clock_gettime(2) and the TSC chapter named in the manual entry.
- r5.7 [How to get consistent results when benchmarking on Linux?](https://easyperf.net/blog/2019/08/02/Perf-measurement-environment-on-Linux): Rule 1: a personal blog post that compiles settings from other sources (the Systems Performance book, LLVM's benchmarking notes, the Intel paper above) rather than reporting an original mechanism or measurement; the LLVM page it draws on is listed instead, and the same author's "Code alignment issues" is the first-hand measurement and is held on the watchlist below.
- r5.8 [Code alignment issues](https://easyperf.net/blog/2018/01/18/Code_alignment_issues): Rule 3, watchlist: a first-hand reproduction of layout bias on Skylake with the clang alignment options that expose it, pending the frequency, SMT state and run count its figures lack; the mechanism itself is carried by the Producing Wrong Data paper in the section.
- r5.9 [C2C - False Sharing Detection in Linux Perf](https://joemario.github.io/blog/2016/09/01/c2c-blog/): Size cap: an implementer's own report with measurements that would otherwise qualify, but the subsection is at the hard cap and the perf-c2c man page sits in the kernel tree directory that the perf-arm-spe(1) entry links into; listed in the memory hierarchy section, where the line size is fixed.
- r5.10 [Parca](https://github.com/parca-dev/parca): Size cap: the continuous profiling design document it descends from is listed; no primary measurement compares the open source continuous profilers, so picking one repository would be preference rather than evidence.
- r5.11 [Arm Statistical Profiling Extension: Performance Analysis Methodology](https://support.arm.com/documentation/109429/latest/): Size cap: the SPE record format is specified in the architecture manual, listed in sections 4, 7 and 13, and the perf SPE document carries the filters and record fields as perf exposes them; this is methodology on top of both and is the next Arm candidate.
- r5.12 [Arm Neoverse V2 Core Telemetry Specification](https://support.arm.com/documentation/109528/latest/): Not a rule failure: it defines top-down metric groups for the core, which is the models section's territory rather than measurement; passed to that owner.
- r5.13 [Coz repository](https://github.com/plasma-umass/coz): Size cap: the Coz paper carries its claim and links the code, and the Stabilizer paper, trimmed below, links its own; a second line each would spend two slots on the same evidence.
- r5.14 [intel_pstate CPU Performance Scaling Driver](https://docs.kernel.org/admin-guide/pm/intel_pstate.html): Size cap: the two driver documents define no_turbo, HWP and the AMD boost interface, but the cpufreq document, listed in sections 3 and 11, links both and defines the governor and boost interfaces they implement.
- r5.15 AMD64 Architecture Programmer's Manual Volume 2, IBS chapter: Size cap: it defines the IBS MSRs architecturally, but the PPR entry carries the per-core event and IBS register tables and the original IBS paper carries the mechanism.
- r5.16 ACM Digital Library copies of the three papers listed: Rule 4: paywalled and return 403 to every automated client; the authors' own copies (a research group page, an author's university page, arXiv) are listed instead.
- r5.17 IEEE Xplore copy of the ISPASS counter paper: Rule 4: paywalled; the first author's project page carries the PDF and is linked instead.
- r5.18 Course mirror of the ASPLOS paper at `users.cs.northwestern.edu/~robby/courses/322-2013-spring/mytkowicz-wrong-data.pdf`: Rule 4: a third-party mirror; used only to read the paper text for the Claims block, the co-author's group page is the entry.
- r5.19 perf design notes in the kernel tree: Size cap: the same counting and sampling split is defined in the perf_event_open(2) man page, which is the maintained interface description and is listed instead.
- r5.20 [Systems Performance: Enterprise and the Cloud, 2nd Edition](https://www.brendangregg.com/systems-performance-2nd-edition-book.html): Trimmed for length: listed in Start here, and the USE Method page carries the checklist the book works through each Linux resource.
- r5.21 [Arm Architecture Reference Manual for A-profile architecture](https://support.arm.com/documentation/ddi0487/latest/): Trimmed for length: one home per URL, and its homes are sections 4, 7 and 13. Arm sampling in this section is carried by perf-arm-spe(1), and the per-core Neoverse events by the telemetry sources in section 6.
- r5.22 [Intel perfmon](https://github.com/intel/perfmon): Trimmed for length: the SDM entry defines the architectural counters and PEBS, and section 6 links the same repository for its TMA metric release, one directory from the event files.
- r5.23 [perf man pages in the kernel tree](https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/tree/tools/perf/Documentation): Trimmed for length: the perf wiki tutorial carries the stat, record and annotate walk, and perf-arm-spe(1) links into the same Documentation directory, one level up from perf-stat, perf-record, perf-mem and perf-c2c.
- r5.24 [FlameGraph](https://github.com/brendangregg/FlameGraph): Trimmed for length: The Flame Graph paper records the design and names the repository, whose stack collapsers define the folded format.
- r5.25 [CPU Performance Scaling](https://docs.kernel.org/admin-guide/pm/cpufreq.html): Trimmed for length: one home per URL, with entries in sections 3 and 11, and the LLVM benchmarking page sets the governor and boost switch in this section.
- r5.26 [Stabilizer: Statistically Sound Performance Evaluation](https://people.cs.umass.edu/~emery/pubs/stabilizer-asplos13.pdf): Trimmed for length: Producing Wrong Data carries layout bias, and the remedy, re-randomising layout each run so a significance test is valid, is recorded under Claims below with its figures. Section 15 also declined it as listed in section 5, so it now has no entry in the README.

### Numbers examined

- c5.1 "frequently by about 33% and once by almost 300%" (run time change from UNIX environment size alone) from https://sape.inf.usi.ch/publications/asplos09: fields present: 1 CPU Intel Core 2 workstation, Core microarchitecture (also Pentium 4, NetBurst, and the m5 O3CPU simulator), 2 cores not stated (single-threaded SPEC programs), 3 freq 2.4 GHz, turbo and SMT not stated (neither exists on that part), 4 compiler gcc 4.1.3 and icc 10.1, -O2 against -O3, 5 workload SPEC CPU2006 C programs, perlbench for the figure, 6 baseline empty environment and default link order, 7 method cycles via papi-3.5.1 and perfmon-2.8, mean of five runs with a 95% confidence interval. Missing: cores used, explicit turbo and SMT statement. Verdict: watchlist for the figures; the entry stays on the mechanism and quotes none of them. Verdict: watchlist.
- c5.2 "the effect of -O3 versus -O2 optimizations is indistinguishable from random noise" (F-test p-value 0.264; -O2 over -O1 p-value 0.0898) from https://people.cs.umass.edu/~emery/pubs/stabilizer-asplos13.pdf: fields present: 1 CPU Intel Core i3-550, dual-core, Westmere (Clarkdale), 2 cores used not stated (single-threaded SPEC programs), 3 freq 3.2 GHz, turbo and SMT state not stated, 4 compiler gcc 4.6.3 front end with dragonegg and LLVM 3.1, -O1 against -O2 against -O3, 5 workload eighteen SPEC CPU2006 C benchmarks, 6 baseline -O2 for the -O3 test and -O1 for the -O2 test, 7 method thirty runs per configuration with layout re-randomised each run, one-way analysis of variance within subjects. Missing: cores used, turbo and SMT state. Verdict: watchlist for the figures; the entry is trimmed for length and the mechanism is carried by the measurement-bias paper. Verdict: watchlist.
- c5.3 "Most x86_64 events are incremented an extra time for every hardware interrupt that occurs" from https://web.eece.maine.edu/~vweaver/projects/deterministic/ispass2013_deterministic.pdf: fields present: 1 CPU eleven x86-64 parts named by microarchitecture and model in the paper's Table III (Atom 230, Core2 X5355, Nehalem X5570, Nehalem-EX X7550, Westmere-EX, SandyBridge-EP, IvyBridge, Pentium D, Phenom, Istanbul, Bobcat), 2 cores one, a single-threaded assembly benchmark, 3 freq, turbo and SMT state not stated, 4 compiler none, hand-written assembly with an instruction count fixed by inspection, 5 workload the retired-instructions assembly microbenchmark, 6 baseline the expected count from code inspection, 7 method ten runs per machine of user-only counts through perf stat or pfmon, compared with the expected total. Missing: frequency, turbo and SMT state. Verdict: core for the mechanism, which is one extra count per interrupt and per fault; the entry quotes no figure. Verdict: core.
- c5.4 "perf variations of less than 0.1%" from https://llvm.org/docs/Benchmarking.html: fields present: 4 flags static linking, 7 method perf stat with ten repetitions inside a cpuset shield, governor set to performance, SMT siblings offline, address space randomisation off, intel_pstate no_turbo set, program and data on tmpfs. Missing: CPU, cores, frequency, workload and baseline. Verdict: the entry rests on the recipe, not on the figure, which is quoted nowhere. Verdict: other.

## 6. Models

### Rejected candidates

- r6.1 [The Roofline Model: Visualizing and Optimizing Performance](https://amcr.lbl.gov/departments/computer-science-department/ppan/roofline-performance-model/): Rule 1 borderline and size cap: a link hub, not a document. Its one substantive list, the below-roof causes paired with fixes, restates the ceilings the listed roofline article defines, and the toolkit, LIKWID and Advisor it links are listed on their own.
- r6.2 [Top-Down Analysis (perf wiki)](https://perfwiki.github.io/main/top-down-analysis/): Size cap: the perf maintainers' page on the TopdownL1 metric groups for Intel cores, written for Linux 6.1. The listed pmu-tools carries the Linux path for Intel and the drill to perf record, and the maintainers' tutorial in the measurement section carries the stat to record walk. Cut to make room for the AMD formulas and the Arm worked case.
- r6.3 [Pipeline Utilization (AMD uProf User Guide)](https://docs.amd.com/r/en-US/57368-uProf-user-guide/Pipeline-Utilization): Size cap: the vendor's prose definition of the same level-one and level-two categories, with the AMDuProfPcm option that reports them, one page from the uProf guide listed in the measurement section. The listed kernel file carries the formulas in events. Next candidate for the top-down subsection.
- r6.4 [Measuring Parallel Processor Performance](https://cacm.acm.org/research/measuring-parallel-processor-performance/): Rule 4: the experimentally determined serial fraction that links a measured speedup curve back to Amdahl's input, but the publisher page carries only the abstract, the DL copy is paywalled and no author copy was found. Watchlist until an open copy appears.
- r6.5 [Roofline: An Insightful Visual Performance Model for Floating-Point Programs and Multicore Architectures, Berkeley technical report EECS-2008-134](https://www2.eecs.berkeley.edu/Pubs/TechRpts/2008/EECS-2008-134.html): Rule 4: the earlier long draft of the CACM article; one canonical copy per source, and the published article is free at the publisher. Same verdict as the Start here section.
- r6.6 [Roofline, LBNL eScholarship copy](https://escholarship.org/uc/item/78h8v7mr): Rule 4: an institutional repository mirror of the CACM article, which is free at the publisher.
- r6.7 [Empirical Roofline Tool (ERT) page at LBNL](https://amcr.lbl.gov/departments/computer-science-department/ppan/roofline-performance-model/empirical-roofline-tool-ert/): Size cap and rule 4: a sub-page of the roofline hub, itself in Rejected, describing the toolkit whose canonical home is the listed repository.
- r6.8 [Introducing a Performance Model for Bandwidth-Limited Loop Kernels](https://arxiv.org/abs/0905.0792): Size cap: the origin of the ECM model, but its overlap rules and notation were only fixed in the listed stencil paper, which is the statement of the model a reader should apply.
- r6.9 [Exploring performance and power properties of modern multicore chips via simple machine models](https://arxiv.org/abs/1208.2908): Size cap: the refinement of ECM to multicore scaling, where the saturation core count originates, restated with the overlap rules in the listed stencil paper, so one entry carries the model.
- r6.10 [Bridging the Architecture Gap: Abstracting Performance-Relevant Properties of Modern Server Processors](https://arxiv.org/abs/1907.00048): Size cap: the refinement of ECM to AMD, IBM and Marvell server cores with a validation on a full solver; first candidate if a roofline slot opens, since it is the only ECM paper past Intel.
- r6.11 [Kerncraft](https://github.com/RRZE-HPC/kerncraft): Size cap: the creators' implementation of both roofline and ECM predictions from loop source; the listed paper carries the model and links the tool.
- r6.12 [The CARM Tool](https://github.com/champ-hub/carm-roofline): Size cap: the model authors' own benchmarking and plotting implementation; the Advisor entry already shows a cache-level roofline in a shipped tool, and the paper carries the definition.
- r6.13 [Applying the Roofline Model, IEEE Xplore](https://ieeexplore.ieee.org/document/6844463): Rule 4: the paywalled copy of the listed paper, which the authors serve from their own Spiral project page.
- r6.14 [A Top-Down Method for Performance Analysis and Counters Architecture, IEEE Xplore](https://ieeexplore.ieee.org/document/6844459): Rule 4: paywalled copy where the author's own page carries the PDF.
- r6.15 [Top-down Microarchitecture Analysis Method, VTune cookbook](https://www.intel.com/content/www/us/en/docs/vtune-profiler/cookbook/2025-0/top-down-microarchitecture-analysis-method.html): Rule 1 and size cap: a tool recipe that restates the method whose definitions are the listed paper and the listed perfmon release; a reader with VTune reaches it from the perfmon README.
- r6.16 [perf topdown.txt in the kernel tree](https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/tree/tools/perf/Documentation/topdown.txt): Size cap: documents the fixed SLOTS counter and the RDPMC path for reading TMA level one from user space on Intel cores; the listed pmu-tools covers the interface a newcomer uses first.
- r6.17 [Arm Neoverse V2 Core Telemetry Specification](https://support.arm.com/documentation/109528/latest/): Size cap: the per-core metric definitions, passed from the measurement section; the listed telemetry-solution repository carries the same per-core JSON for every supported core, so one entry covers all of them.
- r6.18 [perf-tools](https://github.com/aayasin/perf-tools): Size cap: the method author's own scripts layered on toplev and perf; the listed pmu-tools carries the interface.
- r6.19 [How to Quantify Scalability](https://www.perfdynamics.com/Manifesto/USLscalability.html): Size cap: the author's synopsis of the universal scalability law, written as a supplement to chapters of the listed book by the same author.
- r6.20 [Guerrilla Capacity Planning, Springer page](https://link.springer.com/book/10.1007/978-3-540-31010-5): Rule 4: the publisher's paywalled page where the author's own book page exists with the contents and corrigenda.
- r6.21 [usl on CRAN](https://cran.r-project.org/web/packages/usl/index.html): Rule 4: the distribution page; the repository is the canonical home for software.
- r6.22 [Little's Law as Viewed on Its 50th Anniversary](https://pubsonline.informs.org/doi/10.1287/opre.1110.0940): Size cap: the same author's retrospective on the listed proof; paywalled like the original, with no author copy found.
- r6.23 [The Operational Analysis of Queueing Network Models](https://dl.acm.org/doi/10.1145/356733.356735): Rule 1 borderline (published in a surveys journal) and rule 4 (paywalled, 403 to every client): the operational laws it states are in the listed free book by their co-developers.
- r6.24 Course mirror of the Amdahl paper at `www3.cs.stonybrook.edu/~rezaul/Spring-2012/CSE613/reading/Amdahl-1967.pdf`: Rule 4: a third-party mirror of a paper whose publisher copy is open.
- r6.25 ResearchGate and Semantic Scholar copies of the cache-aware roofline paper and of the Amdahl paper: Rule 4: aggregator mirrors.
- r6.26 [CS Roofline Toolkit](https://bitbucket.org/berkeleylab/cs-roofline-toolkit/): Trimmed for length: the listed LIKWID entry carries the measured roofs, with likwid-bench in place of the toolkit's microkernels, and the LBNL roofline hub under Rejected still links the toolkit.
- r6.27 [Analyze CPU Roofline](https://www.intel.com/content/www/us/en/docs/advisor/user-guide/2025-1/analyze-cpu-roofline.html): Trimmed for length: the listed cache-aware roofline paper defines the roof per cache level the tool draws, and the listed LIKWID entry carries the measured point and roofs on Linux.
- r6.28 [Arm Telemetry Solution](https://gitlab.arm.com/telemetry-solution/telemetry-solution): Trimmed for length: the listed Topdown Methodology Specification defines the stall accounting that topdown-tool and the per-core telemetry JSON implement, so the repository is reached from the method rather than listed beside it.
- r6.29 [Arm Neoverse V1 Performance Analysis Methodology](https://support.arm.com/documentation/109199/latest/): Trimmed for length: a worked case of the method the listed Topdown Methodology Specification defines, from the first-stage metrics to the code change, and the first candidate if a top-down slot opens.
- r6.30 [usl](https://github.com/smoeding/usl): Trimmed for length: the listed Guerrilla Capacity Planning states the fitting procedure the package implements by regression, and the book page links the software.

### Numbers examined

- c6.1 "1021 for beam stress analysis using conjugate gradients, 1020 for baffled surface wave simulation using explicit finite differences, and 1016 for unstable fluid flow using flux-corrected transport" (speedups on a 1024-processor hypercube) from http://www.johngustafson.net/pubs/pub13/amdahl.htm: fields present: 1 CPU a 1024-processor hypercube, model and node processor not named on the page, 2 cores 1024 processors, 3 freq, turbo and SMT not stated, 4 compiler and flags not stated, 5 workload the three named simulations, 6 baseline a single processor on the scaled problem, inferred from the serial and parallel fractions rather than run, 7 method not stated on the page, deferred to a Sandia report. Missing: 1 in part, 3, 4, 6 in part, 7. Verdict: core for the definition of scaled speedup, on which the reason rests, with the figures kept out of the annotation. Verdict: core.
- c6.2 "curve fitness above 90%" from https://ieeexplore.ieee.org/document/6506838: fields present: none on the abstract page, and the body is paywalled. Missing: all seven. Verdict: core for the definition of the cache-level roofs, watchlist for the validation claim, which is not quoted. Verdict: core.
- c6.3 "a predicted single-core performance of 2.1 Gflop/s at nominal clock speed and it will saturate at three cores, while the naive, non-pipelined code with 0.9 Gflop/s will require six" from https://arxiv.org/abs/1410.5010: fields present: 1 CPU Intel Xeon E5-2680, Sandy Bridge-EP, 2 cores single-core prediction, scaling across the eight cores of one socket, 3 freq fixed at 2.7 GHz, turbo off by that statement, SMT state not stated (eight cores, sixteen threads listed), 4 compiler Intel C Compiler 13.1.3 for the stencil codes, flags not stated, the summation measured with LIKWID-bench, 5 workload vector summation in AVX and naive form, 6 baseline the model's own prediction against measurement, 7 method LIKWID-bench, run count and statistic not stated. Missing: flags, SMT state, run count and statistic. Verdict: core for the model, with the figures kept out of the annotation. Verdict: core.
- c6.4 The roofline article's kernel results for its four machines are not carried, as the Start here section already records; the entry rests on the definitions of operational intensity and the ceilings.
- c6.5 The measured roofline plots for BLAS and FFT in Applying the Roofline Model are not carried; the entry rests on the counter set and the measurement strategy, and the validation section states the machine but the plots are read, not quoted.
- c6.6 The top-down paper's SPEC CPU2006 characterisation on Ivy Bridge is not carried; the entry rests on the slot hierarchy and the counter proposal.
- c6.7 The Neoverse V1 white paper's case-study speedup for the Arrow CSV parser after the NEON rewrite is not carried; the paper now sits under Rejected, trimmed for length, and the machine, compiler and run method of the case study are stated in the paper for a reader who wants the figure.

## 7. Single-thread optimisation

### Rejected candidates

- r7.1 [Controlling Floating Point Behavior (Clang)](https://clang.llvm.org/docs/UsersManual.html#controlling-floating-point-behavior): Trimmed for length: section 8 is its home and defines -fassociative-math and -ffp-contract there, and the LLVM vectoriser entry above carries the reduction case.
- r7.2 [Optimizing software in C++](https://www.agner.org/optimize/optimizing_cpp.pdf): Trimmed for length: Start here is its home, and the LLVM vectoriser entry states what the vectoriser accepts, so what stops it.
- r7.3 [Fast Multidimensional Matrix Multiplication on CPU from Scratch](https://siboehm.com/articles/22/Fast-MMM-on-CPU): Trimmed for length: the data locality paper defines the interchange and tiling it applies, and the GEMM and BLAS entries in section 13 carry register blocking against a vendor BLAS.
- r7.4 [Filtering numbers quickly with SVE on Amazon Graviton 3 processors](https://lemire.me/blog/2022/06/23/filtering-numbers-quickly-with-sve-on-amazon-graviton-3-processors/): Trimmed for length: Introduction to SVE carries the predication model the measured compaction loop rests on.
- r7.5 [simdutf](https://github.com/simdutf/simdutf): Trimmed for length: Highway carries run-time dispatch of a kernel per ISA, and the base64 and JSON papers carry the measured byte-loop replacements.
- r7.6 [Vectorized and performance-portable Quicksort](https://arxiv.org/abs/2205.05982): Trimmed for length: Highway carries the sizeless vector model the sort is written against, and the sort ships inside the Highway repository.
- r7.7 [SIMD-friendly algorithms for substring searching](http://0x80.pl/notesen/2016-11-28-simd-strfind.html): Trimmed for length: the base64 and JSON papers carry vector compares replacing a per-byte branch, and the sorted-union post carries the branchless case.
- r7.8 [Sheng, a small but fast Deterministic Finite Automaton](https://branchfree.org/2018/05/25/say-hello-to-my-little-friend-sheng-a-small-but-fast-deterministic-finite-automaton/): Trimmed for length: the Hyperscan paper carries the automaton decomposition, and the base64 paper carries shuffle-based lookup replacing loads and branches.
- r7.9 [Software Pipelining: An Effective Scheduling Technique for VLIW Machines](https://dl.acm.org/doi/10.1145/53990.54022): No rule failed; cut on review. A VLIW scheduling paper: on the out-of-order cores in scope the hardware does that scheduling, GCC's modulo scheduler is off at every -O level and LLVM's is not on for x86 or AArch64 by default, so a reader has nothing to act on.
- r7.10 [RISC-V Vector Extension, Version 1.0](https://docs.riscv.org/reference/isa/unpriv/v-st-ext.html): Scope. The list covers x86 and Arm server parts, and the same URL sits on the watchlist pending a seven-field run on a socketed part, so it cannot also be core.
- r7.11 [Intel AVX10.2 Architecture Specification](https://www.intel.com/content/www/us/en/content-details/828965/intel-advanced-vector-extensions-10-2-intel-avx10-2-architecture-specification.html): Watchlist by the README's own rule: no shipped part carries it, and the watchlist's instruction set extensions entry already ties AVX10.2 to Diamond Rapids. Intel has withdrawn the separate AVX10.1 specification (document 784267 now redirects to the 404 page), so no shipped-subset entry replaces it.
- r7.12 [Vector Class Library](https://github.com/vectorclass/version2): No rule failed; cut under the cap. A third fixed-width wrapper, x86 only, whose manual is the reason to read it; the slot went to the measured sort paper that exercises the sizeless model.
- r7.13 [Bit Twiddling Hacks](https://graphics.stanford.edu/~seander/bithacks.html): Rule 1: a catalogue of tricks contributed by many people and credited to them, an aggregation rather than an original report; the single-author book that derives them took the slot.
- r7.14 [Improving the Ratio of Memory Operations to Floating-Point Operations in Loops](https://dl.acm.org/doi/10.1145/197320.197366): No rule failed; cut under the cap. Defines unroll-and-jam and scalar replacement, the register-tiling step the matrix multiplication walkthrough performs and benchmarks/03 measures as independent accumulators.
- r7.15 [Automatic Translation of FORTRAN Programs to Vector Form](https://dl.acm.org/doi/10.1145/29873.29875): No rule failed; cut under the cap. The origin of the dependence tests behind vectorisation legality; the data locality entry states the legality condition the section needs.
- r7.16 [Division by Invariant Integers using Multiplication](https://gmplib.org/~tege/divcnst-pldi94.pdf): No rule failed; cut under the cap. The origin of the multiply-and-shift sequence compilers emit for division by a constant; the book entry derives the same sequence and its magic numbers.
- r7.17 [Floating Point Math (GCC wiki)](https://gcc.gnu.org/wiki/FloatingPointMath): No rule failed; cut under the cap. The GCC statement of the same flags the Clang page defines, -fassociative-math and -ffp-contract among them.
- r7.18 [target_clones function attribute (GCC)](https://gcc.gnu.org/onlinedocs/gcc/Common-Attributes.html#index-target_005fclones): No rule failed; a compiler attribute, so the compilers section is its home. The compiler-defined route to one binary with several ISA clones dispatched through an ifunc; Highway's entry names the dispatch and this subsection is at the cap.
- r7.19 [Fast Random Integer Generation in an Interval](https://arxiv.org/abs/1805.10941): No rule failed; kept out for scope. The mechanism is division avoidance in rejection sampling, which sits with instruction costs rather than with the five topics of the section.
- r7.20 [GCC Optimize Options](https://gcc.gnu.org/onlinedocs/gcc/Optimize-Options.html): No rule failed; cut for depth. It names the unroll, interchange, blocking and modulo-scheduling flags but explains none of them; the LLVM vectoriser page in section 8 explains its transforms.
- r7.21 [The Converged Vector ISA: Intel Advanced Vector Extensions 10 Technical Paper](https://www.intel.com/content/www/us/en/content-details/784343/the-converged-vector-isa-intel-advanced-vector-extensions-10-technical-paper.html): No rule failed; the same watchlist condition as the specification, since no shipped part carries the versioned ISA it describes.
- r7.22 [SVE Programming Examples](https://support.arm.com/documentation/dai0548/latest/): No rule failed; cut for depth. The introduction guide teaches the model; this one is code samples that assume it.
- r7.23 [The ARM Scalable Vector Extension](https://arxiv.org/abs/1803.06185): No rule failed; cut for depth. The architects' paper is the origin of the vector-length-agnostic model, but the introduction guide holds the corpus slot for that model and the measured SVE report took the remaining one.
- r7.24 [Introducing SVE2](https://support.arm.com/documentation/102340/latest/): No rule failed; cut for depth. The architecture manual entry names SVE2, and the introduction guide teaches the scalable model that SVE2 extends.
- r7.25 [Optimizing C code with Neon intrinsics](https://support.arm.com/documentation/102467/latest/): No rule failed; cut for depth. Its interleaved-load treatment of array-of-structures data duplicates the Intel manual's guidance for another ISA.
- r7.26 [RISC-V Vector C Intrinsic API](https://github.com/riscv-non-isa/riscv-rvv-intrinsic-doc): Scope, with the ratified specification above. Redirected from the old repository name.
- r7.27 [hyperscan](https://github.com/intel/hyperscan): No rule failed. One slot per project: the paper is the named corpus item and defines the decomposition; the repository is linked from it.
- r7.28 [Validating UTF-8 In Less Than One Instruction Per Byte](https://arxiv.org/abs/2010.03090): No rule failed. One slot per project: the simdutf repository entry carries this paper and its successor in its references.
- r7.29 [SIMDization of switch statements](http://0x80.pl/notesen/2019-02-03-simd-switch-implementation.html): Rule 1: the note gives no measurement and says of its own dispatch code that little can be said about its performance, so the measured substring note took the author's slot.
- r7.30 [Mispredicted branches can multiply your running times](https://lemire.me/blog/2019/10/15/mispredicted-branches-can-multiply-your-running-times/): No rule failed; the sorted-union post shows the same effect and also shows the case where branchless loses, which a newcomer needs.
- r7.31 [fastbase64](https://github.com/lemire/fastbase64): No rule failed. One slot per project: the paper carries the method and the numbers; it links this repository.
- r7.32 [Vectorscan](https://github.com/VectorCamp/vectorscan): Rule 1: a portability fork, not the original implementation or report; the paper entry covers the mechanism.

### Numbers examined

- c7.1 "speeds up by a factor of 4.3 on a DECstation 3100, and a factor of 3.0 on an IBM RS/6000" from https://dl.acm.org/doi/10.1145/106972.106981: fields present: 1 DECstation 3100 and IBM RS/6000, named by machine, microarchitecture not stated; 2 one core; 3 frequency not stated, no turbo or SMT on either part; 4 compiler and flags not stated; 5 a 300 by 300 matrix multiplication blocked at the register and cache levels, MFLOPS by block size; 6 the unblocked matrix multiplication; 7 MFLOPS measured on the machine with miss rates from simulation, repetitions not stated. Missing: frequency; compiler and flags; repetitions and statistic. Verdict: watchlist for the number; the entry stays for the interference-miss mechanism and the block-size rule. Verdict: watchlist.
- c7.2 "4481 ms" naive against "8 ms" for the vendor BLAS, and "9 FLOPS / core / cycle" against "18" from https://siboehm.com/articles/22/Fast-MMM-on-CPU: fields present: 1 Intel Core i7-6700, which the report calls Haswell (the part is Skylake); 2 one core for every step but the last, four cores and eight threads for the OpenMP step; 3 3.40 GHz, turbo not stated, SMT implied on by the thread count but not stated; 4 Clang 14.0, -O3 -march=native -ffast-math; 5 a 1024 by 1024 single-precision matrix multiplication with dimensions fixed at compile time; 6 NumPy calling Intel MKL; 7 Google Benchmark wall time in milliseconds, run count and statistic not stated, results checked against PyTorch. Missing: turbo and SMT state; run count and statistic. Verdict: watchlist for the number; the entry is trimmed for length, see Rejected. Verdict: watchlist.
- c7.3 "cycles/integer 9.0, 1.8, 0.7" for scalar, branchless scalar and SVE from https://lemire.me/blog/2022/06/23/filtering-numbers-quickly-with-sve-on-amazon-graviton-3-processors/: fields present: 1 Graviton 3, instance size not named, microarchitecture not named in the post (Neoverse V1 per the vendor table in section 14); 2 one core; 3 not stated; 4 GCC 11, flags in the linked repository; 5 removing negative integers from an array, cycles and instructions per input integer; 6 the branchy scalar loop and the branchless scalar loop; 7 hardware counters, repetitions and statistic not stated, code published. Missing: instance and frequency; turbo and SMT; flags in the post; repetitions and statistic. Verdict: watchlist for the number; the entry is trimmed for length, see Rejected. Verdict: watchlist.
- c7.4 "up to 20 times as fast as the sorting algorithms implemented in standard libraries" and "a geometric mean speedup of 1.59" from https://arxiv.org/abs/2205.05982: fields present: 1 Intel Xeon Gold 6154, Skylake-SP, and Apple M1 Max; 2 one core for the single-core tables, 16 threads on one socket for the parallel table; 3 3 GHz on the Xeon with turbo disabled, 3.2 GHz on the M1 Max, SMT not stated; 4 clang++ -O2 with -mavx512f named for the partition benchmark, version not stated; 5 sorting one million and one hundred million uniform random keys of 32-bit float, 32-bit and 64-bit integer and 128-bit types, MB/s; 6 the LLVM C++ library std::sort, heapsort and ips4o; 7 throughput in MB/s, repetitions and statistic not stated, code published in the Highway repository. Missing: SMT state; compiler version and full flags; repetitions and statistic. Verdict: watchlist for the number; the entry is trimmed for length, see Rejected. Verdict: watchlist.
- c7.5 "the Eytzinger layout is usually the fastest" for large arrays, with branch-free binary search "unbeaten by any other strategy" below the L2 size, from https://arxiv.org/abs/1509.05053: fields present: 1 Intel Core i7-4790K, Haswell, plus an Atom 330 and reader-contributed machines; 2 one core for the main results, two to eight threads in one experiment; 3 4 GHz, turbo and SMT not stated; 4 GCC 4.8.4, -std=c++11 -Wall -O4 -march=native; 5 two million uniform random searches over arrays of 32-bit unsigned integers at lengths spaced on a log scale, with 64-bit and 128-bit keys in one experiment; 6 branchy binary search and the standard library lower_bound; 7 wall-clock running time of the two million searches, repetitions and statistic not stated, code and scripts published. Missing: turbo and SMT; repetitions and statistic. Verdict: watchlist for the number; the entry stays for the finding that the winning layout and the branch-free choice change with array size. Verdict: watchlist.
- c7.6 "0.52 versus 1.02 cycles per 8 bytes" from https://arxiv.org/abs/1611.07612: fields present: 1 Intel Core i7-4770, Haswell; 2 one core, single-threaded; 3 3.4 GHz, Turbo Boost disabled, SMT not stated; 4 GCC 5.3, -O3 -march=native; 5 population count over random arrays of the stated sizes, cycles per 64-bit word; 6 an optimised popcnt-instruction loop; 7 each test repeated 500 times, minimum reported, minimum and mean within one percent. Missing: SMT state. Verdict: watchlist for the number (SMT state not stated); the entry stays and its reason quotes no number. Verdict: watchlist.
- c7.7 "encoding ~10x faster, decoding ~7x faster" from https://arxiv.org/abs/1704.00605: fields present: 1 Intel Core i7-6700, Skylake; 2 one core, single-threaded; 3 3.4 GHz, Turbo Boost disabled, SMT not stated; 4 GCC 5.3, -O3 -march=native; 5 base64 encode and decode of long in-memory inputs, cycles per input byte via rdtsc; 6 the Chromium codec, the Linux kernel codec and the QuickTime codec; 7 each test repeated 500 times, minimum reported, minimum and mean within one percent. Missing: SMT state. Verdict: watchlist for the number (SMT state not stated); the entry stays and its reason quotes no number. Verdict: watchlist.
- c7.8 "gigabytes of data per second on a single core" and "a quarter or fewer instructions than RapidJSON" from https://arxiv.org/abs/1902.08318: fields present: 1 Intel Core i7-6700 Skylake and Core i3-8121U Cannon Lake; 2 one core, single-threaded; 3 base 3.4 and 2.2 GHz, maximum 3.7 and 3.2 GHz, clock behaviour verified with avx-turbo, SMT disabled (the paper says hyper-threading); 4 GCC 9.1, -O3 -march=native, no profile-guided optimisation; 5 the listed JSON files parsed from memory; 6 RapidJSON in insitu mode and sajson with static allocation; 7 cycles and instructions per input byte and GB/s from hardware counters, code and raw results published. Missing: whether turbo was disabled (maximum frequency is reported, not the state); number of repetitions and the statistic. Verdict: watchlist for the number; the entry stays for the mechanism. Verdict: watchlist.
- c7.9 "improves the performance of Snort by a factor of 8.7" from https://www.usenix.org/conference/nsdi19/presentation/wang-xiang: fields present: 1 Intel Xeon Platinum 8180, Skylake-SP; 2 one core, packets fed from memory; 3 2.50 GHz base, turbo and SMT not stated; 4 GCC 5.4, flags not stated; 5 Snort with the ET-Open and Talos rulesets over a real Web traffic trace and random packets; 6 Snort with its stock prefilter-based matching; 7 throughput on packets from memory, repetitions and statistic not stated. Missing: turbo and SMT state; flags; repetitions and statistic. Verdict: watchlist for the number; the entry stays for the mechanism. Verdict: watchlist.
- c7.10 "the SOA version was 1.25x faster than the AOS one" from https://pharr.org/matt/assets/ispc.pdf: fields present: 1 four-core Intel Core i7 in an iMac running AVX code, model and microarchitecture not named; 2 not stated for this experiment; 3 3.4 GHz, turbo and SMT not stated; 4 ispc on LLVM for both layouts, flags not stated; 5 collision detection between two groups of spheres, written once in array-of-structures and once in structure-of-arrays layout; 6 the array-of-structures version; 7 minimum of several runs, run count not stated. Missing: exact CPU, core count, turbo and SMT, flags, run count. Verdict: watchlist for the number; the entry stays for the mechanism and the layout keyword. Verdict: watchlist.
- c7.11 "over 10%" from https://lemire.me/blog/2021/07/14/faster-sorted-array-unions-by-reducing-branches/: fields present: 1 Apple M1 and AMD Rome (Zen 2), Zen 2 model not named; 2 one core; 3 not stated; 4 LLVM 12, GCC 10 and LLVM 11, flags in the linked repository; 5 union of two sorted arrays of random integers, nanoseconds per produced element; 6 the conventional branchy union; 7 average nanoseconds per produced element, repetitions not stated, code published. Missing: exact Zen 2 model; frequency, turbo and SMT; flags in the post; repetitions. Verdict: watchlist for the number; the entry stays for the mechanism and the counter-example. Verdict: watchlist.
- c7.12 "SIMD version is 3.1 times faster" than std::string::find on the Armv7 Raspberry Pi 3 from http://0x80.pl/notesen/2016-11-28-simd-strfind.html: fields present: 1 Westmere i5 M540, Bulldozer FX-8150, Haswell i7-4770, Skylake i7-6700, Knights Landing 7210, Armv7 Raspberry Pi 3, Cortex-A57 Opteron A1100; 2 one core; 3 not stated; 4 GCC 6.2.0, 4.8.4, 5.4.1, 5.3.0, 4.9.2 and Clang 3.8.0 per machine, flags in the linked repository; 5 substring search over a fixed text, total seconds; 6 strstr and std::string::find from the GNU library, plus a SWAR version; 7 wall time in seconds, three runs, statistic not stated. Missing: frequency, turbo and SMT; flags in the note; statistic. Verdict: watchlist for the number; the entry is trimmed for length, see Rejected. Verdict: watchlist.
- c7.13 "3.92 bytes per nanosecond against 0.60" from https://branchfree.org/2018/05/25/say-hello-to-my-little-friend-sheng-a-small-but-fast-deterministic-finite-automaton/: fields present: 1 Skylake, model not named; 2 one core; 3 4 GHz, turbo and SMT not stated; 4 not stated; 5 a small automaton run over an input buffer; 6 a table-driven automaton; 7 not stated, code published. Missing: CPU model; compiler and flags; turbo and SMT; method. Verdict: watchlist for the number; the entry is trimmed for length, see Rejected. Verdict: watchlist.
- c7.14 "12-20 MPKI on Nehalem to 0.5-2 MPKI on Haswell" from https://inria.hal.science/hal-01100647: fields present: 1 Xeon W3550 Nehalem, Core i7-2620M Sandy Bridge, Core i7-4770 Haswell; 2 one core; 3 3.07, 2.70 and 3.40 GHz, turbo and SMT not stated; 4 Intel icc 13 for the interpreters, GCC4CLI for the CLI benchmarks, flags not stated; 5 Python 3.3.2 on the Unladen Swallow benchmarks, a Javascript interpreter and a CLI interpreter on SPEC 2000; 6 switch-based dispatch against threaded code; 7 mispredictions per thousand instructions from the PMU, repetitions not stated. Missing: turbo and SMT; flags; repetitions and statistic. Verdict: watchlist for the number; the entry stays for the finding. Verdict: watchlist.

## 8. Compilers and codegen

### Rejected candidates

- r8.1 [What Has My Compiler Done for Me Lately? Unbolting the Compiler's Lid](https://www.youtube.com/watch?v=bSkpMdDe4g4): Trimmed for length: the Start here section is its home, as its last entry, and the Compiler Explorer entry carries the tool the talk teaches.
- r8.2 [Tuning C++: Benchmarks, and CPUs, and Compilers! Oh My! (CppCon 2015)](https://www.youtube.com/watch?v=nXaxk27zwlk): Trimmed for length: the Google Benchmark User Guide entry in the measurement section carries the barriers that stop the optimiser deleting a microbenchmark, and the llvm-objdump entry carries the source-annotated read of a shipped binary.
- r8.3 [-O levels (clang Command Guide)](https://clang.llvm.org/docs/CommandGuide/clang.html#cmdoption-O0): Trimmed for length: the GCC Optimize Options entry carries what each -O level turns on and the warning that -Ofast admits transforms invalid for conforming code, the reason Clang deprecated it, and the Clang floating-point entry names the -ffast-math permissions -Ofast implies.
- r8.4 [AArch64 Options (GCC)](https://gcc.gnu.org/onlinedocs/gcc/AArch64-Options.html): Trimmed for length: the x86 Options entry carries the -march against -mtune rule, which -mcpu sets together on Arm, and -msve-vector-bits, the switch between vector-length agnostic and fixed-width SVE code, stays on the AArch64 page of the same GCC machine-dependent options chapter.
- r8.5 [-fopt-info (GCC)](https://gcc.gnu.org/onlinedocs/gcc/Developer-Options.html#index-fopt-info): Trimmed for length: the Clang optimisation remarks entry carries the report that names each loop left scalar and why, and -fopt-info-vec-missed, the GCC switch for the same report, stays documented on the Developer Options page.
- r8.6 [-fprofile-use and -fauto-profile (GCC)](https://gcc.gnu.org/onlinedocs/gcc/Optimize-Options.html#index-fprofile-use): Trimmed for length: the Clang PGO entry carries the instrumented and sampled workflows, and both GCC switches sit on the Optimize Options page already listed, where the passes a profile turns on are enumerated.
- r8.7 [autofdo](https://github.com/google/autofdo): Trimmed for length: the Clang PGO page names llvm-profgen and create_llvm_prof, the converters that turn LBR or Arm SPE samples in perf.data into a sampled profile, and the AutoFDO paper entry carries the mapping they implement.
- r8.8 [AutoFDO at the ACM Digital Library](https://dl.acm.org/doi/10.1145/2854038.2854044): Rule 4: returns 403 to curl and to WebFetch and shows a bot check in a browser; the authors' own page is listed instead.
- r8.9 [Profile Guided Code Positioning](https://dl.acm.org/doi/10.1145/93542.93550): Rule 4: the origin paper for profile-driven function and basic-block layout, but the DOI returns 403 to curl and to WebFetch and shows a bot check in a browser, and no author copy exists; left out on the same grounds as the AutoFDO ACM copy.
- r8.10 [Propeller: A Profile Guided, Relinking Optimizer for Warehouse-Scale Applications](https://research.google/pubs/propeller-a-profile-guided-relinking-optimizer-for-warehouse-scale-applications/): No rule failed. The page carries only the abstract and links to a copy behind the ACM bot check; the LLVM RFC is the self-contained design statement and took the one Propeller slot.
- r8.11 [AutoFDO at IEEE Xplore](https://ieeexplore.ieee.org/document/7559528/): Rule 4: paywalled copy where an author page exists.
- r8.12 [AutoFDO on Semantic Scholar](https://www.semanticscholar.org/paper/AutoFDO:-Automatic-feedback-directed-optimization-Chen-Li/e8d25bdb1f2cdca5e2d23bb0d2ea6baa91649005): Rule 4: third-party mirror of a document with an official home.
- r8.13 [Profile-guided optimization (Wikipedia)](https://en.wikipedia.org/wiki/Profile-guided_optimization): Rule 1: Wikipedia is excluded outright.
- r8.14 [Intel Architecture Code Analyzer](https://www.intel.com/content/www/us/en/developer/articles/tool/architecture-code-analyzer.html): Rule 4: returns 403 to curl and redirects a fetch into Intel's 404 handler; the tool is withdrawn, and llvm-mca plus uops.info cover the same ground.
- r8.15 [Optimizations in C++ Compilers: A practical journey (ACM Queue)](https://queue.acm.org/detail.cfm?id=3372264): Rule 1: a tutorial explainer of compiler transformations, not the original report of a mechanism or a measurement; the site also returns 403 to curl.
- r8.16 [uops.info](https://uops.info/): No rule failed. Cut from this section on review: it was the third listing of one URL in the README with a reason generic to the site; the measured table lives in the microarchitecture section, and llvm-exegesis is the tool in this section that produces the same measurement.
- r8.17 [uops.info: Characterizing Latency, Throughput, and Port Usage of Instructions on Intel Microarchitectures](https://arxiv.org/abs/1810.04610): No rule failed. Listed in the microarchitecture section beside the site it describes.
- r8.18 [Auto-vectorization in GCC](https://gcc.gnu.org/projects/tree-ssa/vectorization.html): No rule failed. Cut on review: the page's latest news is dated 2011 and its worked examples are driven by -ftree-vectorizer-verbose, which GCC removed in favour of -fopt-info, so a reader who follows it gets an error.
- r8.19 [Clang command line argument reference](https://clang.llvm.org/docs/ClangCommandLineReference.html): No rule failed. Cut on review: a generated flag index whose reason said only that the flags exist, and the third Clang reference page in one subsection; the slot went to the ABI.
- r8.20 [Going Nowhere Faster (CppCon 2017)](https://www.youtube.com/watch?v=2EWejmkKlxs): No rule failed. Cut on review: the fourth CppCon talk in the section and the third by one speaker, whose ground (the loop is memory bound, not compiler bound) is the memory hierarchy section's, and Tuning C++ already covers reading profiler-annotated assembly.
- r8.21 [LLVM Link Time Optimization: Design and Implementation](https://llvm.org/docs/LinkTimeOptimization.html): No rule failed. Cut under the cap on review: the ThinLTO page links it and defines the same linker handshake in its thin form, and the slot went to GCC's LTO Overview so both compilers have a design statement.
- r8.22 [System V ABI: AMD64 Architecture Processor Supplement](https://gitlab.com/x86-psABIs/x86-64-ABI): No rule failed. Cut under the cap: the calling convention any x86 listing is read against, but the Itanium C++ ABI carries the rule the Zero-cost Abstractions entry rests on and the x86 Options page carries the micro-architecture levels; the Arm counterpart is https://github.com/ARM-software/abi-aa/blob/main/aapcs64/aapcs64.rst (curl 200), left out for the same reason.
- r8.23 [C11 committee draft N1570](https://www.open-std.org/jtc1/sc22/wg14/www/docs/n1570.pdf): No rule failed. The formal definition of restrict, already listed in the concurrency section; cut here under the cap because the section's benchmark shows clang vectorises the loop behind a runtime alias check with or without the qualifier, which the Auto-Vectorization in LLVM entry states.
- r8.24 [TSVC_2](https://github.com/UoB-HPC/TSVC_2): No rule failed. The C port of the Test Suite for Vectorising Compilers, one loop per pattern with a checksum; cut under the cap because the section's own benchmark is its measured source and the two remark switches give the same per-loop verdict on the reader's compiler.
- r8.25 [Optimizing real world applications with GCC Link Time Optimization](https://arxiv.org/abs/1010.2196): No rule failed. The implementer report on WHOPR with compile time, memory and code size measured; the GCC LTO Overview page defines the same stages and is what a build is configured from, so it kept the slot under the cap.
- r8.26 [objdump (GNU Binutils)](https://sourceware.org/binutils/docs/binutils/objdump.html): No rule failed. Left out under the cap in favour of llvm-objdump, which reads Mach-O as well as ELF and matches the Clang toolchain used throughout the section.
- r8.27 [An Inline Function is As Fast As a Macro (GCC)](https://gcc.gnu.org/onlinedocs/gcc/Inline.html): No rule failed. It defines the keyword's semantics rather than the optimiser's inlining decisions, which the Optimize Options page already covers; left out under the cap.
- r8.28 [An analysis of inline substitution for a structured programming language](https://dl.acm.org/doi/10.1145/359810.359830): Rule 4: no author copy exists and the publisher copy is paywalled and blocks automated clients; the compiler documentation is the current primary statement of inlining.
- r8.29 [llvm-propeller](https://github.com/google/llvm-propeller): No rule failed. The implementation repository; the RFC entry links it, and the cap left no second Propeller slot.
- r8.30 [Efficiency with Algorithms, Performance with Data Structures (CppCon 2014)](https://www.youtube.com/watch?v=fHNmRkzxHWs): No rule failed. Its subject is algorithm choice and contiguous data layout, which is the memory hierarchy section's ground, not optimisation levels or LTO; cut from this section on review, and the slot went to the Clang -O level text.
- r8.31 [compiler-explorer](https://github.com/compiler-explorer/compiler-explorer): No rule failed. The repository of the project already listed as Compiler Explorer; the hosted site is what a reader opens, so one project keeps one slot and the repository was cut on review.
- r8.32 [ThinLTO: Scalable and Incremental LTO](https://research.google/pubs/thinlto-scalable-and-incremental-lto/): No rule failed. The origin paper, found on review; the Clang usage page defines the same thin-link design and is what a build is configured from, so it kept the slot under the cap.

### Numbers examined

- c8.1 "up to 20.4% on top of FDO and LTO" and "up to 52.1% if the binaries are built without FDO and LTO" from https://arxiv.org/abs/1807.06735: fields present: 1 CPU Intel Xeon E5-2680 v2, Ivy Bridge; 2 cores dual-node, 20 cores, 40 threads with SMT, builds run with -j40; 3 freq 2.80 GHz nominal, SMT on, turbo state not stated; 4 compiler/flags Clang release_60 bootstrapped, PGO+LTO build with -DLLVM_ENABLE_LTO=Full and an instrumented profile, BOLT flags listed in the paper; 5 workload a full ninja build of Clang plus three preprocessed LLVM source files compiled with -std=c++11 -O2; 6 baseline stage1 release Clang, and Clang with PGO+LTO; 7 method wall-clock build and compile time, files compiled "multiple times", no run count or statistic stated. Missing: turbo state, run count, statistic. Verdict: watchlist. Verdict: watchlist.
- c8.2 "up to 8.0% performance speedups on top of profile-guided function reordering and LTO" from https://arxiv.org/abs/1807.06735: fields present: 5 workload five Facebook services including HHVM; 6 baseline GCC builds with HFSort function reordering, HHVM also with LTO; 7 method production traffic with perf counters for i-cache, i-TLB, branch and LLC misses. Missing: 1 CPU, 2 cores, 3 freq/turbo/SMT, 4 compiler flags, run count and statistic. Verdict: watchlist. Verdict: watchlist.
- c8.3 "geomean of 10.5% improvement on benchmarks" and "85% of the gains of traditional FDO" from https://research.google/pubs/autofdo-automatic-feedback-directed-optimization-for-warehouse-scale-applications/ (full text at https://storage.googleapis.com/gweb-research2023-media/pubtools/pdf/45290.pdf): fields present: 4 compiler GCC r231122 from the google/gcc-4_9 branch, baseline built at -O2, other flags not stated; 5 workload internal Google application benchmarks (Table 1) and SPEC CPU2006 integer (Table 2); 6 baseline the -O2 binary without FDO, and instrumented FDO for the 85% ratio; 7 method speedup against the -O2 binary, run count and statistic not stated. Missing: 1 CPU (only "Intel Westmere platform", and only for the SPEC table), 2 cores, 3 freq/turbo/SMT, most of 4, run count. Verdict: watchlist. Verdict: watchlist.
- c8.4 "improves key performance metrics of these benchmarks by 2% to 6%" and, for BOLT on a binary with a "~300M text segment", "Memory foot-print is 70G" and "more than 10 minutes to rewrite" from https://lists.llvm.org/pipermail/llvm-dev/2019-September/135393.html: fields present: 5 workload "large google benchmarks" and SPEC; 6 baseline ThinLTO + PGO. Missing: 1 CPU, 2 cores, 3 freq/turbo/SMT, 4 compiler flags, 7 method. Verdict: watchlist. Verdict: watchlist.

## 9. Concurrency

### Rejected candidates

- r9.1 Threads Cannot Be Implemented as a Library at HP Labs: Rule 4: the host no longer answers at all, and the author's own site still points at it. The same report at the successor host, which serves the paper, is listed.
- r9.2 Foundations of the C++ Concurrency Memory Model at HP Labs: Rule 4: the host no longer answers at all. The co-author's group copy is listed.
- r9.3 Foundations of the C++ Concurrency Memory Model at the successor HPE host: No rule failed. The originating lab's own copy, but the server omits its intermediate certificate and every OpenSSL-based client, the repository link checker included, rejects it. The co-author's copy, which verifies everywhere, is listed.
- r9.4 [Threads Cannot Be Implemented as a Library at the ACM Digital Library](https://dl.acm.org/doi/10.1145/1065010.1065042): Rule 4: returns 403 to curl and to a fetch, and shows a bot check in a browser. The author's report is listed instead.
- r9.5 [C11 committee draft N1570](https://www.open-std.org/jtc1/sc22/wg14/www/docs/n1570.pdf): No rule failed. Clause 7.17 transplants the C++ orders and fences unchanged, so a second normative text for the same model adds nothing the mappings table needs. Left out under the cap in favour of herdtools7, which the subsection lacked.
- r9.6 [Mathematizing C++ Concurrency](https://www.cl.cam.ac.uk/~pes20/cpp/popl085ap-sewell.pdf): No rule failed. The formalisation that found and fixed defects in the draft model. Left out under the cap, and the mappings page links it.
- r9.7 [Repairing Sequential Consistency in C/C++11](https://plv.mpi-sws.org/scfix/paper.pdf): No rule failed. The correction to sequentially consistent fences that C++20 adopted. Left out under the cap.
- r9.8 [Synchronising C/C++ and POWER](https://www.cl.cam.ac.uk/~pes20/ppc-supplemental/pldi105-sarkar.pdf): No rule failed. The proof that the listed mappings are sound for POWER and Arm. Left out under the cap, and the mappings page carries the result.
- r9.9 [x86-TSO: A Rigorous and Usable Programmer's Model for x86 Multiprocessors](https://www.cl.cam.ac.uk/~pes20/weakmemory/cacm.pdf): No rule failed. Listed in section 4 with the same URL, the preamble routes a reader there, and the memory-models subsection is at the hard cap.
- r9.10 [Arm Architecture Reference Manual for A-profile architecture](https://support.arm.com/documentation/ddi0487/latest/): No rule failed. Listed in section 4, and the mappings page carries the part a programmer meets first.
- r9.11 [Simplifying ARM Concurrency: Multicopy-Atomic Axiomatic and Operational Models for ARMv8](https://www.cl.cam.ac.uk/~pes20/armv8-mca/armv8-mca-draft.pdf): No rule failed. Listed in section 4 with the same URL, the preamble routes a reader there, and the memory-models subsection is at the hard cap.
- r9.12 [Memory Barriers: a Hardware View for Software Hackers](http://www.rdrop.com/users/paulmck/scalability/paper/whymb.2010.07.23a.pdf): No rule failed. Listed in section 4, and the same material is an appendix of the perfbook entry.
- r9.13 [diy.inria.fr](https://diy.inria.fr/): No rule failed. The documentation home of herdtools7, whose repository is listed because the tools are what the entry is for and the README points at the documentation.
- r9.14 [std::memory_order at cppreference](https://en.cppreference.com/cpp/atomic/memory_order): Rule 1: a secondary paraphrase of the standard. The draft itself is listed.
- r9.15 [Acquire and Release Semantics](https://preshing.com/20120913/acquire-and-release-semantics/): Rule 1: an explainer, not the original report of a mechanism or of a measurement.
- r9.16 [Everything You Always Wanted to Know About Synchronization but Were Afraid to Ask](https://sigops.org/s/conferences/sosp/2013/papers/p33-david.pdf): No rule failed. Listed in section 4 as the measured cost of moving a line between cores. Its cross-machine lock comparison would repeat the URL, and the locks subsection is at the cap.
- r9.17 [C2C - False Sharing Detection in Linux Perf](https://joemario.github.io/blog/2016/09/01/c2c-blog/): No rule failed. Listed in section 4, where false sharing is measured and the line size is fixed.
- r9.18 [False Sharing and its Effect on Shared Memory Performance](https://www.cs.rochester.edu/u/scott/papers/1993_SEDMS_false_sharing.pdf): No rule failed. The original definition and measurement of false sharing. Left out under the hard cap, and section 4 carries the mechanism and its measurement through the perf c2c and SOSP entries.
- r9.19 [Lock types and their rules](https://docs.kernel.org/locking/locktypes.html): No rule failed. States which lock may be taken in which context, a correctness rule rather than a cost. Left out under the cap, with the mutex design and seqlock documents taking the kernel locking slots.
- r9.20 [A Scalable Concurrent malloc(3) Implementation for FreeBSD](https://people.freebsd.org/~jasone/jemalloc/bsdcan2006/jemalloc.pdf): No rule failed. The arena split restates the per-processor heaps that Hoard defines, for a shipping allocator, and its scaling figure is watchlist under rule 3. Left out under the cap, with Hoard and TCMalloc carrying the allocator mechanisms.
- r9.21 [Mimalloc: Free List Sharding in Action](https://www.microsoft.com/en-us/research/publication/mimalloc-free-list-sharding-in-action/): No rule failed. A fourth allocator in a subsection titled for locks and contention, whose reason rested on a measurement that is watchlist under rule 3. Left out under the cap in favour of Non-scalable locks are dangerous, the modern contention measurement the subsection lacked.
- r9.22 [mimalloc](https://github.com/microsoft/mimalloc): No rule failed. The repository, whose README links the technical report as its design document. Both are left out under the cap, the report for the reason above.
- r9.23 [Wait-Free Synchronization](https://cs.brown.edu/people/mph/Herlihy91/p124-herlihy.pdf): No rule failed. The consensus hierarchy, which the listed textbook restates in full. Left out under the cap.
- r9.24 Systems Programming: Coping with Parallelism: Rule 4: the IBM Research archive that hosted the report answers neither curl nor a fetch, and no other primary home exists. The stack it introduces is given in the listed textbook and in the queue paper's related work.
- r9.25 [Hazard Pointers: Safe Memory Reclamation for Lock-Free Objects](https://ieeexplore.ieee.org/document/1291819): Rule 4: paywalled, the ACM copy sits behind a bot check, and no author copy is online. The inventor's standard paper is listed instead.
- r9.26 [Hazard Pointers at Otago](https://www.cs.otago.ac.nz/cosc440/readings/hazard-pointers.pdf): Rule 4: a third-party course mirror of a document with an official home.
- r9.27 [A Pragmatic Implementation of Non-Blocking Linked-Lists](https://www.cl.cam.ac.uk/research/srg/netos/papers/2001-caslists.pdf): No rule failed. The marked-pointer deletion every lock-free list uses. Left out under the cap, and the textbook carries it.
- r9.28 [Practical lock-freedom](https://www.cl.cam.ac.uk/techreports/UCAM-CL-TR-579.pdf): No rule failed. The origin of epoch-based reclamation. Left out under the cap.
- r9.29 [What is RCU? Part 2: Usage](https://lwn.net/Articles/263130/): No rule failed. Left out under the cap, and the kernel document listed carries the usage patterns.
- r9.30 [RCU part 3: the RCU API](https://lwn.net/Articles/264090/): No rule failed. Left out under the cap, and the kernel document listed carries the full API table.
- r9.31 [Read-Copy Update: Using Execution History to Solve Concurrency Problems](http://www.rdrop.com/users/paulmck/RCU/rclockpdcsproof.pdf): No rule failed. The original RCU paper. Left out under the cap, and the perfbook entry cites it.
- r9.32 [liburcu](https://liburcu.org/): No rule failed. The project home of the user-space RCU library, canonical at git.liburcu.org with GitHub as a mirror. Left out under the cap, and the listed paper defines its flavours and measures them.
- r9.33 [Scheduling Multithreaded Computations by Work Stealing](https://dl.acm.org/doi/10.1145/324133.324234): Rule 4: the ACM copy sits behind a bot check, the group's paper host returns 404 for its former copy, and the group's bibliography page carries no file. The group's own memo, which proves the same bounds, is listed.
- r9.34 Scheduling Multithreaded Computations by Work Stealing at Supertech: Rule 4: redirects to a 404 on the group's new host.
- r9.35 [Executing Multithreaded Programs Efficiently](https://dspace.mit.edu/handle/1721.1/149816): No rule failed. The thesis form of the work-stealing proof. Left out under the cap.
- r9.36 The Implementation of the Cilk-5 Multithreaded Language at Supertech: Rule 4: redirects to a 404 on the group's new host. The first author's own copy is listed.
- r9.37 Thread Scheduling for Multiprogrammed Multiprocessors: Rule 4: the ACM copy sits behind a bot check, the Springer copy is paywalled, and the remaining copies are course mirrors. Considered for the thread-pool subsection and dropped.
- r9.38 [Dynamic Circular Work-Stealing Deque](https://dl.acm.org/doi/10.1145/1073970.1073974): Rule 4: the ACM copy sits behind a bot check and no author copy exists, the originating lab having closed. The listed PPoPP paper restates the algorithm with a proof.
- r9.39 [Dynamic Circular Work-Stealing Deque at Vanderbilt](https://www.dre.vanderbilt.edu/~schmidt/PDF/work-stealing-dequeue.pdf): Rule 4: a third-party course mirror.
- r9.40 [Treiber stack on Wikipedia](https://en.wikipedia.org/wiki/Treiber_stack): Rule 1: Wikipedia is excluded outright.
- r9.41 [Threads Cannot Be Implemented as a Library](https://www.labs.hpe.com/techreports/2004/HPL-2004-209.pdf): Trimmed for length: the argument that the language must own the model, because a thread-unaware compiler introduces races, is carried by Foundations of the C++ Concurrency Memory Model, which states the contract that replaced the library approach.
- r9.42 [Explanation of the Linux-Kernel Memory Consistency Model](https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/tree/tools/memory-model/Documentation/explanation.txt): Trimmed for length: memory-barriers.txt states the kernel's ordering rules and points at the model file, and herdtools7 carries the kernel model as the target klitmus7 runs a litmus test against.
- r9.43 [Generic Mutex Subsystem](https://docs.kernel.org/locking/mutex-design.html): Trimmed for length: the fast and slow paths of a mutex are derived in Futexes Are Tricky, the queue each waiter spins in is the MCS entry, and the spin-before-sleep policy is carried by the oneTBB entry.
- r9.44 [Non-scalable locks are dangerous](https://pdos.csail.mit.edu/papers/linux:lock.pdf): Trimmed for length: the collapse of a non-queue lock as processors are added is measured in the MCS entry, and the cross-machine lock comparison is Everything You Always Wanted to Know About Synchronization in section 4.
- r9.45 [Sequence counters and sequential locks](https://docs.kernel.org/locking/seqlock.html): Trimmed for length: the retrying read side that writes no shared line is derived, with code, in the deferral chapter of the perfbook entry that opens the locks subsection.
- r9.46 [What is RCU, Fundamentally?](https://lwn.net/Articles/262464/): Trimmed for length: the kernel's What is RCU? document carries the split into publish, wait for readers and keep old versions, and the read side that costs no instructions.
- r9.47 [Heartbeat Scheduling: Provable Efficiency for Nested Parallelism](https://www.chargueraud.org/research/2018/heartbeat/heartbeat.pdf): Trimmed for length: the work-first principle it sharpens into a granularity rule is the Cilk-5 entry, and the grain size a shipping runtime applies is carried by the oneTBB entry.

### Numbers examined

- c9.1 "the time to acquire and release the test-and-set lock increases 7.4 µs per additional processor" from https://www.cs.rochester.edu/u/scott/papers/1991_TOCS_synch.pdf: fields present: 1 CPU BBN Butterfly 1, 8 MHz MC68000 (the Sequent Symmetry, 16 MHz Intel 80386 with a 64 KB two-way cache, is the second machine in the paper but not the source of this figure); 2 cores one to the full machine, each point a processor count; 3 freq 8 MHz, no turbo or SMT on the part; 4 compiler hand-written assembler for the spin locks; 5 workload lock acquire and release in a loop, interrupts disabled; 6 baseline the simple test-and-set lock; 7 method average over one hundred thousand acquisitions, loop overhead included, slope by linear least squares. Missing: run count beyond the single averaged loop. Verdict: core for the mechanism and the slope, the value belongs to a machine no longer in use. Verdict: core.
- c9.2 "Throughput peaks at nine cores, at which point it is 4.7× higher than with one core" and "At 48 cores throughput is about 35% of the throughput achieved on one core" for MEMPOP from https://pdos.csail.mit.edu/papers/linux:lock.pdf: fields present: 1 CPU AMD Opteron, eight six-core chips, microarchitecture not named; 2 cores one to forty-eight; 3 freq 2.4 GHz, turbo and SMT state not stated; 4 compiler not stated, kernel 2.6.39; 5 workload one process per core repeatedly mapping and unmapping 64 kB with MAP_POPULATE; 6 baseline the same workload on one core under the kernel's ticket spinlock; 7 method throughput against core count, run count and statistic not stated. Missing: 4, turbo and SMT state, run count and statistic. Verdict: watchlist; the entry is now under Rejected, trimmed for length. Verdict: watchlist.
- c9.3 "Hoard improves performance over the standard Solaris allocator by up to a factor of 60" from https://people.cs.umass.edu/~emery/pubs/berger-asplos2000.pdf: fields present: 1 CPU 400 MHz UltraSPARC with 4 MB level-two cache in a Sun Enterprise 5000; 2 cores 14 processors; 3 freq 400 MHz, turbo and SMT not stated and absent from the part; 4 compiler GNU C++ at -O6 for all but one benchmark, Sun Workshop 5.0 named as rejected; 5 workload threadtest, shbench, Larson, active-false, passive-false, BEMengine and Barnes-Hut; 6 baseline the Solaris 7 allocator, with Ptmalloc and MTmalloc also run; 7 method average of three runs, variation reported as negligible. Missing: none of the seven, though the run count applies to the runtime table and is not restated for the scaling graphs. Verdict: core; the entry does not quote the figure, because its reason rests on the definitions and the ratio belongs to a machine no longer in use. Verdict: core.
- c9.4 "jemalloc scales almost perfectly up to four threads, then remains relatively constant above four threads" from https://people.freebsd.org/~jasone/jemalloc/bsdcan2006/jemalloc.pdf: fields present: 1 CPU AMD Opteron 275, K8; 2 cores four, two dual-core sockets; 5 workload malloc-test, forty million allocation and deallocation cycles of one 512-byte object, divided among the threads; 6 baseline phkmalloc and dlmalloc 2.8.3 through LD_PRELOAD; 7 method three replications per configuration, box plots with the median marked. Missing: 3 frequency and turbo state, 4 compiler and flags beyond FreeBSD-CURRENT/amd64 of April 2006. Verdict: cut with the entry, which is now under Rejected. Verdict: cut.
- c9.5 "speed improvements of 7% and 14%, respectively, on redis" over tcmalloc and jemalloc from https://www.microsoft.com/en-us/research/publication/mimalloc-free-list-sharding-in-action/: fields present: 1 CPU AMD EPYC 7000 series, Zen, in an Amazon EC2 r5a.4xlarge instance; 2 cores sixteen as presented to the virtual machine; 3 freq 2.5 GHz stated, turbo and SMT state not stated; 4 compiler GCC 7.3.0 on Ubuntu 18.04.1 with glibc 2.27, flags not stated; 5 workload redis 5.0.3 serving one million requests that push ten list elements and read back the head ten; 6 baseline mimalloc v1.0.0 with every time reported relative to it, against jemalloc 5.2.0, tcmalloc 2.5, Hoard 3.13, snmalloc, rpmalloc, the TBB allocator and glibc, each through LD_PRELOAD; 7 method wall-clock time and peak resident memory from the time program, average of five runs. Missing: turbo and SMT state, compiler flags, and the instance is virtualised so the physical core count is not stated. Verdict: cut with the entry, which is now under Rejected. Verdict: cut.
- c9.6 "the performance of the QSBR and the signal-based-RCU implementations are more than an order of magnitude greater than that of the per-thread mutex" (Figure 4) from https://www.efficios.com/pub/rcu/urcu-main.pdf: fields present: 1 CPU IBM POWER5+ for the figure, with an Intel Core2 Xeon E5405 reported as behaving alike and not plotted; 2 cores one to sixty-four readers, one hardware thread per core; 3 freq 1.9 GHz POWER5+ and 2.0 GHz Xeon, the second hardware thread of each POWER5+ core left idle, turbo not stated; 4 compiler not stated, glibc 2.7 and 2.5 named for the pthread baselines; 5 workload each reader takes its read-side lock, reads one variable and releases it in a tight loop for ten seconds with no updater; 6 baseline the pthread mutex, the pthread reader-writer lock and a per-thread mutex; 7 method reads per second, run count and statistic not stated. Missing: 4, run count and statistic. Verdict: watchlist; the entry rests on the definitions of the three flavours and on the ordering of their read-side costs. Verdict: watchlist.
- c9.7 "Heartbeat overheads exceed 3% in only 6 cases" against "PBBS programs can incur significant overheads, sometimes over 25% of the execution time" from https://www.chargueraud.org/research/2018/heartbeat/heartbeat.pdf: fields present: 1 CPU Intel Xeon E7-4870, Westmere-EX; 2 cores forty, four ten-core chips; 3 freq 2.4 GHz, turbo and SMT state not stated; 4 compiler GCC 6.3 at -O2 -march=native, with -fcilkplus for the baseline; 5 workload the PBBS benchmarks on their published inputs; 6 baseline the authors' hand-tuned Cilk Plus PBBS codes; 7 method average of thirty runs, standard deviation stated as a few percent. Missing: turbo and SMT state. Verdict: watchlist; the entry is now under Rejected, trimmed for length. Verdict: watchlist.
- c9.8 "at least 1.5 times better than seqcst on both x86 and ARM" from https://inria.hal.science/hal-00802885: fields present: 1 CPU Tegra 3 Armv7, Intel Core i7-2720QM, AMD Opteron 6164HE; 2 cores four, four with SMT off, and two by twelve; 3 freq 1.3 GHz, 2.2 GHz and 1.7 GHz, SMT off on the Core i7, turbo state not stated; 4 compiler GCC 4.7.0, flags not stated; 5 workload a push and take loop with steals injected at a set rate, on a broad tree and a comb-shaped tree; 6 baseline the same deque with every access at seq_cst; 7 method push and take throughput against steal throughput, run count and statistic not stated. Missing: turbo state, compiler flags, run count and statistic. Verdict: watchlist. Verdict: watchlist.
- c9.9 Net execution time for "one million enqueues and dequeues" from https://www.cs.rochester.edu/u/scott/papers/1996_PODC_queues.pdf: fields present: 1 CPU MIPS R4000 in a Silicon Graphics Challenge; 2 cores twelve, with processes pinned to represent dedicated and multiprogrammed loads; 4 compiler the highest optimisation level, compiler not named, code hand-tuned; 5 workload each process enqueues, spins, dequeues and spins; 6 baseline a single-lock queue and four published non-blocking queues; 7 method net time with the single-processor loop cost subtracted, run count not stated. Missing: 3 frequency, turbo and SMT state, the compiler's name and flags, run count and statistic. Verdict: watchlist. Verdict: watchlist.

## 10. NUMA and multi-socket

### Rejected candidates

- r10.1 [The Stanford Dash Multiprocessor](https://ieeexplore.ieee.org/document/121510): Trimmed for length: the directory that keeps caches coherent across nodes is carried by the Intel Xeon Scalable Family overview (the home agent) and the CMN-700 manual (the home nodes), and the origin of cache-coherent NUMA is not a mechanism the placement subsection promises.
- r10.2 [An NUMA API for Linux](http://halobates.de/numaapi3.pdf): Trimmed for length: NUMA Memory Policy states every mode, scope and system call the paper introduced, in its current normative form, and numactl is the API's reference implementation.
- r10.3 [hwloc](https://github.com/open-mpi/hwloc): Trimmed for length: numactl -H prints the node-to-CPU map and distance table that placement needs, and the SNC and NPS entries under Topology and interconnects say what changes the tree hwloc draws.
- r10.4 [sysfs-devices-node ABI](https://www.kernel.org/doc/Documentation/ABI/stable/sysfs-devices-node): Trimmed for length: numactl -H prints the node maps and distance table, Numa policy hit/miss statistics defines the per-node counters, and NUMA Memory Performance explains the access and memory-side cache files, so every file the ABI lists has its meaning in another entry.
- r10.5 [Arm's Neoverse V2, in AWS's Graviton 4](https://chipsandcheese.com/p/arms-neoverse-v2-in-awss-graviton-4): Trimmed for length: its home is section 14, and the socket link it measures is defined by the CMN-700 manual's gateway ports.
- r10.6 [Page migration](https://docs.kernel.org/mm/page_migration.html): Trimmed for length: move_pages(2) carries per-page migration of a running process, and sysctl kernel numa_balancing carries the automatic form, so the kernel's step-by-step internals were the least a newcomer needs.
- r10.7 [A Look into Intel Xeon 6's Memory Subsystem](https://chipsandcheese.com/p/a-look-into-intel-xeon-6s-memory): Trimmed for length: its home is section 14, and the SNC node boundary it measures across is defined by the Intel Xeon 6 HPC tuning guide.
- r10.8 [ACPI Specification 6.5, chapter 5 (SRAT, SLIT, HMAT)](https://uefi.org/specs/ACPI/6.5/05_ACPI_Software_Programming_Model.html): Rule 4: could not be verified live; curl and the fetch tool both get 403 and the site fronts a human-verification challenge, which was not attempted. numactl -H prints the distance table as the kernel exposes it, and NUMA Memory Performance explains the HMAT-derived attributes.
- r10.9 [NUMA Topology for AMD EPYC Naples Family Processors](https://docs.amd.com/v/u/en-US/NUMA-Topology-for-AMD-EPYC-Naples-Family-Processors): No rule failed; it describes a retired package with no measurements, and the 9004 BIOS guide covers the same interleave modes on shipping parts. Left out for size.
- r10.10 [AMD EPYC 9005 Processor Architecture Overview](https://docs.amd.com/v/u/en-US/58462_amd-epyc-9005-tg-architecture-overview): No rule failed; its fabric and NPS text duplicates the 9004 BIOS guide, which also carries the LLC-as-NUMA and xGMI link settings. Left out for the size cap.
- r10.11 [Intel Xeon CPU Max Series Configuration and Tuning Guide](https://www.intel.com/content/www/us/en/content-details/780889/intel-xeon-cpu-max-series-configuration-and-tuning-guide.html): No rule failed; SNC4 and HBM-as-node material overlaps the Xeon 6 guide and numaperf, and its lower-latency claim for SNC4 comes without a measurement. Left out for size.
- r10.12 [Intel Xeon 6 Processor BIOS NUMA Tuning Guide for Microsoft SQL Server](https://cdrdv2-public.intel.com/845720/845720_1p0.pdf): Workload-specific; the SNC content is a settings table for one database product, and the HPC guide defines SNC properly.
- r10.13 [Technical Overview of the Intel Xeon Scalable processor Max Series](https://www.intel.com/content/www/us/en/developer/articles/technical/xeon-scalable-processor-max-series.html): No rule failed; repeats UPI from the Scalable Family overview and SNC4 from the CPU Max guide. Left out for size.
- r10.14 [Effective Synchronization on Linux/NUMA Systems](https://lameter.com/gelato2005.pdf): No rule failed; it is about lock behaviour on Itanium Altix systems, not placement, and the policy and migration entries cover the section. Left out for scope.
- r10.15 [Thread and Memory Placement on NUMA Systems: Asymmetry Matters](https://www.usenix.org/conference/atc15/technical-session/presentation/lepers): No rule failed; it measures on eight-node parts that link width and direction, not hop count, set the cost between nodes, and is the one candidate that measures the stock migrate_pages cost, but it states no compiler or frequency, and the congestion paper plus the latency checker already carry the lesson that placement is decided by measurement rather than by the distance table. Left out for size.
- r10.16 [Traffic management (ACM Digital Library copy)](https://dl.acm.org/doi/10.1145/2451116.2451157): Rule 4: paywalled copy that also returns 403 to curl; the author's own PDF is used instead.
- r10.17 [NUMA: An Overview (ACM Digital Library copy)](https://dl.acm.org/doi/10.1145/2508834.2513149): Rule 4: duplicate of the magazine's own page, which is the canonical location and the entry used; 403 to curl as well.
- r10.18 [Sub-NUMA Clustering, frankdenneman.nl, `](https://frankdenneman.nl/2022/09/21/sub-numa-clustering/): Rule 1: a secondary explainer, neither the vendor's definition nor a measurement by the writer; the URL now redirects to a domain where the page returns 404.
- r10.19 [Intel Ultra Path Interconnect](https://en.wikipedia.org/wiki/Intel_Ultra_Path_Interconnect): Rule 1: Wikipedia.
- r10.20 [Toward better NUMA scheduling](https://lwn.net/Articles/486858/): Rule 1: reporting on kernel work rather than the work itself; the sysctl document is the primary for automatic balancing.
- r10.21 [Optimizing Applications for NUMA](https://www.intel.com/content/www/us/en/developer/articles/technical/optimizing-applications-for-numa.html): Rules 1 and 4: the URL now redirects to the developer overview page, and the article was a vendor explainer rather than the definition of a mechanism or a measurement.
- r10.22 [mbind(2)](https://man7.org/linux/man-pages/man2/mbind.2.html): No rule failed; the rule that a range policy binds only pages allocated after it is set, unless pages are moved, is stated in the policy document, which defers to this page only for flag detail. Left out for size.
- r10.23 [migrate_pages(2)](https://man7.org/linux/man-pages/man2/migrate_pages.2.html): No rule failed; the whole-process form is covered by move_pages(2) and the numa_balancing document. Left out for size.
- r10.24 [numa(7)](https://man7.org/linux/man-pages/man7/numa.7.html): No rule failed; the one thing it documents that the policy document and numactl do not is the /proc/PID/numa_maps format, the per-mapping page count by node that answers where a running process's pages landed without code. Left out for size, and the first candidate if a slot opens.
- r10.25 [set_mempolicy(2)](https://man7.org/linux/man-pages/man2/set_mempolicy.2.html): No rule failed; task-scope policy is covered by the policy document and numactl. Left out for size.

### Numbers examined

- c10.1 "Latency for a memory access (random access) is about 100 ns. Access to memory on a remote node adds another 50 percent to that number" from https://queue.acm.org/doi/10.1145/2508834.2513149: fields present: 5 workload (random access latency), 6 baseline (local access). Missing: 1 CPU (only "a typical high-end business-class server" with two sockets), 2 cores, 3 freq/turbo/SMT, 4 compiler/flags, 7 method. Verdict: cut. The entry is kept for the mechanisms it ties together, which benchmark 10 supports, and the number is not carried anywhere in the section. Verdict: cut.
- c10.2 "The performance degrades by at most 20% under remote-memory" from https://people.ece.ubc.ca/sasha/papers/asplos284-dashti.pdf (Figure 2a): fields present: 1 CPU (four AMD Opteron 8385, Machine A; the model is given, the microarchitecture is not named), 2 cores (one thread per application, one application at a time, on a sixteen-core box), 3 freq (2.3 GHz), 5 workload (NAS, PARSEC and Metis programs with CPU utilisation above thirty percent), 6 baseline (thread and data co-located on one node), 7 method (relative completion time). Missing: 3 turbo and SMT state (not stated; the part has neither, but the paper does not say so), 4 compiler/flags (not stated anywhere in the paper). The paired thirty percent local-versus-remote latency figure in the same sentence is cited from the paper's reference 7 (Corey, OSDI 2008), not measured in the paper, and is not carried. Verdict: watchlist. The entry's value rests on the measured comparison of congestion against wire delay, not on this figure. Verdict: watchlist.
- c10.3 "Memory latency 476 (Streamcluster under interleaving) against 1197 cycles per request (under first touch), with memory-controller imbalance 8% against 170%" from https://people.ece.ubc.ca/sasha/papers/asplos284-dashti.pdf (Table 1): fields present: 1 CPU (Machine A as above), 2 cores (as many threads as cores), 3 freq (2.3 GHz), 5 workload (Streamcluster and PCA under first touch and under interleave), 6 baseline (the interleaved run), 7 method (hardware counters, average cycles per memory request and standard deviation of controller load as a percentage of the mean). Missing: 3 turbo and SMT state, 4 compiler/flags, 7 repeat count for this table (the ten-run statement belongs to the evaluation in Section 4, not to Table 1). The introduction's rounded sentence about 1000 cycles against 200 is a general statement, not a measurement, and is not carried. Verdict: watchlist. Verdict: watchlist.
- c10.4 The evaluation of https://people.ece.ubc.ca/sasha/papers/asplos284-dashti.pdf (Section 4) runs each experiment ten times on Linux v3.6 under default first touch, manual interleave, AutoNUMA v27 and the paper's own placement, on Machines A and B, with the standard deviation across runs stated per configuration; that is why the reason names the kernel balancer as a baseline. The per-program speedups in Figures 5, 6 and 9 carry the same missing fields as Table 1 (turbo and SMT state, compiler and flags) and no speedup figure is quoted. Verdict: watchlist. Verdict: watchlist.
- c10.5 "STREAM Triad 770343 MB/s at NPS=4 against 726995 MB/s at NPS=1 (two cores per L3, 48 threads, SMT off, boost on)" from https://docs.amd.com/v/u/en-US/58002_amd-epyc-9004-tg-hpc (Tables 8-3 to 8-5): fields present: 1 CPU (two AMD EPYC 9654, Zen 4), 2 cores (24 to 192 threads per row), 3 freq (2.4 GHz base, 3.7 GHz boost stated elsewhere in the guide; boost on or off and SMT on or off per column), 4 compiler/flags (AOCC 4.0.0 with the flags given in the Spack line), 5 workload (STREAM Triad, array size 430080000 doubles, NTIMES 10, DDR5-4800 one DIMM per channel, transparent huge pages always, NUMA balancing off), 6 baseline (the NPS=1 and NPS=2 tables under the same settings), 7 method (STREAM reports the best of NTIMES iterations). Missing: 7 number of repeated runs and the statistic across runs; the guide itself warns of run-to-run and system-to-system variation. Verdict: watchlist. Verdict: watchlist.
- c10.6 "should yield approximately 770GB/s with AOCC or Intel compiler" from https://docs.amd.com/v/u/en-US/58002_amd-epyc-9004-tg-hpc: fields present: 1 CPU (two 96-core EPYC 9654), 5 workload (STREAM, DDR5-4800 one DIMM per channel, eight cores per L3). Missing: 2 cores actually used, 3 turbo and SMT, 4 exact flags, 6 baseline, 7 method. Verdict: cut. The tabulated figures above supersede this rounded statement. Verdict: cut.
- c10.7 "UPI at an operational speed of up to 10.4 GT/s" from https://www.intel.com/content/www/us/en/developer/articles/technical/xeon-processor-scalable-family-technical-overview.html: a link specification, not a measurement; none of the seven fields apply. Verdict: cut. Not carried; the entry is used for the mechanism descriptions only. Verdict: cut.
- c10.8 "xGMI links with speeds up to 32Gbps" from https://docs.amd.com/v/u/en-US/58011-epyc-9004-tg-bios-and-workload: a link specification, not a measurement. Verdict: cut. Not carried. Verdict: cut.
- c10.9 The SNC section of the Xeon 6 HPC guide (https://www.intel.com/content/www/us/en/content-details/858491/intel-xeon-6-with-p-cores-configuration-and-tuning-guide-for-hpc-applications.html, revision 1.4) says SNC "reduces memory access latency and improves memory bandwidth" with no table and none of the seven fields. Verdict: cut. The entry rests on the definition of SNC and the placement checks; the measured cost sits with the Xeon 6 memory subsystem article, whose home is section 14. Verdict: cut.
- c10.10 Local and remote DRAM latency across die boundaries in SNC3 from https://chipsandcheese.com/p/a-look-into-intel-xeon-6s-memory: recorded in section 14 with the fields present (Xeon 6 6985P-C, Granite Rapids, a 96-core cloud instance, peak clock named, latency microbenchmarks against Emerald Rapids and Turin, own method) and missing (4 compiler and flags, 7 run count and statistic, SMT state). Verdict: watchlist. The article is no longer an entry in this section, and no figure is quoted here. Verdict: watchlist.
- c10.11 Remote-socket DRAM and cross-socket cache-line latency from https://chipsandcheese.com/p/arms-neoverse-v2-in-awss-graviton-4: recorded in section 14 with the fields present (Graviton 4, Neoverse V2, 96 cores per socket, single- and dual-socket clocks named, no SMT on the part, latency and core-to-core microbenchmarks against Graviton 3, own method) and missing (4 compiler and flags, 7 run count and statistic). Verdict: watchlist. The article is no longer an entry in this section, and no figure is quoted here. Verdict: watchlist.
- c10.12 https://support.arm.com/documentation/102308/latest/ contributes no performance number; its reason rests on the mechanisms it defines. The same held for https://ieeexplore.ieee.org/document/121510 while it was an entry.

## 11. OS and I/O

### Rejected candidates

- r11.1 [Linux multi-core scalability](http://halobates.de/lk09-scalability.pdf): Rule 1: an overview, not a report. Six pages whose abstract says it discusses known bottlenecks, with no hardware named, no benchmark run and no number, and a lock list as of Linux 2.6.31; the survey shape the rule excludes. Its slot went to the storage API study.
- r11.2 [An Analysis of Linux Scalability to Many Cores](https://www.usenix.org/conference/osdi10/analysis-linux-scalability-many-cores): Fails no source rule; cut for depth. The measured record of kernel lock contention under parallel syscalls on 48 cores with MOSBENCH and sixteen named fixes, kernel 2.6.35; the subsection is at the hard cap and the slot went to the storage API study, which measures the interface a newcomer uses today. Author copy at https://pdos.csail.mit.edu/papers/linux:osdi10.pdf.
- r11.3 [Quantifying The Cost of Context Switch](https://www.usenix.org/legacy/events/expcs07/papers/2-li.pdf): Fails no source rule; cut for depth. Separates the direct cost of a switch from the cache refill after it, but on a Pentium-era Xeon, Linux 2.6.17 and an unoptimised gcc 3.2.2 build; the SOSP paper kept above measures context switch on current kernels.
- r11.4 [SPDK](https://github.com/spdk/spdk): Fails no source rule; cut for depth. The user-space polled NVMe driver, the storage analogue of DPDK; the SYSTOR study kept above measures it against io_uring and libaio, and the bypass subsection is at the hard cap.
- r11.5 [SMP IRQ affinity](https://docs.kernel.org/core-api/irq/irq-affinity.html): Fails no source rule; cut for depth. Its whole content, the /proc/irq smp_affinity file, irqaffinity= and the managed_irq flag, is restated in the CPU Isolation guide's IRQ section, which links it.
- r11.6 [NO_HZ: Reducing Scheduling-Clock Ticks](https://docs.kernel.org/timers/no_hz.html): Fails no source rule; cut for depth. The CPU Isolation guide restates full dynticks, the residual housekeeping tick and the RCU offload overhead and links the document; cut with SMP IRQ affinity to make room for the frequency and idle-state documents, which nothing else in the list covered.
- r11.7 [Earliest Eligible Virtual Deadline First: A Flexible and Accurate Mechanism for Proportional Share Resource Allocation](https://people.eecs.berkeley.edu/~istoica/papers/eevdf-tr-95.pdf): Fails no source rule; cut for depth. The origin of eligible time, virtual deadline and the lag bound; the kernel document kept as the entry restates all three, and the subsection is at the hard cap.
- r11.8 [CFS Bandwidth Control](https://docs.kernel.org/scheduler/sched-bwc.html): Fails no source rule; cut for depth. Defines quota, period, the per-CPU slice and burst behind cpu.max; the Control Group v2 entry names the file, and the subsection is at the hard cap.
- r11.9 [rdma-core](https://github.com/linux-rdma/rdma-core): Fails no source rule; scope. The list excludes networking beyond the kernel boundary, and verbs are a NIC data path with no kernel on it, so the entry taught nothing about the CPU and had no measured source beside it.
- r11.10 [Cache QoS: From concept to reality in the Intel Xeon processor E5-2600 v3 product family](https://ieeexplore.ieee.org/document/7446102): Fails no source rule; cut for depth. The implementers' measured report of cache monitoring and allocation on shipped silicon, paywalled with no author copy found; the subsection is at the hard cap with the MPAM document in the slot.
- r11.11 [Efficient IO with io_uring at the bare host](https://kernel.dk/io_uring.pdf): Rule 4: not live; the bare host serves a certificate for another domain over TLS and returns 404 over plain HTTP. Replaced by the www host, which serves the PDF.
- r11.12 [AMD64 Technology Platform Quality of Service Extensions (56375)](https://www.amd.com/content/dam/amd/en/documents/processor-tech-docs/other/56375_1_03_PUB.pdf): Rule 4: 404; the vendor withdrew the standalone specification and folded it into APM Volume 2 chapter 19, itself now under Rejected for length. The third-party mirror at kib.kiev.ua fails the official-copy rule.
- r11.13 [Arm MPAM System Component Specification (IHI0099)](https://support.arm.com/documentation/ihi0099/latest/): Rule 4: 403 to curl and a single-page shell to fetch, so the page cannot be confirmed as the document. Arm's page for the older supplement DDI0598 says it is retired, with the CPU side folded into the Arm ARM already in the list; the kernel's arm64 MPAM document is the entry.
- r11.14 [AMD Server PQOS White Paper for AMD EPYC 9004 and 9005 Series Processors](https://docs.amd.com/v/u/en-US/69127_1.00_PQOSWP): Rule 1: vendor enablement paper that applies the APM Volume 2 chapter, itself now under Rejected for length. Its own reference table names that manual as the architectural specification, and its use-case chapter is marketing.
- r11.15 [io_uring(7)](https://man7.org/linux/man-pages/man7/io_uring.7.html): Fails no source rule; cut for depth. The page is maintained in the liburing repository, itself now under Rejected for length, beside io_uring_setup(2), io_uring_enter(2) and io_uring_register(2), and the subsection is at the hard cap.
- r11.16 [Page Table Isolation (PTI)](https://docs.kernel.org/arch/x86/pti.html): Fails no source rule; cut for depth. Its overhead section is a paragraph, and the SOSP paper listed above measures the same mechanism with the seven fields nearly complete.
- r11.17 [CFS Scheduler](https://docs.kernel.org/scheduler/sched-design-CFS.html): Fails no source rule; cut for depth. The document itself says it is making room for EEVDF and links to the EEVDF page, which is the entry kept.
- r11.18 [sched(7)](https://man7.org/linux/man-pages/man7/sched.7.html): Fails no source rule; cut for depth. Policy overview only; the isolation and cgroup entries cover what a performance engineer acts on.
- r11.19 [CPUSETS (cgroup v1)](https://docs.kernel.org/admin-guide/cgroup-v1/cpusets.html): Fails no source rule; superseded. The CPU Isolation guide recommends cgroup v2 cpuset partitions, covered by the Control Group v2 entry.
- r11.20 [Scaling in the Linux Networking Stack](https://docs.kernel.org/networking/scaling.html): Fails no source rule; cut for depth. RSS, RPS, RFS and XPS matter, but NAPI and the CPU Isolation guide's IRQ section are what a newcomer needs first.
- r11.21 [Meltdown](https://meltdownattack.com/meltdown.pdf): Fails no source rule; cut for depth. Its evaluation is about the attack; the cost of the mitigation is measured better by the SOSP paper kept above.
- r11.22 [Mental Models for modern program tuning](http://halobates.de/applicative-mental-models.pdf): Rule 1: slide deck without a paper.
- r11.23 [Modern Locking](http://halobates.de/modern-locking.pdf): Rule 1: slide deck without a paper; the lock scaling white paper it extracts from belongs to the concurrency section.
- r11.24 [pmu-tools](https://github.com/andikleen/pmu-tools): Fails no source rule; section boundary. It is a measurement tool and belongs with perf and top-down analysis, not the kernel boundary.
- r11.25 [A NUMA API for Linux](http://halobates.de/numaapi3.pdf): Fails no source rule; section boundary. Memory placement is section 10.
- r11.26 [The Linux scheduler: a decade of wasted cores (ACM DL)](https://dl.acm.org/doi/10.1145/2901318.2901326): Rule 4: paywalled copy where the author's own copy exists; the author copy is the entry.
- r11.27 [The eXpress data path (ACM DL)](https://dl.acm.org/doi/10.1145/3281411.3281443): Rule 4: returns 403 to both curl and fetch so the copy cannot be verified, although it is open access there under CC BY-NC-ND; the authors' repository with the PDF and the experimental data is the entry.
- r11.28 [An analysis of performance evolution of Linux's core operations (ACM DL)](https://dl.acm.org/doi/10.1145/3341301.3359640): Rule 4: paywalled copy where the author's own copy exists.
- r11.29 [Understanding Modern Storage APIs (ACM DL)](https://dl.acm.org/doi/10.1145/3534056.3534945): Rule 4: publisher copy behind 403; the author group's own copy is the entry.
- r11.30 [Kernel vs. User-Level Networking (ACM DL)](https://dl.acm.org/doi/10.1145/3626780): Rule 4: publisher copy behind 403; the author's own paper page, which links the preprint and the supplementary code and data, is the entry.
- r11.31 [Heracles: Improving Resource Efficiency at Scale (Google Research)](https://research.google/pubs/heracles-improving-resource-efficiency-at-scale/): Rule 4: landing page rather than the paper; the author's own copy is the entry.
- r11.32 [CPI2 (Google Research)](https://research.google/pubs/cpi2-cpu-performance-isolation-for-shared-compute-clusters/): Rule 4: landing page rather than the paper; the author's own copy is the entry.
- r11.33 [The Linux Scheduler: a Decade of Wasted Cores (the morning paper)](https://blog.acolyer.org/2016/04/26/the-linux-scheduler-a-decade-of-wasted-cores/): Rule 1: secondary summary of a paper that is itself listed.
- r11.34 [Ringing in a new asynchronous I/O API (LWN)](https://lwn.net/Articles/776703/): Rule 1: journalism about the interface; the design note it reports on is the entry, and it says to prefer itself where the two differ.
- r11.35 [lmbench](https://sourceforge.net/projects/lmbench/): Fails no source rule; section boundary. The tool that made lat_syscall a standard number is listed with its paper in the benchmarks section; the project's own host at bitmover.com does not answer, and the SourceForge page is the live home.
- r11.36 [liburing](https://github.com/axboe/liburing): Trimmed for length: the io_uring design note states the SQPOLL and IOPOLL modes and the storage API study measures them, and the repository stays the home of the io_uring man pages, io_uring(7) among them.
- r11.37 [Hardware vulnerabilities](https://docs.kernel.org/admin-guide/hw-vuln/index.html): Trimmed for length: the preamble carries the warning to record the mitigation state, and the Linux core operations paper names each mitigation and measures its cost. The microarchitecture section defers the page to section 11, so it is the first to restore if the cap lifts.
- r11.38 [sched_setaffinity(2)](https://man7.org/linux/man-pages/man2/sched_setaffinity.2.html): Trimmed for length: pinning now rests on the CPU Isolation guide, which builds its recipe on cpuset partitions and isolcpus, and on the Control Group v2 entry, which states how a cpuset bounds a task's affinity.
- r11.39 [CPU Idle Time Management](https://docs.kernel.org/admin-guide/pm/cpuidle.html): Trimmed for length: the preamble records the idle state beside any number, the CPU Isolation guide links the page from its checklist, and section 12 measures the exit cost of idle states with Tales of the Tail and the timerlat tracer.
- r11.40 [Netlink interface for ethtool](https://docs.kernel.org/networking/ethtool-netlink.html#coalesce-set): Trimmed for length: the NAPI entry states that batching normally comes from the device's own interrupt coalescing and defines the software half, and the Kernel vs. User-Level Networking paper measures the interrupt cost that coalescing trades against latency.
- r11.41 [Onload](https://github.com/Xilinx-CNS/onload): Trimmed for length: the bypass points survive in AF_XDP, which hands frames to user space through the kernel, and DPDK, which takes the NIC away from it, and the Kernel vs. User-Level Networking paper measures a user-space stack against the kernel's.
- r11.42 [AMD64 Architecture Programmer's Manual Volume 2](https://docs.amd.com/v/u/en-US/24593_3.45_APM_Vol2_PUB): Trimmed for length: the Intel RDT specification defines the same model of classes of service, masks and bandwidth limits, and the resctrl document states where AMD's semantics differ. Chapter 19 remains the home of AMD's MSRs, and the ISA section is its home if the manual is listed again.
- r11.43 [intel-cmt-cat](https://github.com/intel/intel-cmt-cat): Trimmed for length: the resctrl document defines the schemata and monitoring files the tool writes and reads, and the Intel RDT specification defines the MSRs it drives without resctrl.

### Numbers examined

- c11.1 "peak per-core performance with io_uring is now approximately 1700K 4k IOPS" and "aio reaches a performance cliff much lower than that, at 608K" from https://www.kernel.dk/io_uring.pdf: fields present: 5 workload random 4k reads with polled I/O from a block device (t/io_uring in fio), 6 baseline Linux native aio on the same test, 7 method partial (fio engine, no run count or statistic). Missing: 1 CPU model and microarchitecture, 2 cores used, 3 frequency, turbo and SMT, 4 compiler and flags, most of 7. Verdict: watchlist. The entry earns its place on the design; the number is not used. Verdict: watchlist.
- c11.2 "With just one core, SPDK achieves 305 KIOPS versus the 171 KIOPS and 145 KIOPS of the best io_uring alternative and libaio" and "each polling kernel thread needs a dedicated CPU to achieve the best performance" from https://atlarge-research.com/pdfs/2022-systor-apis.pdf: fields present: 1 two Intel Xeon E5-2630 as stated, ten cores per socket at 2.2 GHz, which matches the v4 Broadwell part, though the paper names neither suffix nor microarchitecture, 2 one and two cores per fio job, up to twenty across the sweep, 3 2.2 GHz with intel_pstate=disable, intel_idle.max_cstate=1 and hyperthreading disabled, 5 fio 3.28 4 KiB unbuffered random reads on Intel DC P3600 NVMe drives, kernel 5.13, Spectre and Meltdown patches on, SPDK at a named commit, 6 libaio on the same drives, 7 IOPS and median latency against queue depth, io_uring in default, SQPOLL and IOPOLL modes. Missing: microarchitecture in 1, 4 compiler and flags, run count and statistic in 7. Verdict: watchlist. Verdict: watchlist.
- c11.3 "user-mode IPC degrades by up to 65% when executing a pwrite every 1,000-2,000 instructions" and "total round-trip time for the gettsc system call is 150 cycles" from https://www.usenix.org/conference/osdi10/flexsc-flexible-system-call-scheduling-exception-less-system-calls: fields present: 1 Intel Core i7, Nehalem (exact model not given), 2 four cores, 3 2.3 GHz with TurboBoost and Hyper-Threading disabled, 5 Xalan from SPEC CPU2006 and SPEC JBB with a pwrite injected at a controlled instruction interval, 6 the same workload with no injected syscalls, 7 hardware performance counters reading user-mode IPC, average of five runs, Linux 2.6.33. Missing: exact CPU model, 4 compiler and flags. Verdict: watchlist. Verdict: watchlist.
- c11.4 "With KPTI, the lower-bound of the constant cost is on the order of 400-500 cycles, whereas without KPTI, the kernel entry and exit overhead is less than 100 cycles" from https://www.eecg.toronto.edu/~stumm/Papers/Ren-sosp-19.pdf: fields present: 1 Intel Xeon E5-2630 v3, Haswell, 3 2.40 GHz nominal, 5 LEBench empty system call, 6 the same kernel version without the KPTI patch, PCID enabled, 7 cycles around the call, 10,000 repetitions, K-best with K of 5 and 5% tolerance, Ubuntu kernel configurations. Missing: 2 cores used, turbo and SMT state in 3, 4 compiler and flags for the benchmark (only the kernel build toolchain is named). Verdict: watchlist. Verdict: watchlist.
- c11.5 "13% higher latency for kernel make, and a 14-23% decrease in TPC-H throughput" from https://people.ece.ubc.ca/sasha/papers/eurosys16-final29.pdf: fields present: 1 AMD Opteron 6272, Bulldozer, 2 64 cores on 8 NUMA nodes, 3 2.1 GHz nominal, 5 kernel make alongside R processes, and TPC-H on a commercial DBMS, 6 the unpatched scheduler on Linux 3.17 through 4.3, 7 completion time before and after each fix. Missing: turbo and SMT state in 3, 4 compiler and flags, run count and statistic in 7. Verdict: watchlist. The entry stands on the bugs and the checker, not the percentages. Verdict: watchlist.
- c11.6 "The baseline performance of XDP for a single core is 24 Mpps, while for DPDK it is 43.5 Mpps" from https://github.com/tohojo/xdp-paper: fields present: 1 Intel Xeon E5-1650 v4, Broadwell, 2 one to five physical cores, 3 3.60 GHz with Hyper-Threading disabled, 5 packet drop under load from a TRex generator through Mellanox ConnectX-5 100 Gbps adapters, pre-release Linux 4.18 with retpoline and full preemption disabled, 6 DPDK and the Linux stack on the same machine, 7 a single run per point, stated repeatable, with the full configuration in the repository. Missing: turbo state in 3, 4 compiler and flags, run count and statistic in 7. Verdict: watchlist. Verdict: watchlist.
- c11.7 "up to 45% increased throughput without compromising tail latency" from https://cs.uwaterloo.ca/~mkarsten/papers/sigmetrics2024.html (preprint PDF linked from the page): fields present: 1 two Intel Xeon E5-2680, eight cores each, Sandy Bridge by the model number though the paper does not name the microarchitecture, 2 one core for the Memcached experiments because F-Stack is single-threaded, more for the web server runs, 3 fixed at 2.7 GHz with Turbo Boost disabled and SMT avoided by placing threads on the first hardware thread of each core, 5 Memcached under Mutilate and a web server under wrk, kernel 5.15 with the polling patch, Mellanox ConnectX-3 10 GbE, seven identical client machines, 6 unmodified Linux and F-Stack on DPDK on the same server, 7 throughput at a tail-latency bound with the run scripts and data published on the supplementary page. Missing: microarchitecture in 1, 4 compiler and flags, run count and statistic in 7. Verdict: watchlist. Verdict: watchlist.
- c11.8 "Heracles achieves an effective machine utilization of 90% averaged across all colocation combinations" from https://csl.stanford.edu/~christos/publications/2015.heracles.isca.pdf: fields present: 1 dual-socket Intel Xeon, Haswell, with way-partitioned LLC, 3 2.3 GHz nominal with Hyper-Threading on, 5 production websearch, ml_cluster and memkeyval as latency-critical jobs with batch antagonists, 6 the latency-critical job alone at each load point, 7 controller run across a load sweep with tail latency against the SLO. Missing: exact CPU model and core count in 1 and 2, turbo state in 3, 4 compiler and flags. Verdict: watchlist. Not a CPU microbenchmark; the entry stands on the control structure. Verdict: watchlist.

## 12. Tail latency and production systems

### Rejected candidates

- r12.1 [Smart Batching](https://mechanical-sympathy.blogspot.com/2011/10/smart-batching.html): Rule 1 borderline: a pattern post whose only figures are a worked arithmetic example on an assumed write cost, with no code and no measurement; the pattern is implemented and measured in the kept Disruptor paper and carried through the kept Aeron, so the entry added a name, not a source.
- r12.2 [Simple Binary Encoding](https://github.com/aeron-io/simple-binary-encoding): No rule failed; cut for scope. A serialisation codec whose mechanism the rest of the list does not explain, and whose reason could only justify it by reference to the transport it serves; the slot went to a measured C++ queue for an audience that reads C.
- r12.3 [Optimizing RHEL 9 for Real Time for low latency operation](https://docs.redhat.com/en/documentation/red_hat_enterprise_linux_for_real_time/9/html/optimizing_rhel_9_for_real_time_for_low_latency_operation/index): No rule failed; cut for depth under the entry cap. Its isolation and interrupt settings are the CPU Isolation, SMP IRQ affinity and NO_HZ entries in section 11, the per-CPU kthreads document, trimmed below, names what those leave, the kept Tales of the Tail measures the power-state effect it warns of, and the firmware and wakeup tests it prescribes ship in the kept rt-tests.
- r12.4 [Heracles: Improving Resource Efficiency at Scale](https://csl.stanford.edu/~christos/publications/2015.heracles.isca.pdf): No rule failed; cross-section duplicate. Listed in section 11 under Noisy neighbours, where its reason names the per-resource controls; the copy here restated the same finding from the service's side.
- r12.5 [Profiling a warehouse-scale computer](https://static.googleusercontent.com/media/research.google.com/en//pubs/archive/44271.pdf): No rule failed; cross-section duplicate. Listed in section 15 under Methodology, where its fleet profile is the point; a fleet cycle profile is neither a tail-latency nor a load-generation source, and the copy here restated the same finding.
- r12.6 [perf-sched](https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/tree/tools/perf/Documentation/perf-sched.txt): No rule failed; left out under the entry cap. Its per-wakeup report of wait time and scheduling delay for an application's own threads is the one measurement the jitter subsection lacks, and the bcc collection kept in section 5 carries runqlat, which reports the same quantity.
- r12.7 [clock_gettime(2)](https://man7.org/linux/man-pages/man2/clock_gettime.2.html): No rule failed; section boundary. Which clock a recorder reads and what the read costs belong beside vdso(7) in section 11, which defines the path clock_gettime takes without a mode switch.
- r12.8 [Low Latency Tuning Guide](https://rigtorp.se/low-latency-guide/): Rule 1: a tuning checklist that cites kernel and vendor documents rather than reporting a mechanism or measurement of its own; its one jitter figure has none of the seven fields (rule 3). The same author's virtual memory and ring buffer reports are kept because they carry code and measurements, and the kernel documents the checklist cites are entries in their own right, in section 11 and under Where jitter comes from.
- r12.9 [HdrHistogram project site](https://hdrhistogram.org/): Rule 4: not the canonical home, a frameset that embeds the hdrhistogram.github.io landing page, and the histogram log format it was credited with is defined in the kept repository, so the entry duplicated one already kept; the plotter's own page is https://hdrhistogram.github.io/HdrHistogram/plotFiles.html.
- r12.10 [jHiccup](https://github.com/giltene/jHiccup): No rule failed; left out under the entry cap because the kept osnoise tracer measures the same pauses in-kernel and names their source, and LatencyUtils, trimmed below, carries the same pause detector for in-process use.
- r12.11 [cyclictest howto on the Linux Foundation RT wiki](https://wiki.linuxfoundation.org/realtime/documentation/howto/tools/cyclictest/start): Rule 4: cannot be verified; curl and fetch both receive a Cloudflare challenge page. The rt-tests README carries the same usage and interpretation guidance.
- r12.12 [Why Mechanical Sympathy?](https://mechanical-sympathy.blogspot.com/2011/07/why-mechanical-sympathy.html): Rule 1: the post that names the idea, but it reports no mechanism and no measurement; the posts kept do.
- r12.13 [Achieving Rapid Response Times in Large Online Services](https://static.googleusercontent.com/media/research.google.com/en//people/jeff/Berkeley-Latency-Mar2012.pdf): Rule 1: slide deck without a paper; the same material became the CACM paper that is kept.
- r12.14 [Latency Numbers Every Programmer Should Know](https://gist.github.com/jboner/2841832): Rule 1 and rule 3: an unsourced aggregate table with no CPU, method or baseline for any figure.
- r12.15 [Your Load Generator is Probably Lying to You](https://highscalability.com/your-load-generator-is-probably-lying-to-you-take-the-red-pi/): Rule 1: aggregator write-up of the kept talk.
- r12.16 [Everything You Know About Latency Is Wrong](https://bravenewgeek.com/everything-you-know-about-latency-is-wrong/): Rule 1: secondary summary of the kept talk.
- r12.17 [Principles of Mechanical Sympathy](https://martinfowler.com/articles/mechanical-sympathy-principles.html): Rule 1: secondary explainer of the original posts, which are kept.
- r12.18 [On Coordinated Omission](https://www.scylladb.com/2021/04/22/on-coordinated-omission/): Rule 1: vendor blog restating a definition whose original source is kept.
- r12.19 [How NOT to Measure Latency, QCon San Francisco 2015](https://www.infoq.com/presentations/latency-response-time/): Duplicate of the kept talk, recorded at another conference; one recording is kept and it is the conference channel's own.
- r12.20 [How NOT to Measure Latency, QCon London 2013](https://www.infoq.com/presentations/latency-pitfalls/): Duplicate of the kept talk, an earlier and shorter recording.
- r12.21 [Understanding Latency](https://www.youtube.com/watch?v=9MKY4KypBzg): Rule 1 borderline: a talk with no paper or code that repeats the measurement material of the kept talk.
- r12.22 [The tail at scale on ACM DL](https://dl.acm.org/doi/10.1145/2408776.2408794): Rule 4: publisher copy that answers 403 to automated clients; the author's own copy is kept.
- r12.23 [Attack of the killer microseconds on ACM DL](https://dl.acm.org/doi/10.1145/3015146): Rule 4: as above; the author's own copy is kept.
- r12.24 [Profiling a warehouse-scale computer on ACM DL](https://dl.acm.org/doi/10.1145/2749469.2750392): Rule 4: publisher copy behind 403; the copy hosted by the authors' employer is the entry in section 15.
- r12.25 [Reconciling high server utilization and sub-millisecond quality-of-service on ACM DL](https://dl.acm.org/doi/10.1145/2592798.2592821): Rule 4: publisher copy behind 403; the authors' own copy is kept.
- r12.26 [Heracles on ACM DL](https://dl.acm.org/doi/10.1145/2749469.2749475): Rule 4: publisher copy behind 403; the authors' own copy is the entry in section 11.
- r12.27 [TailAtScale.pdf on a course site](https://cseweb.ucsd.edu/classes/fa16/cse291-g/applications/ln/TailAtScale.pdf): Rule 4: third-party mirror of a paper the author hosts.
- r12.28 [2015-kanev.pdf on gwern.net](https://gwern.net/doc/cs/hardware/2015-kanev.pdf): Rule 4: third-party mirror of a paper the authors' employer hosts.
- r12.29 [Disruptor 1.0 PDF](https://lmax-exchange.github.io/disruptor/files/Disruptor-1.0.pdf): Duplicate of the kept paper page, which is the project's maintained copy of the same document.
- r12.30 [LMAX Disruptor repository](https://github.com/LMAX-Exchange/disruptor): No rule failed; left out under the entry cap because the kept paper page is the design's primary document and links the repository.
- r12.31 [hiccups](https://github.com/rigtorp/hiccups): No rule failed; left out under the entry cap because the kept osnoise tracer measures the same quantity in-kernel and names the source of each event.
- r12.32 [hwlat detector](https://docs.kernel.org/trace/hwlat_detector.html): No rule failed; left out under the entry cap because firmware stalls are counted by the kept osnoise tracer and hwlatdetect ships inside the kept rt-tests.
- r12.33 [NO_HZ: Reducing Scheduling-Clock Ticks](https://docs.kernel.org/timers/no_hz.html): No rule failed; left out under the entry cap because the per-CPU kthreads document, trimmed below, sets out the same nohz_full requirements, and it is an entry in section 11.
- r12.34 [CPU Idle Time Management](https://docs.kernel.org/admin-guide/pm/cpuidle.html): No rule failed; left out under the entry cap because the kept Tales of the Tail measures the effect of the idle states it defines and the timerlat tracer, trimmed below, exposes their exit cost in the IRQ part of a wakeup.
- r12.35 [Transparent Hugepage Support](https://docs.kernel.org/admin-guide/mm/transhuge.html): No rule failed; left out under the entry cap because the kept virtual memory report measures the compaction and fault stalls the document describes.
- r12.36 [OSADL latency plots](https://www.osadl.org/Latency-plots.latency-plots.0.html): No rule failed; a primary and continuously measured cyclictest dataset, left out under the entry cap.
- r12.37 [LatencyUtils](https://github.com/LatencyUtils/LatencyUtils): Trimmed for length: the pause it corrects for is the recording error Coordinated Omission defines, and HdrHistogram's expected-interval recording carries the same correction for in-process use.
- r12.38 [Timerlat tracer](https://docs.kernel.org/trace/timerlat-tracer.html): Trimmed for length: rt-tests measures the same periodic wakeup latency from user space, and the osnoise tracer names the IRQ, softirq and thread events that lengthen it.
- r12.39 [Reducing OS jitter due to per-cpu kthreads](https://docs.kernel.org/admin-guide/kernel-per-CPU-kthreads.html): Trimmed for length: the osnoise tracer names each kthread or interrupt that wakes an isolated CPU at runtime, and the boot parameters it lists belong with the CPU Isolation entry in section 11, the home of the isolation recipe.
- r12.40 [mutilate](https://github.com/leverich/mutilate): Trimmed for length: the paper listed beside it, Reconciling High Server Utilization and Sub-millisecond Quality-of-Service, describes the same master and agent design that keeps client queueing out of the measured tail, and wrk2 carries the open-loop timing.

### Numbers examined

- c12.1 "63% of user requests will take more than one second" from https://www.barroso.org/publications/TheTailAtScale.pdf: arithmetic on a hypothetical, not a measurement: one server in a hundred at its 99th percentile, fanned out across a hundred servers. Fields present: none of the seven apply. Verdict: not quoted in the annotation; the entry rests on the argument, so core. Verdict: not_quoted.
- c12.2 "99.9th-percentile latency ... from 1,800ms to 74ms while sending just 2% more requests" from https://www.barroso.org/publications/TheTailAtScale.pdf: fields present: 5 workload (read of a thousand keys from a BigTable spread over a hundred servers, hedge after a fixed delay), 6 baseline (no hedging). Missing: 1 CPU, 2 cores, 3 freq/turbo/SMT, 4 compiler/flags, 7 method (runs, statistic). Verdict: watchlist; number not quoted. Verdict: watchlist.
- c12.3 "reduces median latency by 16% ... nearly 40% reduction at the 99.9th-percentile latency" from https://www.barroso.org/publications/TheTailAtScale.pdf: fields present: 5 workload (small uncached BigTable read, three replicas, tied request after a fixed delay), 6 baseline (no tied requests). Missing: 1, 2, 3, 4, 7. Verdict: watchlist; number not quoted. Verdict: watchlist.
- c12.4 "Table 1 ... datacenter networking: O(1us), high-end flash: O(10us), new NVM memories: O(1us)" from https://www.barroso.org/publications/AttackoftheKillerMicroseconds.pdf: orders of magnitude, not measurements; no machine. Missing: 1, 2, 3, 4, 5, 6, 7. Verdict: not quoted; the entry rests on the argument, so core. Verdict: not_quoted.
- c12.5 "for each of those ~536msec freezes, there are ~535 missing observations ... the 99.99%'ile is at least 582 msec" from https://groups.google.com/g/mechanical-sympathy/c/icNZJejUHfE: arithmetic on a third party's published log4j2 benchmark table, not a measurement of the author's own. Fields present: 5 workload (async logging benchmark), 7 method (recount of samples per stall). Missing: 1, 2, 3, 4, 6. Verdict: not quoted; the entry rests on the definition, so core. Verdict: not_quoted.
- c12.6 "value recording times as low as 3-6 nanoseconds on modern (circa 2012) Intel CPUs" from https://github.com/HdrHistogram/HdrHistogram: fields present: none beyond a vendor and a year. Missing: 1, 2, 3, 4, 5, 6, 7. Verdict: cut; the entry rests on the data structure's guarantees, not on this number. Verdict: cut.
- c12.7 "worst-case latency of 39.242 milliseconds ... on a realtime-enabled system ... 18 microseconds" from https://git.kernel.org/pub/scm/utils/rt-tests/rt-tests.git/: illustrative output in the README, no machine named. Missing: 1, 2, 3, 4, 5, 6, 7 (the README itself says a run without stress is meaningless). Verdict: cut; entry stays core as the tool. Verdict: cut.
- c12.8 "median latency of 11 us and a 99.9th percentile latency of 32 us at 80% utilization on a four-core system ... single-core system has a median latency of 100 us and a 99.9th percentile latency of 5 ms" from https://drkp.net/papers/latency-socc14.pdf: fields present: 1 CPU (two Intel Xeon L5640, Westmere-EP, in a Dell PowerEdge R610), 2 cores (four for the tuned case, one for the baseline, others offlined), 3 freq (2.27 GHz; SMT disabled for the single-core runs; turbo not stated), 5 workload (Memcached 1.4.15, 64-byte keys, 1024-byte values, 90/10 read/write, open-loop Poisson clients, Ubuntu 12.04, kernel 3.2.0), 6 baseline (default Linux configuration), 7 method (NIC-to-NIC timestamps written into each packet by a modified kernel and driver, compared against an M/D/c simulation fed the measured arrival trace). Missing: 4 compiler and flags, turbo state. Verdict: watchlist for the number; entry core on the method. Verdict: watchlist.
- c12.9 "latency spikes ... up to ~900 us" from https://rigtorp.se/virtual-memory/: fields present: 1 CPU (AMD Ryzen 3900X, Zen 2), 5 workload (writes to a file-backed mapping while page cache writeback runs; a munmap loop for TLB shootdowns; Linux 5.7, Samsung 970 EVO NVMe), 6 baseline (anonymous memory without writeback), 7 method (timestamp loop, code published). Missing: 2 cores used, 3 frequency, turbo and SMT state, 4 compiler and flags. Verdict: watchlist; number not quoted, entry rests on the code and the mechanism. Verdict: watchlist.
- c12.10 "in one example ... the 99'th percentile ... 200x larger" from https://github.com/giltene/wrk2: an illustrative comparison in the README with a stated stall length but no server named. Missing: 1, 2, 3, 4, 6, 7 (5 partly: constant-rate HTTP load). Verdict: cut; entry stays core as the tool. Verdict: cut.
- c12.11 "Total QPS = 318710.8 ... read ... 99th 1170.6" from https://github.com/leverich/mutilate: illustrative output in the README with the client thread and connection counts stated but no server, CPU or kernel named. Missing: 1, 2, 3, 4, 6, 7 (5 partly: 24 threads and 8 connections against one memcached). Verdict: cut; the entry is trimmed for length. Verdict: cut.
- c12.12 "Principle (i): For a given load, mean response times are significantly lower in closed systems than in open systems" from https://www.usenix.org/conference/nsdi-06/open-versus-closed-cautionary-tale: a qualitative principle backed by implementation runs. Fields present: 1 CPU (Intel Pentium III 700 MHz for the Apache case, 2.4 GHz Pentium 4 for PostgreSQL), 5 workload (static HTTP, e-commerce database, auction site), 6 baseline (open versus closed generator at equal load), 7 method (measured response time versus load, plus simulation). Missing: 2 cores (single-core parts, not stated), 3 turbo and SMT (not applicable, not stated), 4 compiler and flags. Verdict: no number quoted; entry core on the principles. Verdict: other.
- c12.13 "substantial queuing delay is observed at loads above 70% ... the maximum memcached load we can provision for co-located servers is 60% of peak" from https://csl.stanford.edu/~christos/publications/2014.mutilate.eurosys.pdf: fields present: 1 CPU (two Intel Xeon L5640, Westmere-EP), 2 cores (twelve across two sockets, the share given to memcached varies by experiment), 3 freq (2.27 GHz; DVFS and the C6 idle state disabled; SMT state not stated), 5 workload (memcached 1.4.15 on Linux 3.5.0, exponentially paced requests from the mutilate generator on twenty clients, 30-byte keys and 200-byte values, no pipelining, SPEC CPU2006 and cache microbenchmarks as co-located antagonists), 6 baseline (memcached alone on the server), 7 method (95th percentile latency against load, with a QoS bound of five times the low-load value). Missing: 4 compiler and flags, SMT state, runs and statistic. Verdict: watchlist; number not quoted, entry rests on the method and the attribution. Verdict: watchlist.
- c12.14 The tail latency against load curves in Section V of https://people.csail.mit.edu/sanchez/papers/2016.tailbench.iiswc.pdf, linked from https://tailbench.csail.mit.edu/: fields present: 1 CPU (Intel Xeon E5-2670, Sandy Bridge), 2 cores (eight), 3 freq (2.4 GHz nominal; TurboBoost and deep sleep states disabled; SMT state not stated), 5 workload (the eight suite applications under the harness on Ubuntu 14.04 with Linux 4.2.3, exponentially distributed inter-arrivals), 7 method (sojourn time recorded per request by the harness, 95th percentile against offered load). Missing: 4 compiler and flags, 6 baseline (a characterisation, no comparison), SMT state, runs and statistic. Verdict: watchlist; no number quoted, entry rests on the suite and the harness. Verdict: watchlist.
- c12.15 "300 ms one thread, 10,000 ms one thread with lock, 118,000 ms two threads with lock, 500 million increments on a 2.4GHz Westmere" from https://mechanical-sympathy.blogspot.com/2011/09/single-writer-principle.html: fields present: 1 CPU (Westmere, model not named), 3 freq (2.4 GHz; turbo and SMT not stated), 5 workload (64-bit counter increment), 6 baseline (uncontended single thread). Missing: 2 cores, 4 JVM version and flags, 7 method (runs, statistic). Verdict: watchlist; number not quoted, entry rests on the rule. Verdict: watchlist.
- c12.16 "~45ns ... C++ ... 50ns ... Java ... between 2 cores on a 2.2 GHz machine" from https://mechanical-sympathy.blogspot.com/2011/08/inter-thread-latency.html: fields present: 1 CPU (Sandy Bridge laptop, model not named), 2 cores (two, pinned with taskset), 3 freq (2.2 GHz; turbo and whether the two logical CPUs share a core not stated), 4 compiler (Oracle JDK 1.7.0; g++ -O3), 5 workload (ping-pong on two counters, five hundred million exchanges), 7 method (average over the run). Missing: 6 baseline (none), exact model, turbo, SMT layout. Verdict: watchlist; number not quoted, entry rests on the code and the mechanism. Verdict: watchlist.
- c12.17 "5513850 ops/s ... 112287037 ops/s" from https://rigtorp.se/ringbuffer/: fields present: 1 CPU (AMD Ryzen 9 3900X, Zen 2), 2 cores (two threads pinned to different core complexes), 4 compiler (g++ -O3 -march=native -std=c++20), 5 workload (single-producer single-consumer ring buffer push and pop of integers), 6 baseline (the same queue with aligned indices but no cached copies), 7 method (perf stat with cache-miss and coherence counters). Missing: 3 frequency, turbo and SMT state, runs and statistic. Verdict: watchlist; number not quoted, entry rests on the code and the counters. Verdict: watchlist.
- c12.18 "Serial 100us best, 500us average, 1,000us worst; smart batching 100us, 150us, 200us" from https://mechanical-sympathy.blogspot.com/2011/10/smart-batching.html: a worked arithmetic example with an assumed write cost, not a measurement. Missing: 1, 2, 3, 4, 6, 7. Verdict: cut; the entry is rejected on the same ground. Verdict: cut.
- c12.19 "mean latency ... 3 orders of magnitude lower ... approximately 8 times more throughput" from https://lmax-exchange.github.io/disruptor/disruptor.html: fields present: 1 CPU (Intel Core i7-2720QM, Sandy Bridge, and a Nehalem part for throughput; an AMD EPYC 9374F added in the maintained copy), 3 freq (2.2 GHz and 2.8 GHz; turbo and SMT not stated), 4 runtime (Java 1.6.0_25 64-bit; OpenJDK 11.0.24 for the newer table), 5 workload (unicast and pipeline topologies, five hundred million messages; fifty million events at one microsecond spacing for latency), 6 baseline (ArrayBlockingQueue), 7 method (best of three runs). Missing: 2 cores used, turbo and SMT state, heap and GC settings. Verdict: watchlist; number not quoted, entry rests on the design. Verdict: watchlist.

## 13. Inference on CPU

### Rejected candidates

- r13.1 [ruy](https://github.com/google/ruy): Scope and the cap: its README targets the small and rectangular shapes of mobile inference under TensorFlow Lite on Arm, and the framework path on Arm server parts is carried by the listed oneDNN entry, which calls Compute Library there.
- r13.2 [Apple AMX Instruction Set](https://github.com/corsix/amx): Scope: independent reverse engineering of the undocumented matrix unit in M-series parts, which are neither x86 nor Neoverse server cores. The benchmark machine's Accelerate SGEMM runs on that unit, which the benchmark README can say without an entry.
- r13.3 [OpenMP Specifications](https://www.openmp.org/specifications/): Rule 4, one copy per source: listed in section 9 for the same fork-join, places, binding and wait-policy controls, and the inference runtimes built on OpenMP add no mechanism of their own to them.
- r13.4 [llamafile](https://github.com/mozilla-ai/llamafile): Left out under the cap once the attention pattern entered: the report on it is held below, and the tinyBLAS kernels it holds were upstreamed into the listed llama.cpp.
- r13.5 [GGUF](https://github.com/ggml-org/ggml/blob/master/docs/gguf.md): Left out under the cap once the bf16 definition entered: a file format rather than a mechanism, whose tensor-type table is reached from the listed k-quants and ggml entries.
- r13.6 [ZenDNN](https://github.com/amd/ZenDNN): Left out under the cap: AMD's primitive library for EPYC, with the LowOHA matmul and a llama.cpp backend documented at docs/backend/ZenDNN.md in the listed llama.cpp, competed with the Arm Compute Library for the one runtime slot and the Arm framework path was the larger gap. First candidate if a runtime slot opens.
- r13.7 [Optimize Weight Rounding via Signed Gradient Descent for the Quantization of LLMs](https://arxiv.org/abs/2309.05516): Left out under the cap: the weight-only int4 method (AutoRound, repository at intel/auto-round) behind the Intel MLPerf CPU submissions the last subsection links, held behind the kernel-side entries because the list is about mechanisms. First candidate if a quantization slot opens.
- r13.8 [STREAM Benchmark Reference Information](https://www.cs.virginia.edu/stream/ref.html): Rule 4, one copy per source: listed in section 15 for the run rules, and the measured roof a decode rate times bytes per token is compared with is the same figure there, with no mechanism specific to inference.
- r13.9 [Understanding Memory Formats](https://uxlfoundation.github.io/oneDNN/dev_guide_understanding_memory_formats.html): Left out under the seven-entry cap: a sub-page of the listed oneDNN that explains blocked layouts and ties block width to vector width, and the reorder itself is defined on the reorder primitive page rather than there.
- r13.10 [bitnet.cpp](https://github.com/microsoft/BitNet): Left out under the cap, and rule 3: the speedups in the repository and in its companion report (arXiv 2410.16144) state neither CPU model, core count, frequency, compiler nor run statistic. Its lookup-table kernels descend from T-MAC, which stays.
- r13.11 [Efficient LLM Inference on CPUs](https://arxiv.org/abs/2311.00502): Rule 3: the only measurement names a processor family without model, core count, frequency, compiler or run statistic, and the runtime it describes (intel/neural-speed) was archived in 2024, so the entry would rest on a number it cannot carry.
- r13.12 [Intel SDM Volume 1, versioned PDF](https://cdrdv2-public.intel.com/922477/253665-092-sdm-vol-1.pdf): Rule 4: a revision-numbered PDF whose address changes at every release, so the manual index in Intel's document library, which reaches every volume, is linked instead.
- r13.13 [Intel SDM Volume 2, VPDPBUSD](https://cdrdv2-public.intel.com/922478/325383-092-sdm-vol-2abcd.pdf): Rule 4 and the cap: the same revision-numbered form, and the VPDPBUSD page is reached from the listed manual index, so no second SDM entry is needed.
- r13.14 [Intel Optimization Reference Manual, versioned PDF](https://cdrdv2-public.intel.com/821612/248966-Optimization-Reference-Manual-V1-050.pdf): Rule 4: a revision-numbered PDF, replaced by the document-library page for the manual that sections 1 and 7 link.
- r13.15 [Tuning Guide for AI on the 4th Generation Intel Xeon Scalable Processors](https://www.intel.com/content/www/us/en/developer/articles/technical/tuning-guide-for-ai-on-the-4th-generation.html): Left out under the cap: its BIOS, SMT and NUMA advice is carried by the OpenVINO scheduling page, and its AMX guidance by the Optimization Reference Manual in sections 1, 2, 4 and 7. Returns 403 to curl.
- r13.16 [Intel Extension for PyTorch Performance Tuning Guide](https://intel.github.io/intel-extension-for-pytorch/cpu/latest/tutorials/performance_tuning/tuning_guide.html): Primary, but left out under the cap: its cores-versus-logical-processors, numactl and OpenMP affinity advice is carried by the listed OpenVINO scheduling page and the listed Thread management entry. Only its allocator advice (jemalloc, tcmalloc) is lost.
- r13.17 [CPU Performance](https://github.com/ggml-org/llama.cpp/discussions/3167): Rule 3: the numbers carry a CPU model and thread count but no frequency, compiler, flags or run statistic, and the bandwidth argument in the thread is asserted rather than measured.
- r13.18 [Performance of llama.cpp on Apple Silicon M-series](https://github.com/ggml-org/llama.cpp/discussions/4167): Rule 5: the tables are Metal GPU measurements.
- r13.19 [ONNX Runtime Execution Providers](https://onnxruntime.ai/docs/execution-providers/): Rule 4: no dedicated CPU provider page exists and the overview only names it, so the MLAS source tree, which is the provider's kernel library, is linked instead.
- r13.20 [Accelerate Matrix Multiplication Performance with SME2](https://learn.arm.com/learning-paths/cross-platform/multiplying-matrices-with-sme2/): Rule 1: a tutorial that restates the kernels the SME Programmer's Guide already shows, and the guide is the more primary source.
- r13.21 [KleidiAI on GitHub](https://github.com/Arm-software/kleidiai): Rule 4: the README states it is a read-only mirror of the Arm GitLab repository, which is the copy held below.
- r13.22 [oneDNN Matrix Multiplication Primitive on oneapi-src](https://oneapi-src.github.io/oneDNN/dev_guide_matmul.html): Rule 4: the project moved to the UXL Foundation and the old site is a mirror that still serves the page.
- r13.23 [Anatomy of High-Performance Matrix Multiplication at ACM](https://dl.acm.org/doi/10.1145/1356052.1356053): Rule 4: paywalled and 403 to curl, and the authors' own copy exists.
- r13.24 [BLIS: A Framework for Rapidly Instantiating BLAS Functionality at ACM](https://dl.acm.org/doi/10.1145/2764454): Rule 4: paywalled and 403 to curl, and the authors' own copy exists.
- r13.25 [Roofline: An Insightful Visual Performance Model, Berkeley technical report EECS-2008-134](https://www2.eecs.berkeley.edu/Pubs/TechRpts/2008/EECS-2008-134.html): Rule 4: the earlier long draft of the CACM article, which is free at the publisher and is the copy sections 1 and 6 carry.
- r13.26 [MLPerf Inference Benchmark](https://arxiv.org/abs/1911.02549): Rule 4, one copy per source: listed in section 15, whose reason already covers the load generator and the scenarios.
- r13.27 [T-MAC](https://github.com/microsoft/T-MAC): Rule 3: the README's Jetson AGX Orin table is the nearest measured CPU-against-GPU comparison, but it states no CPU frequency, compiler or run statistic, so it stays on the section 16 watchlist pending those fields. The paper stays here on the lookup-table mechanism, which is a definition, not a number.
- r13.28 [MLPerf Inference: Datacenter results page](https://mlcommons.org/benchmarks/inference-datacenter/): Rule 4: 403 to curl and the page is a viewer over the results repository, which holds the system descriptions and logs and is linked instead.
- r13.29 [Flash attention in llama.cpp](https://github.com/ggml-org/llama.cpp/pull/5021): Left out for the attention slot: the pull request describes its CPU path as slow and for testing, so the oneDNN pattern definition is the attention entry instead.
- r13.30 [BLIS](https://github.com/flame/blis): Trimmed for length: a repository whose reason said where the code lives, and the micro-kernel contract and blocksize rules it holds are the ones the listed BLIS paper defines.
- r13.31 [Arm Compute Library](https://github.com/ARM-software/ComputeLibrary): Trimmed for length: the listed oneDNN entry carries the framework path on Arm, which ends in Compute Library's NEON and SVE kernels.
- r13.32 [XNNPACK](https://github.com/google/XNNPACK): Trimmed for length: the per-ISA micro-kernel table dispatched at run time is carried by the listed ONNX Runtime MLAS entry, which reaches AMX and SME as well.
- r13.33 [bfloat16 Hardware Numerics Definition](https://www.intel.com/content/www/us/en/content-details/671279/bfloat16-hardware-numerics-definition.html): Trimmed for length: the bf16 operand format is defined with the AVX512_BF16 and AMX-BF16 instructions in the listed Intel Software Developer Manuals, and the listed attention entry carries f32 accumulation under bf16 inputs.
- r13.34 [FBGEMM: Enabling High-Performance Low-Precision Deep Learning Inference](https://arxiv.org/abs/2101.05615): Trimmed for length: the packed weight format and fused requantization are carried by the listed oneDNN Matrix Multiplication Primitive, run-time kernel generation by the listed oneDNN entry, and VNNI saturation by the listed Nuances of int8 Computations.
- r13.35 [Intel Optimization Reference Manual, Intel AMX and Int8 Deep Learning Inference chapters](https://www.intel.com/content/www/us/en/content-details/671488/intel-64-and-ia-32-architectures-optimization-reference-manual-volume-1.html): Trimmed for length: its homes are sections 1, 2, 4 and 7, and the listed Intel Software Developer Manuals entry defines AMX, so tile throughput and the AMX frequency level are read from the manual in those sections.
- r13.36 [KleidiAI](https://gitlab.arm.com/kleidi/kleidiai): Trimmed for length: the SME2 int8 matmul, gemv and lookup-table kernel forms are in the listed SME Programmer's Guide, and run-time SME dispatch on Arm is carried by the listed ONNX Runtime MLAS entry.
- r13.37 [PyTorch Performance Tuning Guide, CPU specific optimizations](https://docs.pytorch.org/tutorials/recipes/recipes/tuning_guide.html): Trimmed for length: one thread per physical core with affinity is carried by the listed Thread management entry, and one instance per socket by the listed Performance Hints and Thread Scheduling entry.
- r13.38 [oneTBB](https://github.com/uxlfoundation/oneTBB): Trimmed for length: its home is section 9, which carries the arena model and the spin policy, and the listed OpenVINO CPU Device entry carries the streams built on it.
- r13.39 [Roofline: An Insightful Visual Performance Model for Multicore Architectures](https://cacm.acm.org/research/roofline-an-insightful-visual-performance-model-for-multicore-architectures/): Trimmed for length: its homes are sections 1 and 6, and the preamble carries the batch-one decode case, every weight read once for a few flops under the bandwidth roof.
- r13.40 [Cache-aware Roofline model: Upgrading the loft](https://ieeexplore.ieee.org/document/6506838): Trimmed for length: its home is section 6, whose reason carries the roof per cache level that covers weights fitting in L2 or L3.
- r13.41 [MLPerf Inference Rules](https://github.com/mlcommons/inference_policies/blob/master/inference_rules.adoc): Trimmed for length: its home is section 15, and the clause that inputs start in host memory, so a GPU's timed region holds a transfer a CPU never pays, is read from the same document there.

### Numbers examined

- c13.1 "810 gigaflops" against "MKL processes this matrix size at 295 gigaflops" from http://justine.lol/matmul/: fields present: 1 CPU Intel Core i9-14900K (the report calls it Alderlake), 2 cores not stated (the report says llamafile avoids efficiency cores but gives no thread count), 3 freq/turbo/SMT overclocked with no clock stated, performance governor set, SMT state not stated, 4 compiler/flags the example kernels are compiled with -O3 -ffast-math -march=native under GCC, the llamafile build flags are not stated, 5 workload SGEMM of 513 by 512 times 512 by 512 f32 matrices, 6 baseline MKL on the same machine, 7 method not stated (no run count, no statistic). Missing: 2, 3, 4 (for the shipped build), 7. Verdict: watchlist, the entry is now under Rejected, trimmed for length, and the numbers do not enter. Verdict: watchlist.
- c13.2 "63 prompt tok/sec" for llamafile-0.7 against "40" for llama.cpp 2024-03-26, Mistral 7B q8_0 from http://justine.lol/matmul/: fields present: 1 CPU Intel Core i9-14900K, 2 cores not stated, 3 freq/turbo/SMT overclocked, performance governor, no clock or SMT state, 4 compiler/flags not stated for the shipped binaries, 5 workload Mistral 7B q8_0 prompt evaluation, prompt length not stated, 6 baseline llama.cpp at a dated commit and llamafile-0.6.2, 7 method not stated. Missing: 2, 3, 4, 7, and the prompt length in 5. Verdict: watchlist. Verdict: watchlist.
- c13.3 "ms/tok@4th, Ryzen7950X: F16 214, Q4_K_S 68, Q4_K_M 71" from https://github.com/ggml-org/llama.cpp/pull/1684: fields present: 1 CPU AMD Ryzen 9 7950X (Zen 4), 2 cores 4 threads, 3 freq/turbo/SMT not stated, 4 compiler/flags not stated (AVX2 path implied), 5 workload LLaMA 7B single-token prediction after a short prompt, perplexity on wikitext, 6 baseline F16 and the earlier Q4_0 table on the project page, 7 method not stated (no run count, no statistic). Missing: 3, 4, 7. Verdict: watchlist, the entry stays on the format definitions and the perplexity curve. Verdict: watchlist.
- c13.4 "tg128 1.25 ± 0.00 t/s" on one NUMA node against "2.38 ± 0.00 t/s" with --numa distribute (190.4 percent) from https://github.com/ggml-org/llama.cpp/discussions/11733: fields present: 1 CPU AMD EPYC 9374F (Zen 4, Genoa) in NPS2 to emulate two sockets, plus 2x EPYC 9175F (Zen 5, Turin) and 2x EPYC 9654 (Genoa) on other platforms, 2 cores 16 threads per node and 32 across both, 3 freq/turbo/SMT not stated, 4 compiler/flags cc (Ubuntu 13.2.0-23ubuntu4) 13.2.0, build flags not stated, 5 workload Llama-3.1-70B-Instruct F16, pp512 and tg128, 6 baseline the single-node run, 7 method llama-bench at build c026ba3c (4663), -r 3, mean and standard deviation. Missing: 3, and the flags in 4. Verdict: watchlist, the entry stays on the NUMA placement finding and the numatop remote-access evidence. Verdict: watchlist.
- c13.5 "eval time ... 21.32 tokens per second" before against "42.30 tokens per second" after, and "benchmark-matmult 32 cores 263.7 to 3694.85 gFlops" from https://github.com/ggml-org/llama.cpp/pull/7707: fields present: 1 CPU Intel Xeon CPU Max 9480 (Sapphire Rapids with HBM), 2 cores 1, 4, 16 and 32 for the matmult table, not stated for the model run, 3 freq/turbo/SMT not stated, 4 compiler/flags not stated, 5 workload llama2-7b q4_0 with a six-token prompt, and the ggml benchmark-matmult, 6 baseline the AVX-512 path before the change, 7 method a single llama_print_timings run. Missing: 2 (model run), 3, 4, 7. Verdict: watchlist, the entry stays on the kernel design and the reason no longer claims when tile GEMM wins over VNNI. Verdict: watchlist.
- c13.6 "30 tokens/s on M2-Ultra with a single core and 71 tokens/s with eight cores" and "up to 4x increase in throughput and 70% reduction in energy consumption compared to llama.cpp" from https://arxiv.org/abs/2407.00088: fields present: 1 CPU Apple M2 Ultra, also Raspberry Pi 5 and Snapdragon X Elite, 2 cores 1 and 8, 3 freq/turbo/SMT not stated, 4 compiler/flags not stated in the abstract (TVM-generated kernels in the paper), 5 workload BitNet-3B at 1.58 bits, token generation, 6 baseline llama.cpp, 7 method not stated in the abstract. Missing: 3, 4, 7. Verdict: watchlist, the entry stays on the lookup-table mechanism. Verdict: watchlist.
- c13.7 "Tokens per second: 1229.56" (Offline, Llama 3.1 8B) from https://github.com/mlcommons/inference_results_v6.0 (closed/Intel/results/1-node-2S-GNR_128C/llama3_1-8b/Offline/performance/run_1/mlperf_log_summary.txt): fields present: 1 CPU Intel Xeon 6980P (Granite Rapids), two sockets, 2 cores 128 per socket, physical cores only (the launch script enumerates unique core and socket pairs), 3 freq/turbo not stated, the systems JSON records host_processor_frequency as N/A, SMT siblings unused, 4 compiler/flags pinned inside the docker image intel/mlperf:mlperf-inference-6.0-llama3.1-8b with PyTorch and SGLang, flags not written down, 5 workload Llama 3.1 8B Instruct with AutoRound int4 weights (w4g128), CNN/DailyMail, Offline scenario, 6 baseline none, MLPerf reports absolute throughput against an accuracy floor, 7 method LoadGen, 600 s minimum duration, early stopping satisfied, VALID. Missing: 3 (frequency), 4 (flags as text), 6. Verdict: watchlist, the entry stays because the rules and logs make the run repeatable. The direction the reason states is read from the same repository: the same model's Offline log under closed/NVIDIA/results/B300-SXM-270GBx8_TRT reads "Tokens per second: 165432", and the Intel systems JSON records accelerators_per_node 0; Dell, Lenovo and Supermicro hold the other CPU-only Granite Rapids entries. Verdict: watchlist.
- c13.8 "125 ms/token" for GPT-J 6B f16 from https://github.com/ggml-org/ggml/blob/master/examples/gpt-j/README.md: fields present: 1 CPU Apple M1 Pro, 2 cores 8 threads, 3 freq/turbo/SMT not stated, 4 compiler/flags not stated, 5 workload GPT-J 6B f16, 200 generated tokens after a short prompt, 6 baseline the same run with half the network on the GPU through Metal Performance Shaders, 7 method a single timed run. Missing: 3, 4, 7. Verdict: watchlist, the entry stays on the observation that a GPU sharing the memory bus gained nothing, not on the figure. Verdict: watchlist.
- c13.9 "just 5 threads are enough to fully utilize the memory bandwidth" and "~3 t/s" expected against "2 t/s" for 7B q4_0 from https://johannesgaessler.github.io/llamacpp_performance: fields present: 1 CPU AMD Ryzen 7 3700X (Zen 2) with dual-channel DDR4 at 3200 MHz, and Intel Xeon E5-2683 v4 (Broadwell) with quad-channel DDR4 at 2133 MHz, 2 cores swept by thread count, 3 freq 2.1 GHz stated for the Xeon only, turbo and SMT state not stated, 4 compiler/flags not stated, 5 workload 128 generated tokens from an empty prompt at 2048 context, q4_0, 6 baseline the same machine at other memory frequencies and thread counts, 7 method not stated, the page is marked WIP. Missing: 3 (in part), 4, 7. Verdict: watchlist, the entry stays on the experiment design, memory clock varied with all else fixed. Verdict: watchlist.
- c13.10 "1.42x reduction in end-to-end latency" and Table 1 "Dense: memory bound 100%, DRAM bound 87.5%" from https://arxiv.org/abs/2502.12444: fields present: 1 CPU Intel Xeon Gold 6430L (Sapphire Rapids), 2 cores swept in the scaling figure, not stated for the headline, 3 freq/turbo/SMT not stated, 4 compiler/flags not stated, PyTorch with custom kernels, 5 workload Llama 3 8B decode at 512 context, and 32 repeated linear layers of 4192 by 14336 for the VTune table, 6 baseline stock PyTorch, 7 method VTune pipeline-slot accounting for the table, run statistic not stated for the speedup. Missing: 2 (headline), 3, 4, 7 (speedup). Verdict: watchlist, the entry stays on the counter evidence and the released kernels, the speedup does not enter. Verdict: watchlist.

## 14. Hardware generations

### Rejected candidates

- r14.1 [Intel Optimization Reference Manual](https://www.intel.com/content/www/us/en/content-details/671488/intel-64-and-ia-32-architectures-optimization-reference-manual-volume-1.html): Trimmed for length: its homes are sections 1, 2, 4 and 7, where the Golden Cove, Redwood Cove and Crestmont chapters are read from, and the Sapphire Rapids overview and the generations paper carry the server generations here.
- r14.2 [Intel Xeon 6 with P-cores Configuration and Tuning Guide for HPC Applications](https://www.intel.com/content/www/us/en/content-details/858491/intel-xeon-6-with-p-cores-configuration-and-tuning-guide-for-hpc-applications.html): Trimmed for length: its home is section 10, which holds the sub-NUMA clustering and uncore mode chapters, and the Xeon 6 memory article here measures the per-die L3 and the die crossing those modes create.
- r14.3 [High Performance Computing Tuning Guide for AMD EPYC 9004 Series Processors](https://docs.amd.com/v/u/en-US/58002_amd-epyc-9004-tg-hpc): Trimmed for length: its home is section 10, which holds the NUMA modes and boost settings, and the Zen 4 paper here states how one core yields Genoa and Bergamo as distinct packages.
- r14.4 [Arm Neoverse V2 Core Technical Reference Manual](https://support.arm.com/documentation/102375/latest/): Trimmed for length: its home is section 4, and the V2 optimisation guide here carries the pipeline widths and timings of the core Graviton 4, Grace and Axion share.
- r14.5 [Arm Neoverse V1 Core Technical Reference Manual](https://support.arm.com/documentation/101427/latest/): Trimmed for length: the Graviton repository's generation table names Neoverse V1 as the Graviton 3 core and links the manual, and the SVE paper measures the part with its configuration stated.
- r14.6 [uops.info](https://uops.info/): Trimmed for length: its home is section 3, and the comparison paper here measures Golden Cove and Zen 4 in-core at fixed clock.
- r14.7 [Hot Chips 2023 BHS and Granite Rapids Xeon Architecture Presentation](https://www.intel.com/content/www/us/en/content-details/787386/hot-chips-2023-bhs-and-granite-rapids-xeon-architecture-presentation.html): Rule 4: the landing page answers 200 but the file it offers answers 404 from Intel's download endpoint and AccessDenied from the public bucket, so the deck cannot be read; the Xeon 6 HPC guide, in section 10, carries the package description instead.
- r14.8 [Hot Chips 2023 Sierra Forest Xeon Architecture Presentation](https://www.intel.com/content/www/us/en/content-details/787431/hot-chips-2023-sierra-forest-xeon-architecture-presentation.html): Rule 4: the same dead download behind a live landing page; the Crestmont chapter of the optimisation manual, listed in sections 1, 2, 4 and 7, is the vendor document left for the E-core part, and the generations paper measures two Sierra Forest sockets with their configuration stated.
- r14.9 [Hot Chips 2023 Xeon Press Briefing](https://download.intel.com/newsroom/2023/data-center-hpc/Hot_Chips_23_Granite_Rapids_Sierra_Forest_Xeon_Press_Briefing.pdf): Rule 3: a press deck whose multiples are labelled architectural projections against unnamed prior platforms, with none of the seven fields; its one page of E-core pipeline widths is restated in the optimisation manual.
- r14.10 [Sapphire Rapids: The Next-Generation Intel Xeon Scalable Processor](https://ieeexplore.ieee.org/document/9731107): No rule failed; the ISSCC disclosure of the four-die package, cut for size because the free developer overview states the same platform changes and the listed measurement shows their cost.
- r14.11 [4th Generation Intel Xeon Scalable Processor Family Instruction Throughput and Latency](https://www.intel.com/content/www/us/en/content-details/765484/4th-generation-intel-xeon-scalable-processor-family-based-on-sapphire-rapids-architecture-instruction-throughput-and-latency.html): No rule failed; the vendor's machine-readable per-instruction table for the Sapphire Rapids core, cut for size because uops.info, in section 3, measures the same core family and the optimisation manual gives the port map.
- r14.12 [5th Gen Intel Xeon Scalable Processors](https://www.intel.com/content/www/us/en/products/details/processors/xeon/5th-gen-xeon-scalable-processors.html): Rule 1: a product page; the ISSCC paper is the designers' statement of the generation and is listed instead.
- r14.13 [Intel Xeon 6 Processors Performance and Power Profiles](https://builders.intel.com/solutionslibrary/intel-xeon-6-processors-performance-and-power-profiles): Rule 3 and rule 1: a solutions-library paper whose own text says its results may be estimated or simulated, and the Xeon 6 HPC guide states the same latency-optimised uncore mode with the BIOS menu, the script and the sysfs check that set and verify it, so the guide, in section 10, is listed instead.
- r14.14 [Intel Xeon 6 with E-Core Processors](https://www.intel.com/content/www/us/en/products/details/processors/xeon/6-e-core-series.html): Rule 3: the page carries an integer-throughput multiple against a two-generation-old part with the configuration off page; section 16 holds it on the watchlist and it is not repeated here.
- r14.15 [Tackling Throughput Computing with Sierra Forest](https://www.intel.com/content/www/us/en/newsroom/news/tackling-throughput-computing-sierra-forest.html): Rule 1: a newsroom editorial that defines no mechanism.
- r14.16 [Meteor Lake's E-Cores: Crestmont Makes Incremental Progress](https://chipsandcheese.com/p/meteor-lakes-e-cores-crestmont-makes-incremental-progress): No rule failed; the only measured Crestmont core, on a client part outside the section's server boundary, held until an independent microbenchmark of a Xeon 6 E-core socket with its configuration stated appears; the generations paper measures the socket on application codes, not the core.
- r14.17 [The Microarchitecture of Intel, AMD, and VIA CPUs](https://www.agner.org/optimize/microarchitecture.pdf): No rule failed; listed in sections 2 and 3, and its chapters cover Alder Lake, Zen 4 and Zen 5 as client cores, with no Sapphire Rapids, Redwood Cove or Crestmont chapter, so a third listing would add nothing about a server generation.
- r14.18 [Software Optimization Guide for the AMD Zen4 Microarchitecture](https://docs.amd.com/v/u/en-US/57647_zen4_sog_1.01): Rule 4: the viewer renders 401 in a browser and the portal search no longer lists publication 57647, as section 2 found; the IEEE Micro paper is the vendor's remaining description of the core.
- r14.19 [AMD EPYC 9004 Series Architecture Overview](https://docs.amd.com/v/u/en-US/58015-epyc-9004-tg-architecture-overview): Rule 4: the amd.com PDF path redirects here and the viewer renders 404; the portal index no longer carries 58015, so the HPC tuning guide, whose first chapter is the same layout for Genoa and Bergamo, is listed in section 10.
- r14.20 4th Gen AMD EPYC Processor Architecture whitepaper, publication 70351: Rule 4 and rule 3: the portal search indexes it and its API serves the PDF, but every viewer address for it renders 404; its generation-on-generation IPC and per-core multiples sit in footnotes with no fields.
- r14.21 [5th Gen AMD EPYC Processor Architecture](https://docs.amd.com/v/u/en-US/5th-gen-amd-epyc-processor-architecture-white-paper): Rule 3: leads with world-record counts and double-digit IPC claims that name no configuration; the technical overview of the same package is listed instead.
- r14.22 [AMD EPYC 9004 Series Processors Memory and CXL Advances](https://docs.amd.com/v/u/en-US/231963000-A_en_AMD-EPYC-9004-Series-Processors-Memory-and-CXL-Advances-White-Paper_pdf): Rule 3: a bandwidth multiple against the prior generation with the configuration in a footnote, on a paper that otherwise restates channel count and speed.
- r14.23 [Zen 4: The AMD 5nm 5.7GHz x86-64 Microprocessor Core](https://ieeexplore.ieee.org/document/10067540): No rule failed; the ISSCC core paper, cut for size because the IEEE Micro paper by the same team covers the core and the server packages together.
- r14.24 [Zen 4c: The AMD 5nm Area-Optimized x86-64 Microprocessor Core](https://ieeexplore.ieee.org/document/10454507): No rule failed; the origin of the statement that Zen 4c keeps the Zen 4 design with half the L3 per core, cut for size because the tuning guide, in section 10, states the layout and the Bergamo measurement tests it.
- r14.25 [Zen 5: The AMD High-Performance 4nm x86-64 Microprocessor Core](https://ieeexplore.ieee.org/document/10904529): No rule failed; cut for size because the Zen 5 optimisation guide is the fuller vendor account and is free.
- r14.26 [Zen 5 Variants and More, Clock for Clock](https://chipsandcheese.com/p/zen-5-variants-and-more-clock-for-clock): No rule failed; the only measured Zen 5c core, on a client part outside the section's server boundary, held until an independent measurement of a dense Turin socket with its configuration stated appears, so the Zen 5c dies the 9005 overview names have no measured source in the section.
- r14.27 [Genoa-X: Server V-Cache Round 2](https://chipsandcheese.com/p/genoa-x-server-v-cache-round-2): No rule failed; measures the stacked-cache variant, cut for size because the section covers the base parts and the Bergamo and Turin articles carry the same tests.
- r14.28 [AMD's EPYC 9355P: Inside a 32 Core Zen 5 Server Chip](https://chipsandcheese.com/p/amds-epyc-9355p-inside-a-32-core): No rule failed; a second Turin part on the same tests, cut for size in favour of the launch article.
- r14.29 [Arm Neoverse V3 Core Software Optimization Guide](https://support.arm.com/documentation/109678/latest/): No rule failed; the timing guide for the core in Graviton 5 and Cobalt 200, parts with no public measurement yet, held on the watchlist in section 16 until one exists, so it is not listed in the core and on the watchlist at once.
- r14.30 [Arm Neoverse N3 Core Software Optimization Guide](https://support.arm.com/documentation/109637/latest/): No rule failed; the timing guide for the core in the Axion N4A series, shipped with no public measurement, held on the watchlist in section 16 pending a seven-field run, for the same reason.
- r14.31 [Arm Neoverse V1 Software Optimization Guide](https://support.arm.com/documentation/109897/latest/): Rule 4: the address answers 200 for the site's script shell, but Arm's documentation service reports that document 109897 does not exist, a browser is sent to the site's not-found page, and the older PJDOC address the Graviton repository links for the guide answers the same, so the Graviton repository, whose generation table links the V1 technical reference manual, is the document left for the Graviton 3 core.
- r14.32 [Azure Cobalt Processor-based Virtual Machines](https://learn.microsoft.com/en-us/azure/virtual-machines/sizes/cobalt-overview): No rule failed; Microsoft's statement that Cobalt 100 is an N2 part at a stated clock with a whole physical core behind every vCPU, cut at the cap because the N2 guide names the part and the Yitian measurement covers the same core.
- r14.33 [General-purpose Machine Family for Compute Engine](https://docs.cloud.google.com/compute/docs/general-purpose-machines): No rule failed; Google's statement that C4A carries Neoverse V2 and N4A carries Neoverse N3, cut at the cap because the V2 guide names the series, the SVE paper measures C4A with its configuration stated, and N4A sits on the watchlist in section 16 with the N3 guide.
- r14.34 [AArch64SchedAmpere1B.td](https://github.com/llvm/llvm-project/blob/main/llvm/lib/Target/AArch64/AArch64SchedAmpere1B.td): No rule failed; the public timing model LLVM schedules AmpereOne by, written at a compiler contractor rather than by Ampere, cut at the cap because the Hot Chips article puts the disclosed structure beside measurements and the Altra datasheet is the Ampere document.
- r14.35 [AmpereOne Product Brief](https://amperecomputing.com/briefs/ampereone-family-product-brief): Rule 3: a cost multiple with no fields on a page whose technical content is a spec list; the Hot Chips article carries the public description of the core beside measurements.
- r14.36 [Ampere Altra Family Product Brief](https://amperecomputing.com/briefs/ampere-altra-family-product-brief): Rule 3: power and space multiples against an unnamed baseline; the datasheet is the technical document and is listed.
- r14.37 [AmpereOne AC03 Device Documentation](https://amperecomputing.com/customer-connect/products/AmpereOne-device-documentation): Rule 4: every file on the page downloads through a login, and no datasheet is listed even then.
- r14.38 [NVIDIA Grace CPU Benchmarking Guide](https://nvidia.github.io/grace-cpu-benchmarking-guide/): No rule failed; the vendor's build and run recipes for standard benchmarks on Grace, cut for size because section 15 owns the suites and the tuning guide is the architecture document and is listed.
- r14.39 [NVIDIA Grace CPU Superchip Whitepaper](https://resources.nvidia.com/en-us-grace-cpu/nvidia-grace-cpu-superchip): Rule 3 and rule 1: a marketing flipbook whose comparisons name no configuration; the tuning guide states the same architecture without them.
- r14.40 [Grace Hopper, Nvidia's Halfway APU](https://chipsandcheese.com/p/grace-hopper-nvidias-halfway-apu): Rule 5: the Grace measurements share the article with the GPU half of the module, and the list carries no GPU content; the arXiv comparison measures Grace as a CPU.
- r14.41 [Graviton 3: First Impressions](https://chipsandcheese.com/p/graviton-3-first-impressions): No rule failed; a preview by its own statement, without the instance size or clock used, cut for size in favour of the Graviton 4 article and the SVE paper, which both measure Graviton 3.
- r14.42 [Core to Core Latency Data on Large Systems](https://chipsandcheese.com/p/core-to-core-latency-data-on-large-systems): No rule failed; one test across many parts, cut for size because each listed article runs the same test on its part.
- r14.43 [Introducing Google's New Arm-based CPU](https://cloud.google.com/blog/products/compute/introducing-googles-new-arm-based-cpu): Rule 3 and rule 1: an announcement whose price-performance multiples name no configuration; the machine family documentation states the cores without them and is held for size.
- r14.44 [Azure Cobalt 100-based Virtual Machines Are Now Generally Available](https://azure.microsoft.com/en-us/blog/azure-cobalt-100-based-virtual-machines-are-now-generally-available/): Rule 1: a launch post; the documentation page, held for size, carries the same core, clock and vCPU facts.
- r14.45 [New Azure Cobalt 200 VMs Deliver 50% Performance Improvement](https://azure.microsoft.com/en-us/blog/new-azure-cobalt-200-vms-deliver-50-performance-improvement-fully-optimized-for-modern-agentic-ai-workloads/): Rule 3: per-workload multiples against Cobalt 100 with no fields; the V3 guide sits on the watchlist in section 16 with the core.
- r14.46 [Announcing Cobalt 200: Azure's next cloud-native CPU](https://techcommunity.microsoft.com/blog/azureinfrastructureblog/announcing-cobalt-200-azure%E2%80%99s-next-cloud-native-cpu/4469807): Rule 4 and rule 3: answers 403 to curl and 400 with a browser agent, and the fetch tool sees only the title; the design facts in it are repeated in the Azure post above with the same unfielded multiple.
- r14.47 [Evaluating HPC-Style CPU Performance and Cost in Virtualized Cloud Infrastructures](https://arxiv.org/abs/2511.08948): No rule failed; a cost study across cloud instances with Graviton 3 among them, cut for size because the SVE paper covers the Arm parts with the fuller configuration statement.

### Numbers examined

- c14.1 "GNR in absolute terms is still on average 2.71× faster than SPR DDR and 1.93× faster than EMR", "SRF is on average 5% and 10% faster compared to the increase in bandwidth for the 96 core and 192 core variants" and "87% less energy consumed" from https://real.mtak.hu/215594/: fields present: 1 CPU (Xeon MAX 9480 and Platinum 8480+, Sapphire Rapids with HBM and with DDR5; Platinum 8592+, Emerald Rapids; Platinum 6960P, Granite Rapids; Xeon 6740E and an unnamed 192-core part, Sierra Forest; EPYC 9B14, Zen 4), 2 cores (two sockets each, cores per socket stated per system), 3 frequency (base and all-core turbo per system, SMT on for the P-core Xeons, off for Sierra Forest and the Genoa instance), 4 compiler (oneAPI 2024.1 on the Intel systems, GCC 11.3 on Genoa, flags in the makefiles of the linked test repository), 5 workload (BabelStream Triad and seven structured- and unstructured-mesh codes with problem sizes stated), 6 baseline (Sapphire Rapids with DDR5), 7 method (four runs averaged, wall time and RAPL energy per run). Missing: the flags in the paper text itself, held in the repository it links. Verdict: core, the figures are not quoted. Verdict: core.
- c14.2 "SPR and Genoa eventually fall down to a frequency of 2.0 GHz and 3.1 GHz for AVX-512-heavy code, which results in 53% and 84% of their respective single-core turbo limit" and "GCS exhibits a constant frequency of 3.4 GHz" from https://arxiv.org/abs/2409.08108: fields present: 1 CPU (Intel Xeon Platinum 8470, Golden Cove; AMD EPYC 9684X, Zen 4; NVIDIA Grace CPU Superchip, Neoverse V2), 2 cores (scaled across one socket, 52, 96 and 72), 3 frequency (turbo allowed for this test, clock of every active core tracked with LIKWID counters), 4 compiler (GCC 12.1, oneAPI 2023.2 and Clang 17.0.6 on x86; Arm Compiler for Linux 23.10 and GCC 13.2 on Grace), 5 workload (arithmetic-heavy microbenchmarks per ISA extension, each run for several minutes), 6 baseline (each part's single-core turbo limit), 7 method (counter-tracked sustained clock). Missing: SMT state, the flags for this test, run count and statistic. Verdict: core for the method, the figures are not quoted. Verdict: core.
- c14.3 "Genoa only achieves 78% of its theoretical memory bandwidth peak, while GCS and SPR reach 87% and 90%" from https://arxiv.org/abs/2409.08108: the same systems, with the paper itself saying the comparison depends on the memory fitted. Missing: 3 SMT state, 4 flags, 7 run count and statistic. Verdict: watchlist, not quoted. Verdict: watchlist.
- c14.4 "Graviton4 and Axion are 1.6X and 1.4X slower" than the x86 instances, and "Axion exhibits the fastest execution time, being 9.4% and 36% faster than Graviton4 and Yitian" from https://arxiv.org/abs/2506.09505: fields present: 1 CPU (Graviton3 c7g.metal, Neoverse V1; Graviton4 c8g.metal-24xl, Neoverse V2; Yitian 710, Neoverse N2; Axion c4a-standard-72, Neoverse V2; EPYC 9R14, Zen 4; Xeon 8488C, Sapphire Rapids), 2 cores (whole instances, sizes named), 3 frequency (per part, with turbo noted for the Xeon and Axion's clock estimated because the vendor states none), 4 compiler (GCC 14.2), 5 workload (Merkle tree of Poseidon hashes over the Goldilocks field, leaf count stated), 6 baseline (AVX-512 build on the x86 instances), 7 method (at least five runs, average reported, spread under a stated bound). Missing: 3 SMT state, 4 flags. Verdict: core for the configuration statement, the multiples are not quoted. Verdict: core.
- c14.5 "This generation delivers an 18% performance improvement for general integer compute workloads and a 24% improvement for floating-point workloads at iso power vs. the 4th-Gen Xeon processors" from https://ieeexplore.ieee.org/document/10454434: fields present: 1 CPU (the 5th and 4th generation families, no models), 6 baseline (4th generation). Missing: 2, 3, 4, 5, 7 (held in a reference the abstract cites). Verdict: cut; the entry rests on the die count and cache statement. Verdict: cut.
- c14.6 "13% IPC improvement over the previous generation on an average of single-threaded desktop applications" and "more than 29% increase generationally in single-threaded performance" from https://ieeexplore.ieee.org/document/10067540: fields present: 6 baseline (Zen 3). Missing: 1 model, 2, 3, 4, 5 (unnamed desktop applications), 7. Verdict: cut. Verdict: cut.
- c14.7 "35% smaller", "more than 25% improvement in performance/mm2, and 9% improvement in performance/W on SPECrate2017_int_base" and "up to 3.1GHz in frequency in a server configuration" from https://ieeexplore.ieee.org/document/10454507: fields present: 1 CPU (Zen 4c against Zen 4), 5 workload (SPECrate2017_int_base for the efficiency figure), 6 baseline (Zen 4). Missing: 2, 3, 4, 7. Verdict: cut; the area and clock statements are design facts, the efficiency figure is not quoted. Verdict: cut.
- c14.8 "on a few Intel AMX-based Deep Learning (DL) workloads, up to ~7% performance degradation was observed with the feature enabled" for the L2P AMP prefetcher from https://www.intel.com/content/www/us/en/content-details/671488/intel-64-and-ia-32-architectures-optimization-reference-manual-volume-1.html: fields present: 1 CPU (5th Generation Xeon Scalable, no model), 5 workload (AMX deep learning, unnamed), 6 baseline (the prefetcher disabled). Missing: 2, 3, 4, 7. Verdict: cut; the manual is under Rejected here with its homes named. Verdict: cut.
- c14.9 Workload tables comparing the default and latency-optimised profiles (Tables 6-1 and 7-1) from https://builders.intel.com/solutionslibrary/intel-xeon-6-processors-performance-and-power-profiles: fields present: 1 CPU (Xeon 6900-series with P-cores and 6700-series with E- and P-cores, models in the backup), 5 workload (named suites per row), 6 baseline (the latency-optimised profile). Missing on the page: 2, 3, 4, 7 (the paper says configurations sit in backup and results may be estimated or simulated). Verdict: cut; the paper is under Rejected and the HPC guide, in section 10, carries the mode and the knob. Verdict: cut.
- c14.10 "L3 latency also regressed by about 33% compared to Ice Lake SP" and "around 33 ns" from https://chipsandcheese.com/p/a-peek-at-sapphire-rapids: fields present: 1 CPU (Xeon Platinum 8480, Sapphire Rapids), 2 cores (a 56-core part, threads per test not stated), 3 frequency (idle, ramp and peak clocks observed on Intel Developer Cloud, a lower fixed clock on GCP, SMT state not stated), 5 workload (pointer-chasing latency and bandwidth microbenchmarks), 6 baseline (Ice Lake SP), 7 method (own microbenchmarks). Missing: 4 compiler and flags, 7 run count and statistic. Verdict: watchlist; the entry stays for the shape of the latency ladder, the figures are not quoted. Verdict: watchlist.
- c14.11 "L3 cache latency of just over 33 ns" in SNC3, "Accessing the L3 on an adjacent die increases latency by about 24 ns", local DRAM "131.5 ns" and "691.62 GB/s" from https://chipsandcheese.com/p/a-look-into-intel-xeon-6s-memory: fields present: 1 CPU (Xeon 6 6985P-C, Granite Rapids; Xeon Platinum 8559C, Emerald Rapids; EPYC 9355P, Turin), 2 cores (96, in a cloud instance), 3 frequency (peak clock named, SMT state not stated), 5 workload (latency, bandwidth and core-to-core microbenchmarks), 6 baseline (Emerald Rapids and Turin on the same tests), 7 method (own microbenchmarks). Missing: 4 compiler and flags, 7 run count and statistic. Verdict: watchlist; the entry stays for the die-crossing structure, the figures are not quoted. Verdict: watchlist.
- c14.12 "16.89 ns" against "13.35 ns" L3 latency and "just short of 360 GB/s" from https://chipsandcheese.com/p/testing-amds-bergamo-zen-4c-spam: fields present: 1 CPU (Bergamo, Zen 4c, model not named; a desktop Ryzen 9 7950X3D capped to Bergamo's clock as the Zen 4 comparison), 2 cores (128 per socket, two sockets, NPS1), 3 frequency (capped at the part's ceiling for the comparison, SMT state not stated), 5 workload (latency, bandwidth, core-to-core, libx264 and 7-Zip), 6 baseline (the desktop Zen 4 part), 7 method (own microbenchmarks). Missing: 1 model, 4 compiler and flags, 7 run count and statistic. Verdict: watchlist; the entry stays for the same-core finding, the figures are not quoted. Verdict: watchlist.
- c14.13 "nearly 99% of the theoretical 576 GB/s" read bandwidth from https://chipsandcheese.com/p/amds-turin-5th-gen-epyc-launched: fields present: 1 CPU (EPYC 9575F, Zen 5), 2 cores (64), 3 frequency (single-thread and all-core clocks named, SMT state not stated), 5 workload (bandwidth and latency microbenchmarks), 6 baseline (theoretical peak of the fitted DDR5), 7 method (own microbenchmarks). Missing: 4 compiler and flags, 7 run count and statistic. Verdict: watchlist, not quoted. Verdict: watchlist.
- c14.14 "114.08 ns" local DRAM latency and "138.6 ns on average" cross-socket from https://chipsandcheese.com/p/arms-neoverse-v2-in-awss-graviton-4: fields present: 1 CPU (Graviton 4, Neoverse V2), 2 cores (96 per socket), 3 frequency (single- and dual-socket clocks named, no SMT on the part), 5 workload (latency, core-to-core and structure-size microbenchmarks), 6 baseline (Graviton 3), 7 method (own microbenchmarks). Missing: 4 compiler and flags, 7 run count and statistic. Verdict: watchlist; the entry stays for the rename-width and mesh findings, the figures are not quoted. Verdict: watchlist.
- c14.15 "35.48 ns on Yitian 710" L3 latency and "141 ns" DRAM latency from https://chipsandcheese.com/p/arms-neoverse-n2-cortex-a710-for-servers: fields present: 1 CPU (Alibaba Yitian 710, Neoverse N2), 2 cores (an eight-core instance), 3 frequency (locked clock stated, no SMT on the core), 5 workload (structure-size and latency microbenchmarks), 6 baseline (Ampere Altra and Graviton 2 on Neoverse N1, and Zen 4), 7 method (own microbenchmarks). Missing: 4 compiler and flags, 7 run count and statistic. Verdict: watchlist, not quoted. Verdict: watchlist.
- c14.16 "3.68 nanoseconds" L2 latency and "about 166 ns" DRAM latency from https://chipsandcheese.com/p/ampereone-at-hot-chips-2024-maximizing-density: fields present: 1 CPU (AmpereOne on an Oracle instance, SKU not named), 2 cores (an instance under a sixteen-core quota), 3 frequency (the instance's clock ceiling stated, no SMT on the core), 5 workload (pointer-chasing and structure-size microbenchmarks, 7-Zip and libx264), 6 baseline (Zen 4, Zen 5, Crestmont, Neoverse V2 in Grace and Graviton 4 on the same tests; Altra is named in prose only), 7 method (own microbenchmarks). Missing: 1 SKU, 4 compiler and flags, 7 run count and statistic. Verdict: watchlist, not quoted. Verdict: watchlist.
- c14.17 "Est. SPECrate 2017_int_base (SKU: AC-108025002): 301 at Usage Power: 187 W" from https://amperecomputing.com/assets/Altra_Rev_A1_DS_v1_50_20240130_3375c3dec5_1c5d4604fa.pdf: an estimate by the vendor's own label. Fields present: 1 CPU (one Altra SKU), 5 workload. Missing: 2, 3, 4, 6, 7. Verdict: cut; the entry rests on the cache, mesh and memory statements, not on the estimate. Verdict: cut.
- c14.18 "up to 34% better price performance" for Lambda and "up to 8x" for PCRE2 from https://github.com/aws/aws-graviton-getting-started: fields present: none beyond the product line. Verdict: cut; the entry rests on the generation table and the runbook. Verdict: cut.
- c14.19 "up to 50% better CPU performance" and per-workload multiples from https://azure.microsoft.com/en-us/blog/new-azure-cobalt-200-vms-deliver-50-performance-improvement-fully-optimized-for-modern-agentic-ai-workloads/: fields present: 6 baseline (Cobalt 100). Missing: 1 to 5, 7. Verdict: cut. Verdict: cut.
- c14.20 "cuts TCO by 40%" from https://amperecomputing.com/briefs/ampereone-family-product-brief and "reduce power consumption by up to 3X over the legacy x86 ecosystem" from https://amperecomputing.com/briefs/ampere-altra-family-product-brief: fields present: none. Verdict: cut. Verdict: cut.
- c14.21 "up to two times the performance per watt of conventional x86 servers" from https://docs.nvidia.com/dccpu/grace-perf-tuning-guide/index.html: fields present: none beyond the product line. Verdict: cut; the entry rests on the fabric, memory and MPAM statements and the multiple is not repeated. Verdict: cut.
- c14.22 Bisection bandwidth, memory bandwidth and memory power figures from https://docs.nvidia.com/dccpu/grace-perf-tuning-guide/index.html: specification figures for the fabric and the LPDDR5X fit, not measurements, and no comparison is drawn from them. Verdict: not subject to the seven-field test; nothing is quoted. Verdict: other.
- c14.23 STREAM Triad rows per NPS setting from https://docs.amd.com/v/u/en-US/58002_amd-epyc-9004-tg-hpc: recorded in section 10, the guide's home, with the fields present and missing there; the guide is under Rejected here.

## 15. Benchmarks

### Rejected candidates

- r15.1 [SPEC CPU 2017 Run and Reporting Rules](https://www.spec.org/cpu2017/Docs/runrules.html): Trimmed for length: the listed 2026 rules carry the same base against peak and rate against speed definitions and the same run and disclosure requirements, so a result published under either suite reads the same way, with only the reference machine changed.
- r15.2 [DCPerf](https://github.com/facebookresearch/DCPerf): Trimmed for length: the listed DCPerf paper carries the misprojection finding and the fleet-matching method, and cites the repository, which keeps the multi-process and microservice structure of the services on x86 and Arm.
- r15.3 [LIKWID](https://github.com/RRZE-HPC/likwid): Trimmed for length: its home is section 6, where the same repository is listed for the measured roofs and counter groups, and likwid-bench with its pinned kernels and per-node placement is reached from there.
- r15.4 [nanoBench](https://github.com/andreas-abel/nanoBench): Trimmed for length: the listed nanoBench paper explains the kernel-mode, interrupts-off timing the harness implements and names the repository in its first footnote, so sections 3 and 5, which defer the harness to this section, land on the paper.
- r15.5 [Twelve Ways to Fool the Masses When Giving Performance Results on Parallel Computers](https://www.davidhbailey.com/dhbpapers/twelve-ways.pdf): Trimmed for length: the listed crimes checklist carries the inflation faults, from sub-setting and microbenchmarks standing for the whole to improper baselines, and the listed scientific benchmarking rules carry the speedup base case that scaled problem sizes hide.
- r15.6 [Fleetbench](https://github.com/google/fleetbench): Trimmed for length: the listed fleet profile establishes the shared routines the suite benchmarks over fleet-sampled inputs, and the DCPerf paper under Standard suites is the listed fleet-derived suite.
- r15.7 [SPEC CPU2017: Next-Generation Compute Benchmark](https://research.spec.org/icpe_proceedings/2018/companion/p41.pdf): No rule failed. Superseded: the committee's paper on the 2026 suite explains the same speed and rate design and adds the rolling round-robin rate, and the listed 2026 rules carry the base, peak, rate and speed definitions that results published under either suite need.
- r15.8 [SPEC CPU2017: Next-Generation Compute Benchmark in the ACM Digital Library](https://dl.acm.org/doi/10.1145/3185768.3185771): Rule 4: returns 403 to automated clients. The SPEC Research Group's own ICPE proceedings archive serves the same two pages, itself now cut as superseded.
- r15.9 [SPEC CPU 2017 Overview](https://www.spec.org/cpu2017/Docs/overview.html): No rule failed. A question-and-answer introduction that restates the run rules' definitions of rate, speed, base and peak, cut for size in favour of the rules themselves.
- r15.10 [SPEC CPU 2017 Results](https://www.spec.org/cpu2017/results/): Leaderboard: results move and the list is not a results index. The run rules say how to read any result found there.
- r15.11 [Wait of a Decade: Did SPEC CPU 2017 Broaden the Performance Horizon?](https://lca.ece.utexas.edu/pubs/HPCA_SPEC17_ShuangSong.pdf): No rule failed. The counter characterisation of the 2017 suite alone; the listed 2026 characterisation measures both suites on current x86 and Arm parts, so one entry carries the measurement.
- r15.12 [MLPerf Inference Benchmark](https://arxiv.org/abs/1911.02549): Rule 5: the rules document above carries the load-generator placement and the scenario definitions a CPU reader needs, so the paper adds nothing here.
- r15.13 [MLPerf Inference: Datacenter results](https://mlcommons.org/benchmarks/inference-datacenter/): Leaderboard: results move and the list is not a results index. The rules define what a row on it means.
- r15.14 [MLPerf Inference reference implementations](https://github.com/mlcommons/inference): No rule failed. The LoadGen and reference model source, cut for size because the rules document names it.
- r15.15 [Geekbench 6 Benchmark Internals](https://www.geekbench.com/doc/geekbench6-benchmark-internals.pdf): No rule failed. Scope: a closed-source client benchmark with no repository behind its score, and it says nothing about the server parts the list is for, so the slot went to the current SPEC suite.
- r15.16 [Geekbench 6 CPU Workloads](https://www.geekbench.com/doc/geekbench6-cpu-workloads.pdf): No rule failed. A subset of the Benchmark Internals document, itself cut for scope.
- r15.17 [CoreMark](https://github.com/eembc/coremark): No rule failed. Scope: an embedded core test whose working set fits any first-level cache, quoted for microcontroller cores rather than the server parts the list is for, so the slot went to a datacenter suite.
- r15.18 [CoreMark at EEMBC](https://www.eembc.org/coremark/): No rule failed. The organisation's product page for a test that is out of scope for a server list; the repository carries the run rules and the reporting line if one is needed.
- r15.19 [DCPerf in the ACM Digital Library](https://dl.acm.org/doi/10.1145/3695053.3731411): Rule 4: returns 403 to automated clients. The camera-ready copy on the authors' team page is listed instead.
- r15.20 [Sustainable Memory Bandwidth in Current High Performance Computers](https://www.cs.virginia.edu/~mccalpin/papers/bandwidth/bandwidth.html): No rule failed. The author's later revision of the same survey. The balance paper is the one the STREAM FAQ cites, so it is listed instead.
- r15.21 [STREAM home page](https://www.cs.virginia.edu/stream/): No rule failed. A results portal with the source directory. The reference page under it holds the run rules and the counting convention, so that page is listed.
- r15.22 [lmbench](https://sourceforge.net/projects/lmbench/): No rule failed. The only live home of the source, but a tarball download page rather than a browsable repository, so the paper carries the tests and the slot went to the nanoBench paper that sections 3 and 5 defer to this section for.
- r15.23 [lmbench project home](http://www.bitmover.com/lmbench/): Rule 4: the host does not answer. The SourceForge release page is live and is recorded above.
- r15.24 [perf-bench](https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/tree/tools/perf/Documentation/perf-bench.txt): No rule failed. The kernel tree's scheduler, futex, NUMA and memory microbenchmarks, cut at the cap because the lmbench entry covers system calls and IPC and LIKWID, listed in section 6, covers pinned bandwidth.
- r15.25 [Benchmarking Crimes: An Emerging Threat in Systems Security](https://arxiv.org/abs/1801.02381): No rule failed. A literature audit showing that the crimes list is needed, a fact about publishing rather than a rule a reader applies, cut for size in favour of the geometric-mean origin and the fleet profile.
- r15.26 [A Profiling-Based Benchmark Suite for Warehouse-Scale Computers](https://ieeexplore.ieee.org/document/10590038/): Rule 4: the IEEE record answers an interstitial rather than the page and no author copy was found. The Fleetbench repository, itself now trimmed for length, states the fleet-derived input distributions its reason rested on.
- r15.27 [Profiling a warehouse-scale computer in the ACM Digital Library](https://dl.acm.org/doi/10.1145/2749469.2750392): Rule 4: returns 403 to automated clients. The copy hosted by the authors' employer is listed, the address section 12 also records.
- r15.28 [Profiling a warehouse-scale computer on gwern.net](https://gwern.net/doc/cs/hardware/2015-kanev.pdf): Rule 4: a third-party mirror of a paper the authors' employer hosts.
- r15.29 [Producing Wrong Data Without Doing Anything Obviously Wrong!](https://sape.inf.usi.ch/publications/asplos09): No rule failed. Listed in section 5, where measurement bias belongs. The methodology entries above start from its conclusion rather than repeat it.
- r15.30 [Stabilizer: Statistically Sound Performance Evaluation](https://people.cs.umass.edu/~emery/pubs/stabilizer-asplos13.pdf): No rule failed. Listed in section 5 as the remedy for layout bias, so it is not repeated.
- r15.31 [Statistically Rigorous Java Performance Evaluation](https://dri.es/files/oopsla07-georges.pdf): No rule failed. Its repetition and confidence-interval method is specific to managed runtimes and is generalised by the Kent paper listed, cut for size.
- r15.32 [Clearing the Clouds: A Study of Emerging Scale-out Workloads on Modern Hardware](https://infoscience.epfl.ch/entities/publication/674caf70-abd5-40a6-b096-713741ee689c): No rule failed. Measured mismatch between scale-out services and the standard suites on real cores, cut for size because the fleet profile and the DCPerf paper make the point at larger scale on current cores.
- r15.33 [Scientific Benchmarking of Parallel Computing Systems in the ACM Digital Library](https://dl.acm.org/doi/10.1145/2807591.2807644): Rule 4: returns 403 to automated clients. The authors' own copy is listed instead.
- r15.34 [Rigorous Benchmarking in Reasonable Time in the ACM Digital Library](https://dl.acm.org/doi/10.1145/2464157.2464160): Rule 4: returns 403 to automated clients. The authors' institutional copy, which also corrects the printed version, is listed instead.

### Numbers examined

- c15.1 "nearly 30% of cycles" for the datacenter tax and "15-30% of all pipeline slots" for front-end stalls from https://static.googleusercontent.com/media/research.google.com/en//pubs/archive/44271.pdf: fields present: 1 CPU Intel Ivy Bridge (model not named); 2 cores about twenty thousand machines sampled per collection, whole machines; 3 SMT enabled, frequency and turbo state not stated; 5 workload live production jobs across the fleet; 6 baseline the fleet itself, no external baseline; 7 method random sampling of machines and time through Google-Wide Profiling, performance counters reduced with the top-down method. Missing: 4 compiler and flags, CPU model, frequency and turbo state. Verdict: watchlist; the entry stays for the shape of the profile, the percentages are not quoted. Verdict: watchlist.
- c15.2 "overestimate the performance of Meta's latest server SKU by 28%" for SPEC CPU 2017 and "within a 3.3% error margin" for DCPerf from https://aisystemcodesign.github.io/papers/DCPerf-ISCA25.pdf: fields present: 2 cores the logical core count of each of four server SKUs, from thirty-six to one hundred and seventy-six; 5 workload DCPerf, SPEC CPU 2006 and SPEC CPU 2017 against the production services each benchmark models; 6 baseline the oldest SKU, with projections compared against production performance measured across thousands of servers; 7 method geometric mean of per-benchmark scores normalised to the first SKU, production side weighted by each workload's share of fleet power. Missing: 1 CPU model and microarchitecture (the SKUs are anonymised x86 parts), 3 frequency, turbo and SMT state (SMT on is stated for one figure only), 4 compiler and flags. Verdict: watchlist for the percentages; the entry stays for the direction of the finding and the fleet-matching method, no figure is quoted. Verdict: watchlist.
- c15.3 Per-workload IPC, miss-rate and stall tables from https://arxiv.org/abs/2605.03713: fields present: 1 CPU nine named parts (Xeon Platinum 8160, 8380 and 8468, Xeon Max 9480, EPYC 7763, 9454 and 9555, Ampere Altra on Neoverse N1, Nvidia Grace on Neoverse V2); 2 cores per-socket core counts stated; 4 compiler gcc -O3, with gcc 13 to 15 in the compiler sensitivity study; 5 workload SPEC CPU 2026 and 2017 rate and speed, DCPerf and MLPerf; 6 baseline SPEC CPU 2017 on the same machines; 7 method Linux perf counters under default deployment settings. Missing: 3 turbo and SMT state (dynamic frequency scaling left on is stated, SMT is not). Verdict: core; no figure is quoted in the annotation. Verdict: core.
- c15.4 "4-5% of their rated peak speeds" for out-of-cache arithmetic kernels from https://www.cs.virginia.edu/stream/ref.html: fields present: 5 workload simple vector kernels on out-of-cache operands; 6 baseline the vendor's rated peak. Missing: 1 CPU model, 2 cores, 3 frequency and SMT state, 4 compiler and flags, 7 method. Verdict: cut; the entry rests on the run rules and the counting convention, not on the figure. Verdict: cut.
- c15.5 "77.38 Tflop/s" for High-Performance Linpack on 64 nodes of Piz Daint from https://htor.inf.ethz.ch/publications/img/hoefler-scientific-benchmarking.pdf: presented by the paper as an example of an uninterpretable number, not as a result. Fields present: 5 workload HPL with N stated; 6 baseline theoretical peak. Missing: 1, 2, 3, 4, 7 by the paper's own design. Verdict: cut; the entry rests on the twelve rules. Verdict: cut.

## 16. Watchlist

### Rejected candidates

- r16.1 [Intel APX Architecture Specification](https://www.intel.com/content/www/us/en/content-details/784266/intel-advanced-performance-extensions-intel-apx-architecture-specification.html): Trimmed for length: the ISE programming reference line carries APX, its doubled registers and the tie to Diamond Rapids that the specification itself does not name.
- r16.2 [Arm Neoverse N3 Core Software Optimization Guide](https://support.arm.com/documentation/109637/latest/): Trimmed for length: the Neoverse V3 guide line carries the watch on a Neoverse core shipped in a cloud part with vendor timing tables and no public run, and the N3 behind Axion N4A needs the same run.
- r16.3 Arm Neoverse N4: Trimmed for length: an announced core with no guide or shipped part to link, carried by the Neoverse V3 guide line until Arm publishes its optimisation guide, when it returns.
- r16.4 [io_uring zero copy rx, PATCH v14 00/11](https://lore.kernel.org/io-uring/20250215000947.789731-1-dw@davidwei.uk/): Trimmed for length: the kernel document line carries the mechanism and names the implementer's epoll run, whose figures and missing fields stay under Claims.
- r16.5 [BitNet](https://github.com/microsoft/BitNet): Trimmed for length: the T-MAC line carries the table-lookup kernels BitNet descends from, and its speedups lack the same frequency, compiler and run count.
- r16.6 [SME Programmer's Guide](https://support.arm.com/documentation/109246/latest/): No rule failed. Listed in section 13 under Matrix extensions as the SME2 kernel document, so it is not repeated; the watched object is a Neoverse core that carries SME, for which no document exists, so the watchlist carries it as an unlinked line.
- r16.7 [Intel Xeon 6 with E-Core Processors](https://www.intel.com/content/www/us/en/products/details/processors/xeon/6-e-core-series.html): Rule 3, and superseded: the page's integer-throughput multiple holds its configuration off page, and the successor Xeon 6+ series launched in June 2026 with the same shape of claim and no public measurement either, so the successor's page is listed as the E-core part being watched.
- r16.8 [ParEval](https://github.com/parallelcodefoundry/ParEval): No rule failed. Left out under the entry cap: the benchmark for LLM-written parallel code (paper: Can Large Language Models Write Parallel Code?, arXiv 2401.12554) scores correctness and speedup across OpenMP, MPI and Kokkos tasks, but its speedups carry none of the seven fields and no routine from it has shipped, while the libc++ sorts have. It returns when a report of its speedups states all seven fields.
- r16.9 [D118029: Introduce new sorting algorithms for libc++](https://reviews.llvm.org/D118029): No rule failed. Left out because the Nature paper is the primary report of the generated sorts; the review confirms the merge and that the routines were run on isolated Skylake, Arm and AMD machines, and adds no field.
- r16.10 [Amazon EC2 M9g and M9gd general purpose instances](https://aws.amazon.com/about-aws/whats-new/2026/06/ec2-m9g-m9gd-instances-graviton5-processors-available/): No rule failed. Left out because the V3 optimisation guide is the primary document for the core and its line records that Graviton5 has shipped, while the notice adds percentages against M8g with none of the seven fields. It stood in for the guide while section 14 carried it, and the guide moved here once that section ceded it to the watchlist.
- r16.11 [N4A machine series](https://docs.cloud.google.com/compute/docs/general-purpose-machines#n4a_series): No rule failed. Left out because the N3 optimisation guide is the primary document for the core and its line names the series; the page states which core the series carries and no performance number, and section 14 rejects the same page at its cap.
- r16.12 [AWS Graviton Processors](https://aws.amazon.com/ec2/graviton/): No rule failed. Left out because the family page changes with each generation and carries no date, while the V3 guide line records the fact the watch rests on, that a Graviton5 part has shipped.
- r16.13 [Intel Outlines Architectures for Agentic AI at Hot Chips 2026](https://www.intel.com/content/www/us/en/newsroom/news/client-computing/intel-outlines-architectures-for-agentic-ai-at-hot-chips-2026.html): No rule failed. Left out because the announcement of Diamond Rapids says only "enhanced" AMX and "new" APX with no ship date, while the ISE programming reference is where AMX-FP8, APX and AVX10.2 are tied to the part by name, so the reference is listed instead.
- r16.14 [Intel Xeon 6 Performance Index](https://edc.intel.com/content/www/us/en/products/performance/benchmarks/intel-xeon-6/): Rule 3: the vendor's own disclosures behind its generation-on-generation percentages give CPU, cores, SMT and turbo state, workload and baseline, but not compiler flags, run count or the statistic reported, so no percentage from it can be quoted. The most complete vendor claims page found, and still short of seven.
- r16.15 [AMD Announces Production Ramp of Next-Generation AMD EPYC Processor Venice](https://newsroom.amd.com/news/amd-announces-production-ramp-of-next-generation-a/): No rule failed. Superseded as the vendor announcement by the July launch release of the same family, which names the series, core counts and frequencies the ramp release does not.
- r16.16 [AMD EPYC 9006 Server CPUs](https://www.amd.com/en/products/processors/server/epyc/9006-series.html): Rule 3: the product page's rack-level multiples against Zen 5, Xeon 6 and Vera are labelled performance projections, and the model tables are marked subject to change, so it is neither a measurement nor a stable specification. The launch release is listed instead.
- r16.17 [AMD EPYC Claims](https://www.amd.com/en/legal/claims/epyc.html): Rule 3 and rule 4: the footnote page behind the vendor's percentages renders only through script (a browser shows an empty footnote list until each claim is searched) and, where a footnote is reached, it names systems and workloads but no compiler flags, run count or statistic.
- r16.18 [AAI 2026: AMD Delivers Full-Stack Compute for the Agentic AI Era](https://newsroom.amd.com/news/aai-2026-full-stack-compute-agentic-ai/): Rule 3: a core-count ratio per rack presented as capacity, with the competitor part and every field beyond model and core count held in off-page footnotes. The EPYC launch release from the same event is listed as the vendor announcement.
- r16.19 [Intel Xeon 6 Product Brief](https://www.intel.com/content/www/us/en/products/docs/xeon-6-product-brief.html): Rule 3: multiples against prior-generation parts with the configurations held in off-page footnotes. The Xeon 6+ page is listed instead as the vendor page for the current E-core part.
- r16.20 [Intel Xeon 6 Processors with MRDIMM Solution Brief](https://www.intel.com/content/www/us/en/content-details/919018/intel-xeon-6-processors-with-mrdimm-solution-brief.html): Rule 3: a solution brief whose landing page carries "up to 39% higher bandwidth", "40% lower latency" and "2x faster AI performance than competitors" with every configuration held in the downloadable brief's footnotes, and rule 1 as marketing collateral rather than a specification. The standards body's published data buffer standard is listed instead.
- r16.21 [New Ultrafast Memory Boosts Intel Xeon Chips](https://www.intel.com/content/www/us/en/newsroom/news/new-ultrafast-memory-boosts-intel-xeon-chips.html): Rule 3: a transfer-rate ratio presented as a bandwidth gain and a "33% faster" figure credited to a third-party review, with no model, cores, frequency, compiler, workload or method. The data buffer standard is listed instead.
- r16.22 [JEDEC Advances DDR5 MRDIMM Ecosystem with New Memory Interface Logic and Expanded MRDIMM Roadmap](https://www.jedec.org/news/pressreleases/jedec%C2%AE-advances-ddr5-mrdimm-ecosystem-new-memory-interface-logic-and-expanded): Rule 1: the standards body's press release rather than the standard, stating that the data buffer standard is published, the clock driver standard is pending and the Gen2 module standard is still in progress, so the published standard page is listed instead.
- r16.23 [Arm Neoverse V3](https://www.arm.com/products/silicon-ip-cpu/neoverse/neoverse-v3): Rule 3: "double-digit performance improvements over Neoverse V2" with none of the seven fields. The core's software optimisation guide is the watchlist line; the page names SVE and not SME, which the SME line relies on.
- r16.24 [Arm AGI CPU](https://www.arm.com/products/cloud-datacenter/arm-agi-cpu): Rule 3: Arm's own Neoverse V3 server part, still listed as not yet available, with a per-rack multiple over x86 that the page's footnote marks as an estimate. It confirms a V3 part without SME and changes nothing on the SME line.
- r16.25 [Arm Neoverse CSS N4](https://www.arm.com/products/cloud-datacenter/neoverse-compute-subsystems/css-n4): Rule 3: per-socket, per-watt and bandwidth multiples against the prior subsystem with none of the seven fields, and no SME on the feature list, so it changes nothing on the SME line. The N4 core it names has no optimisation guide or shipped part yet, so the watchlist carries the core as an unlinked line rather than the subsystem page.
- r16.26 [SME2 on arm.com](https://www.arm.com/technologies/sme2): Rule 1: a marketing page whose shipped examples are all client parts. The programmer's guide in Arm's document library is listed in section 13, and the watchlist carries the Neoverse core as an unlinked line.
- r16.27 [The Scalable Matrix Extension (SME), for Armv9-A](https://support.arm.com/documentation/ddi0616/latest/): Rule 4: Arm marks the supplement retired and points to the Architecture Reference Manual, which the memory and single-thread sections already list. The programmer's guide, in section 13, is the SME-specific document kept.
- r16.28 [Powering Microsoft's Azure Cobalt 200 with Arm Neoverse CSS V3](https://newsroom.arm.com/blog/microsoft-azure-cobalt-200-arm-neoverse-css-v3): Rule 1: a partner announcement with no specification and no measurement, for a part whose instances are in preview rather than shipped; the V3 guide line carries the core, shipped in Graviton5, instead.
- r16.29 [riscv-v-spec repository](https://github.com/riscvarchive/riscv-v-spec): Rule 4: archived, and https://github.com/riscv/riscv-v-spec now answers 301 to it. The ratified text lives in the RISC-V specifications library, which the watchlist links, so the archive adds nothing.
- r16.30 [Tenstorrent Announces Availability of TT-Ascalon](https://tenstorrent.com/en/newsroom/tenstorrent-announces-availability-of-tt-ascalon): Rule 3: SPEC-per-GHz figures with no statement of silicon versus simulation, no core count, compiler, baseline or method. The RVV line counts it only as an announced core.
- r16.31 [SiFive Performance P870](https://www.sifive.com/cores/performance-p870): Rule 4: answers 307 to a consolidated product page, which carries no measurement either way, so there is nothing to promote.
- r16.32 [Ventana Introduces Veyron V2](https://riscv.org/blog/ventana-introduces-veyron-v2-worlds-highest-performance-data-center-class-risc-v-processor-and-platform/): Rule 3: a superlative performance claim with none of the seven fields, for IP and chiplets rather than a socketed part, reached through RISC-V International's republication because ventanamicro.com refused every connection during verification.
- r16.33 [Pond: CXL-Based Memory Pooling Systems for Cloud Platforms](https://arxiv.org/abs/2203.00241): Rule 3: the added-latency figures are estimates from topology analysis, and the workload results emulate them with a remote NUMA node on named Skylake and Zen 2 parts rather than measuring a shipped pool. The measurement on shipped expansion devices is the Demystifying CXL Memory line; Pond returns to the core once a multi-host pool is measured.
- r16.34 [Compute Express Link Linux driver documentation](https://docs.kernel.org/driver-api/cxl/index.html): No rule failed. Left out because the specification page is the promotion target and the kernel side is not what is unproven.
- r16.35 [scx](https://github.com/sched-ext/scx): No rule failed. Left out because the kernel document is the specification and links the shipped schedulers.
- r16.36 [Private Cloud Compute](https://security.apple.com/blog/private-cloud-compute/): No rule failed. Left out under the entry cap because it is the one candidate with none of the three promotion requirements in public, no part name, no specification and no measurement, so there is nothing to watch beyond the announcement. It returns once Apple names the part or publishes a specification for it.
- r16.37 [1-bit AI Infra: Part 1.1, Fast and Lossless BitNet b1.58 Inference on CPUs](https://arxiv.org/abs/2410.16144): Rule 3: the technical report behind the BitNet line's speedups names CPUs and thread counts but no frequency, compiler, baseline version or run statistic. The repository is listed as the implementation and the report is where the missing fields would be published.
- r16.38 [Bitnet.cpp: Efficient Edge Inference for Ternary LLMs](https://arxiv.org/abs/2502.11880): Rule 3: the later report on the same kernels states speedups over full-precision and low-bit baselines with no CPU model in the abstract and the same fields missing. The repository stays as the single BitNet line.
- r16.39 [The Era of 1-bit LLMs](https://arxiv.org/abs/2402.17764): No rule failed as a model paper. Left out because the watchlist question is the CPU kernel measurement, not the model.
- r16.40 [T-MAC: CPU Renaissance via Table Lookup for Low-Bit LLM Deployment on Edge](https://arxiv.org/abs/2407.00088): No rule failed as the mechanism paper, which section 13 lists under Quantization. Its numbers state warm-up, run count, mean and baseline version but not frequency, turbo, compiler or flags, so no number is quoted anywhere, and the repository, where the missing fields would be published, is the watchlist line.

### Numbers examined

- c16.1 "1.37x to 5.07x on ARM CPUs" and "2.37x to 6.17x on x86 CPUs" from https://github.com/microsoft/BitNet, restated from https://arxiv.org/abs/2410.16144: fields present: 1 CPU (Apple M2 Ultra; Intel Core i7-13700H, Raptor Lake), 2 cores (two threads, or unrestricted with the best speed reported), 5 workload (BitNet b1.58 models from 125M to 100B parameters, token generation), 6 baseline (llama.cpp at fp16, version not stated). Missing: 3 frequency, turbo and SMT state, 4 compiler and flags, 7 method (runs, statistic). Verdict: watchlist. Verdict: watchlist.
- c16.2 "5-7 tokens per second" for a 100B model "on a single CPU" from https://github.com/microsoft/BitNet: fields present: 5 workload (100B BitNet b1.58, token generation). Missing: 1, 2, 3, 4, 6, 7. Verdict: watchlist; number not quoted. Verdict: watchlist.
- c16.3 "up to a 6.25x increase in speed over full-precision baselines and up to 2.32x over low-bit baselines" from https://arxiv.org/abs/2502.11880: fields present in the abstract: 6 baseline (full-precision and low-bit llama.cpp, versions not stated). Missing in the abstract: 1, 2, 3, 4, 5, 7. Verdict: watchlist; number not quoted. Verdict: watchlist.
- c16.4 "up to 4x increase in throughput and 70% reduction in energy consumption compared to llama.cpp" and "30 tokens/s with a single core and 71 tokens/s with eight cores on M2-Ultra" from https://arxiv.org/abs/2407.00088, repeated at https://github.com/microsoft/T-MAC: fields present: 1 CPU (Apple M2 Ultra; Cortex-A76 in a Raspberry Pi 5; Cortex-A78AE in a Jetson AGX Orin; Intel Core i5-1035G7, Ice Lake), 2 cores (one and eight for the headline, per-device counts in a table), 5 workload (BitNet-3B and low-bit Llama, sixty-four tokens generated over twenty iterations), 6 baseline (llama.cpp b2794), 7 method (ten warm-up iterations then a hundred runs, mean; energy sampled with powermetrics on the Mac). Missing: 3 frequency, turbo and SMT state, 4 compiler, flags and TVM version. Verdict: watchlist; the closest of the inference candidates to seven. Verdict: watchlist.
- c16.5 "up to 70% for sequences of a length of five and roughly 1.7% for sequences exceeding 250,000 elements" from https://www.nature.com/articles/s41586-023-06004-9: fields present: 1 CPU as families only (ARMv8, Intel Skylake, AMD Zen 2, no model), 5 workload (the libc++ sort benchmarks over uint32, uint64 and float, ten thousand random inputs per measurement), 6 baseline (the previous libc++ routines), 7 partly (fifth percentile of latency across a hundred machines, read from the unhalted core cycles counter). Missing: 1 CPU model, 2 cores, 3 frequency, turbo and SMT state, 4 compiler and flags. Verdict: watchlist; number not quoted. Verdict: watchlist.
- c16.6 "82.2 Gbps" epoll against "116.2 Gbps (+41%)" io_uring with the application thread and receive softirq on different cores, and "62.6 Gbps" against "80.9 Gbps (+29%)" on the same core, from https://lore.kernel.org/io-uring/20250215000947.789731-1-dw@davidwei.uk/: fields present: 1 CPU (AMD EPYC 9454, Zen 4), 2 cores (one application thread and the softirq, pinned apart or together), 5 workload (a single TCP flow over a Broadcom BCM957508 at 4K MTU, driven by kperf on a v6.11 kernel base), 6 baseline (epoll on the copying receive path), 7 partly (the tool and its fork are named). Missing: 3 frequency, turbo and SMT state, 4 compiler and flags, 7 run count and statistic. The letter itself notes that the comparison with TCP_ZEROCOPY_RECEIVE was not yet done. Verdict: watchlist. Verdict: watchlist.
- c16.7 Memory access latency and bandwidth for three CXL expansion devices against local and remote DDR5 from https://arxiv.org/abs/2303.15375 (Table 1 and section 3 of the paper): fields present: 1 CPU (two Intel Xeon 6430, Sapphire Rapids), 2 cores (32 per socket), 3 frequency fixed at 2.1 GHz with Hyper-Threading disabled, 5 workload (Intel MLC and the authors' own pointer-chasing tool, per instruction type), 6 baseline (local DDR5-4800 and a one-channel remote NUMA node emulating CXL), 7 method (ten thousand repeats, median; Ubuntu 22.04.2 with kernel 6.2). Missing: 4 compiler and flags. Verdict: watchlist; the nearest CXL measurement to seven fields, and it covers expansion devices, not a multi-host pool. Verdict: watchlist.
- c16.8 "Up to 2.5x higher performance" and "Up to 79% greater performance per watt" against Intel Xeon 6 with E-cores, and "Up to 30% better performance per thread" against an unnamed competitor, from https://www.intel.com/content/www/us/en/products/details/processors/xeon/6-plus-series.html: fields present: 1 CPU (the Xeon 6+ E-core series, no model on the claim), 6 baseline (the prior E-core series; the competitor is not named). Missing: 2, 3, 4, 5, 7 (each footnote points to the vendor claims site). Verdict: watchlist; number not quoted, the entry rests on the shipped part. Verdict: watchlist.
- c16.9 "1.8x faster agentic sandbox performance" and "2x memory bandwidth" against "leading x86 CPUs" from https://www.nvidia.com/en-us/data-center/vera-cpu/: fields present: 1 CPU (Vera, Olympus cores), 2 cores (88). Missing: 3, 4, 5 beyond the phrase "agentic sandbox", 6 (the x86 part is not named), 7 (the page says "relative performance based on measured data, subject to change"). Verdict: watchlist; number not quoted. Verdict: watchlist.
- c16.10 "up to 25% better compute performance" against M8g instances, with "up to 30%" for databases and "up to 35%" for web applications and machine learning, from https://aws.amazon.com/about-aws/whats-new/2026/06/ec2-m9g-m9gd-instances-graviton5-processors-available/: fields present: 1 CPU (Graviton5, Neoverse V3), 6 baseline (M8g on Graviton4). Missing: 2, 3, 4, 5, 7. Verdict: cut for the notice; the shipped part is recorded on the V3 guide line. Verdict: cut.
- c16.11 https://docs.cloud.google.com/compute/docs/general-purpose-machines#n4a_series carries no performance number for N4A; it names the Neoverse N3 core, the vCPU and memory ceilings and the network bandwidth. Verdict: nothing to quote; the series is named on the N3 guide line, and the page is Rejected only. Verdict: not_quoted.
- c16.12 "peak bandwidth rises by almost 40%, from 6,400 megatransfers per second (MT/s) to 8,800 MT/s" and "completed jobs as much as 33% faster" from https://www.intel.com/content/www/us/en/newsroom/news/new-ultrafast-memory-boosts-intel-xeon-chips.html: fields present: 1 CPU (Xeon 6 family, no model), 6 baseline (the same system with RDIMMs at 6400 MT/s). Missing: 2 cores, 3 frequency, turbo and SMT, 4 compiler and flags, 5 workload, 7 method (the second figure is credited to a third-party review). The first figure is a transfer-rate ratio, not a measurement. Verdict: cut for the page; the module stays on the watchlist. Verdict: cut.
- c16.13 "up to 39% higher bandwidth", "40% lower latency" and "2x faster AI performance than competitors" from https://www.intel.com/content/www/us/en/content-details/919018/intel-xeon-6-processors-with-mrdimm-solution-brief.html: fields present on the landing page: none (configurations sit in the downloadable brief's footnotes). Missing: 1 to 7. Verdict: cut for the brief; the module stays on the watchlist through the JEDEC data buffer standard, which carries no performance number. Verdict: cut.
- c16.14 "double-digit performance improvements over Neoverse V2 on cloud and ML applications" from https://www.arm.com/products/silicon-ip-cpu/neoverse/neoverse-v3: fields present: 6 baseline (Neoverse V2). Missing: 1, 2, 3, 4, 5, 7. Verdict: cut. Verdict: cut.
- c16.15 "Greater than 2x performance per rack on Arm" from https://www.arm.com/products/cloud-datacenter/arm-agi-cpu: the page's footnote marks it as an estimate for an unshipped part. Fields present: 1 CPU (Arm AGI CPU, Neoverse V3), 2 cores (136 per socket). Missing: 3, 4, 5, 6 (the x86 rack is not named), 7. Verdict: cut. Verdict: cut.
- c16.16 "up to 2x performance per socket", "1.25x performance per watt" and "1.75x memory bandwidth" from https://www.arm.com/products/cloud-datacenter/neoverse-compute-subsystems/css-n4: fields present: 6 baseline (the prior Neoverse CSS generation). Missing: 1, 2, 3, 4, 5, 7. Verdict: cut for the page; the N4 core stays on the watchlist as an unlinked line. Verdict: cut.
- c16.17 "more than 22 SPECint 2006/GHz, >2.3 SPECint 2017/GHz and >3.6 SPECfp 2017/GHz" and "operates at >2.5 GHz on the Samsung SF4X process node" from https://tenstorrent.com/en/newsroom/tenstorrent-announces-availability-of-tt-ascalon: fields present: 1 CPU (TT-Ascalon core on Samsung SF4X), 3 partly (a frequency floor, no turbo or SMT statement), 5 workload (SPEC CPU). Missing: 2 cores, 4 compiler and flags, 6 baseline, 7 method (silicon or simulation, runs, statistic). Verdict: cut. Verdict: cut.
- c16.18 "The AMD EPYC 9996 (256C) provides 2.08x the cores (and threads with SMT) per rack" and "Highest Agents / CPU W" from https://newsroom.amd.com/news/aai-2026-6th-gen-epyc/ and https://newsroom.amd.com/news/aai-2026-full-stack-compute-agentic-ai/: the footnotes state that agent counts are "estimates derived from available CPU thread resources used as a proxy under a consistent theoretical workload", so these are core-count and power ratios, not measurements. Fields present: 1 CPU (EPYC 9996, Zen 6c; Nvidia Vera; Xeon 6980P), 2 cores (256, 88 and 128), 3 partly (SMT on, no frequency or turbo state). Missing: 4, 5, 6 (a rack power envelope, not a run), 7. Verdict: cut; the family stays on the watchlist as an announced part. Verdict: cut.
- c16.19 "Up to 3.3x" rack performance for EPYC 9996 against "NVIDIA Vera 1.0" from https://www.amd.com/en/products/processors/server/epyc/9006-series.html: labelled on the page as a performance projection within a rack power budget. Fields present: 1 CPU (EPYC 9996; EPYC 9965; Xeon 6980P; Nvidia Vera). Missing: 2, 3, 4, 5, 6 beyond the part name, 7 (projection, not measurement). Verdict: cut. Verdict: cut.
- c16.20 "Up to 256 new cores with 1.28 GB LLC" and "16 memory channels with 12800 MT/s" from https://www.intel.com/content/www/us/en/newsroom/news/client-computing/intel-outlines-architectures-for-agentic-ai-at-hot-chips-2026.html: specification figures for an unshipped part, not measurements; no performance number appears on the page. Verdict: nothing to quote; the part is carried by the ISE reference line. Verdict: not_quoted.
- c16.21 "Up to 3.6 x higher integer throughput performance with Intel Xeon 6780E vs. 2nd Gen Intel Xeon 8280" from https://www.intel.com/content/www/us/en/products/details/processors/xeon/6-e-core-series.html: fields present: 1 CPU (Xeon 6780E, Sierra Forest; Xeon Platinum 8280, Cascade Lake), 6 baseline. Missing on the page: 2, 3, 4, 5, 7 (the footnote points to the vendor claims site). Verdict: cut for the page, which is superseded by the Xeon 6+ series; the E-core part stays on the watchlist through the successor's page. Verdict: cut.
- c16.22 "2x Intel(R) 6980P, 128 cores, 500W TDP, HT On, Turbo On" behind a BERT-large inference claim on https://edc.intel.com/content/www/us/en/products/performance/benchmarks/intel-xeon-6/: fields present: 1 CPU (Xeon 6980P, Granite Rapids), 2 cores, 3 SMT and turbo state (frequency itself not fixed), 5 workload (BERT-large int8, batch one, with an SLA), 6 baseline (Xeon Platinum 8592+). Missing: 4 flags (framework versions only), 7 method (runs, statistic). Verdict: cut; the nearest vendor disclosure to seven fields, and still short. Verdict: cut.
- c16.23 "we estimate that CXL will add 70-90ns to access latencies over same-NUMA-node DRAM with a pool size of 8-16 sockets, and add more than 180ns for rack-scale pooling" from https://arxiv.org/abs/2203.00241: an estimate from topology analysis; the workload study emulates it with a remote NUMA node. Fields present for the emulation: 1 CPU (two Intel Xeon Platinum 8157M, Skylake; two AMD EPYC 7452, Zen 2), 5 workload (158 workloads), 6 baseline (same-NUMA-node DRAM), 7 method (measured local and remote latency and bandwidth, then slowdown under the emulated latency). Missing: 2 cores, 3 partly (SMT, turbo and C-states off, frequency not stated), 4 compiler and flags, and no shipped pool was measured. Verdict: watchlist, on the specification line; the paper itself is rejected. Verdict: watchlist.

