Sources
20 entries in 4 parts
In the MCP server
cpuperf://section/12

Measuring the tail

  1. 01
    The Tail at Scale paper

    Shows why fan-out makes a rare slow server a common slow request, and names the techniques that tolerate variance.

  2. 02
    Attack of the Killer Microseconds paper

    Defines the stall band that out-of-order hardware cannot hide and a context switch cannot amortise.

  3. 03
    How NOT to Measure Latency talk

    Shows that a summary without a max discards the samples that define the tail, and closed-loop load never records them.

  4. 04
    Coordinated Omission report

    The original definition of the recording error, with arithmetic for how far a reported percentile sits from the truth.

  5. 05
    HdrHistogram repository

    Keeps the whole distribution at fixed relative precision in constant time, so the far percentiles and max survive.

Reproduce it

closed-loop against open-loop p99 under the same stalls.

Reproduce it · 12-coordinated-omission Coordinated omission A closed-loop load generator hides stalls that an open-loop one reports

Where jitter comes from

  1. 01
    rt-tests repository

    The reference wakeup-latency measurement for Linux, whose README states that an unloaded run proves nothing.

  2. 02
    osnoise tracer manual

    Counts the noise a spinning thread suffers and attributes each event to NMI, IRQ, softirq, thread or hardware.

  3. 03
    Tales of the Tail paper

    Derives the queueing-ideal tail and attributes the excess to scheduling, interrupt placement, power saving and NUMA.

  4. 04
    Latency Implications of Virtual Memory report

    Measures with code the page-fault, TLB-shootdown and writeback stalls that memory mapping hides from the caller.

  5. 05
    The KVM halt polling system manual

    Defines the host-side polling after a vCPU halt that trades idle host CPU for guest wakeup time, unseen by the guest.

Load generation and production workloads

  1. 01
    Open Versus Closed: A Cautionary Tale paper

    Shows that open and closed load models disagree on response time and scheduling gains, with rules for choosing one.

  2. 02
    wrk2 repository

    Issues requests on a fixed schedule and times each from when it was due, so server stalls reach the percentiles.

  3. 03
    Reconciling High Server Utilization and Sub-millisecond Quality-of-Service paper

    Shows the tail, not throughput, caps a latency-critical server's utilisation, and how far co-located work lowers it.

  4. 04
    TailBench report

    Pairs latency-critical services with an open-loop harness that records sojourn against service time per request.

  5. 05
    Workload Analysis of a Large-Scale Key-Value Store paper

    Measures the key, value and inter-arrival distributions of live key-value traffic, the shape load generators imitate.

Mechanical sympathy

  1. 01
    Inter Thread Latency report

    Measures with code the floor for handing a cache line between cores, which every queue and lock is built on.

  2. 02
    Single Writer Principle report

    States the design rule that removes write contention outright, using a contended increment's cost as the argument.

  3. 03
    Optimizing a Ring Buffer for Throughput report

    Adds cached indices to a single-producer single-consumer ring and shows with counters the coherence traffic removed.

  4. 04
    LMAX Disruptor manual

    Applies the single writer rule and cache-line padding to a ring buffer, with the queue comparison that motivated it.

  5. 05
    Aeron repository

    Carries the single writer and batching rules through a whole transport, the reference beyond one in-process queue.