---
title: "12. Tail latency and production systems"
url: https://cpuperf.com/learn/tail-latency-and-production-systems/
source: https://github.com/usamahz/cpu-performance-engineering/blob/deb5a0bac46760503b6f4a2608bdfed470c8532e/README.md#L521
commit: deb5a0bac46760503b6f4a2608bdfed470c8532e
---
## 12. Tail latency and production systems

A latency figure means nothing without its percentile, its load model and the way it was recorded.

### Measuring the tail

- [The Tail at Scale](https://www.barroso.org/publications/TheTailAtScale.pdf) - Shows why fan-out makes a rare slow server a common slow request, and names the techniques that tolerate variance.
- [Attack of the Killer Microseconds](https://www.barroso.org/publications/AttackoftheKillerMicroseconds.pdf) - Defines the stall band that out-of-order hardware cannot hide and a context switch cannot amortise.
- [How NOT to Measure Latency](https://www.youtube.com/watch?v=lJ8ydIuPFeU) - Shows that a summary without a max discards the samples that define the tail, and closed-loop load never records them.
- [Coordinated Omission](https://groups.google.com/g/mechanical-sympathy/c/icNZJejUHfE) - The original definition of the recording error, with arithmetic for how far a reported percentile sits from the truth.
- [HdrHistogram](https://github.com/HdrHistogram/HdrHistogram) - Keeps the whole distribution at fixed relative precision in constant time, so the far percentiles and max survive.

Reproduce it: [misc/benchmarks/12-coordinated-omission](https://cpuperf.com/benchmarks/12-coordinated-omission/), closed-loop against open-loop p99 under the same stalls.

### Where jitter comes from

- [rt-tests](https://git.kernel.org/pub/scm/utils/rt-tests/rt-tests.git/) - The reference wakeup-latency measurement for Linux, whose README states that an unloaded run proves nothing.
- [osnoise tracer](https://docs.kernel.org/trace/osnoise-tracer.html) - Counts the noise a spinning thread suffers and attributes each event to NMI, IRQ, softirq, thread or hardware.
- [Tales of the Tail](https://syslab.cs.washington.edu/papers/latency-socc14.pdf) - Derives the queueing-ideal tail and attributes the excess to scheduling, interrupt placement, power saving and NUMA.
- [Latency Implications of Virtual Memory](https://rigtorp.se/virtual-memory/) - Measures with code the page-fault, TLB-shootdown and writeback stalls that memory mapping hides from the caller.
- [The KVM halt polling system](https://docs.kernel.org/virt/kvm/halt-polling.html) - Defines the host-side polling after a vCPU halt that trades idle host CPU for guest wakeup time, unseen by the guest.

### Load generation and production workloads

- [Open Versus Closed: A Cautionary Tale](https://www.usenix.org/conference/nsdi-06/open-versus-closed-cautionary-tale) - Shows that open and closed load models disagree on response time and scheduling gains, with rules for choosing one.
- [wrk2](https://github.com/giltene/wrk2) - Issues requests on a fixed schedule and times each from when it was due, so server stalls reach the percentiles.
- [Reconciling High Server Utilization and Sub-millisecond Quality-of-Service](https://csl.stanford.edu/~christos/publications/2014.mutilate.eurosys.pdf) - Shows the tail, not throughput, caps a latency-critical server's utilisation, and how far co-located work lowers it.
- [TailBench](https://tailbench.csail.mit.edu/) - Pairs latency-critical services with an open-loop harness that records sojourn against service time per request.
- [Workload Analysis of a Large-Scale Key-Value Store](https://jiangs.utasites.cloud/pubs/papers/atikoglu12-memcached.pdf) - Measures the key, value and inter-arrival distributions of live key-value traffic, the shape load generators imitate.

### Mechanical sympathy

- [Inter Thread Latency](https://mechanical-sympathy.blogspot.com/2011/08/inter-thread-latency.html) - Measures with code the floor for handing a cache line between cores, which every queue and lock is built on.
- [Single Writer Principle](https://mechanical-sympathy.blogspot.com/2011/09/single-writer-principle.html) - States the design rule that removes write contention outright, using a contended increment's cost as the argument.
- [Optimizing a Ring Buffer for Throughput](https://rigtorp.se/ringbuffer/) - Adds cached indices to a single-producer single-consumer ring and shows with counters the coherence traffic removed.
- [LMAX Disruptor](https://lmax-exchange.github.io/disruptor/disruptor.html) - Applies the single writer rule and cache-line padding to a ring buffer, with the queue comparison that motivated it.
- [Aeron](https://github.com/aeron-io/aeron) - Carries the single writer and batching rules through a whole transport, the reference beyond one in-process queue.
