Learn / Out
Tail latency and production systems
Read first
A latency figure means nothing without its percentile, its load model and the way it was recorded.
- Sources
- 20 entries in 4 parts
- Reproduce it
- 12-coordinated-omission
- Related sections
- §11 OS and I/O, §15 Benchmarks
- In the MCP server
cpuperf://section/12
Measuring the tail
-
01 The Tail at Scale paper
Shows why fan-out makes a rare slow server a common slow request, and names the techniques that tolerate variance.
-
02
Defines the stall band that out-of-order hardware cannot hide and a context switch cannot amortise.
-
03
Shows that a summary without a max discards the samples that define the tail, and closed-loop load never records them.
-
04 Coordinated Omission report
The original definition of the recording error, with arithmetic for how far a reported percentile sits from the truth.
-
05 HdrHistogram repository
Keeps the whole distribution at fixed relative precision in constant time, so the far percentiles and max survive.
Reproduce it
closed-loop against open-loop p99 under the same stalls.
Reproduce it · 12-coordinated-omission Coordinated omission A closed-loop load generator hides stalls that an open-loop one reportsWhere jitter comes from
-
01 rt-tests repository
The reference wakeup-latency measurement for Linux, whose README states that an unloaded run proves nothing.
-
02 osnoise tracer manual
Counts the noise a spinning thread suffers and attributes each event to NMI, IRQ, softirq, thread or hardware.
-
03 Tales of the Tail paper
Derives the queueing-ideal tail and attributes the excess to scheduling, interrupt placement, power saving and NUMA.
-
04
Measures with code the page-fault, TLB-shootdown and writeback stalls that memory mapping hides from the caller.
-
05 The KVM halt polling system manual
Defines the host-side polling after a vCPU halt that trades idle host CPU for guest wakeup time, unseen by the guest.
Load generation and production workloads
-
01
Shows that open and closed load models disagree on response time and scheduling gains, with rules for choosing one.
-
02 wrk2 repository
Issues requests on a fixed schedule and times each from when it was due, so server stalls reach the percentiles.
-
03
Shows the tail, not throughput, caps a latency-critical server's utilisation, and how far co-located work lowers it.
-
04 TailBench report
Pairs latency-critical services with an open-loop harness that records sojourn against service time per request.
-
05
Measures the key, value and inter-arrival distributions of live key-value traffic, the shape load generators imitate.
Mechanical sympathy
-
01 Inter Thread Latency report
Measures with code the floor for handing a cache line between cores, which every queue and lock is built on.
-
02 Single Writer Principle report
States the design rule that removes write contention outright, using a contended increment's cost as the argument.
-
03
Adds cached indices to a single-producer single-consumer ring and shows with counters the coherence traffic removed.
-
04 LMAX Disruptor manual
Applies the single writer rule and cache-line padding to a ring buffer, with the queue comparison that motivated it.
-
05 Aeron repository
Carries the single writer and batching rules through a whole transport, the reference beyond one in-process queue.