Learn / Out
OS and I/O
Read first
A syscall's cost depends on the mitigation state, the governor and the idle state the core was in, three sysfs settings that change after boot, so each is recorded beside any number below.
- Sources
- 20 entries in 4 parts
- Reproduce it
- 11-syscall-cost
- Related sections
- §12 Tail latency and production systems
- In the MCP server
cpuperf://section/11
Syscalls and asynchronous I/O
-
01 vdso(7) manual
Defines the calls the kernel answers without a mode switch, and the clocksource condition for skipping the trap.
-
02
Separates a syscall's trap cost from its cache and TLB pollution, and shows the pollution can dominate.
-
03
Measures syscall and context switch cost across kernel releases and traces each slowdown to a named mitigation.
-
04
States the goals aio failed, and which io_uring features remove a syscall and which remove a copy.
-
05
Measures io_uring's polling modes against libaio and SPDK, and shows the kernel poller needs its own core.
Reproduce it
the fixed cost of a kernel crossing across request sizes.
Reproduce it · 11-syscall-cost Syscall cost A kernel crossing has a fixed cost that request size does not amortiseScheduling, affinity and isolation
-
01 EEVDF Scheduler manual
Defines lag and virtual deadline, which the default class schedules by, and the slice request in sched_setattr.
-
02
Proves cores sit idle while runnable threads queue, and gives the invariant checker that found the load-balancer bugs.
-
03 Control Group v2 manual
Defines cpu.max throttling, cpu.weight and the cpusets that bound affinity, the controls behind every container limit.
-
04 CPU Performance Scaling manual
Defines the governors, driver and boost switch that set a core's frequency, the sysfs state a measurement records.
-
05 CPU Isolation manual
Ties isolcpus, nohz_full, IRQ affinity, RCU offload and cpusets into one recipe, and lists the jitter it leaves.
Interrupts and kernel bypass
-
01 NAPI manual
Defines the polling, software coalescing, busy polling and IRQ suspension knobs that trade interrupts against latency.
-
02 DPDK Programmer's Guide manual
Defines the full bypass model, pinned poll-mode cores with no interrupts, that every kernel path is measured against.
-
03 The eXpress Data Path repository
Measures an in-kernel programmable path against DPDK and the stack per core, with the full configuration published.
-
04
Separates direct and indirect NIC interrupt cost, measures the stack against bypass, and is where IRQ suspension began.
-
05 AF_XDP manual
Defines the socket and UMEM rings handing XDP frames to user space, and the zero-copy and need-wakeup modes.
Cache and bandwidth partitioning
-
01
Defines classes of service, cache masks, bandwidth allocation and monitoring IDs, the model resctrl exposes.
-
02
Defines the filesystem through which Linux exposes Intel, AMD and Arm partitioning, and the schemata format.
-
03 MPAM manual
Maps Arm's cache portion and bandwidth controls onto resctrl's schemata, and states which platform limits apply.
-
04
Shows at fleet scale that cycles per instruction alone finds an interfering neighbour and the one to throttle.
-
05
Shows cache ways, cores, bandwidth and power must be partitioned together, or batch work reaches the tail.