---
title: "11. OS and I/O"
url: https://cpuperf.com/learn/os-and-io/
source: https://github.com/usamahz/cpu-performance-engineering/blob/deb5a0bac46760503b6f4a2608bdfed470c8532e/README.md#L483
commit: deb5a0bac46760503b6f4a2608bdfed470c8532e
---
## 11. OS and I/O

A syscall's cost depends on the mitigation state, the governor and the idle state the core was in, three sysfs settings that change after boot, so each is recorded beside any number below.

### Syscalls and asynchronous I/O

- [vdso(7)](https://man7.org/linux/man-pages/man7/vdso.7.html) - Defines the calls the kernel answers without a mode switch, and the clocksource condition for skipping the trap.
- [FlexSC: Flexible System Call Scheduling with Exception-Less System Calls](https://www.usenix.org/conference/osdi10/flexsc-flexible-system-call-scheduling-exception-less-system-calls) - Separates a syscall's trap cost from its cache and TLB pollution, and shows the pollution can dominate.
- [An Analysis of Performance Evolution of Linux's Core Operations](https://www.eecg.toronto.edu/~stumm/Papers/Ren-sosp-19.pdf) - Measures syscall and context switch cost across kernel releases and traces each slowdown to a named mitigation.
- [Efficient IO with io_uring](https://www.kernel.dk/io_uring.pdf) - States the goals aio failed, and which io_uring features remove a syscall and which remove a copy.
- [Understanding Modern Storage APIs: A systematic study of libaio, SPDK, and io_uring](https://atlarge-research.com/pdfs/2022-systor-apis.pdf) - Measures io_uring's polling modes against libaio and SPDK, and shows the kernel poller needs its own core.

Reproduce it: [misc/benchmarks/11-syscall-cost](https://cpuperf.com/benchmarks/11-syscall-cost/), the fixed cost of a kernel crossing across request sizes.

### Scheduling, affinity and isolation

- [EEVDF Scheduler](https://docs.kernel.org/scheduler/sched-eevdf.html) - Defines lag and virtual deadline, which the default class schedules by, and the slice request in sched_setattr.
- [The Linux Scheduler: a Decade of Wasted Cores](https://people.ece.ubc.ca/sasha/papers/eurosys16-final29.pdf) - Proves cores sit idle while runnable threads queue, and gives the invariant checker that found the load-balancer bugs.
- [Control Group v2](https://docs.kernel.org/admin-guide/cgroup-v2.html) - Defines cpu.max throttling, cpu.weight and the cpusets that bound affinity, the controls behind every container limit.
- [CPU Performance Scaling](https://docs.kernel.org/admin-guide/pm/cpufreq.html) - Defines the governors, driver and boost switch that set a core's frequency, the sysfs state a measurement records.
- [CPU Isolation](https://docs.kernel.org/admin-guide/cpu-isolation.html) - Ties isolcpus, nohz_full, IRQ affinity, RCU offload and cpusets into one recipe, and lists the jitter it leaves.

### Interrupts and kernel bypass

- [NAPI](https://docs.kernel.org/networking/napi.html) - Defines the polling, software coalescing, busy polling and IRQ suspension knobs that trade interrupts against latency.
- [DPDK Programmer's Guide](https://doc.dpdk.org/guides/prog_guide/) - Defines the full bypass model, pinned poll-mode cores with no interrupts, that every kernel path is measured against.
- [The eXpress Data Path](https://github.com/tohojo/xdp-paper) - Measures an in-kernel programmable path against DPDK and the stack per core, with the full configuration published.
- [Kernel vs. User-Level Networking: Don't Throw Out the Stack with the Interrupts](https://cs.uwaterloo.ca/~mkarsten/papers/sigmetrics2024.html) - Separates direct and indirect NIC interrupt cost, measures the stack against bypass, and is where IRQ suspension began.
- [AF_XDP](https://docs.kernel.org/networking/af_xdp.html) - Defines the socket and UMEM rings handing XDP frames to user space, and the zero-copy and need-wakeup modes.

### Cache and bandwidth partitioning

- [Intel Resource Director Technology Architecture Specification](https://www.intel.com/content/www/us/en/content-details/789566/intel-resource-director-technology-intel-rdt-architecture-specification.html) - Defines classes of service, cache masks, bandwidth allocation and monitoring IDs, the model resctrl exposes.
- [User Interface for Resource Control feature (resctrl)](https://docs.kernel.org/filesystems/resctrl.html) - Defines the filesystem through which Linux exposes Intel, AMD and Arm partitioning, and the schemata format.
- [MPAM](https://docs.kernel.org/arch/arm64/mpam.html) - Maps Arm's cache portion and bandwidth controls onto resctrl's schemata, and states which platform limits apply.
- [CPI2: CPU performance isolation for shared compute clusters](https://john.e-wilkes.com/papers/2013-EuroSys-CPI2.pdf) - Shows at fleet scale that cycles per instruction alone finds an interfering neighbour and the one to throttle.
- [Heracles: Improving Resource Efficiency at Scale](https://csl.stanford.edu/~christos/publications/2015.heracles.isca.pdf) - Shows cache ways, cores, bandwidth and power must be partitioned together, or batch work reaches the tail.
