Sources
20 entries in 4 parts
Reproduce it
11-syscall-cost
In the MCP server
cpuperf://section/11

Syscalls and asynchronous I/O

  1. 01
    vdso(7) manual

    Defines the calls the kernel answers without a mode switch, and the clocksource condition for skipping the trap.

  2. 02
    FlexSC: Flexible System Call Scheduling with Exception-Less System Calls paper

    Separates a syscall's trap cost from its cache and TLB pollution, and shows the pollution can dominate.

  3. 03
    An Analysis of Performance Evolution of Linux's Core Operations paper

    Measures syscall and context switch cost across kernel releases and traces each slowdown to a named mitigation.

  4. 04
    Efficient IO with io_uring paper

    States the goals aio failed, and which io_uring features remove a syscall and which remove a copy.

  5. 05
    Understanding Modern Storage APIs: A systematic study of libaio, SPDK, and io_uring paper

    Measures io_uring's polling modes against libaio and SPDK, and shows the kernel poller needs its own core.

Reproduce it

the fixed cost of a kernel crossing across request sizes.

Reproduce it · 11-syscall-cost Syscall cost A kernel crossing has a fixed cost that request size does not amortise

Scheduling, affinity and isolation

  1. 01
    EEVDF Scheduler manual

    Defines lag and virtual deadline, which the default class schedules by, and the slice request in sched_setattr.

  2. 02
    The Linux Scheduler: a Decade of Wasted Cores paper

    Proves cores sit idle while runnable threads queue, and gives the invariant checker that found the load-balancer bugs.

  3. 03
    Control Group v2 manual

    Defines cpu.max throttling, cpu.weight and the cpusets that bound affinity, the controls behind every container limit.

  4. 04
    CPU Performance Scaling manual

    Defines the governors, driver and boost switch that set a core's frequency, the sysfs state a measurement records.

  5. 05
    CPU Isolation manual

    Ties isolcpus, nohz_full, IRQ affinity, RCU offload and cpusets into one recipe, and lists the jitter it leaves.

Interrupts and kernel bypass

  1. 01
    NAPI manual

    Defines the polling, software coalescing, busy polling and IRQ suspension knobs that trade interrupts against latency.

  2. 02
    DPDK Programmer's Guide manual

    Defines the full bypass model, pinned poll-mode cores with no interrupts, that every kernel path is measured against.

  3. 03
    The eXpress Data Path repository

    Measures an in-kernel programmable path against DPDK and the stack per core, with the full configuration published.

  4. 04
    Kernel vs. User-Level Networking: Don't Throw Out the Stack with the Interrupts paper

    Separates direct and indirect NIC interrupt cost, measures the stack against bypass, and is where IRQ suspension began.

  5. 05
    AF_XDP manual

    Defines the socket and UMEM rings handing XDP frames to user space, and the zero-copy and need-wakeup modes.

Cache and bandwidth partitioning

  1. 01
    Intel Resource Director Technology Architecture Specification manual

    Defines classes of service, cache masks, bandwidth allocation and monitoring IDs, the model resctrl exposes.

  2. 02
    User Interface for Resource Control feature (resctrl) manual

    Defines the filesystem through which Linux exposes Intel, AMD and Arm partitioning, and the schemata format.

  3. 03
    MPAM manual

    Maps Arm's cache portion and bandwidth controls onto resctrl's schemata, and states which platform limits apply.

  4. 04
    CPI2: CPU performance isolation for shared compute clusters paper

    Shows at fleet scale that cycles per instruction alone finds an interfering neighbour and the one to throttle.

  5. 05
    Heracles: Improving Resource Efficiency at Scale paper

    Shows cache ways, cores, bandwidth and power must be partitioned together, or batch work reaches the tail.