Linux scheduler gains steal-time preferred CPUs for VM guests
A steal governor lets paravirtualized guests shrink their usable CPU set under host contention, cutting preemption costs beyond lost cycles.
The Linux kernel scheduler has gained a way for paravirtualized guests to back off from contended physical CPUs, aimed at workloads where vCPU preemption costs more than the CPU time stolen alone.
Shrikanth Hegde of IBM has landed a series that pairs a core preferred-CPU concept with a steal_governor driver. On dense shared-processor hosts, guests already account for steal time: cycles the hypervisor took while a vCPU was runnable. For databases and similar mixes of OLTP and OLAP work, that preemption can also stall lock holders, stretch critical sections, and thrash TLBs and caches. High steal now signals the guest to shrink the set of CPUs it prefers for running tasks; low steal grows the set again.
Once a CPU is non-preferred, fair-class load balancing stops pulling work onto it and the running task is pushed off at the next tick. User affinities still win, so a task pinned only to non-preferred CPUs is not stranded. The preferred mask is exposed in sysfs so user-space tools such as irqbalance can follow it. The governor polls on a tunable interval (default one second) against low and high steal thresholds (defaults about 2% and 5%). Building it as a module is recommended so parameters can be retuned by reload and so administrators enable it deliberately on cooperating VMs.
Hegde reports solid results on PowerPC shared-processor LPARs and earlier gains under KVM on s390 and x86. Pure CPU-bound work may regress slightly. The design is generic for any architecture with paravirtualization and steal-time accounting; architecture-specific hooks and a test framework are planned later. Peter Zijlstra merged the work after extensive review, notably from Yury Norov.