On Fri, Jul 24, 2026 at 07:37:22PM +0530, Shrikanth Hegde wrote: > Add documentation for new cpumask called cpu_preferred_mask. This could > help users in understanding what this mask is and the concept behind it. > > Document how to enable it and implementation aspects of it. > > Reported-by: kernel test robot <[email protected]> > Closes: > https://lore.kernel.org/oe-kbuild-all/[email protected]/ > Signed-off-by: Shrikanth Hegde <[email protected]> > --- > Documentation/scheduler/sched-arch.rst | 58 ++++++++++++++++++++++++++
And there's really nothing arch-specific in your series. Can you find a better place for the docs? > 1 file changed, 58 insertions(+) > > diff --git a/Documentation/scheduler/sched-arch.rst > b/Documentation/scheduler/sched-arch.rst > index ed07efea7d02..0a0a8eadbfd6 100644 > --- a/Documentation/scheduler/sched-arch.rst > +++ b/Documentation/scheduler/sched-arch.rst > @@ -62,6 +62,64 @@ Your cpu_idle routines need to obey the following rules: > arch/x86/kernel/process.c has examples of both polling and > sleeping idle functions. > > +Preferred CPUs > +============== > + > +In virtualised environments it is possible to overcommit CPU resources. i.e. > +the sum of virtual CPUs (vCPUs) of all VMs is greater than number of physical > +CPUs (pCPUs). Under such conditions when all or many VMs have high > utilization, > +hypervisor won't be able to satisfy the CPU requirement and has to context > +switch within or across VMs. The hypervisor needs to preempt one vCPU to run > +another. This is called vCPU preemption. This is more expensive compared to > +task context switch within a vCPU. > + > +In such cases it is better that combined vCPU ask from all VMs is reduced > +by not using some of the vCPUs in each VM. vCPUs where workload can be safely > +scheduled which won't increase any contention for pCPU are called as > +"Preferred CPUs". > + > +Main design construct is preferred CPUs are always a subset of active CPUs. > +In most cases preferred CPUs will be same as active CPUs, when there is pCPU > +contention, Preferred CPUs will reduce based on the amount of steal time. > +When the pCPU contention goes away as indicated by steal time, Preferred CPUs > +will become same as active CPUs again. These policy decisions are taken by > +steal_governor driver available at drivers/virt/steal_governor.c > + > +Scheduling decisions such as wakeup, pushing the task etc, need this > +CPU state info. This is maintained in cpu_preferred_mask. > +vCPUs which are not in cpu_preferred_mask should be treated as vCPUs which > +should not be used at this moment provided it doesn't break user affinity. > + > +This is achieved by: > + > +1. Selecting a preferred CPU at wakeup using fallback mechanism. > +2. Push the task away from non-preferred CPU at tick. > +3. Only select preferred CPUs for load balance. > + > +/sys/devices/system/cpu/preferred prints the current cpu_preferred_mask in > +cpulist format. > + > +Notes: > + > +1. This feature is available under CONFIG_PREFERRED_CPU. It is selected > + by steal_governor driver (CONFIG_STEAL_GOVERNOR). On enabling the > + driver, CPU preferred state can change based on steal time. Without that > + driver, preferred CPUs is same as active CPUs. > + > +2. This feature works for FAIR class only. > + > +3. A task pinned, which can't be moved to preferred CPUs will continue > + to run based on its affinity. But no load balancing happens. > + > +4. Decision to change the preferred CPU state is driven by kernel. > + Hence it shouldn't break user affinities. One of the main reasons why > + CPU hotplug or Isolated cpuset partitions was not a solution. > + > +5. This feature works best only when all the VMs enable the feature as > + it is a co-operative scheme. If a specific VM doesn't enable this feature > + it may end up with more CPUs than others, still should lead to better > + performance when seen from system view. > + Users who enable this driver must ensure it is enabled in all VMs. > > Possible arch/ problems > ======================= > -- > 2.47.3

