Hi Yury,

On 7/25/26 3:10 AM, Yury Norov wrote:
On Fri, Jul 24, 2026 at 07:37:26PM +0530, Shrikanth Hegde wrote:
When cpu is marked as non preferred, any load pulled towards it is
pointless since in the next tick task will be pushed out again.
So, Consider only preferred CPUs for load balance.

This makes it not fight against the push task mechanism which happens
at tick. Also, this stops active balance to happen on non-preferred CPU
pulling the load.

This means there is no load balancing if the task is pinned only to
non-preferred CPUs. They will continue to run where they were previously
running before the CPUs was marked as non-preferred.

Bailout early for NEWIDLE and IDLE balance as load balancing is done
only on preferred CPUs.

Signed-off-by: Shrikanth Hegde <[email protected]>
---
  kernel/sched/fair.c | 11 +++++------
  1 file changed, 5 insertions(+), 6 deletions(-)

diff --git a/kernel/sched/fair.c b/kernel/sched/fair.c
index df8c9c2c7918..12f5b7de28d2 100644
--- a/kernel/sched/fair.c
+++ b/kernel/sched/fair.c
@@ -13399,7 +13399,7 @@ static int sched_balance_rq(int this_cpu, struct rq 
*this_rq,
        };
        bool need_unlock = false;
- cpumask_and(cpus, sched_domain_span(sd), cpu_active_mask);
+       cpumask_and(cpus, sched_domain_span(sd), cpu_preferred_mask);
schedstat_inc(sd->lb_count[idle]); @@ -14337,7 +14337,8 @@ static void _nohz_idle_balance(struct rq *this_rq, unsigned int flags)
                        update_rq_clock(rq);
                        rq_unlock_irqrestore(rq, &rf);
- if (flags & NOHZ_BALANCE_KICK)
+                       if (flags & NOHZ_BALANCE_KICK &&
+                           cpu_preferred(balance_cpu))
                                sched_balance_domains(rq, CPU_IDLE);

Here you skip the sched_balance_domains() for idle non-preferred CPUs,
which means you don't re-calculate the rq->next_balance.

In the following code, we update the global nohz.next_balance
depending on the new rq->next_balance, which doesn't happen if
the sched_balance_domains() is not invoked.

So, you propagate an already-expired per-CPU deadline back into the
global NOHZ deadline. It may lead to unneeded asynchronous IPI in the
nohz_balancer_kick() -> kick_ilb() path.

I think, you need another helper, something like:

         void sched_balance_domains(rq, idle)
         {
                 __sched_balance_domains(rq, idle);
                 advance_rq_next_balance(rq, idle);
         }

And then in the code above:

         if (flags & NOHZ_BALANCE_KICK) {
                 if (cpu_preferred(balance_cpu))
                         __sched_balance_domains(rq, CPU_IDLE);

                 advance_rq_next_balance(rq, idle);
         }


I looked more into it. Removing this check will solve it naturally.
This happens because:
- In sched_balance_rq - env->cpus is masked for preferred CPUs.
- In should_we_balance - fails since dst_cpu isn;t part of env->cpus.
- In sched_balance_domains - updates the rq->next_balance based on intervals.
So the above check is not needed. I will remove it.


While there, it makes me think it is probably a good idea to put check in
kick_ilb path in find_new_ilb. So that it find a idle CPU which is preferred. 
If all
idle CPUs are non-preferred then chose first idle cpu which is non-preferred 
such that
time elapsing still happens correctly. performance numbers are similar.

Something like below?
---

diff --git a/kernel/sched/fair.c b/kernel/sched/fair.c
index 12f5b7de28d2..ccb18834fbfa 100644
--- a/kernel/sched/fair.c
+++ b/kernel/sched/fair.c
@@ -13964,6 +13964,7 @@ static inline int find_new_ilb(void)
 {
        int this_cpu = smp_processor_id();
        const struct cpumask *hk_mask;
+       int first_idle_cpu = -1;
        int ilb_cpu;
hk_mask = housekeeping_cpumask(HK_TYPE_KERNEL_NOISE);
@@ -13972,11 +13973,16 @@ static inline int find_new_ilb(void)
                if (ilb_cpu == this_cpu)
                        continue;
- if (idle_cpu(ilb_cpu))
-                       return ilb_cpu;
+               if (idle_cpu(ilb_cpu)) {
+                       if (cpu_preferred(ilb_cpu))
+                               return ilb_cpu;
+
+                       if (first_idle_cpu == -1)
+                               first_idle_cpu = ilb_cpu;
+               }
        }
- return -1;
+       return first_idle_cpu;
 }
/*


Reply via email to