Track A (cpuset_update_sd_hk_unlock) updates the HK_TYPE_KERNEL_NOISE
and HK_TYPE_MANAGED_IRQ cpumasks but performs no per-CPU reconfiguration.
Tick suppression, RCU callback offloading and managed-IRQ remapping only
take effect when the affected CPUs pass through the CPU hotplug machinery.

Implement dhm_cycle_isolated_cpus() and call it from
cpuset_update_sd_hk_unlock() with cpuset_top_mutex still held: per the
locking convention at the top of this file, cpuset_top_mutex is the
outermost lock, so remove_cpu()/add_cpu() may freely acquire
cpus_write_lock() underneath it.  Keeping the mutex held across the
whole cycle also matches dhm_prev_isolated's existing "protected by
cpuset_top_mutex" comment and prevents a second isolated-partition
update from entering cpuset_update_sd_hk_unlock() while a cycle for
the previous one is still in flight.

On isolation, for each newly-isolated CPU:
  1. remove_cpu()              - offline; dying callbacks migrate IRQs
  2. housekeeping_update_types() - publish this CPU as kernel-noise
                                 isolated now that it is actually offline
  3. tick_nohz_cpu_isolate()  - enable context tracking so the tick
                                 is suppressed for this CPU (B0/B3)
  4. rcu_nocb_cpu_isolate()   - lazy nocb init, spawn kthreads, offload
                                 callbacks (B1)
  5. add_cpu()                 - online; tick and IRQ online callbacks
                                 reconfigure against the now-published
                                 HK masks

On de-isolation, the reverse order is applied, publishing the CPU as
no longer kernel-noise isolated right after its own remove_cpu()
succeeds instead of before any CPU in the batch is touched.

The managed-IRQ remapping requires no explicit call:
irq_migrate_all_off_this_cpu() (dying callback) and
irq_affinity_online_cpu() (online callback) already consult the
updated HK_TYPE_MANAGED_IRQ mask.

dhm_prev_isolated tracks the previous isolation set so that only CPUs
whose state changed are cycled rather than the full isolation set.
lockup_detector_hk_update() (B2) is called once after all CPUs are
cycled to update the watchdog mask.

Some CPUs cannot be taken offline: cpu_is_hotpluggable() rejects a
CPU with hotplug disabled in the architecture (e.g. the x86-64 boot
CPU), and remove_cpu() itself can still fail for a CPU that passed
that filter (e.g. it turns out to be the last CPU in its sched
domain, or it is the current tick_do_timer_cpu holder and
tick_nohz_cpu_hotpluggable() rejects it).  Publishing
housekeeping_update_types() per CPU, strictly after that CPU's own
remove_cpu() has already succeeded, makes both cases handle
themselves: a CPU this function ends up skipping is simply never
added to the published mask, so HK_TYPE_KERNEL_NOISE never claims a
CPU is isolated while it keeps ticking and running RCU callbacks
normally, without needing a separate pre-filter pass or a
republish-on-failure step afterwards.  Also free
newly_isolated/newly_deisolated/cur_isolated on the early return for
a no-op update, which previously leaked the allocations.

Signed-off-by: Qiliang Yuan <[email protected]>
---
 kernel/cgroup/cpuset.c | 155 ++++++++++++++++++++++++++++++++++++++++++++-----
 1 file changed, 140 insertions(+), 15 deletions(-)

diff --git a/kernel/cgroup/cpuset.c b/kernel/cgroup/cpuset.c
index 6213e63cf348d..9fa4bf504c3d9 100644
--- a/kernel/cgroup/cpuset.c
+++ b/kernel/cgroup/cpuset.c
@@ -20,6 +20,8 @@
  */
 #include "cpuset-internal.h"
 
+#include <linux/cpu.h>
+#include <linux/cpuhplock.h>
 #include <linux/init.h>
 #include <linux/interrupt.h>
 #include <linux/kernel.h>
@@ -33,7 +35,9 @@
 #include <linux/sched/task.h>
 #include <linux/security.h>
 #include <linux/oom.h>
+#include <linux/nmi.h>
 #include <linux/sched/isolation.h>
+#include <linux/tick.h>
 #include <linux/wait.h>
 #include <linux/workqueue.h>
 #include <linux/task_work.h>
@@ -165,6 +169,14 @@ static cpumask_var_t       isolated_hk_cpus;       /* T */
 static DEFINE_SPINLOCK(dhm_cycling_lock);
 static cpumask_var_t   dhm_cycling_cpus;
 
+/*
+ * Snapshot of the isolated CPUs from the previous housekeeping update.
+ * Used to compute the delta (newly isolated / newly de-isolated) so that
+ * only the changed CPUs are cycled rather than the full isolation set.
+ * Protected by cpuset_top_mutex.
+ */
+static cpumask_var_t   dhm_prev_isolated;
+
 /*
  * A flag to force sched domain rebuild at the end of an operation.
  * It can be set in
@@ -1423,6 +1435,108 @@ static bool prstate_housekeeping_conflict(int prstate, 
struct cpumask *new_cpus)
        return false;
 }
 
+/*
+ * dhm_cycle_isolated_cpus - Apply kernel-noise isolation via hotplug cycling
+ *
+ * For each CPU newly entering isolation: cycle it offline, configure tick
+ * suppression and RCU callback offloading while it is offline, then bring
+ * it back online.  The managed-IRQ state is handled automatically by the
+ * existing irq_migrate_all_off_this_cpu() dying callback and the
+ * irq_affinity_online_cpu() online callback which both consult the
+ * already-updated HK_TYPE_MANAGED_IRQ mask.
+ *
+ * For each CPU leaving isolation: cycle it offline, de-offload RCU and
+ * restore the tick, then bring it back online.
+ *
+ * Called with cpuset_top_mutex held and no other cpuset or hotplug locks
+ * held: cpuset_top_mutex is the outermost lock, so remove_cpu()/add_cpu()
+ * may freely take cpus_write_lock() underneath it.
+ */
+static void dhm_cycle_isolated_cpus(const struct cpumask *new_isolated)
+{
+       static const unsigned long noise_types =
+               BIT(HK_TYPE_KERNEL_NOISE) | BIT(HK_TYPE_MANAGED_IRQ);
+       cpumask_var_t newly_isolated, newly_deisolated, cur_isolated;
+       int cpu;
+
+       if (!alloc_cpumask_var(&newly_isolated, GFP_KERNEL) ||
+           !alloc_cpumask_var(&newly_deisolated, GFP_KERNEL) ||
+           !alloc_cpumask_var(&cur_isolated, GFP_KERNEL)) {
+               free_cpumask_var(newly_isolated);
+               free_cpumask_var(newly_deisolated);
+               return;
+       }
+
+       cpumask_andnot(newly_isolated, new_isolated, dhm_prev_isolated);
+       cpumask_andnot(newly_deisolated, dhm_prev_isolated, new_isolated);
+       /* cur_isolated tracks the mask actually published so far. */
+       cpumask_copy(cur_isolated, dhm_prev_isolated);
+       cpumask_copy(dhm_prev_isolated, new_isolated);
+
+       if (cpumask_empty(newly_isolated) && cpumask_empty(newly_deisolated))
+               goto out_free;
+
+       /* Mark cycling CPUs so cpuset_hotplug_update_tasks skips invalidation 
*/
+       spin_lock(&dhm_cycling_lock);
+       cpumask_or(dhm_cycling_cpus, newly_isolated, newly_deisolated);
+       spin_unlock(&dhm_cycling_lock);
+
+       /*
+        * Publish each CPU's kernel-noise/managed-IRQ housekeeping state
+        * strictly between its own remove_cpu() succeeding and add_cpu()
+        * bringing it back, never as a batch before the whole cycle.
+        * housekeeping_update_types() then always describes exactly which
+        * CPUs are actually ticking and running RCU callbacks normally at
+        * that instant, including the current tick_do_timer_cpu holder,
+        * which tick_nohz_cpu_hotpluggable() must see as still a
+        * housekeeping CPU for as long as it is still online and un-cycled.
+        */
+       for_each_cpu(cpu, newly_isolated) {
+               if (!cpu_is_hotpluggable(cpu)) {
+                       pr_warn_once("cpuset: CPU%d cannot be isolated (hotplug 
disabled)\n",
+                                    cpu);
+                       cpumask_clear_cpu(cpu, dhm_prev_isolated);
+                       continue;
+               }
+               if (remove_cpu(cpu)) {
+                       pr_warn_once("cpuset: failed to offline CPU%d for 
isolation\n",
+                                    cpu);
+                       cpumask_clear_cpu(cpu, dhm_prev_isolated);
+                       continue;
+               }
+               cpumask_set_cpu(cpu, cur_isolated);
+               WARN_ON_ONCE(housekeeping_update_types(noise_types, 
cur_isolated));
+               WARN_ON_ONCE(tick_nohz_cpu_isolate(cpu));
+               WARN_ON_ONCE(rcu_nocb_cpu_isolate(cpu));
+               WARN_ON_ONCE(add_cpu(cpu));
+       }
+
+       for_each_cpu(cpu, newly_deisolated) {
+               if (remove_cpu(cpu)) {
+                       pr_warn_once("cpuset: failed to offline CPU%d for 
de-isolation\n",
+                                    cpu);
+                       cpumask_set_cpu(cpu, dhm_prev_isolated);
+                       continue;
+               }
+               cpumask_clear_cpu(cpu, cur_isolated);
+               WARN_ON_ONCE(housekeeping_update_types(noise_types, 
cur_isolated));
+               WARN_ON_ONCE(rcu_nocb_cpu_deoffload(cpu));
+               tick_nohz_cpu_deisolate(cpu);
+               WARN_ON_ONCE(add_cpu(cpu));
+       }
+
+       spin_lock(&dhm_cycling_lock);
+       cpumask_clear(dhm_cycling_cpus);
+       spin_unlock(&dhm_cycling_lock);
+
+       lockup_detector_hk_update();
+
+out_free:
+       free_cpumask_var(newly_isolated);
+       free_cpumask_var(newly_deisolated);
+       free_cpumask_var(cur_isolated);
+}
+
 /*
  * cpuset_update_sd_hk_unlock - Rebuild sched domains, update HK & unlock
  *
@@ -1439,8 +1553,6 @@ static void cpuset_update_sd_hk_unlock(void)
                rebuild_sched_domains_locked();
 
        if (update_housekeeping) {
-               static const unsigned long noise_types =
-                       BIT(HK_TYPE_KERNEL_NOISE) | BIT(HK_TYPE_MANAGED_IRQ);
                int ret;
 
                update_housekeeping = false;
@@ -1450,9 +1562,16 @@ static void cpuset_update_sd_hk_unlock(void)
                cpus_read_unlock();
 
                /*
-                * housekeeping_update() is now called without holding
-                * cpus_read_lock and cpuset_mutex. Only cpuset_top_mutex
-                * is still being held for mutual exclusion.
+                * housekeeping_update() and housekeeping_update_types() are
+                * now called without holding cpus_read_lock and cpuset_mutex.
+                * cpuset_top_mutex stays held all the way through
+                * dhm_cycle_isolated_cpus() below: per the locking convention
+                * at the top of this file it is the outermost lock, so it may
+                * legally nest around the cpus_write_lock() that remove_cpu()/
+                * add_cpu() take. Holding it here is also what makes
+                * dhm_prev_isolated's "protected by cpuset_top_mutex" comment
+                * true, and keeps a second isolated-partition update from
+                * entering this function while a cycle is still in flight.
                 */
 
                /*
@@ -1464,18 +1583,23 @@ static void cpuset_update_sd_hk_unlock(void)
                WARN_ON_ONCE(ret);
 
                /*
-                * Only touch the kernel-noise housekeeping masks
-                * (HK_TYPE_KERNEL_NOISE and HK_TYPE_MANAGED_IRQ) once the
-                * sched domain update above actually succeeded: HK_TYPE_
-                * KERNEL_NOISE must stay a subset of HK_TYPE_DOMAIN, so a CPU
-                * can never end up tick-suppressed while still scheduled as
-                * part of the normal (non-isolated) sched domain.  The tick,
-                * RCU and managed-interrupt state is reconfigured as the
-                * affected CPUs are cycled through the CPU hotplug machinery.
+                * Only cycle CPUs through hotplug, applying the kernel-noise
+                * types along the way, when the sched domain update above
+                * actually succeeded.  HK_TYPE_KERNEL_NOISE must stay a
+                * subset of HK_TYPE_DOMAIN, so a CPU can never end up
+                * tick-suppressed while still scheduled as part of the
+                * normal (non-isolated) sched domain.  The tick, RCU and
+                * managed-interrupt state is reconfigured as the affected
+                * CPUs are cycled through the CPU hotplug machinery;
+                * dhm_cycle_isolated_cpus() publishes HK_TYPE_KERNEL_NOISE
+                * and HK_TYPE_MANAGED_IRQ itself, one CPU at a time, strictly
+                * between that CPU's own remove_cpu() and add_cpu(), so a
+                * CPU that cpu_is_hotpluggable() rejects (e.g. the current
+                * tick_do_timer_cpu) is simply never published as isolated.
                 */
                if (!ret)
-                       WARN_ON_ONCE(housekeeping_update_types(noise_types,
-                                                              
isolated_hk_cpus));
+                       dhm_cycle_isolated_cpus(isolated_hk_cpus);
+
                mutex_unlock(&cpuset_top_mutex);
        } else {
                cpuset_full_unlock();
@@ -3915,6 +4039,7 @@ int __init cpuset_init(void)
        BUG_ON(!zalloc_cpumask_var(&isolated_cpus, GFP_KERNEL));
        BUG_ON(!zalloc_cpumask_var(&isolated_hk_cpus, GFP_KERNEL));
        BUG_ON(!zalloc_cpumask_var(&dhm_cycling_cpus, GFP_KERNEL));
+       BUG_ON(!zalloc_cpumask_var(&dhm_prev_isolated, GFP_KERNEL));
 
        cpumask_setall(top_cpuset.cpus_allowed);
        nodes_setall(top_cpuset.mems_allowed);

-- 
2.43.0


Reply via email to