Oh, I forgot to add the actual results from the LPAR: Model name: POWER10 (architected), altivec supported Model: 2.0 (pvr 0080 0200) Type : Dedicated Mode : Capped Entitled Capacity : 1.00 Memory Mode : Dedicated
Results (median bogo-ops/sec, N=15 per cell): Mode 7.1 7.2-rc4 Delta Enforcing 68,275 70,154 +2.7% Permissive 68,082 67,166 -1.3% Disabled 99,654 100,510 +0.9% On Thu, Sep 10, 2026 at 9:54 PM Jirka Hladky <[email protected]> wrote: > > Hi Mukesh, > > Following up on the ppc64le SELinux/kill() regression - it doesn't reproduce > with proper multi-iteration sampling. stress-ng --kill has ~30% > run-to-run variance > on our LPARs. Part of the variance comes from LPARs being much less > isolated than I'd assumed - we saw contention bleed across LPARs on the same > physical server, even between "Dedicated" partitions. With 15 iterations per > config, 7.1 and 7.2-rc4 are statistically even (within ~3%) across enforcing/ > permissive/disabled, confirmed on a second, physically separate LPAR (7.1 vs > 7.2-rc6). This matches what you said - you couldn't reproduce it either. > > I also built 7.2-rc4 from source with your proposed > rcu_ct_kernel_exit_qs() patch and re-ran the same sweep - it shows no > consistent effect vs. stock 7.2-rc4 (differences within the same noise > band), which is consistent with there being no real regression here for it > to address. > > Sorry for the false alarm, and thank you for the patch and for taking the > time to look into this - I don't think it's needed for this issue, though > it may still be a reasonable efficiency improvement on its own merits (it > does avoid the unconditional rcu_preempt_deferred_qs() call on the common > fast path). I'll update the internal RHEL tracking ticket (RHEL-216740) to > reflect that this doesn't reproduce. > > Thank you! > Jirka > > > On Tue, Sep 8, 2026 at 6:01 PM Jirka Hladky <[email protected]> wrote: > > > > Hi Mukesh, > > > > Re: policy details - I'm running an ELN kernel on RHEL-10.2, with SELinux > > in enforcing mode using the default (targeted) policy, nothing custom > > or hardened on top of it. Let me know if you need anything more specific > > (policy version, booleans, etc.) and I'll pull it. > > > > Re: the rcu_ct_kernel_exit_qs() diff — thanks, that guard against the > > unconditional rcu_preempt_deferred_qs() call lines up well with what I > > was seeing on the RCU-barrier side. I'll apply it on top of 7.2 on > > POWER10 box and re-run the stress-ng --kill numbers > > (enforcing/permissive/disabled) to see how much of the SELinux tax it > > recovers. No need to prepare anything separately. > > > > I'm OOO tomorrow, but I'll get to this by the end of the week and report > > back with results. > > > > Thanks > > Jirka > > > > > > On Tue, Sep 8, 2026 at 3:26 PM Mukesh Kumar Chaurasiya > > <[email protected]> wrote: > > > > > > On Thu, Jul 30, 2026 at 01:24:46PM +0200, Jirka Hladky wrote: > > > > On Thu, Jul 30, 2026 at 11:36 AM Mukesh Kumar Chaurasiya > > > > <[email protected]> wrote: > > > > > > > > > > On Mon, Jul 27, 2026 at 11:41:07AM +0200, Jirka Hladky wrote: > > > > > > Hi Mukesh, Shrikanth, Madhavan, > > > > > > > > > > > > I'm seeing a ~19% throughput regression on ppc64le (POWER10) in > > > > > > syscall-heavy workloads between kernel 7.1 and 7.2-rc4, which I've > > > > > > narrowed down to the GENERIC_ENTRY enablement for powerpc (commit > > > > > > bee25f97ad24 "powerpc: Enable GENERIC_ENTRY feature"). > > > > > > > > > > > > The regression manifests specifically when SELinux is enabled. > > > > > > Regression > > > > > > disappears with selinux=0. With SELinux disabled, 7.2 is actually > > > > > > slightly *faster* than 7.1. > > > > > > > > > > > > Test: stress-ng --kill 1 -t 23 (single-thread kill() syscall loop) > > > > > > Machine: IBM,9080-HEX POWER10, 8 vCPUs (1 core SMT-8), 64 GiB > > > > > > > > > > > > Results (bogo-ops/sec, higher is better): > > > > > > > > > > > > Kernel Enforcing Permissive Disabled SELinux tax > > > > > > ------ --------- ---------- -------- ----------- > > > > > > 7.1 69,107 70,248 93,566 -26.1% > > > > > > 7.2-rc4 56,247 59,075 99,240 -43.3% > > > > > > Delta -18.6% -15.9% +6.1% > > > > > > > > > > > > Key observations: > > > > > > - SELinux permissive ~ enforcing on both kernels (overhead is in the > > > > > > code path, not policy evaluation) > > > > > > - SELinux disabled: 7.2 is 6% faster than 7.1 -- the generic entry > > > > > > path itself is not slower for the base syscall > > > > > > - The SELinux "tax" nearly doubles from 26% to 43% between 7.1->7.2 > > > > > > - No changes to security/selinux/avc.c between v7.1 and v7.2-rc4 > > > > > > > > > > > > The kill() hot path (sys_kill -> check_kill_permission -> > > > > > > security_task_kill > > > > > > -> selinux_task_kill -> avc_has_perm -> avc_lookup) wraps every AVC > > > > > > lookup > > > > > > in rcu_read_lock()/rcu_read_unlock(). On ppc64le with PREEMPT_RCU, > > > > > > each > > > > > > pair requires lwsync/isync barriers. I suspect the generic entry > > > > > > path > > > > > > changes something in how context tracking, tracing, or RCU interacts > > > > > > with the syscall fast path that amplifies the per-call RCU cost. > > > > > > > > > > > > Note: this is on top of a separate ~36% regression from 6.12->7.1 > > > > > > caused > > > > > > by HAVE_PREEMPT_DYNAMIC_KEY enabling CONFIG_PREEMPT_RCU on ppc64le > > > > > > (commit 6ad7751537e8), which I reported separately on the v4 thread. > > > > > > > > > > > > Thank you > > > > > > Jirka > > > > > [...] > > > > > Hey Jirka, > > > > > > > > > > Thanks for reporting this. > > > > > I'll take a look. > > > > > > > > > > Thanks, > > > > > Mukesh > > > > > > > > > > > > > Thank you, Mukesh! I'm happy to run additional tests on the POWER10 > > > > machine > > > > if you have patches to try or specific experiments you'd like me to do. > > > > > > > > Jirka > > > > > > > > > > can you try this with the below diff? > > > > > > diff --git a/include/linux/rcutiny.h b/include/linux/rcutiny.h > > > index e56ded733b1b..dcad641eb2c2 100644 > > > --- a/include/linux/rcutiny.h > > > +++ b/include/linux/rcutiny.h > > > @@ -120,6 +120,7 @@ static inline bool > > > rcu_preempt_need_deferred_qs(struct task_struct *t) > > > return false; > > > } > > > static inline void rcu_preempt_deferred_qs(struct task_struct *t) { } > > > +static inline void rcu_ct_kernel_exit_qs(void) { } > > > void rcu_scheduler_starting(void); > > > static inline void rcu_end_inkernel_boot(void) { } > > > static inline bool rcu_inkernel_boot_has_ended(void) { return true; } > > > diff --git a/include/linux/rcutree.h b/include/linux/rcutree.h > > > index 16a04202888b..d623f2a7d3fc 100644 > > > --- a/include/linux/rcutree.h > > > +++ b/include/linux/rcutree.h > > > @@ -87,6 +87,7 @@ static inline void rcu_irq_exit_check_preempt(void) { } > > > > > > struct task_struct; > > > void rcu_preempt_deferred_qs(struct task_struct *t); > > > +void rcu_ct_kernel_exit_qs(void); > > > > > > void exit_rcu(void); > > > > > > diff --git a/kernel/context_tracking.c b/kernel/context_tracking.c > > > index a743e7ffa6c0..011018214c6d 100644 > > > --- a/kernel/context_tracking.c > > > +++ b/kernel/context_tracking.c > > > @@ -118,7 +118,7 @@ static void noinstr ct_kernel_exit(bool user, int > > > offset) > > > lockdep_assert_irqs_disabled(); > > > trace_rcu_watching(TPS("End"), ct_nesting(), 0, > > > ct_rcu_watching()); > > > WARN_ON_ONCE(IS_ENABLED(CONFIG_RCU_EQS_DEBUG) && !user && > > > !is_idle_task(current)); > > > - rcu_preempt_deferred_qs(current); > > > + rcu_ct_kernel_exit_qs(); > > > > > > // instrumentation for the noinstr ct_kernel_exit_state() > > > instrument_atomic_write(&ct->state, sizeof(ct->state)); > > > diff --git a/kernel/rcu/tree.c b/kernel/rcu/tree.c > > > index 96848fc1f02b..c23478f70c17 100644 > > > --- a/kernel/rcu/tree.c > > > +++ b/kernel/rcu/tree.c > > > @@ -368,6 +368,25 @@ notrace void rcu_momentary_eqs(void) > > > } > > > EXPORT_SYMBOL_GPL(rcu_momentary_eqs); > > > > > > +/** > > > + * rcu_ct_kernel_exit_qs - report deferred QS on syscall/exception exit > > > if needed > > > + * > > > + * Called from ct_kernel_exit() on every return to userspace. Guards the > > > + * rcu_preempt_deferred_qs() call with rcu_preempt_need_deferred_qs() so > > > that > > > + * on the common fast path -- where nothing is deferred -- we avoid the > > > + * cache-line traffic on current->rcu_read_unlock_special that the > > > unconditional > > > + * call causes. This is particularly significant on weakly-ordered > > > architectures > > > + * (e.g. ppc64le) where rcu_read_lock/unlock issue lwsync/isync barriers > > > and > > > + * already touch that cache line in the syscall body. > > > + * > > > + * Follows the same pattern used by rcu_flavor_sched_clock_irq(). > > > + */ > > > +notrace void rcu_ct_kernel_exit_qs(void) > > > +{ > > > + if (rcu_preempt_need_deferred_qs(current)) > > > + rcu_preempt_deferred_qs(current); > > > +} > > > + > > > /** > > > * rcu_is_cpu_rrupt_from_idle - see if 'interrupted' from idle > > > * > > > > > > > > > > > > -- > > -Jirka > > > > -- > -Jirka -- -Jirka
