> On Oct 6, 2026, at 7:55 AM, Cunlong Li <[email protected]> wrote:
> 
> Hi Joel, thanks for the review!
> 
>> On Mon, Oct 05, 2026 at 12:56:30PM +0000, Joel Fernandes wrote:
>> 
>> 
>>>> On Oct 3, 2026, at 9:14 PM, Cunlong Li <[email protected]> wrote:
>>> 
>>> The kernel.panic_on_rcu_stall and kernel.max_rcu_stall_to_panic sysctls
>>> count RCU CPU stalls since system boot, and invoke panic() once
>>> max_rcu_stall_to_panic stalls have elapsed.  This count is never reset,
>>> so stalls caused by transient incidents keep consuming the budget of a
>>> long-running system, and a later unrelated stall can
>> 
>> Please provide examples of specific transient events you ran into, and 
>> recovered from?
>> 
>>> immediately
>>> trigger panic() instead of allowing the intended fresh window of
>>> stalls.
>>> 
>>> This commit therefore introduces the kernel.rcu_stall_panic_count
>>> sysctl.  Reading this file reports the number of stalls counted so far,
>>> and writing 0 to it resets the count.  This allows the cumulative stall
>>> history to be cleared after an incident has been resolved, and also
>>> allows userspace to clear the count periodically, so that panic() is
>>> only triggered by a burst of stalls occurring within a short period.
>> 
>> Could you provide details of such an incident that got resolved without 
>> requiring a reboot?
> 
> The motivation comes from our customer systems running codex on Ubuntu
> 24.04, with a memory limit of ~8G and zram enabled for swap.  When the
> agent's memory footprint grows, the box spends long stretches in direct
> reclaim (zram makes reclaim CPU-heavy), and we get a continuous stream
> of RCU stall warnings.  In the incidents we have seen so far, the
> warnings kept coming and the systems stayed hung, so we would like to
> enable panic_on_rcu_stall so that the box reboots and the service
> recovers automatically.  However, we do not actually know whether some
> of those stalls were transient, i.e. whether the systems would have
> recovered on their own once the memory pressure subsided.  If they
> were, a panic would kill an otherwise recoverable system, which is
> what we would like to avoid.
> 
> We have not yet been able to reproduce this in the lab.  But transient,
> self-recovering stalls are not a hypothetical: the commit that
> introduced max_rcu_stall_to_panic, dfe564045c65 ("rcu: Panic after
> fixed number of stalls"), states the premise outright:
> 
>    Some stalls are transient, so that system fully recovers.  This
>    commit therefore allows users to configure the number of stalls
>    that must happen in order to trigger kernel panic.
> 
> As described in the commit message, since the counter is never reset,
> stalls caused by transient incidents keep consuming the budget of
> max_rcu_stall_to_panic, and a later unrelated stall can immediately
> trigger panic() instead of allowing the intended fresh window of
> stalls.  This is what motivated the patch.

Ok, add all this to commit message. Use case justification is a must as we 
weigh in additional complexity.

> 
>> 
>> 
>>> 
>>> Signed-off-by: Cunlong Li <[email protected]>
>> 
>> If the commit is AI assisted, please add an assisted tag.
> 
> Yes, it is AI-assisted, and I will state that in the v2 commit
> message.

OK.

Thanks,

Joel 

> 
> Thanks,
> Cunlong
> 
>> 
>> Thanks,
>> 
>> Joel
>> 
>>> ---
>>> Documentation/admin-guide/sysctl/kernel.rst | 15 ++++++++++++++-
>>> kernel/rcu/tree_stall.h                     | 28 
>>> ++++++++++++++++++++++++++--
>>> 2 files changed, 40 insertions(+), 3 deletions(-)
>>> 
>>> diff --git a/Documentation/admin-guide/sysctl/kernel.rst 
>>> b/Documentation/admin-guide/sysctl/kernel.rst
>>> index ffea61d448eb..64fe2985e358 100644
>>> --- a/Documentation/admin-guide/sysctl/kernel.rst
>>> +++ b/Documentation/admin-guide/sysctl/kernel.rst
>>> @@ -959,7 +959,20 @@ max_rcu_stall_to_panic
>>> When ``panic_on_rcu_stall`` is set to 1, this value determines the
>>> number of times that RCU can stall before panic() is called.
>>> 
>>> -When ``panic_on_rcu_stall`` is set to 0, this value is has no effect.
>>> +When ``panic_on_rcu_stall`` is set to 0, this value has no effect.
>>> +
>>> +rcu_stall_panic_count
>>> +=====================
>>> +
>>> +Indicates the number of RCU CPU stalls that have been counted since
>>> +system boot or since the counter was reset. When ``panic_on_rcu_stall``
>>> +is set to 1, this count is compared against ``max_rcu_stall_to_panic``
>>> +to decide whether panic() should be called.
>>> +
>>> +Writing 0 to this file resets the counter to zero, which restarts the
>>> +``max_rcu_stall_to_panic`` window of stalls. This allows system
>>> +administrators to clear the cumulative stall count after an incident
>>> +has been resolved, without requiring a system restart.
>>> 
>>> perf_cpu_time_max_percent
>>> =========================
>>> diff --git a/kernel/rcu/tree_stall.h b/kernel/rcu/tree_stall.h
>>> index 091e7850ab6e..a80f03e1c7ea 100644
>>> --- a/kernel/rcu/tree_stall.h
>>> +++ b/kernel/rcu/tree_stall.h
>>> @@ -19,6 +19,20 @@
>>> /* panic() on RCU Stall sysctl. */
>>> static int sysctl_panic_on_rcu_stall __read_mostly;
>>> static int sysctl_max_rcu_stall_to_panic __read_mostly;
>>> +static unsigned long sysctl_rcu_stall_panic_count;
>>> +
>>> +/* Reset the RCU stall panic count when written to. */
>>> +static int proc_do_rcu_stall_panic_count(const struct ctl_table *table, 
>>> int write,
>>> +                     void *buffer, size_t *lenp, loff_t *ppos)
>>> +{
>>> +    if (!write)
>>> +        return proc_doulongvec_minmax(table, write, buffer, lenp, ppos);
>>> +
>>> +    WRITE_ONCE(sysctl_rcu_stall_panic_count, 0);
>>> +    *ppos += *lenp;
>>> +
>>> +    return 0;
>>> +}
>>> 
>>> static const struct ctl_table rcu_stall_sysctl_table[] = {
>>>   {
>>> @@ -39,6 +53,13 @@ static const struct ctl_table rcu_stall_sysctl_table[] = 
>>> {
>>>       .extra1        = SYSCTL_ONE,
>>>       .extra2        = SYSCTL_INT_MAX,
>>>   },
>>> +    {
>>> +        .procname    = "rcu_stall_panic_count",
>>> +        .data        = &sysctl_rcu_stall_panic_count,
>>> +        .maxlen        = sizeof(sysctl_rcu_stall_panic_count),
>>> +        .mode        = 0644,
>>> +        .proc_handler    = proc_do_rcu_stall_panic_count,
>>> +    },
>>> };
>>> 
>>> static int __init init_rcu_stall_sysctl(void)
>>> @@ -161,7 +182,7 @@ early_initcall(check_cpu_stall_init);
>>> /* If so specified via sysctl, panic, yielding cleaner stall-warning 
>>> output. */
>>> static void panic_on_rcu_stall(const struct cpumask *stalled_mask)
>>> {
>>> -    static int cpu_stall;
>>> +    unsigned long count;
>>> 
>>>   /*
>>>    * Attempt to kick out the BPF scheduler if it's installed and defer
>>> @@ -170,7 +191,10 @@ static void panic_on_rcu_stall(const struct cpumask 
>>> *stalled_mask)
>>>   if (scx_rcu_cpu_stall(stalled_mask))
>>>       return;
>>> 
>>> -    if (++cpu_stall < sysctl_max_rcu_stall_to_panic)
>>> +    /* A lost RMW update only delays the panic by one stall. */
>>> +    count = READ_ONCE(sysctl_rcu_stall_panic_count) + 1;
>>> +    WRITE_ONCE(sysctl_rcu_stall_panic_count, count);
>>> +    if (count < (unsigned long)READ_ONCE(sysctl_max_rcu_stall_to_panic))
>>>       return;
>>> 
>>>   if (sysctl_panic_on_rcu_stall)
>>> 
>>> ---
>>> base-commit: ce1e0223d8ad4211275c82a17ed6d43ab81e13d9
>>> change-id: 20261003-rcu-375496d7704e
>>> 
>>> Best regards,
>>> --
>>> Cunlong Li <[email protected]>
>>> 

Reply via email to