Re: [mainline][BUG] Observed Workqueue lockups on offline CPUs.

2026-04-30 Thread Samir M



On 30/04/26 8:30 pm, Paul E. McKenney wrote:

On Thu, Apr 30, 2026 at 11:49:02AM +0530, Samir M wrote:

On 29/04/26 11:21 pm, Shrikanth Hegde wrote:

Hi Samir.

On 4/29/26 3:46 PM, Samir M wrote:

Hi Boqun,

Thank you for pointing me to the existing patches. I have tested
both Paul's patch [1] and TJ's workqueue patch [2] on my PowerPC
system (80 CPUs), and can confirm that the workqueue lockup issue is
not observed.


Can you try only paul's patch and confirm if the issue is fixed?


Hi Shrikanth,

I have verified Paul’s “alone” patch, and it resolves the issue no workqueue
lockups are observed.

Thank you, Samir!  May I add your Tested-by?

Thanx, Paul


Hi Paul,

Yes, you may add:
Tested-by: Samir 

Thanks,
Samir

Regards,
Samir.



Test Environment:
- System: PowerPC with 80 CPUs ( e.g. PowerPC LPARs with 80 online
and 384 possible CPUs)
- Kernel version: Latest upstream (7.1-rc1)

Regression Testing Results:
All tests completed successfully with no issues observed:
- Hackbench
- Kernel selftests
- LTP scheduler tests

The workqueue lockup that was previously occurring is no longer
present with the patches applied.

References:
[1]: https://lore.kernel.org/rcu/ed1fa6cd-7343-4ca3-8b9d-
d699ca496f83@paulmck-laptop/
[2]: https://lore.kernel.org/rcu/[email protected]/

Best regards,
Samir




Re: [mainline][BUG] Observed Workqueue lockups on offline CPUs.

2026-04-30 Thread Paul E. McKenney
On Thu, Apr 30, 2026 at 11:49:02AM +0530, Samir M wrote:
> 
> On 29/04/26 11:21 pm, Shrikanth Hegde wrote:
> > Hi Samir.
> > 
> > On 4/29/26 3:46 PM, Samir M wrote:
> > > 
> > > Hi Boqun,
> > > 
> > > Thank you for pointing me to the existing patches. I have tested
> > > both Paul's patch [1] and TJ's workqueue patch [2] on my PowerPC
> > > system (80 CPUs), and can confirm that the workqueue lockup issue is
> > > not observed.
> > > 
> > 
> > Can you try only paul's patch and confirm if the issue is fixed?
> > 
> 
> Hi Shrikanth,
> 
> I have verified Paul’s “alone” patch, and it resolves the issue no workqueue
> lockups are observed.

Thank you, Samir!  May I add your Tested-by?

Thanx, Paul

> Regards,
> Samir.
> 
> 
> > > Test Environment:
> > > - System: PowerPC with 80 CPUs ( e.g. PowerPC LPARs with 80 online
> > > and 384 possible CPUs)
> > > - Kernel version: Latest upstream (7.1-rc1)
> > > 
> > > Regression Testing Results:
> > > All tests completed successfully with no issues observed:
> > > - Hackbench
> > > - Kernel selftests
> > > - LTP scheduler tests
> > > 
> > > The workqueue lockup that was previously occurring is no longer
> > > present with the patches applied.
> > > 
> > > References:
> > > [1]: https://lore.kernel.org/rcu/ed1fa6cd-7343-4ca3-8b9d-
> > > d699ca496f83@paulmck-laptop/
> > > [2]: https://lore.kernel.org/rcu/[email protected]/
> > > 
> > > Best regards,
> > > Samir
> > 



Re: [mainline][BUG] Observed Workqueue lockups on offline CPUs.

2026-04-29 Thread Samir M



On 29/04/26 11:21 pm, Shrikanth Hegde wrote:

Hi Samir.

On 4/29/26 3:46 PM, Samir M wrote:


Hi Boqun,

Thank you for pointing me to the existing patches. I have tested both 
Paul's patch [1] and TJ's workqueue patch [2] on my PowerPC system 
(80 CPUs), and can confirm that the workqueue lockup issue is not 
observed.




Can you try only paul's patch and confirm if the issue is fixed?



Hi Shrikanth,

I have verified Paul’s “alone” patch, and it resolves the issue no 
workqueue lockups are observed.


Regards,
Samir.



Test Environment:
- System: PowerPC with 80 CPUs ( e.g. PowerPC LPARs with 80 online 
and 384 possible CPUs)

- Kernel version: Latest upstream (7.1-rc1)

Regression Testing Results:
All tests completed successfully with no issues observed:
- Hackbench
- Kernel selftests
- LTP scheduler tests

The workqueue lockup that was previously occurring is no longer 
present with the patches applied.


References:
[1]: https://lore.kernel.org/rcu/ed1fa6cd-7343-4ca3-8b9d- 
d699ca496f83@paulmck-laptop/

[2]: https://lore.kernel.org/rcu/[email protected]/

Best regards,
Samir






Re: [mainline][BUG] Observed Workqueue lockups on offline CPUs.

2026-04-29 Thread Shrikanth Hegde

Hi Samir.

On 4/29/26 3:46 PM, Samir M wrote:


Hi Boqun,

Thank you for pointing me to the existing patches. I have tested both 
Paul's patch [1] and TJ's workqueue patch [2] on my PowerPC system (80 
CPUs), and can confirm that the workqueue lockup issue is not observed.




Can you try only paul's patch and confirm if the issue is fixed?


Test Environment:
- System: PowerPC with 80 CPUs ( e.g. PowerPC LPARs with 80 online and 
384 possible CPUs)

- Kernel version: Latest upstream (7.1-rc1)

Regression Testing Results:
All tests completed successfully with no issues observed:
- Hackbench
- Kernel selftests
- LTP scheduler tests

The workqueue lockup that was previously occurring is no longer present 
with the patches applied.


References:
[1]: https://lore.kernel.org/rcu/ed1fa6cd-7343-4ca3-8b9d- 
d699ca496f83@paulmck-laptop/

[2]: https://lore.kernel.org/rcu/[email protected]/

Best regards,
Samir





Re: [mainline][BUG] Observed Workqueue lockups on offline CPUs.

2026-04-29 Thread Samir M



On 27/04/26 9:13 pm, Boqun Feng wrote:

On Mon, Apr 27, 2026 at 05:00:10PM +0530, Samir M wrote:
Hi Samir,


On 27/04/26 3:32 pm, Samir M wrote:

Hi Paul,

I've been testing the latest upstream kernel on a PowerPC system and
encountered workqueue lockup issues that I've bisected to commit
61bbcfb50514 ("srcu: Push srcu_node allocation to GP when
non-preemptible").
After booting, I'm seeing workqueue lockup warnings for CPUs 81-96,
which are offline on my system. The workqueues remain stuck for over 237
seconds:

[  243.309302][    C0] BUG: workqueue lockup - pool cpus=81 node=0
flags=0x4 nice=0 stuck for 237s!
[  243.309311][    C0] BUG: workqueue lockup - pool cpus=82 node=0
flags=0x4 nice=0 stuck for 237s!
[  243.309318][    C0] BUG: workqueue lockup - pool cpus=83 node=0
flags=0x4 nice=0 stuck for 237s!
[  243.309326][    C0] BUG: workqueue lockup - pool cpus=84 node=0
flags=0x4 nice=0 stuck for 237s!
[  243.309333][    C0] BUG: workqueue lockup - pool cpus=85 node=0
flags=0x4 nice=0 stuck for 237s!
[  243.309341][    C0] BUG: workqueue lockup - pool cpus=86 node=0
flags=0x4 nice=0 stuck for 237s!
[  243.309348][    C0] BUG: workqueue lockup - pool cpus=87 node=0
flags=0x4 nice=0 stuck for 237s!
[  243.309355][    C0] BUG: workqueue lockup - pool cpus=88 node=0
flags=0x4 nice=0 stuck for 237s!
[  243.309363][    C0] BUG: workqueue lockup - pool cpus=89 node=0
flags=0x4 nice=0 stuck for 237s!
[  243.309370][    C0] BUG: workqueue lockup - pool cpus=90 node=0
flags=0x4 nice=0 stuck for 237s!
[  243.309377][    C0] BUG: workqueue lockup - pool cpus=91 node=0
flags=0x4 nice=0 stuck for 237s!
[  243.309384][    C0] BUG: workqueue lockup - pool cpus=92 node=0
flags=0x4 nice=0 stuck for 237s!
[  243.309392][    C0] BUG: workqueue lockup - pool cpus=93 node=0
flags=0x4 nice=0 stuck for 237s!
[  243.309399][    C0] BUG: workqueue lockup - pool cpus=94 node=0
flags=0x4 nice=0 stuck for 237s!
[  243.309406][    C0] BUG: workqueue lockup - pool cpus=95 node=0
flags=0x4 nice=0 stuck for 237s!
[  243.309413][    C0] BUG: workqueue lockup - pool cpus=96 node=0
flags=0x4 nice=0 stuck for 237s!

Git bisect identified this as the first bad commit:

commit 61bbcfb50514a8a94e035a7349697a3790ab4783
Author: Paul E. McKenney 
Date:   Fri Mar 20 20:29:20 2026 -0700

     srcu: Push srcu_node allocation to GP when non-preemptible

     When the srcutree.convert_to_big and srcutree.big_cpu_lim kernel boot
     parameters specify initialization-time allocation of the srcu_node
     tree for statically allocated srcu_struct structures (for example, in
     DEFINE_SRCU() at build time instead of init_srcu_struct() at
runtime),
     init_srcu_struct_nodes() will attempt to dynamically allocate this
tree
     at the first run-time update-side use of this srcu_struct structure,
     but while holding a raw spinlock. Because the memory allocator can
     acquire non-raw spinlocks, this can result in lockdep splats.

     This commit therefore uses the same SRCU_SIZE_ALLOC trick that is
used
     when the first run-time update-side use of this srcu_struct structure
     happens before srcu_init() is called. The actual allocation then
takes
     place from workqueue context at the ends of upcoming SRCU grace
periods.

     [boqun: Adjust the sha1 of the Fixes tag]

     Fixes: 175b45ed343a ("srcu: Use raw spinlocks so call_srcu() can be
used under preempt_disable()")
     Signed-off-by: Paul E. McKenney 
     Signed-off-by: Boqun Feng 

  kernel/rcu/srcutree.c | 7 +--
  1 file changed, 5 insertions(+), 2 deletions(-)

Reverting this commit resolves the issue.

The problem appears to be that the workqueue is attempting to execute on
offline CPUs. The commit moves SRCU node allocation to workqueue context
to avoid lockdep issues with memory allocation under raw spinlocks,
which makes sense. However, it seems the workqueue scheduling doesn't
properly account for CPU online/offline state in this code path.

My test environment:
- Architecture: PowerPC
- Kernel version: Latest upstream (7.1-rc1)
- CPUs 81-96 are offline at boot time

I suspect the issue might be related to:
1. Workqueue not checking CPU online status before scheduling SRCU
allocation work
2. Missing CPU hotplug awareness in the new workqueue-based allocation
path
3. Possible race condition with CPU hotplug events

Would it make sense to use queue_work_on() with explicit online CPU
selection, or add CPU hotplug handlers for this workqueue? I'm not
deeply familiar with the workqueue internals, so I might be missing
something.
Please let me know if you need any additional details or if you'd like
me to test any patches.

If you happen to fix the above issue, then please add below tag.
Reported-by: Samir M 


Thanks,
Samir

Hi Paul,


I worked on fixing the issue and introduced the changes below. With these
updates, I no longer observe any workqueue lockup messages for offline CPUs.
Could you please review the changes and share your feedback?

The commit 61bbcfb50514 ("src

Re: [mainline][BUG] Observed Workqueue lockups on offline CPUs.

2026-04-27 Thread Boqun Feng
On Mon, Apr 27, 2026 at 05:00:10PM +0530, Samir M wrote:
> 

Hi Samir,

> On 27/04/26 3:32 pm, Samir M wrote:
> > Hi Paul,
> > 
> > I've been testing the latest upstream kernel on a PowerPC system and
> > encountered workqueue lockup issues that I've bisected to commit
> > 61bbcfb50514 ("srcu: Push srcu_node allocation to GP when
> > non-preemptible").
> > After booting, I'm seeing workqueue lockup warnings for CPUs 81-96,
> > which are offline on my system. The workqueues remain stuck for over 237
> > seconds:
> > 
> > [  243.309302][    C0] BUG: workqueue lockup - pool cpus=81 node=0
> > flags=0x4 nice=0 stuck for 237s!
> > [  243.309311][    C0] BUG: workqueue lockup - pool cpus=82 node=0
> > flags=0x4 nice=0 stuck for 237s!
> > [  243.309318][    C0] BUG: workqueue lockup - pool cpus=83 node=0
> > flags=0x4 nice=0 stuck for 237s!
> > [  243.309326][    C0] BUG: workqueue lockup - pool cpus=84 node=0
> > flags=0x4 nice=0 stuck for 237s!
> > [  243.309333][    C0] BUG: workqueue lockup - pool cpus=85 node=0
> > flags=0x4 nice=0 stuck for 237s!
> > [  243.309341][    C0] BUG: workqueue lockup - pool cpus=86 node=0
> > flags=0x4 nice=0 stuck for 237s!
> > [  243.309348][    C0] BUG: workqueue lockup - pool cpus=87 node=0
> > flags=0x4 nice=0 stuck for 237s!
> > [  243.309355][    C0] BUG: workqueue lockup - pool cpus=88 node=0
> > flags=0x4 nice=0 stuck for 237s!
> > [  243.309363][    C0] BUG: workqueue lockup - pool cpus=89 node=0
> > flags=0x4 nice=0 stuck for 237s!
> > [  243.309370][    C0] BUG: workqueue lockup - pool cpus=90 node=0
> > flags=0x4 nice=0 stuck for 237s!
> > [  243.309377][    C0] BUG: workqueue lockup - pool cpus=91 node=0
> > flags=0x4 nice=0 stuck for 237s!
> > [  243.309384][    C0] BUG: workqueue lockup - pool cpus=92 node=0
> > flags=0x4 nice=0 stuck for 237s!
> > [  243.309392][    C0] BUG: workqueue lockup - pool cpus=93 node=0
> > flags=0x4 nice=0 stuck for 237s!
> > [  243.309399][    C0] BUG: workqueue lockup - pool cpus=94 node=0
> > flags=0x4 nice=0 stuck for 237s!
> > [  243.309406][    C0] BUG: workqueue lockup - pool cpus=95 node=0
> > flags=0x4 nice=0 stuck for 237s!
> > [  243.309413][    C0] BUG: workqueue lockup - pool cpus=96 node=0
> > flags=0x4 nice=0 stuck for 237s!
> > 
> > Git bisect identified this as the first bad commit:
> > 
> > commit 61bbcfb50514a8a94e035a7349697a3790ab4783
> > Author: Paul E. McKenney 
> > Date:   Fri Mar 20 20:29:20 2026 -0700
> > 
> >     srcu: Push srcu_node allocation to GP when non-preemptible
> > 
> >     When the srcutree.convert_to_big and srcutree.big_cpu_lim kernel boot
> >     parameters specify initialization-time allocation of the srcu_node
> >     tree for statically allocated srcu_struct structures (for example, in
> >     DEFINE_SRCU() at build time instead of init_srcu_struct() at
> > runtime),
> >     init_srcu_struct_nodes() will attempt to dynamically allocate this
> > tree
> >     at the first run-time update-side use of this srcu_struct structure,
> >     but while holding a raw spinlock. Because the memory allocator can
> >     acquire non-raw spinlocks, this can result in lockdep splats.
> > 
> >     This commit therefore uses the same SRCU_SIZE_ALLOC trick that is
> > used
> >     when the first run-time update-side use of this srcu_struct structure
> >     happens before srcu_init() is called. The actual allocation then
> > takes
> >     place from workqueue context at the ends of upcoming SRCU grace
> > periods.
> > 
> >     [boqun: Adjust the sha1 of the Fixes tag]
> > 
> >     Fixes: 175b45ed343a ("srcu: Use raw spinlocks so call_srcu() can be
> > used under preempt_disable()")
> >     Signed-off-by: Paul E. McKenney 
> >     Signed-off-by: Boqun Feng 
> > 
> >  kernel/rcu/srcutree.c | 7 +--
> >  1 file changed, 5 insertions(+), 2 deletions(-)
> > 
> > Reverting this commit resolves the issue.
> > 
> > The problem appears to be that the workqueue is attempting to execute on
> > offline CPUs. The commit moves SRCU node allocation to workqueue context
> > to avoid lockdep issues with memory allocation under raw spinlocks,
> > which makes sense. However, it seems the workqueue scheduling doesn't
> > properly account for CPU online/offline state in this code path.
> > 
> > My test environment:
> > - Architecture: PowerPC
> > - Kernel version: Latest upstream (7.1-rc1)
> > - CPUs 81-96 are offline at boot time
> > 
> > I suspect the issue might be related to:
> > 1. Workqueue not checking CPU online status before scheduling SRCU
> > allocation work
> > 2. Missing CPU hotplug awareness in the new workqueue-based allocation
> > path
> > 3. Possible race condition with CPU hotplug events
> > 
> > Would it make sense to use queue_work_on() with explicit online CPU
> > selection, or add CPU hotplug handlers for this workqueue? I'm not
> > deeply familiar with the workqueue internals, so I might be missing
> > something.
> > Please let me know if you need any additional details or if you'd like
> > me to test any p

Re: [mainline][BUG] Observed Workqueue lockups on offline CPUs.

2026-04-27 Thread Samir M



On 27/04/26 3:32 pm, Samir M wrote:

Hi Paul,

I've been testing the latest upstream kernel on a PowerPC system and 
encountered workqueue lockup issues that I've bisected to commit 
61bbcfb50514 ("srcu: Push srcu_node allocation to GP when 
non-preemptible").
After booting, I'm seeing workqueue lockup warnings for CPUs 81-96, 
which are offline on my system. The workqueues remain stuck for over 
237 seconds:


[  243.309302][    C0] BUG: workqueue lockup - pool cpus=81 node=0 
flags=0x4 nice=0 stuck for 237s!
[  243.309311][    C0] BUG: workqueue lockup - pool cpus=82 node=0 
flags=0x4 nice=0 stuck for 237s!
[  243.309318][    C0] BUG: workqueue lockup - pool cpus=83 node=0 
flags=0x4 nice=0 stuck for 237s!
[  243.309326][    C0] BUG: workqueue lockup - pool cpus=84 node=0 
flags=0x4 nice=0 stuck for 237s!
[  243.309333][    C0] BUG: workqueue lockup - pool cpus=85 node=0 
flags=0x4 nice=0 stuck for 237s!
[  243.309341][    C0] BUG: workqueue lockup - pool cpus=86 node=0 
flags=0x4 nice=0 stuck for 237s!
[  243.309348][    C0] BUG: workqueue lockup - pool cpus=87 node=0 
flags=0x4 nice=0 stuck for 237s!
[  243.309355][    C0] BUG: workqueue lockup - pool cpus=88 node=0 
flags=0x4 nice=0 stuck for 237s!
[  243.309363][    C0] BUG: workqueue lockup - pool cpus=89 node=0 
flags=0x4 nice=0 stuck for 237s!
[  243.309370][    C0] BUG: workqueue lockup - pool cpus=90 node=0 
flags=0x4 nice=0 stuck for 237s!
[  243.309377][    C0] BUG: workqueue lockup - pool cpus=91 node=0 
flags=0x4 nice=0 stuck for 237s!
[  243.309384][    C0] BUG: workqueue lockup - pool cpus=92 node=0 
flags=0x4 nice=0 stuck for 237s!
[  243.309392][    C0] BUG: workqueue lockup - pool cpus=93 node=0 
flags=0x4 nice=0 stuck for 237s!
[  243.309399][    C0] BUG: workqueue lockup - pool cpus=94 node=0 
flags=0x4 nice=0 stuck for 237s!
[  243.309406][    C0] BUG: workqueue lockup - pool cpus=95 node=0 
flags=0x4 nice=0 stuck for 237s!
[  243.309413][    C0] BUG: workqueue lockup - pool cpus=96 node=0 
flags=0x4 nice=0 stuck for 237s!


Git bisect identified this as the first bad commit:

commit 61bbcfb50514a8a94e035a7349697a3790ab4783
Author: Paul E. McKenney 
Date:   Fri Mar 20 20:29:20 2026 -0700

    srcu: Push srcu_node allocation to GP when non-preemptible

    When the srcutree.convert_to_big and srcutree.big_cpu_lim kernel boot
    parameters specify initialization-time allocation of the srcu_node
    tree for statically allocated srcu_struct structures (for example, in
    DEFINE_SRCU() at build time instead of init_srcu_struct() at 
runtime),
    init_srcu_struct_nodes() will attempt to dynamically allocate this 
tree

    at the first run-time update-side use of this srcu_struct structure,
    but while holding a raw spinlock. Because the memory allocator can
    acquire non-raw spinlocks, this can result in lockdep splats.

    This commit therefore uses the same SRCU_SIZE_ALLOC trick that is 
used

    when the first run-time update-side use of this srcu_struct structure
    happens before srcu_init() is called. The actual allocation then 
takes
    place from workqueue context at the ends of upcoming SRCU grace 
periods.


    [boqun: Adjust the sha1 of the Fixes tag]

    Fixes: 175b45ed343a ("srcu: Use raw spinlocks so call_srcu() can 
be used under preempt_disable()")

    Signed-off-by: Paul E. McKenney 
    Signed-off-by: Boqun Feng 

 kernel/rcu/srcutree.c | 7 +--
 1 file changed, 5 insertions(+), 2 deletions(-)

Reverting this commit resolves the issue.

The problem appears to be that the workqueue is attempting to execute 
on offline CPUs. The commit moves SRCU node allocation to workqueue 
context to avoid lockdep issues with memory allocation under raw 
spinlocks, which makes sense. However, it seems the workqueue 
scheduling doesn't properly account for CPU online/offline state in 
this code path.


My test environment:
- Architecture: PowerPC
- Kernel version: Latest upstream (7.1-rc1)
- CPUs 81-96 are offline at boot time

I suspect the issue might be related to:
1. Workqueue not checking CPU online status before scheduling SRCU 
allocation work
2. Missing CPU hotplug awareness in the new workqueue-based allocation 
path

3. Possible race condition with CPU hotplug events

Would it make sense to use queue_work_on() with explicit online CPU 
selection, or add CPU hotplug handlers for this workqueue? I'm not 
deeply familiar with the workqueue internals, so I might be missing 
something.
Please let me know if you need any additional details or if you'd like 
me to test any patches.


If you happen to fix the above issue, then please add below tag.
Reported-by: Samir M 


Thanks,
Samir


Hi Paul,


I worked on fixing the issue and introduced the changes below. With 
these updates, I no longer observe any workqueue lockup messages for 
offline CPUs.

Could you please review the changes and share your feedback?

The commit 61bbcfb50514 ("srcu: Push srcu_node allocation to GP when
non-preemptible") introduced workqueu