On Sun, Aug 02, 2026 at 09:56:53PM +0200, Kumar Kartikeya Dwivedi wrote:
> On Thu Jul 30, 2026 at 5:07 AM CEST, Paul E. McKenney wrote:
> > On Wed, Jul 29, 2026 at 09:21:59AM -0700, Puranjay Mohan wrote:
> >> call_rcu() and call_srcu() only ever touch their per-CPU callback lists
> >> with interrupts disabled: the enqueue runs under local_irq_save() (and the
> >> nocb locks when offloaded), and so do callback invocation and grace-period
> >> work.  That is fine as long as call_rcu() itself is invoked with
> >> interrupts enabled, but it is not always.  An NMI handler can call
> >> call_rcu(), and instrumentation can reenter it.  The case that prompted
> >> this is a BPF program attached to rcu_segcblist_enqueue() that frees an
> >> object: the free reaches call_rcu_tasks_trace(), which is call_srcu()
> >> under the hood, back on the same CPU with the srcu_data lock already held,
> >> and it deadlocks on that lock.  Either way, enqueuing directly can corrupt
> >> the list or deadlock.
> >>
> >> Rather than scatter context checks through the enqueue, make it defer
> >> whenever interrupts are disabled: stage the callback on a per-CPU lockless
> >> list and re-issue it from an irq_work once interrupts are back on, going
> >> straight to the enqueue helper so the re-issue cannot defer again.  Only
> >> the drain side takes a lock; the staging is a bare llist_add() and stays
> >> safe from NMI.  This is behind a new hidden CONFIG_RCU_DEFER, which is set
> >> wherever a reentrant enqueue is possible (HAVE_NMI, KPROBES,
> >> FUNCTION_TRACER or TRACEPOINTS); without it call_rcu() enqueues exactly as
> >> before.
> >>
> >> CPU offline is the awkward part.  A callback can be deferred very late in
> >> the outgoing CPU's teardown -- from do_idle() or cpuhp_ap_report_dead(),
> >> past the CPUHP_AP_SMPCFD_DYING flush that would otherwise run the irq_work
> >> -- so the irq_work can no longer run there to re-issue it.  rcu_barrier()
> >> and srcu_barrier() therefore drain the deferred lists themselves before
> >> they wait: for online CPUs they wait the irq_work out, and for offline
> >> ones they drain the list directly, since that irq_work may never run
> >> again.  rcutree_migrate_callbacks() drains the outgoing CPU's list too, so
> >> a late deferral still lands on a callback list even when nobody calls a
> >> barrier.  To keep those three drainers from stepping on each other, the
> >> drain holds a per-CPU raw lock across the llist_del_all() and the
> >> re-issue, so a drainer never returns having pulled callbacks off the
> >> deferred list but not yet put them on a callback list.  Every lock the
> >> re-issue touches (nocb, rcu_node, srcu_data) is already raw, so the
> >> nesting is fine.
> >>
> >> The irq_work is IRQ_WORK_INIT_HARD.  It is not needed for correctness, but
> >> a non-HARD irq_work runs from a kthread on PREEMPT_RT and can be delayed
> >> under load, letting deferred callbacks pile up; running the re-issue in
> >> hard-irq context keeps that from turning into an OOM.
> >>
> >> Patches 1 and 2 do Tree and Tiny RCU, 3 and 4 Tree and Tiny SRCU.  Patch 5
> >> teaches rcutorture to issue ->call() from a perf-overflow NMI -- the
> >> nmi_calls parameter, on by default -- on the flavors that advertise it,
> >> and checks that every callback issued from NMI is later invoked.  Patch 6
> >> adds the BPF reentry reproducer described above.
> >
> > Nice!  I applied this series to -rcu for testing and review.  Patch 6
> > might want to go up a different path, but let's see how it goes.
> > If it goes via -rcu, it will need an appropriate ack.
> >
> 
> I think it makes sense to take it through your tree for now. If there are 
> issues
> when we get these changes on BPF side (with the selftest), we'll fix forward. 
> If
> we take it now, it will end up deadlocking the CI, so should come with the
> relevant changes once trees are synced.
> 
> You can add my ack when taking it in.
> 
> Acked-by: Kumar Kartikeya Dwivedi <[email protected]>

Very good, and I will apply this on my next rebase.

                                                        Thanx, Paul

> >> Puranjay Mohan (6):
> >>   rcu: Make call_rcu() safe to call from any context
> >>   rcu: Make Tiny call_rcu() safe to call from any context
> >>   srcu: Make call_srcu() safe to call from any context
> >>   srcu: Make Tiny call_srcu() safe to call from any context
> >>   rcutorture: Exercise ->call() from NMI context
> >>   selftests/bpf: Add a call_srcu() re-entry reproducer
> >>
> >>  include/linux/srcutiny.h                      |  11 +-
> >>  include/linux/srcutree.h                      |   4 +
> >>  kernel/rcu/Kconfig                            |   6 +
> >>  kernel/rcu/rcu.h                              |  18 +++
> >>  kernel/rcu/rcutorture.c                       | 115 +++++++++++++++
> >>  kernel/rcu/srcutiny.c                         |  53 ++++++-
> >>  kernel/rcu/srcutree.c                         | 138 +++++++++++++++++-
> >>  kernel/rcu/tiny.c                             | 101 ++++++++++---
> >>  kernel/rcu/tree.c                             | 122 ++++++++++++++--
> >>  kernel/rcu/tree.h                             |   5 +
> >>  .../selftests/bpf/prog_tests/rcu_reentry.c    |  58 ++++++++
> >>  .../testing/selftests/bpf/progs/rcu_reentry.c |  45 ++++++
> >>  12 files changed, 636 insertions(+), 40 deletions(-)
> >>  create mode 100644 tools/testing/selftests/bpf/prog_tests/rcu_reentry.c
> >>  create mode 100644 tools/testing/selftests/bpf/progs/rcu_reentry.c
> >>
> >>
> >> base-commit: 9dc303e69bcd49f9668ca090ae45325269531fbb
> >> --
> >> 2.53.0-Meta
> >>
> 

Reply via email to