Hi Mukesh.

On 9/29/26 12:46 PM, Mukesh Kumar Chaurasiya (IBM) wrote:
When running with CONFIG_PPC_IRQ_SOFT_MASK_DEBUG=y and LOCKDEP=y, an
Oops is triggered during NMI exit:

   [  202.258332] Oops: Exception in kernel mode, sig: 4 [#1]
   NIP [000000000003e178] 0x3e178
   LR  [c00000000003e174] arch_local_irq_restore+0x2a4/0x328
   TRAP: 0700  MSR: 0000000000081000 <ME>

The faulting instruction at NIP is a 'trap' (0fe00000) inside the
rfi_flush_fallback trampoline at a userspace-looking low address, reached
after a bad rfid caused by an incorrect SRR0 value.

Root cause
----------
DEFINE_INTERRUPT_HANDLER_NMI used to call arch_interrupt_nmi_exit_prepare()
*before* irqentry_nmi_exit(), giving this sequence:

   arch_interrupt_nmi_enter_prepare()   // saves irq_soft_mask, irq_happened
   irqentry_nmi_enter()                 // __nmi_enter() -> nmi_count++, 
in_nmi()=true
   ____func(regs)                       // NMI handler body
   arch_interrupt_nmi_exit_prepare()    // RESTORES irq_soft_mask=0 
(IRQS_ENABLED)
                                        // RESTORES irq_happened with pending 
bits
   irqentry_nmi_exit()                  // instrumentation block runs here...
     trace_hardirqs_on_prepare()        //   ...with in_nmi()==true still
     lockdep_hardirqs_on_prepare()      //   ...and irq_soft_mask already 
restored to 0
       -> arch_local_irq_restore(0)     // mask=0 -> enable path entered
            WARN_ON_ONCE(in_nmi())      // FIRES: __nmi_exit() hasn't run yet
     __nmi_exit()                       // nmi_count-- (too late)

arch_interrupt_nmi_exit_prepare() restores irq_soft_mask to whatever
the interrupted thread had (IRQS_ENABLED / 0) and restores irq_happened
with any pending deferred-interrupt bits that were present when the NMI
fired.  This is correct and necessary before rfid back to the interrupted
context, but it must not happen until irqentry_nmi_exit() has fully
completed.

With irq_soft_mask=0 and pending bits in irq_happened, the
trace_hardirqs_on_prepare() call inside irqentry_nmi_exit()'s
instrumentation block reaches lockdep_hardirqs_on_prepare() which
calls __trace_hardirqs_on_caller(), which eventually calls
arch_local_irq_restore(0) (the IRQS_ENABLED / enable path).

arch_local_irq_restore() with mask=0 hits the CONFIG_PPC_IRQ_SOFT_MASK_DEBUG
guards at irq_64.c:217:

   WARN_ON_ONCE(in_nmi());   // TRUE: __nmi_exit() not yet called

The WARN_ON_ONCE fires a 'trap' instruction (tw 31,0,0 / 0fe00000) in
kernel mode with MSR_PR=0, MSR_IR=0 (real-mode).  The program check
handler (TRAP 0700) takes the sig:4 (SIGILL) path via do_program_check()
and calls die(), producing the Oops.

The RESTART_TABLE mechanism in arch_local_irq_restore then causes the
bad rfid that lands at the userspace-looking address visible in the NIP.

The bug is invisible with LOCKDEP=n because trace_hardirqs_on_prepare()
is a no-op when CONFIG_TRACE_IRQFLAGS is not set (TRACE_IRQFLAGS is
selected by PROVE_LOCKING which selects LOCKDEP).

Fix
---
Move arch_interrupt_nmi_exit_prepare() to after the irqentry_nmi_exit()
block, so that:

   irqentry_nmi_exit()                 // __nmi_exit() -> nmi_count--, 
in_nmi()=false
   arch_interrupt_nmi_exit_prepare()   // now safe to restore soft mask state

The instrumentation inside irqentry_nmi_exit() now executes while
in_nmi() is still true (nmi_count not yet decremented), but
irq_soft_mask is still IRQS_ALL_DISABLED (NMI-entry value), so
arch_local_irq_restore() is never called with mask=0 from that path.
After __nmi_exit() returns, arch_interrupt_nmi_exit_prepare() restores
the pre-NMI soft mask and irq_happened, which is the correct state for
the rfid back to the interrupted context.

Reported-by: Venkat Rao Bagalkote <[email protected]>
Closes: 
https://lore.kernel.org/all/[email protected]/
Fixes: bee25f97ad24 ("powerpc: Enable GENERIC_ENTRY feature")
Signed-off-by: Mukesh Kumar Chaurasiya (IBM) <[email protected]>
---
  arch/powerpc/include/asm/interrupt.h | 2 +-
  1 file changed, 1 insertion(+), 1 deletion(-)

diff --git a/arch/powerpc/include/asm/interrupt.h 
b/arch/powerpc/include/asm/interrupt.h
index 1b45a49e9bed..56f855c6f07e 100644
--- a/arch/powerpc/include/asm/interrupt.h
+++ b/arch/powerpc/include/asm/interrupt.h
@@ -301,7 +301,6 @@ interrupt_handler long func(struct pt_regs *regs)           
        \
                state = irqentry_nmi_enter(regs);                       \
        }                                                               \
        ret = ____##func (regs);                                        \
-       arch_interrupt_nmi_exit_prepare(regs, &nmi_state);          \
        if (mfmsr() & MSR_DR) {                                             \
                /* nmi_exit if relocations are on */                    \
                irqentry_nmi_exit(regs, state);                         \
@@ -317,6 +316,7 @@ interrupt_handler long func(struct pt_regs *regs)           
        \
        } else {                                                        \
                irqentry_nmi_exit(regs, state);                         \
        }                                                               \
+       arch_interrupt_nmi_exit_prepare(regs, &nmi_state);          \
                                                                        \
        return ret;                                                     \
  }                                                                     \

Thanks for the fix.

Ordering also seems right now visually and it reflects similar to code previous
to generic entry.

arch_interrupt_nmi_enter_prepare
irqentry_nmi_enter
irqentry_nmi_exit
arch_interrupt_nmi_exit_prepare

So,
Reviewed-by: Shrikanth Hegde <[email protected]>

Reply via email to