On Sat, Aug 29, 2026 at 09:49:00AM +0530, Ritesh Harjani (IBM) wrote:
> So the syscall return sequence is as follows:
> A syscall return to userspace is prepared and then a short asm sequence
> that actually does the RFI. Note that this asm range is restartable i.e.
> EE is still on, so an interrupt (e.g. decrementer or external interrupt)
> can hit while SRR/GPRs are being loaded. This is defined via:
> 
> RESTART_TABLE(.Lsyscall_rst_start, .Lsyscall_rst_end, syscall_restart)
> 
> This restart table then sends us to syscall_restart rather than resuming
> in the middle of the RFI. The same stub is also used if irq_happened
> already has a pending bit (soft-masked irq that has not been replayed
> yet (PowerPC special case of local_irq_disable())).
> 
> Here is a bit of a flow of sequence of code to visualize:
>   syscall_exit_prepare
>       decide full-GPR restore (_TIF_RESTOREALL) for signal,
>       rt_sigreturn or syscall trace
>       save that in regs->exit_result and return it in r3
>            |
>            v
>   .Lsyscall_rst_start .. _end     EE still on
>       irq_happened set or interrupt in this range?
>            | no                         | yes
>            v                            v
>       cmpdi r3,0                  syscall_exit_restart
>       restore all / zero            replay irq, try exit again
>       volatiles; RFI                must return flags in r3
>                                     again for the same cmpdi
> 
> Now r3 after prepare is the flags word, not the actual syscall return. A 
> nested
> interrupt clobbers it, so the restart stub reloads RESULT into r3 and the
> C handler (syscall_exit_restart()) should put the flags back (because later 
> asm
> checks whether r3 returned from C has _TIF_RESTOREALL set or not):
>       cmpdi r3, 0
>       bne     .Lsyscall_restore_regs
> 
> Note that syscall_exit_restart() already ORs any new _TIF_RESTOREALL into
> exit_result, but then it only returns the new sample and not the full
> regs->exit_result.
> 
> That sample could be often 0 even when restore-all is still required:
> 
>   - rt_sigreturn / syscall trace set the bit in prepare's local
>     ret and in exit_result. They never set exit_flags, which is
>     what restart samples.
> 
>   - a signal does set exit_flags but restart clears it. A
>     second pass through the stub then returns 0 while
>     exit_result still has the bit.
> 
> The asm as mentioned earlier then treats r3==0 as the fast path and
> zeros r0/r4-r12. That means the userspace that needed the full register
> set could SIGSEGVs, (which could happen often in ld64.so.2 like while
> doing a parallel kernel build as reported by Venkat).
> 
> So we should instead return the accumulated exit_result, like how we do
> in interrupt_exit_user_restart(). Note that prior to this commit
> 263e5159e00a ("powerpc: Fix exit_flags field placement in pt_regs for ptrace")
> we were returning regs->exit_result from syscall_exit_restart(), but
> this commit changed that behaviour.
> 
> Fixes: 263e5159e00a ("powerpc: Fix exit_flags field placement in pt_regs for 
> ptrace")
> Reported-by: Venkat Rao Bagalkote <[email protected]>
> Closes: 
> https://lore.kernel.org/all/[email protected]/
> Signed-off-by: Ritesh Harjani (IBM) <[email protected]>
> ---
> Sorry about the long commit msg. It took sometime for me to fully understand
> that complex path, so I thought I may as well document that properly.
> 
>  arch/powerpc/kernel/interrupt.c | 2 +-
>  1 file changed, 1 insertion(+), 1 deletion(-)
> 
> diff --git a/arch/powerpc/kernel/interrupt.c b/arch/powerpc/kernel/interrupt.c
> index 5b88bf72786c..55f9c0c9922a 100644
> --- a/arch/powerpc/kernel/interrupt.c
> +++ b/arch/powerpc/kernel/interrupt.c
> @@ -175,7 +175,7 @@ notrace unsigned long syscall_exit_restart(unsigned long 
> r3, struct pt_regs *reg
>       current_thread_info()->exit_flags &= ~_TIF_RESTOREALL;
>       regs->exit_result |= ret;
> 
> -     return ret;
> +     return regs->exit_result;
>  }
>  #endif
> 
> --
> 2.39.5
> 
> 
LGTM

Reviewed-by: Mukesh Kumar Chaurasiya (IBM) <[email protected]>

Reply via email to