HI Sourabh,

On 19/07/26 9:59 AM, Sourabh Jain wrote:


On 16/07/26 10:21, Ritesh Harjani (IBM) wrote:
Sourabh Jain <[email protected]> writes:

This issue only occurs with the Radix MMU.
Booting the same kernel with disable_radix=1 does not reproduce the issue.

The system also has pmem.
@Sourabh,

Few requests -
1. Could you also attach the boot logs from the same system with the
    same kernel but with disable_radix=1?

2. I can't really tell on how is MMU related to this issue? But I think
    if you could share the working dmesg logs from Hash that might give
    some more info. I am suspecting what if the firmware negotiates
    different page sizes even for TCEs for Radix v/s Hash.

Attached the dmesg of the kernel where issue is reproducible with and without
disable_radix=1


3. Would it be possible for you to try git bisect and identify the
    culprit commit?

git bisect start
# status: waiting for both good and bad commits
# good: [1471c517cf7dae1a6342fb821d8ed501af956dd0] powerpc/iommu: bypass DMA APIs for coherent allocations for pre-mapped memory
git bisect good 1471c517cf7dae1a6342fb821d8ed501af956dd0
# status: waiting for bad commit, 1 good commit known
# bad: [328335a79487ec38d6b0e1aa807785b0f75e594d] powerpc/powernv/iommu: iommu incorrectly bypass DMA APIs
git bisect bad 328335a79487ec38d6b0e1aa807785b0f75e594d
# bad: [939faf71cf7ca9ab3d1bd2912ac0e203d4d7156a] Merge tag 'drm- next-2026-02-11' of https://gitlab.freedesktop.org/drm/kernel
git bisect bad 939faf71cf7ca9ab3d1bd2912ac0e203d4d7156a
# good: [66bbe4a8ed73f1187a4271c58f0ea30f42debe0d] Merge tag 'irq- core-2026-02-09' of git://git.kernel.org/pub/scm/linux/kernel/git/tip/tip
git bisect good 66bbe4a8ed73f1187a4271c58f0ea30f42debe0d
# bad: [d061251387903e8502843ac983553f0b2e098ef8] Merge tag 'sound-7.0- rc1' of git://git.kernel.org/pub/scm/linux/kernel/git/tiwai/sound
git bisect bad d061251387903e8502843ac983553f0b2e098ef8
# good: [f51d34065de4f29fd237276a49cad06ac7356a55] Merge tag 'socfpga_dts_updates_for_v6.20_v3' of git://git.kernel.org/pub/scm/ linux/kernel/git/dinguyen/linux into soc/dt
git bisect good f51d34065de4f29fd237276a49cad06ac7356a55
# bad: [e86dda7bde8801d32ffe7d1570fe173cab14d1ba] Merge tag 'spi-v6.20' of git://git.kernel.org/pub/scm/linux/kernel/git/broonie/spi
git bisect bad e86dda7bde8801d32ffe7d1570fe173cab14d1ba
# good: [bdbddf72a2ab1cfea699959795d70df3931eefe7] Merge tag 'soc- drivers-7.0' of git://git.kernel.org/pub/scm/linux/kernel/git/soc/soc
git bisect good bdbddf72a2ab1cfea699959795d70df3931eefe7
# good: [8ea39d960c9f890e9213cdcfcbe4b3f281acd12f] spi: add multi-lane support
git bisect good 8ea39d960c9f890e9213cdcfcbe4b3f281acd12f
# bad: [d5cbd9f332c70be9589201474b9477baf9b5a24d] Merge tag 'regmap- v6.20' of git://git.kernel.org/pub/scm/linux/kernel/git/broonie/regmap
git bisect bad d5cbd9f332c70be9589201474b9477baf9b5a24d
# good: [8ad8d24d962804d3c54b5b7542c4cb7cd6bacb3b] Merge tag 'parisc- for-7.0-rc1' of git://git.kernel.org/pub/scm/linux/kernel/git/deller/ parisc-linux
git bisect good 8ad8d24d962804d3c54b5b7542c4cb7cd6bacb3b
# bad: [1d72a02d65a733ad124bacc2db3fb90fa81a612b] docs: trusted- encryped: add PKWM as a new trust source
git bisect bad 1d72a02d65a733ad124bacc2db3fb90fa81a612b
# bad: [11d45eee9f42f7d3f9c9f190a6f79e570d345fe2] powerpc64/bpf: Additional NVR handling for bpf_throw
git bisect bad 11d45eee9f42f7d3f9c9f190a6f79e570d345fe2
# bad: [2ed2d8f6fb3853daa536d79e0799cbac91fac9f5] powerpc64/bpf: Support tailcalls with subprogs
git bisect bad 2ed2d8f6fb3853daa536d79e0799cbac91fac9f5
# bad: [815a8d2feb5615ae7f0b5befd206af0b0160614c] powerpc/eeh: fix recursive pci_lock_rescan_remove locking in EEH event handling
git bisect bad 815a8d2feb5615ae7f0b5befd206af0b0160614c

Reverting commit 815a8d2feb56 ("powerpc/eeh: fix recursive pci_lock_rescan_remove locking in EEH event handling")
on top of commit fce2dfa773ced15f27d resolves the issue.

- Sourabh Jain



Btw, to clarify I think there are two issues here:
1. EEH triggered for the USB adapter - I think you meant that with Hash
    you don't even see EEH error being triggered at all.

2. Second problem is - EEH recovery causing this WARNING.

[   17.375534] [    T118] WARNING: arch/powerpc/kernel/eeh_pe.c:439 at eeh_pe_tree_remove+0xf8/0x2f0, CPU#10: eehd/118
<...>
[   17.375599] [    T118] CPU: 10 UID: 0 PID: 118 Comm: eehd Tainted: G            E       7.2.0-rc3upstream+ #3 PREEMPT(lazy) f29942c1738362234fefd88de0d1ed711f1a557e
[   17.375603] [    T118] Tainted: [E]=UNSIGNED_MODULE
[   17.375604] [    T118] Hardware name: IBM,9080-HEX Power11 (architected) 0x820200 0xf000007 of:IBM,FW1120.00 (NH1120_135) hv:phyp pSeries
<...>
[   17.375650] [    T118] NIP [c00000000004a048] eeh_pe_tree_remove+0xf8/0x2f0 [   17.375654] [    T118] LR [c00000000004ba6c] eeh_pe_detach_dev+0x5c/0xa0
[   17.375658] [    T118] Call Trace:
[   17.375659] [    T118] [c00000000a357c90] [c00000000dd32800] 0xc00000000dd32800 (unreliable) [   17.375668] [    T118] [c00000000a357d30] [c00000000004ba6c] eeh_pe_detach_dev+0x5c/0xa0 [   17.375672] [    T118] [c00000000a357d60] [c0000000000496ec] eeh_pe_traverse+0x6c/0xf0 [   17.375676] [    T118] [c00000000a357da0] [c00000000004e100] eeh_reset_device+0x1f8/0x298 [   17.375680] [    T118] [c00000000a357e50] [c00000000004ce5c] eeh_handle_normal_event+0x61c/0xa60 [   17.375685] [    T118] [c00000000a357f30] [c00000000004e298] eeh_event_handler+0xf8/0x190 [   17.375689] [    T118] [c00000000a357f90] [c00000000025dab4] kthread+0x194/0x1b0


-ritesh

Thanks for the bisect output — that's very helpful.

Commit 1010b4c012b0 ("powerpc/eeh: Make EEH driver device hotplug safe")
restructured EEH driver synchronization with the PCI hotplug layer, but
inadvertently introduced a deadlock that was silently preventing eehd from detecting and recovering from EEH events.

Commit 815a8d2feb56 ("powerpc/eeh: fix recursive pci_lock_rescan_remove
locking in EEH event handling") resolved that deadlock — but as a
consequence, EEH processing now proceeds further and exposes this
pre-existing WARNING in eeh_pe_tree_remove().

This means reverting 815a8d2feb56 makes the issue disappear only because
it re-introduces the deadlock that was masking it — it is not a real fix.

To further validate this, it would be useful to inject an error into the
PE on a kernel with 815a8d2feb56 applied, and confirm whether EEH
detection and recovery can proceed correctly end-to-end once the
WARNING in eeh_pe_tree_remove() is also addressed.

Thanks,
Narayana.


Reply via email to