Hi Michal, Thanks for sharing the host information and the full guest boot log.
On 2026/09/29 02:18 PM, Michal Suchánek wrote: > On Tue, Sep 29, 2026 at 04:39:05PM +0530, Amit Machhiwal wrote: > > On 2026/09/29 01:04 PM, Michal Suchánek wrote: > > > On Tue, Sep 29, 2026 at 09:36:05AM +0530, Amit Machhiwal wrote: > > > > On 2026/09/28 11:43 AM, Amit Machhiwal wrote: > > > > > Hi Michal, > > > > > > > > > > Thanks for the report. We have been looking at the crash and have > > > > > some findings > > > > > and follow-up questions. > > > > > > > > > > On 2026/09/25 10:10 AM, Michal Suchánek wrote: > > > > > > Hello, > > > > > > > > > > > > There appears to be a regression between > > > > > > > > < snip > > > > > > > > > > > > > > > > > Linux 7.0.12 > > > > > > https://github.com/openSUSE/kernel/tree/4c9922875554ec191166ba6c0145f12d8ff9d4fc > > > > > > https://github.com/openSUSE/kernel-source/blob/2ebf0bc/config/ppc64le/default > > > > > > > > I have been trying to recreate the issue with this kernel and config > > > > but I > > > > haven't been able to. The L2 KVM guest boots fine everytime. > > > > > > > > In addition to the requested information, could you please also share > > > > your qemu > > > > cmdline/guest xml you used? > > > > > > The XML of the VM is below. What other information do you require? > > > > Thanks for sharing the guest XML. I had requested some more info [1]. > > Could > > you please share that? > > > > [1] > > https://lore.kernel.org/all/[email protected]/ > > Hello, > > full log below. > > host: > > numactl --hardware > available: 2 nodes (0-1) > node 0 cpus: 0 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 > 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 > 51 52 53 54 55 > node 0 size: 122139 MB > node 0 free: 103635 MB > node 1 cpus: 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 > 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 > 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 > node 1 size: 139516 MB > node 1 free: 128483 MB > node distances: > node 0 1 > 0: 10 20 > 1: 20 10 > > The patch does indeed make it possible to boot when the buffer > allocation fails. Glad the first fix works. That fixes the crash (the NULL deref in __find_event_file()), but the ring buffer allocation failure itself is a separate bug worth fixing too. My suspicion is that with CONFIG_DEFERRED_STRUCT_PAGE_INIT=y, si_mem_available() returns falsely negative during early boot because NR_FREE_PAGES only reflects the non-deferred memory pool at that point — the bulk of RAM hasn't been handed to the buddy allocator yet. The check in __rb_allocate_pages() rejects the allocation prematurely. [ 0.149312][ T743] node 0 deferred pages initialised in 20ms Since the issue is not recreating on my environment, could you please give this patch a try and see if the actual crash goes away? diff --git a/kernel/trace/ring_buffer.c b/kernel/trace/ring_buffer.c index 04bb94c29f58..224cc0e5c066 100644 --- a/kernel/trace/ring_buffer.c +++ b/kernel/trace/ring_buffer.c @@ -2454,7 +2454,7 @@ static int __rb_allocate_pages(struct ring_buffer_per_cpu *cpu_buffer, * not going to succeed. */ i = si_mem_available(); - if (i < nr_pages) + if (system_state != SYSTEM_BOOTING && i < nr_pages) return -ENOMEM; /* Thanks, Amit
