On Fri, Jul 10, 2026 at 05:02:53PM +0100, Peter Maydell wrote: > (cc list is a random mix of maintainers of code that calls this > function, Paolo as "main loop maintainer" and a few others who > I thought might have an opinion.) > > We have a qemu_system_guest_panicked() function which causes QEMU > to report this to the user and do one of a couple of possible options > (shutdown, pause the VM, do nothing). This seems mostly intended for > "the guest OS told us by some mechanism that it just panicked". But > we use it more widely than that... > > Cases which are "the guest told us about a panic": > - accel/kvm/kvm-all.c handling of the KVM_EXIT_SYSTEM_EVENT SEV_TERM > and CRASH subtypes > - the pvpanic device > - the spapr ibm,os-term RTAS call > - x86 kvm: tdx_panicked_on_fatal_error() > - x86 xen: the SHUTDOWN_crash shutdown subtype > > Cases which are not: > - hw/spapr/rtas.c: if the FDT has no RTAS address during system reset > - hw/spapr/spapr_events.c: if we wanted to deliver a machine check > exception to the guest but the FDT has no RTAS address > - hw/spapr/spapr_events.c: if we want to deliver a machine check > but the guest is still dealing with a previous machine check > - arm whpx, if we get an unknown/unexpected exit code trying to run the VM > - x86 nvmm, for an unknown/unexpected VM exit code > - ppc TCG, in powerpc_checkstop(), for a machine check exception I think > - s390_handle_wait(): not sure exactly what this is > - s390 unmanageable_intercept(): again not sure, think this is where > the guest has gone off the rails and we can't keep running > > At least one or two of the above have comments to the effect that > they don't want to use e.g. cpu_abort() because they want to give > the user the ability to examine the VM after this unrecoverable > guest error, rather than just exiting QEMU. > > So I guess my question is, is it OK to mash these two categories of > "we can't keep running the VM" together, or should we define a new > one for the "unrecoverable guest error" case, or do we already have > some better thing to do that I missed?
IMHO we should NOT be abusing "panicked" for cases which are not guest OS panics. Adding new QMP events is cheap and we should do so. We need at least a "MCE" event. IMHO the "unknown / unexpected" VM exits likely deserve a different event again. > > In particular, "do nothing" might be a reasonable response for the > user to configure to a guest panic notification, since the guest will > presumably stick the vcpu into a do-nothing loop, but "do nothing" > doesn't make sense for "unrecoverable guest error" because we'll > probably then sit in QEMU in a tight loop retrying whatever it > was that failed. > > If we had some kind of qemu_system_unrecoverable_guest_error() then > we could maybe convert some uses of cpu_abort() over to that (though > uses of cpu_abort() are a very mixed bunch, some of which should be > LOG_UNIMP or LOG_GUEST_ERROR and continue and some of which should > be straightforward assertions, as well as some which might be this > new exit case). > > (This query was prompted by a patch for arm whpx which adds a new > qemu_system_guest_panicked() call for an unhandled VM exit situation.) > > thanks > -- PMM > With regards, Daniel -- |: https://berrange.com ~~ https://hachyderm.io/@berrange :| |: https://libvirt.org ~~ https://entangle-photo.org :| |: https://pixelfed.art/berrange ~~ https://fstop138.berrange.com :|
