(cc list is a random mix of maintainers of code that calls this function, Paolo as "main loop maintainer" and a few others who I thought might have an opinion.)
We have a qemu_system_guest_panicked() function which causes QEMU to report this to the user and do one of a couple of possible options (shutdown, pause the VM, do nothing). This seems mostly intended for "the guest OS told us by some mechanism that it just panicked". But we use it more widely than that... Cases which are "the guest told us about a panic": - accel/kvm/kvm-all.c handling of the KVM_EXIT_SYSTEM_EVENT SEV_TERM and CRASH subtypes - the pvpanic device - the spapr ibm,os-term RTAS call - x86 kvm: tdx_panicked_on_fatal_error() - x86 xen: the SHUTDOWN_crash shutdown subtype Cases which are not: - hw/spapr/rtas.c: if the FDT has no RTAS address during system reset - hw/spapr/spapr_events.c: if we wanted to deliver a machine check exception to the guest but the FDT has no RTAS address - hw/spapr/spapr_events.c: if we want to deliver a machine check but the guest is still dealing with a previous machine check - arm whpx, if we get an unknown/unexpected exit code trying to run the VM - x86 nvmm, for an unknown/unexpected VM exit code - ppc TCG, in powerpc_checkstop(), for a machine check exception I think - s390_handle_wait(): not sure exactly what this is - s390 unmanageable_intercept(): again not sure, think this is where the guest has gone off the rails and we can't keep running At least one or two of the above have comments to the effect that they don't want to use e.g. cpu_abort() because they want to give the user the ability to examine the VM after this unrecoverable guest error, rather than just exiting QEMU. So I guess my question is, is it OK to mash these two categories of "we can't keep running the VM" together, or should we define a new one for the "unrecoverable guest error" case, or do we already have some better thing to do that I missed? In particular, "do nothing" might be a reasonable response for the user to configure to a guest panic notification, since the guest will presumably stick the vcpu into a do-nothing loop, but "do nothing" doesn't make sense for "unrecoverable guest error" because we'll probably then sit in QEMU in a tight loop retrying whatever it was that failed. If we had some kind of qemu_system_unrecoverable_guest_error() then we could maybe convert some uses of cpu_abort() over to that (though uses of cpu_abort() are a very mixed bunch, some of which should be LOG_UNIMP or LOG_GUEST_ERROR and continue and some of which should be straightforward assertions, as well as some which might be this new exit case). (This query was prompted by a patch for arm whpx which adds a new qemu_system_guest_panicked() call for an unhandled VM exit situation.) thanks -- PMM
