On Mon, Jul 13, 2026 at 03:57:25PM +0200, Markus Armbruster wrote:
> Daniel P. Berrangé <[email protected]> writes:
> 
> > On Mon, Jul 13, 2026 at 01:43:46PM +0200, Markus Armbruster wrote:
> >> Daniel P. Berrangé <[email protected]> writes:
> >> 
> >> > On Mon, Jul 13, 2026 at 10:57:58AM +0200, Markus Armbruster wrote:
> >> >> Daniel P. Berrangé <[email protected]> writes:
> >> >> 
> >> >> > On Fri, Jul 10, 2026 at 05:02:53PM +0100, Peter Maydell wrote:
> >> >> 
> >> >> [...]
> >> >> 
> >> >> >> So I guess my question is, is it OK to mash these two categories of
> >> >> >> "we can't keep running the VM" together, or should we define a new
> >> >> >> one for the "unrecoverable guest error" case, or do we already have
> >> >> >> some better thing to do that I missed?
> >> >> >
> >> >> > IMHO we should NOT be abusing "panicked" for cases which are
> >> >> > not guest OS panics.
> >> >> 
> >> >> Point.
> >> >> 
> >> >> > Adding new QMP events is cheap and we should do so.
> >> >> 
> >> >> Changing the event sent on a certain situation is technically a
> >> >> compatibility break.  Would it matter here?
> >> >
> >> > What wins "compat break" or "bug fix" ? A strict POV prevents almost
> >> > any bug fixes, if you want to remain bug-for-bug compatible with
> >> > old QEMU.
> >> 
> >> There is no hard and fast rule.
> >> 
> >> If a patch changes behavior, and no one is around to observe it, should
> >> we still treat it as compatibility break?  The pragmatic answer is no.
> >> 
> >> It's of course hard to be sure about non-observation.  The pragmatic
> >> answer to that is "we use the best available data, and where it is
> >> lacking, reasonably conservative guesses."
> >> 
> >> How likely is it that the fix breaks something else, and how painful
> >> could such breakage be?  Again, hard to be sure, thus reasonably
> >> conservative guesses.
> >> 
> >> > With my "management app" hat on, I want QEMU to stop sending panic
> >> > events for things that are not panics, as that is triggering incorrect
> >> > actions / admin activities. ie on a panic, I'm going to take a guest
> >> > memory dump and try to analyse what is broken in the guest kernel.
> >> >
> >> > The QAPI spec says:
> >> >
> >> >   ##
> >> >   # @GUEST_PANICKED:
> >> >   #
> >> >   # Emitted when guest OS panic is detected
> >> >
> >> >
> >> > and
> >> >
> >> >   ##
> >> >   # @RunState:
> >> >   #
> >> >   # An enumeration of VM run states.
> >> >   #
> >> >   ..
> >> >   # @guest-panicked: guest has been panicked as a result of guest OS
> >> >   #     panic
> >> >
> >> >
> >> > I don't think "machine check exception" or "unknown VM exit"
> >> > can be said to match either of those docs, and thus fixing
> >> > compliance should trump bug-for-bug compatibility IMHO.
> >> 
> >> I'm not objecting, I just want the compatibility issues considered.
> >> 
> >> What are the known observers of GUEST_PANICKED?  How would they be
> >> affected by the change?
> >> 
> >> What are the use cases for observing GUEST_PANICKED?  How could they be
> >> affected?
> >> 
> >> Reasonably conservative guesses will do.
> >
> > When libvirt see a GUEST_PANICKED, it will transition the state to
> > "CRASHED" and assign a reason of "PANICKED" as the trigger / cause.
> >
> > Then depending on the guest XML config for <on_crash> it will do one
> > of
> >
> >  * Take a core dump of QEMU
> >  * Terminate QEMU
> >  * Restart QEMU
> >  * Take a core dump of QEMU and restart
> >  * Leave it in crashed state (to allow a debugger to attach)
> >
> >
> > The "machine check" and "unknown VM exit" scenarios, would still map
> > to libvirt's "CRASHED" state, but we would want to assign distinct
> > "cause" for each of them.
> 
> Say we add event GUEST_MACHINE_CHECK, then send it instead of
> GUEST_PANICKED on a machine check.
> 
> With an updated version of libvirt, this changes exactly the "cause"
> recorded for a machine check in the "CRASHED" state.  This change is
> desirable.
> 
> Older versions of libvirt ignore the unknown GUEST_MACHINE_CHECK event.
> They therefore no longer take the <on_crash> action.  This is
> undesirable, I'm afraid.  Is it?
> 
> An orderly transition could look like this:
> 
> 1. Add a new event for each distinct cause, and emit the appropriate
>    event in addition to GUEST_PANICKED.
> 
> 2. Deprecate GUEST_PANICKED.
> 
> 3. Remove GUEST_PANICKED after a suitable grace period.

Another option is to accept  "GUEST_PANICKED" as a sub-optimal name
forever and add a "source" field to the event to explain what actually
triggered it, where "source" can be  "guest-os", "machine-check", ...
etc...

Which would probably be better compatible with libvirt too

With regards,
Daniel
-- 
|: https://berrange.com       ~~        https://hachyderm.io/@berrange :|
|: https://libvirt.org          ~~          https://entangle-photo.org :|
|: https://pixelfed.art/berrange   ~~    https://fstop138.berrange.com :|


Reply via email to