> -----Original Message----- > From: Cédric Le Goater <[email protected]> > Sent: 02 September 2026 09:08 > To: Farhan Ali <[email protected]>; [email protected]; qemu- > [email protected] > Cc: [email protected]; [email protected]; [email protected]; > [email protected]; [email protected]; Shameer Kolothum Thodi > <[email protected]> > Subject: Re: [PATCH v4 2/4] vfio/pci: Add an error handler callback > > External email: Use caution opening links or attachments > > > +Shameer, who is looking at forwarding AER errors to guest :
Thanks Cédric. I have sent out a kernel RFC series yesterday: https://lore.kernel.org/all/[email protected]/ Please take a look. Probably this series can be extended later to come with a generic AER error handling incorporating the new eventfd proposed in that RFC. I will go through this one. Thanks, Shameer > > https://lore.kernel.org/qemu- > devel/[email protected] > amprd12.prod.outlook.com/ > > On 8/31/26 20:31, Farhan Ali wrote: > > Provide a vfio error handling callback, that can be used by devices to > > handle PCI errors for passthrough devices. > > > > Signed-off-by: Farhan Ali <[email protected]> > > --- > > hw/vfio/pci.c | 27 +++++++++++++++++++++------ > > hw/vfio/pci.h | 1 + > > 2 files changed, 22 insertions(+), 6 deletions(-) > > > > diff --git a/hw/vfio/pci.c b/hw/vfio/pci.c index > > 428ab2f069..a2f489b34d 100644 > > --- a/hw/vfio/pci.c > > +++ b/hw/vfio/pci.c > > @@ -3244,21 +3244,36 @@ void vfio_pci_put_device(VFIOPCIDevice > *vdev) > > static void vfio_err_notifier_handler(void *opaque) > > { > > VFIOPCIDevice *vdev = opaque; > > + Error *err = NULL; > > > > if (!event_notifier_test_and_clear(&vdev->err_notifier)) { > > return; > > } > > > > /* > > - * TBD. Retrieve the error details and decide what action > > - * needs to be taken. One of the actions could be to pass > > - * the error to the guest and have the guest driver recover > > - * from the error. This requires that PCIe capabilities be > > - * exposed to the guest. For now, we just terminate the > > + * We can retrieve the error details and decide what action > > + * needs to be taken in err_handler(). One of the actions could > > + * be to pass the error to the guest and have the guest driver > > + * recover from the error. This requires that PCIe capabilities be > > + * exposed to the guest. > > + * > > + * If err_handler() is not implemented/fails, we just terminate > > + the > > * guest to contain the error. > > */ > > > > - error_report("%s(%s) Unrecoverable error detected. Please collect any > data possible and then kill the guest", __func__, vdev->vbasedev.name); > > + if (vdev->err_handler && vdev->err_handler(vdev, &err)) { > > + return; > > + } > > + > > + if (err) { > > + error_prepend(&err, "Unrecoverable PCIe error detected for device > %s", > > + vdev->vbasedev.name); > > + error_report_err(err); > > + } else { > > + error_printf("Unrecoverable PCIe error detected for device %s", > > + vdev->vbasedev.name); > > + } > > + error_printf("Please collect any data possible and then kill the > > + guest"); > > how about that instead : > > if (err) { > error_report("Unrecoverable PCIe error detected for device %s: %s", > vdev->vbasedev.name, error_get_pretty(err)); > error_free(err); > } else { > error_report("Unrecoverable PCIe error detected for device %s", > vdev->vbasedev.name); > } > error_printf("Please collect any data possible and then kill the guest\n"); > > > > > vm_stop(RUN_STATE_INTERNAL_ERROR); > > } > > diff --git a/hw/vfio/pci.h b/hw/vfio/pci.h index > > c9ab949870..c067bbbebc 100644 > > --- a/hw/vfio/pci.h > > +++ b/hw/vfio/pci.h > > @@ -146,6 +146,7 @@ struct VFIOPCIDevice { > > EventNotifier err_notifier; > > EventNotifier req_notifier; > > int (*resetfn)(struct VFIOPCIDevice *); > > + bool (*err_handler)(struct VFIOPCIDevice *, Error **); > > Please add documentation, something like : > > /* > * Platform-specific error recovery handler. > * > * Called when the host reports a PCI error via the VFIO error notifier. > * The handler should attempt to recover the device and forward the > * error to the guest if the platform supports it. > * > * @vdev: the VFIO PCI device that triggered the error > * @errp: set with the failure reason on false return > * > * Return true on success, the VM continues running. > * Return false on failure and set @errp, the VM will be stopped. > */ > bool (*err_handler)(struct VFIOPCIDevice *vdev, Error **errp); > > > That said, I'd prefer to see AER forwarding first. > > Thanks, > > C. > > > > uint32_t vendor_id; > > uint32_t device_id; > > uint32_t sub_vendor_id;
