+Shameer, who is looking at forwarding AER errors to guest :
https://lore.kernel.org/qemu-devel/sj0pr12mb8614ddfee3a9575564edef99ab...@sj0pr12mb8614.namprd12.prod.outlook.com/
On 8/31/26 20:31, Farhan Ali wrote:
Provide a vfio error handling callback, that can be used by devices to
handle PCI errors for passthrough devices.
Signed-off-by: Farhan Ali <[email protected]>
---
hw/vfio/pci.c | 27 +++++++++++++++++++++------
hw/vfio/pci.h | 1 +
2 files changed, 22 insertions(+), 6 deletions(-)
diff --git a/hw/vfio/pci.c b/hw/vfio/pci.c
index 428ab2f069..a2f489b34d 100644
--- a/hw/vfio/pci.c
+++ b/hw/vfio/pci.c
@@ -3244,21 +3244,36 @@ void vfio_pci_put_device(VFIOPCIDevice *vdev)
static void vfio_err_notifier_handler(void *opaque)
{
VFIOPCIDevice *vdev = opaque;
+ Error *err = NULL;
if (!event_notifier_test_and_clear(&vdev->err_notifier)) {
return;
}
/*
- * TBD. Retrieve the error details and decide what action
- * needs to be taken. One of the actions could be to pass
- * the error to the guest and have the guest driver recover
- * from the error. This requires that PCIe capabilities be
- * exposed to the guest. For now, we just terminate the
+ * We can retrieve the error details and decide what action
+ * needs to be taken in err_handler(). One of the actions could
+ * be to pass the error to the guest and have the guest driver
+ * recover from the error. This requires that PCIe capabilities be
+ * exposed to the guest.
+ *
+ * If err_handler() is not implemented/fails, we just terminate the
* guest to contain the error.
*/
- error_report("%s(%s) Unrecoverable error detected. Please collect any data possible and then kill the guest", __func__, vdev->vbasedev.name);
+ if (vdev->err_handler && vdev->err_handler(vdev, &err)) {
+ return;
+ }
+
+ if (err) {
+ error_prepend(&err, "Unrecoverable PCIe error detected for device %s",
+ vdev->vbasedev.name);
+ error_report_err(err);
+ } else {
+ error_printf("Unrecoverable PCIe error detected for device %s",
+ vdev->vbasedev.name);
+ }
+ error_printf("Please collect any data possible and then kill the guest");
how about that instead :
if (err) {
error_report("Unrecoverable PCIe error detected for device %s: %s",
vdev->vbasedev.name, error_get_pretty(err));
error_free(err);
} else {
error_report("Unrecoverable PCIe error detected for device %s",
vdev->vbasedev.name);
}
error_printf("Please collect any data possible and then kill the guest\n");
vm_stop(RUN_STATE_INTERNAL_ERROR);
}
diff --git a/hw/vfio/pci.h b/hw/vfio/pci.h
index c9ab949870..c067bbbebc 100644
--- a/hw/vfio/pci.h
+++ b/hw/vfio/pci.h
@@ -146,6 +146,7 @@ struct VFIOPCIDevice {
EventNotifier err_notifier;
EventNotifier req_notifier;
int (*resetfn)(struct VFIOPCIDevice *);
+ bool (*err_handler)(struct VFIOPCIDevice *, Error **);
Please add documentation, something like :
/*
* Platform-specific error recovery handler.
*
* Called when the host reports a PCI error via the VFIO error notifier.
* The handler should attempt to recover the device and forward the
* error to the guest if the platform supports it.
*
* @vdev: the VFIO PCI device that triggered the error
* @errp: set with the failure reason on false return
*
* Return true on success, the VM continues running.
* Return false on failure and set @errp, the VM will be stopped.
*/
bool (*err_handler)(struct VFIOPCIDevice *vdev, Error **errp);
That said, I'd prefer to see AER forwarding first.
Thanks,
C.
uint32_t vendor_id;
uint32_t device_id;
uint32_t sub_vendor_id;