On Mon, 27 Jul 2026 07:39:24 +0200
Cédric Le Goater <[email protected]> wrote:

> Hello,
> 
> Live migration of VFIO-passthrough devices - SR-IOV VFs, vGPUs - is a
> growing requirement, but real hardware with migration support is
> scarce and hard to debug. An emulated device provides a fully
> controlled testbed for developing and validating the entire software
> stack - vfio-pci variant drivers, VFIO core migration v2 framework,
> QEMU, libvirt - and for tuning complex migration policies such as
> downtime convergence. It also serves as an educational reference for
> understanding VFIO migration end-to-end, from device state
> serialization to dirty page tracking.
> 
> This series adds an experimental VF live migration interface to the
> emulated igb (82576) device. It enables a vfio-pci variant driver
> (igb-vfio-pci) to migrate VFs using the standard VFIO migration v2
> protocol with stop-copy and pre-copy support.
> 
> The target scenario is nested virtualization:
> 
>   L0 QEMU (these patches)
>     igb PF with x-vf-migration=on
>     └── VFs with migration BAR + vendor cap
> 
>   L1 kernel
>     igb-vfio-pci variant driver [1]
>     translates VFIO migration v2 ioctls → BAR2 MMIO
> 
>   L1 QEMU (stock, unmodified)
>     vfio-pci device model, standard migration fd
> 
>   L2 guest
>     standard igbvf driver, unaware of migration
> 
> The L1 QEMU is completely unmodified -- it sees a standard VFIO
> migratable device and uses the normal migration fd path.
> 
> * Design
> 
> The migration interface is exposed through a hidden 64KB PCI BAR
> (BAR2) on each VF, discovered via a vendor-specific PCI capability
> ("MIGB", PCI_CAP_ID_VNDR). The BAR exposes a register-based state
> machine that mirrors VFIO migration states (RUNNING, STOP, STOP_COPY,
> RESUMING, PRE_COPY).

I think you're placing the migration BAR on the VF in order to
implement this in a small footprint, QEMU + vfio-pci variant driver,
without PF guest driver changes.  A model that better matches real
world hardware might be to put the migration BAR on the PF, segmented
per VF, and then have the PF driver vend those segments out to the VF
drivers.  That would remove the BAR always mapped problem, but expands
the footprint to include the PF driver.  However, we're not exactly
clean with respect to the PF driver as implemented here when we're
going around the PF driver's back to setup DMA mappings.

Can we take advantage of the fact that this is a virtual device to
avoid all these warts?

For example, do we really need MMIO BAR space for the register set
exposed or can we prune that down to some key registers and doorbells
and move the rest to memory?  We can put the vendor capability in
extended config space to give ourselves more room to work with if
necessary.  We also don't really need to play by the physical rules for
access, the variant driver in the L1 kernel can allocate contiguous
ranges and write GPAs into config space registers.  L0 QEMU can just
write migration data and dirty bitmaps directly to those GPAs,
bypassing any pretense of DMA mapping.

There might be some tricks we can steal from virtio as it seems to
optionally honor things like vIOMMUs as well.  Anyway, if we want to
confine the implementation to the virtual VF, avoiding dependencies on
the PF driver, both at the cross-driver API and device DMA state, I
think we can probably lean harder on QEMU being able to push data into
an arbitrary GPA regardless of the IO topology we're exposing.  Thanks,

Alex

Reply via email to