On Wed, 7 Oct 2026 at 10:10, David Marchand <[email protected]> wrote: > > This patchset introduces a major refactor of the VFIO subsystem in DPDK to > support character device (cdev) interface introduced in Linux kernel, as well > as > make the API more streamlined and useful. The goal is to simplify device > management, improve compatibility, make the code readable, and clarify API. > > The following sections outline the key issues addressed by this patchset and > the > corresponding changes introduced. > > 1. Only group mode is supported > =============================== > > Since kernel version 4.14.327 (LTS), VFIO supports the new character device > (cdev)-based way of working with VFIO devices (otherwise known as IOMMUFD). > This > is a device-centric mode and does away with all the complexity regarding > groups > and IOMMU types, delegating it all to the kernel, and exposes a much simpler > interface to userspace. The old group-based implementation will still be > around, > and will need to be kept in DPDK for compatibility reasons. > > To enable this, VFIO is heavily refactored, so that the code can support both > modes while relying on (mostly) common infrastructure. > > Additionally, new `vfio_get_mode` API is added for those cases that need > some introspection into VFIO's internals, with two modes: group (old-style), > and cdev (the new mode). > > Historically, no-IOMMU mode was technically a variant of group mode, the > distinction is largely irrelevant to the user, as all usages of noiommu checks > in our codebase are for deciding whether to use IOVA or PA, not anything to do > with managing groups. However, now that upcoming kernel versions will support > no-IOMMU for both group, cdev compatibility, and full cdev paths, a new > `vfio_get_iommu_mode` is also added, with two modes: safe (full IOMMU > backing), > and unsafe (no-IOMMU mode). The naming is chosen explicitly to emphasize that > using no-IOMMU mode is not ideal. > > 2. Custom container assignment API does not map to cdev mode > ============================================================ > > The existing `rte_vfio_device_setup/release` model is fundamentally > incompatible > with cdev mode, because for custom container cases, the expected flow is that > the user binds the IOMMU group (and thus, implicitly, the device itself) to a > specific container using `rte_vfio_container_group_bind`, whereas this step is > not needed for cdev as the device fd is assigned to the container straight > away. > > Therefore, what we do instead is introduce a new API for container device > assignment which, semantically, will assign a device to specified container, > so > that when it is mapped using `rte_pci_map_device`, the appropriate container > is > selected. Under the hood though, we essentially transition to getting device > fd > straight away at assign stage, so that by the time the PCI bus attempts to map > the device, it is already mapped and we just return an fd. There is no > "unassign" API because `release_device` already performs that function. > > Because the API is now unified around device assignment, the old > group-specific > API's can be removed and, where appropriate, reimplemented using new API. > There > were other users of VFIO which relied on group API but only for convenience > purposes; no actual VFIO functionality depended on those API's. > > List of removed API's: > > * `rte_vfio_get_group_fd` > * `rte_vfio_clear_group` > * `rte_vfio_container_group_bind` (replaced by container assign API) > * `rte_vfio_container_group_unbind` > * `rte_vfio_noiommu_is_enabled` (replaced by new mode API) > > 3. The API responsibilities aren't clear and bleed into each other > ================================================================== > > Some API's do multiple things at once. In particular: > > * `rte_vfio_get_device_info` will setup the device > * `rte_vfio_setup_device` will get device info > > These API's have been adjusted to do one thing only. > > 4. The API does not need to be public > ===================================== > > The initial idea for exposing VFIO API was to enable userspace applications to > directly map memory for DMA, but it turns out that in practice only drivers > use > this API. Therefore, the entire VFIO API is made internal, driver-only, and is > renamed from `rte_vfio` to `dev_vfio`. > > v19: > > TL;DR: the v18/v19 diff shows no major changes, more could have been done > probably (especially on the huge API revamp patch), but rc1 is coming... > Compilation per patch still works, hopefully my splitting broke nothing at > runtime. > > - Rebased so that CI can run on the series (did not happen with v18) > - Reordered patches: > - moved the uAPI update at the end of the series when needed > - moved small cleanups earlier in the series (vhost DMA capability cleanup, > bus updates like PCI or fslmc...) > - Replaced bus/pci update with the change I prepared when replying on v18: > the most notable change is that the pci bus object is left in common code > - Fixed (well, removed) VFIO header from FreeBSD and Windows exports list > and removed all mentions of "Linux only" in this same header since it is > now only exported by Linux > (iow lib/eal/include/dev_vfio.h -> lib/eal/linux/include/dev_vfio.h) > - Squashed the patches touching kernel module: those are mechanical and the > diff is small. Also fixed one small string truncation issue reported by AI > - Squashed the two patches around making VFIO internal only: the diff is huge, > but the changes are mechanical and I saw no interest in keeping them > separate > - Dropped rte_errno updates in patches that did not document the change in > eal_vfio.h/dev_vfio.h (such updates were dropped later in the series anyway) > - Split the too big "vfio: cleanup and refactor" patch > - Mechanical changes have been isolated in first patches: like renaming or > reorganising structures > - Macros only used in one location were removed from headers > - vfio_containers is not exposed out of eal_vfio.c anymore > - Fixed uapi/linux/vfio.h inclusion (must be included *first*, important > because of stddef dependencies) > - Fixed coding style like indentation, or line wrapping for 80 columns(?) > or log messages split over multiple lines > > v18: > - Adjusted pre-refactor VFIO cleanup to not rely on per-config enabled flag > - Moved loaded module detection from EAL into VFIO > - Added API to check for specific VFIO modules being loaded > - Removed VFIO stubs from non-Linux code by removing PCI bus dependency on > VFIO > - Reworked vhost DMA capability check > - Renamed "IOMMU mode" and "safe/unsafe" terminology to IOVA VA/PA > - Fixed typo in FSLMC bus IOVA mode selection
Known CI issue for https://mails.dpdk.org/archives/test-report/2026-October/1056865.html. The Intel CI failure on doc generation can be ignored, since updates of doc/ are filtered when applying the patches... -- David Marchand

