On Wed, 7 Oct 2026 at 10:10, David Marchand <[email protected]> wrote:
>
> This patchset introduces a major refactor of the VFIO subsystem in DPDK to
> support character device (cdev) interface introduced in Linux kernel, as well 
> as
> make the API more streamlined and useful. The goal is to simplify device
> management, improve compatibility, make the code readable, and clarify API.
>
> The following sections outline the key issues addressed by this patchset and 
> the
> corresponding changes introduced.
>
> 1. Only group mode is supported
> ===============================
>
> Since kernel version 4.14.327 (LTS), VFIO supports the new character device
> (cdev)-based way of working with VFIO devices (otherwise known as IOMMUFD). 
> This
> is a device-centric mode and does away with all the complexity regarding 
> groups
> and IOMMU types, delegating it all to the kernel, and exposes a much simpler
> interface to userspace. The old group-based implementation will still be 
> around,
> and will need to be kept in DPDK for compatibility reasons.
>
> To enable this, VFIO is heavily refactored, so that the code can support both
> modes while relying on (mostly) common infrastructure.
>
> Additionally, new `vfio_get_mode` API is added for those cases that need
> some introspection into VFIO's internals, with two modes: group (old-style),
> and cdev (the new mode).
>
> Historically, no-IOMMU mode was technically a variant of group mode, the
> distinction is largely irrelevant to the user, as all usages of noiommu checks
> in our codebase are for deciding whether to use IOVA or PA, not anything to do
> with managing groups. However, now that upcoming kernel versions will support
> no-IOMMU for both group, cdev compatibility, and full cdev paths, a new
> `vfio_get_iommu_mode` is also added, with two modes: safe (full IOMMU 
> backing),
> and unsafe (no-IOMMU mode). The naming is chosen explicitly to emphasize that
> using no-IOMMU mode is not ideal.
>
> 2. Custom container assignment API does not map to cdev mode
> ============================================================
>
> The existing `rte_vfio_device_setup/release` model is fundamentally 
> incompatible
> with cdev mode, because for custom container cases, the expected flow is that
> the user binds the IOMMU group (and thus, implicitly, the device itself) to a
> specific container using `rte_vfio_container_group_bind`, whereas this step is
> not needed for cdev as the device fd is assigned to the container straight 
> away.
>
> Therefore, what we do instead is introduce a new API for container device
> assignment which, semantically, will assign a device to specified container, 
> so
> that when it is mapped using `rte_pci_map_device`, the appropriate container 
> is
> selected. Under the hood though, we essentially transition to getting device 
> fd
> straight away at assign stage, so that by the time the PCI bus attempts to map
> the device, it is already mapped and we just return an fd. There is no
> "unassign" API because `release_device` already performs that function.
>
> Because the API is now unified around device assignment, the old 
> group-specific
> API's can be removed and, where appropriate, reimplemented using new API. 
> There
> were other users of VFIO which relied on group API but only for convenience
> purposes; no actual VFIO functionality depended on those API's.
>
> List of removed API's:
>
> * `rte_vfio_get_group_fd`
> * `rte_vfio_clear_group`
> * `rte_vfio_container_group_bind` (replaced by container assign API)
> * `rte_vfio_container_group_unbind`
> * `rte_vfio_noiommu_is_enabled` (replaced by new mode API)
>
> 3. The API responsibilities aren't clear and bleed into each other
> ==================================================================
>
> Some API's do multiple things at once. In particular:
>
> * `rte_vfio_get_device_info` will setup the device
> * `rte_vfio_setup_device` will get device info
>
> These API's have been adjusted to do one thing only.
>
> 4. The API does not need to be public
> =====================================
>
> The initial idea for exposing VFIO API was to enable userspace applications to
> directly map memory for DMA, but it turns out that in practice only drivers 
> use
> this API. Therefore, the entire VFIO API is made internal, driver-only, and is
> renamed from `rte_vfio` to `dev_vfio`.
>
> v19:
>
> TL;DR: the v18/v19 diff shows no major changes, more could have been done
> probably (especially on the huge API revamp patch), but rc1 is coming...
> Compilation per patch still works, hopefully my splitting broke nothing at
> runtime.
>
> - Rebased so that CI can run on the series (did not happen with v18)
> - Reordered patches:
>   - moved the uAPI update at the end of the series when needed
>   - moved small cleanups earlier in the series (vhost DMA capability cleanup,
>     bus updates like PCI or fslmc...)
> - Replaced bus/pci update with the change I prepared when replying on v18:
>   the most notable change is that the pci bus object is left in common code
> - Fixed (well, removed) VFIO header from FreeBSD and Windows exports list
>   and removed all mentions of "Linux only" in this same header since it is
>   now only exported by Linux
>   (iow lib/eal/include/dev_vfio.h -> lib/eal/linux/include/dev_vfio.h)
> - Squashed the patches touching kernel module: those are mechanical and the
>   diff is small. Also fixed one small string truncation issue reported by AI
> - Squashed the two patches around making VFIO internal only: the diff is huge,
>   but the changes are mechanical and I saw no interest in keeping them
>   separate
> - Dropped rte_errno updates in patches that did not document the change in
>   eal_vfio.h/dev_vfio.h (such updates were dropped later in the series anyway)
> - Split the too big "vfio: cleanup and refactor" patch
>   - Mechanical changes have been isolated in first patches: like renaming or
>     reorganising structures
>   - Macros only used in one location were removed from headers
>   - vfio_containers is not exposed out of eal_vfio.c anymore
>   - Fixed uapi/linux/vfio.h inclusion (must be included *first*, important
>     because of stddef dependencies)
>   - Fixed coding style like indentation, or line wrapping for 80 columns(?)
>     or log messages split over multiple lines
>
> v18:
> - Adjusted pre-refactor VFIO cleanup to not rely on per-config enabled flag
> - Moved loaded module detection from EAL into VFIO
> - Added API to check for specific VFIO modules being loaded
> - Removed VFIO stubs from non-Linux code by removing PCI bus dependency on 
> VFIO
> - Reworked vhost DMA capability check
> - Renamed "IOMMU mode" and "safe/unsafe" terminology to IOVA VA/PA
> - Fixed typo in FSLMC bus IOVA mode selection

Known CI issue for
https://mails.dpdk.org/archives/test-report/2026-October/1056865.html.
The Intel CI failure on doc generation can be ignored, since updates
of doc/ are filtered when applying the patches...


-- 
David Marchand

Reply via email to