On Thu, 13 Aug 2026 15:06:11 +0530
<[email protected]> wrote:
> From: Manish Honap <[email protected]>
>
> A CXL device needs the vfio-cxl callbacks, but pulling vfio-cxl and the
> CXL core in unconditionally would bloat every vfio-pci setup. At bind,
> detect a CXL device with pcie_is_cxl() and request_module("vfio-cxl")
> only then, and hand the device to the registered ops.
>
> Each bound CXL device pins vfio-cxl through try_module_get() and drops
> the reference at release, so vfio-cxl can unload once no CXL device is
> bound. If vfio-cxl is absent the device is driven as plain vfio-pci.
>
> Signed-off-by: Manish Honap <[email protected]>
> ---
> drivers/vfio/pci/vfio_pci_core.c | 81 ++++++++++++++++++++++++++++++--
> include/linux/vfio_pci_core.h | 3 ++
> 2 files changed, 81 insertions(+), 3 deletions(-)
>
> diff --git a/drivers/vfio/pci/vfio_pci_core.c
> b/drivers/vfio/pci/vfio_pci_core.c
> index 88e68d43af9a..0f9b5dfeea66 100644
> --- a/drivers/vfio/pci/vfio_pci_core.c
> +++ b/drivers/vfio/pci/vfio_pci_core.c
> @@ -2176,6 +2176,44 @@ static void vfio_pci_vga_uninit(struct
> vfio_pci_core_device *vdev)
> VGA_RSRC_LEGACY_MEM);
> }
>
> +static const struct vfio_cxl_ops *vfio_pci_cxl_ops;
> +static DEFINE_MUTEX(vfio_pci_cxl_ops_lock);
> +
> +static const struct vfio_cxl_ops *vfio_pci_get_cxl_ops(void)
> +{
> + const struct vfio_cxl_ops *ops;
> +
> + mutex_lock(&vfio_pci_cxl_ops_lock);
> + ops = vfio_pci_cxl_ops;
> + if (ops && !try_module_get(ops->owner))
> + ops = NULL;
> + mutex_unlock(&vfio_pci_cxl_ops_lock);
> +
> + return ops;
> +}
> +
> +/*
> + * A CXL Type-2 device advertises both CXL.cache and CXL.mem in its CXL
> DVSEC.
> + * pcie_is_cxl() is also true for Type-1 (cache only) and Type-3 (mem only)
> + * devices, which the vfio-cxl provider does not handle, so confirm the
> Type-2
> + * identity before engaging it.
We don't expect to handle type 3 class code compliant devices, but what about
the things referred to sometimes as CXL Type 3+?
No CXL.cache support, but accelerators none the less - typically using back
invalidate to ensure what they are working on isn't held by the host and
CXL.IO (i.e. PCI) for control path.
It is also plausible we'd pass a full compressed RAM device through to the
guest without paravirtualizing like we currently plan to do for class
code Type 3 devices (for DCD, sharing etc). +CC Gregory to point out where
I am wrong on this ;)
Not sure what that means for this checking function.
More generally, why are we controlling usecases? A class code complaint type 3
device 'could' be passed through I think if someone wanted to do that.
I'd not encourage it but why is it a linux policy to not support it?
Jonathan
> + */
> +static bool vfio_pci_is_cxl_type2(struct pci_dev *pdev)
> +{
> + u16 dvsec, cap;
> +
> + dvsec = pci_find_dvsec_capability(pdev, PCI_VENDOR_ID_CXL,
> + PCI_DVSEC_CXL_DEVICE);
> + if (!dvsec)
> + return false;
> +
> + if (pci_read_config_word(pdev, dvsec + PCI_DVSEC_CXL_CAP, &cap))
> + return false;
> +
> + return (cap & PCI_DVSEC_CXL_CACHE_CAPABLE) &&
> + (cap & PCI_DVSEC_CXL_MEM_CAPABLE);
> +}