> -----Original Message-----
> From: Alex Williamson <[email protected]>
> Sent: Saturday, August 29, 2026 3:24 AM
> To: Manish Honap <[email protected]>
> Cc: [email protected]; Ankit Agrawal <[email protected]>; [email protected];
> [email protected]; [email protected]; Srirangan Madhavan
> <[email protected]>; [email protected]; [email protected];
> [email protected]; [email protected]; [email protected];
> [email protected]; [email protected]; Yishai Hadas
> <[email protected]>; Shameer Kolothum Thodi
> <[email protected]>; [email protected]; [email protected];
> [email protected]; [email protected]; [email protected]; Neo Jia
> <[email protected]>; Krishnakant Jaju <[email protected]>; Vikram Sethi
> <[email protected]>; Zhi Wang <[email protected]>; linux-
> [email protected]; [email protected]; [email protected];
> [email protected]; [email protected]; linux-
> [email protected]; [email protected]; [email protected]
> Subject: Re: [PATCH v4 22/27] vfio/cxl: Revoke the HDM mapping on reset and
> power transitions
> 
> External email: Use caution opening links or attachments
> 
> 
> On Thu, 13 Aug 2026 15:06:26 +0530 <[email protected]> wrote:
> 
> > From: Manish Honap <[email protected]>
> >
> > The HDM region is a device region, not a BAR, so vfio_pci_zap_bars()
> > leaves its PTEs in place. A runtime-PM entry, a D3 transition, or a
> > reset would then leave the guest with live mappings into a quiesced device.
> >
> > Add a zap hook, called alongside the BAR zap under memory_lock, that
> > unmaps the window. The fault path already refuses to re-insert PFNs
> > while the device is suspended or its Memory Space is disabled.
> 
> We need to think about what happens in the dmabuf mmap world[1].
> Currently there are no device specific regions supporting mmap.
> Zapping is left as a compatibility interface, but the right solution is 
> probably to
> use dmabuf for mmap where we can.  Otherwise zap should likely be handled
> generically for device specific regions supporting mmap rather than as a CXL
> one-off.
> 
> AFAIK, Matt's series is still in the works and this will conflict.
> Thanks,
> 
> Alex

Agree. I will align it with the dma-buf direction. Kernel-v5 will add a dma-buf 
export
of the HDM region (imported by iommufd) and revoke it on reset via move_notify.

I will coordinate with Matt to resolve this conflict.

Manish

> 
> [1]https://lore.kernel.org/all/[email protected]/
> 
> 
> > Signed-off-by: Manish Honap <[email protected]>
> > ---
> >  drivers/vfio/pci/cxl/vfio_cxl_core.c | 29 ++++++++++++++++++++++++++++
> >  drivers/vfio/pci/vfio_pci_core.c     |  8 ++++++++
> >  include/linux/vfio_pci_core.h        |  2 ++
> >  3 files changed, 39 insertions(+)
> >
> > diff --git a/drivers/vfio/pci/cxl/vfio_cxl_core.c
> > b/drivers/vfio/pci/cxl/vfio_cxl_core.c
> > index 0fb5ed5d86b7..f1c6bf06c408 100644
> > --- a/drivers/vfio/pci/cxl/vfio_cxl_core.c
> > +++ b/drivers/vfio/pci/cxl/vfio_cxl_core.c
> > @@ -530,6 +530,31 @@ static void vfio_cxl_release_device(struct
> vfio_pci_core_device *vdev)
> >       vdev->cxl = NULL;
> >  }
> >
> > +static void vfio_cxl_zap(struct vfio_pci_core_device *vdev) {
> > +     struct vfio_cxl_state *cxl = vdev->cxl;
> > +
> > +     lockdep_assert_held_write(&vdev->memory_lock);
> > +
> > +     if (!cxl)
> > +             return;
> > +
> > +     /*
> > +      * Revoke the mapping so a later access re-faults. Do not touch 
> > hdm_valid
> > +      * here: zap also runs on a plain PCI Memory-Space disable, across 
> > which
> > +      * the committed HDM decoder stays valid (CXL.mem is not gated by PCI
> > +      * Memory-Space). hdm_valid tracks decoder validity and is cleared 
> > only
> by
> > +      * the paths that can leave the decoder unrestored (a failed reset or 
> > PM
> > +      * restore). A reset or D3 transition holds memory_lock for write 
> > while it
> > +      * runs, so no fault races the revoke, and a runtime-suspended device 
> > is
> > +      * caught by the pm_runtime_engaged check on the insert path.
> > +      */
> > +     unmap_mapping_range(vdev->vdev.inode->i_mapping,
> > +                         VFIO_PCI_INDEX_TO_OFFSET(VFIO_PCI_NUM_REGIONS +
> > +                                                  cxl->hdm_region_idx),
> > +                         range_len(&cxl->hpa_range), true); }
> > +
> >  static int vfio_cxl_open_device(struct vfio_pci_core_device *vdev)  {
> >       struct vfio_cxl_state *cxl = vdev->cxl; @@ -594,6 +619,9 @@
> > static int vfio_cxl_open_device(struct vfio_pci_core_device *vdev)
> >       if (ret)
> >               goto err_free_shadows;
> >
> > +     /* Remember where the HDM region landed so it can be zapped by index.
> */
> > +     cxl->hdm_region_idx = vdev->num_regions - 1;
> > +
> >       ret = vfio_pci_core_register_dev_region(vdev, VFIO_REGION_TYPE_CXL,
> >                                               
> > VFIO_REGION_SUBTYPE_CXL_COMP_REGS,
> >                                               &vfio_cxl_comp_regops,
> > cxl->hdm_len, @@ -738,6 +766,7 @@ static const struct vfio_cxl_ops
> vfio_cxl_ops = {
> >       .close_device   = vfio_cxl_close_device,
> >       .config_read    = vfio_cxl_config_read,
> >       .config_write   = vfio_cxl_config_write,
> > +     .zap            = vfio_cxl_zap,
> >       .owner          = THIS_MODULE,
> >  };
> >
> > diff --git a/drivers/vfio/pci/vfio_pci_core.c
> > b/drivers/vfio/pci/vfio_pci_core.c
> > index 77f8f39dd670..1a54f15d1c2c 100644
> > --- a/drivers/vfio/pci/vfio_pci_core.c
> > +++ b/drivers/vfio/pci/vfio_pci_core.c
> > @@ -1832,6 +1832,14 @@ void
> > vfio_pci_zap_and_down_write_memory_lock(struct vfio_pci_core_device
> *vdev)  {
> >       down_write(&vdev->memory_lock);
> >       vfio_pci_zap_bars(vdev);
> > +     /*
> > +      * The HDM region lives in the device-region offset range that
> > +      * vfio_pci_zap_bars() does not cover, so revoke it here too. 
> > Otherwise
> > +      * a runtime-PM entry, D3 transition, or reset would leave the guest
> > +      * with live mappings into a quiesced device.
> > +      */
> > +     if (vdev->cxl_ops && vdev->cxl_ops->zap)
> > +             vdev->cxl_ops->zap(vdev);
> >  }
> >
> >  u16 vfio_pci_memory_lock_and_enable(struct vfio_pci_core_device
> > *vdev) diff --git a/include/linux/vfio_pci_core.h
> > b/include/linux/vfio_pci_core.h index 294e95b5e881..8b93949d4484
> > 100644
> > --- a/include/linux/vfio_pci_core.h
> > +++ b/include/linux/vfio_pci_core.h
> > @@ -76,6 +76,8 @@ struct vfio_cxl_ops {
> >                              int count, __le32 *val);
> >       int     (*config_write)(struct vfio_pci_core_device *vdev, int pos,
> >                               int count, __le32 val);
> > +     /* Revoke the HDM mapping; paired with the BAR zap */
> > +     void    (*zap)(struct vfio_pci_core_device *vdev);
> >
> >       /* Pinned per bound CXL device so vfio-cxl cannot unload under usage 
> > */
> >       struct module *owner;


Reply via email to