Hello Manish,

On 8/13/26 15:06, [email protected] wrote:
From: Manish Honap <[email protected]>

This series adds QEMU support for passing a CXL Type-2 device (an
accelerator with host-managed device memory, e.g. a GPU) to a guest via
vfio-pci. The guest drives its own virtual endpoint HDM decoder and QEMU
maps the device memory at the guest physical address the guest commits,
while the host owns the host physical placement.

It is a full rewrite of the RFC v1 series [1] on clean upstream master,
reworked to address Jonathan's review.

Base: master (post pull-9p-20260725), commit 6333226c2a.


Kernel dependency
-----------------

Pairs with the kernel vfio-cxl series "vfio/cxl: CXL Type-2 device
passthrough" [2].
- The kernel exposes the device memory as an HPA-backed VFIO region, traps
   the HDM decoder block and runs its lock-on-commit FSM, handles the CXL
   DVSEC (including guest-triggered reset), and reports two things through
   VFIO:
   - A device flag (VFIO_DEVICE_FLAGS_CXL) and
   - The component-register geometry (VFIO_REGION_INFO_CAP_CXL_COMP_REGS).


Sample supported topology
-------------------------

Guest disk, network, and system-RAM lines are omitted:

   -machine virt,accel=kvm,gic-version=3,hmat=on,cxl=on,ras=on, \
            highmem-mmio-size=4T
   -object iommufd,id=iommufd0
   -device pxb-cxl,bus_nr=12,bus=pcie.0,id=cxl.1
   -device cxl-rp,port=1,bus=cxl.1,id=rport0.1,chassis=4, \
           pref64-reserve=2G,mem-reserve=1G
   -M cxl-fmw.0.targets.0=cxl.1,cxl-fmw.0.size=256G
   -device arm-smmuv3,primary-bus=cxl.1,id=smmuv3.0,accel=on,ats=on, \
           ril=on,ssidsize=8,oas=48
   -device vfio-pci-nohotplug,host=<BDF>,bus=rport0.1,id=dev0, \
           iommufd=iommufd0
   -object acpi-generic-initiator,id=gi0,pci-dev=dev0,node=2
   ... (one acpi-generic-initiator per guest NUMA node the HDM memory backs)


Address model
-------------

The kernel fixes device memory to a host physical range before the guest
sees the device, and hardware presents a firmware-committed, locked
endpoint HDM decoder whose registers hold that host physical base.

QEMU never exposes Host physical base. It virtualizes the decoder base registers
in the trapped component-register read and returns the base of the device's
CFMWS window, a guest physical address, so the guest only ever sees a GPA.

QEMU maps the RAM-device region (backed by the fixed HPA) at that CFMWS base.

The kernel never learns the GPA and the guest never learns the HPA.

Because the decoder is already committed at boot, no guest commit write
triggers the mapping. QEMU maps once the guest enables memory decoding
(the Command register Memory-Space bit) and re-checks on any decoder
control write, so the region enters the guest address space and the IOAS
while the device is live.

When the guest clears Memory-Space, QEMU withdraws the mapping, matching the
kernel's revoke of the backing PTEs on the same write, so a guest access during
the disabled interval cannot fault a zapped mapping and stop the VM; the enable
path re-installs it.

This also helps to keep the guest-visible base as the CFMWS base by
construction, and the host physical placement stays with the kernel.

The RFC added the pxb-cxl _DSM because OS may treat PCI configuration as
reassignable and a BAR move would break the CXL.mem mapping. This change keeps
it narrower than the RFC v1 _DSM: it applies only to the host bridge that
carries the passed-through CXL device.


Reset
-----

There is no QEMU reset patch. A guest CXL reset is a DVSEC write that
lands in vfio config space and is handled entirely by the host kernel,
which stamps the outcome into DVSEC STATUS2. The kernel re-commits this
firmware-fixed decoder across the reset, and the guest reaches its memory
through the mapping QEMU already installed.


Reviewer feedback addressed
---------------------------

The RFC [1] thread has the full discussion.

Jonathan Cameron
   - The high-MMIO window and the cxl-fmws-base property are removed. The
     guest-visible base is the CFMWS base by construction, so it equals the
     decoder base and stays stable without a hack.
   - The guest programs its virtual (GPA) decoder and QEMU maps the memory at
     commit time.
   - The host owns the HPA, resolved before the guest sees the device.
   - The committed decoder is the fast path this series ships. The
     guest-programmed uncommitted case is delivered by cxl-core resolving the
     range at enumeration, not by a QEMU or vfio dynamic branch, so it is a
     later cxl-core item this series does not depend on.
   - FIRMWARE_COMMITTED is dropped; the cap is no longer exposed.
   - The one-endpoint, non-interleaved, no-switch topology is enforced at
     realize.
   - PCI/BAR configuration and CXL.mem stay independent.


Validation
----------

- Every patch passes scripts/checkpatch.pl --codespell --strict
   (patch 1 carries the expected imported-from-Linux warning)
- The series applies cleanly on the stated base.
- Ran couple of rounds of masoncl/review-prompts on this series before posting.


Pending items
-------------

Future enhancements for the multi-decoder, interleaved devices and switched
topologies.

Trapped CXL RAS registers are planned as a new VFIO region subtype that the same
region-by-subtype detection already handles, not as a change to the
component-register cap.

The bios-tables test refresh for the new _DSM is still to be added.

Before diving into CXL and doing a review of the QEMU patches, could you
provide a status update on the kernel vfio-cxl series [2] ?

The QEMU side is a consumer of the VFIO regions and capabilities the
kernel exposes, and patch 1 notes the capability ID is still provisional.
Knowing where the kernel series stands would help. That said, the UAPI
looks simple enough.

What has become of the CXL Type-2 device emulation effort from Zhi Wang?

  https://lore.kernel.org/qemu-devel/[email protected]/

That series introduced a bare-minimum emulated CXL Type-2 device in
QEMU, so kernel and QEMU developers could test the CXL Type-2 stack
without real hardware. I liked the idea quite a lot.

CXL is still new and complex, and hardware is scarce. Having an
emulated device and passthrough support gives us a full environment to
flush out design issues across the stack (Linux kernel, QEMU, libvirt)
and in associated subsystems like IOMMU and ACPI, without needing
physical devices. It is also useful for education and for running CI
tests. Do you see both efforts as complementary, and is the emulation
work still active?

I will take a closer look at the series after some time off.

Thanks,


C.








Reply via email to