Stefano, Anthony, Paul,

Hi guys, hope this finds you all well. We, in XenServer, have been
doing some back to back Xen VM PV storage backend performance
comparisons between tapdisk + VHD and QEMU + QCOW2. The image
storage format appears to be largely irrelevant to the rest of this
conversation. In doing so we have found a few significant issues in
QEMU, mainly around the xen-block dataplane and adjacent code.

## Testing Environment

Debian 12 VM, 8 vCPUS, 4G Ram on XenServer 9. Three virtual block
devices, one for the OS on host local storage and two, unformatted,
20G test devices, one backed by tapdisk and one backed by QEMU. These
disks themselves are then backed by a BRD based target hosted on a
native installed Rocky 9.8 host. The two hosts are connected over a
25G, 9000 MTU, switched storage network whose switch is rated to run
all ports at wire speed without contention. In the guest
`xen_blkfront` is configured with `max_ring_page_order=3`, which is
honoured by both tapdisk and QEMU. QEMU supports 4 but tapdisk caps at
3 so for consistency 3 was used. Performance data was collected by
`fio` v3.33 from the default Debian 12 repo.

Note: we did find during this that QEMU has a separate problem with
feature publication when hot plugging a device -

https://lists.gnu.org/archive/html/qemu-devel/2026-09/msg03492.html

This was worked around by always cold booting the VM which then
negotiates correctly.

## Results

The following table is the full grid from the testing
performed. Points are expressed as ratios of QEMU/tapdisk so any
points < 1.0 show QEMU having a performance deficit. Each figure is
the median of the per-repeat ratio over 12 repeats, the first discarded
as unstable on a newly booted VM. The tapdisk arm is measured
seconds away from the QEMU arm in every repeat, so that drift in the
rig cancels rather than being attributed to either datapath.

| block size | queue depth | read | write |
|:-----------|------------:|-----:|------:|
| 4k         |           1 | 0.73 |  0.73 |
| 64k        |           1 | 0.81 |  0.83 |
| 256k       |           1 | 0.67 |  0.92 |
| 1m         |           1 | 0.67 |  0.76 |
| 4m         |           1 | 0.81 |  1.01 |
| 4k         |          32 | 1.90 |  1.32 |
| 64k        |          32 | 0.99 |  1.00 |
| 256k       |          32 | 1.14 |  0.84 |
| 1m         |          32 | 0.86 |  0.88 |
| 4m         |          32 | 0.87 |  0.86 |

Median across the 20 cells is 0.86, with 5 of 20 at or above parity.

The deficit is concentrated at queue depth 1, which is what pointed at
per-request latency rather than throughput. Note that 1m QD1 read is the
noisiest cell in the grid at 24% coefficient of variation across
repeats, so treat 0.67 as "much worse" rather than as a precise figure;
the 4k and 64k QD1 cells sit at 3-4% and are solid. 1m QD1 is noisy on
both backends and is probably an artifact of the environment.

## Issues found

1. **Large (4m) blocksize read**. Analysis determined that this was
   impacted by QEMU being overly keen in notifying the frontend of
   completion responses being placed into the ring. In turn this
   results in a Dom0 to DomU context switch and all the associated
   hypervisor flushes and mitigations. This can also be observed in
   the `ctx` count tracker in `fio` itself. Tapdisk mitigates this by
   only sending the notification on the `final` completion, managed by
   a `bool` flag. For QEMU we can use built in primitives such as
   `qemu_bh_schedule` so that the notify is sent at the tail of
   processing the `iothread` events.

2. **Small (4k) Queue Depth 1 read/writes**. Analysis pointed at
   dataplane wakeup latency of ~57us being the culprit. Tapdisk
   mitigates this by implementing polling, for a defined time period,
   after receiving a request from the ring. QEMU contains built in
   support for equivalent polling which can be set a property on the
   `iothread`. However there is no plumbing for this in
   `xen_device_set_event_channel_context` as `io_poll_ready` is passed
   as NULL to `aio_set_fd_handler` so no polling occurs regardless of
   the value set for `poll-max-ns`. Adding an implementation for this
   goes most of the way to addressing this issue. Matching tapdisk we
   used 8,000,000, i.e. 8ms. This ensures that the poll window exceeds
   the backend service time. There should be no `evtchn` unmask in the
   `io_poll_end` as this will introduce a CPU spin in the poller.

3. **Overall performance deficit**. After addressing the above two
   issues, overall comparative performance was still found to be
   deficient relative to tapdisk. Analysis found that QEMU was making
   many more, much smaller, I/O requests to the block layer than
   tapdisk was for equivalent guest I/O patterns. This I/O pattern
   difference was traced to tapdisk performing merges of adjacent,
   same operation, I/O requests at the blkif level before submitting
   to the lower level format drivers. Equivalent behaviour can be
   achieved in QEMU using the `qemu_iovec` function family and issuing
   `blk_aio_pwritev`/`blk_aio_preadv` as necessary. Suitable
   completion fan-out is then needed to respond appropriately to the
   associated request IDs.

## Outcomes

After addressing the above items the resulting comparison table is as
follows, measured the same way on the same rig as the table above,
again 12 repeats, first discarded.

| block size | queue depth | read | write |
|:-----------|------------:|-----:|------:|
| 4k         |           1 | 1.07 |  1.07 |
| 64k        |           1 | 1.02 |  1.03 |
| 256k       |           1 | 0.93 |  1.04 |
| 1m         |           1 | 0.80 |  1.00 |
| 4m         |           1 | 1.03 |  1.01 |
| 4k         |          32 | 2.24 |  1.50 |
| 64k        |          32 | 1.05 |  1.00 |
| 256k       |          32 | 1.32 |  0.99 |
| 1m         |          32 | 1.17 |  0.99 |
| 4m         |          32 | 1.06 |  0.98 |

Median across the 20 cells moves from 0.86 to 1.03, and the number of
cells at or above parity from 5 of 20 to 13 of 20. The cells still
well below parity are 1m QD1 read at 0.80, 256k QD1 read at 0.93, the
other deficient cells are 1-2% down.

4k QD1 operations benefited most from the fixes for issue 2, voluntary
context switches per I/O went from 1.94 to 0.00, corresponding CPU
usage went from 26.5% to 96.7% which is a significant increase but it
was mostly doing useful work.

We believe that the final 3 rows of QD32 operations are probably I/O
bound on the iSCSI transport.

One cell deserves comment so it is not mistaken for one of our fixes.
4k QD32 sits well above parity at 2.24 read and 1.50 write, but it was
already at 1.90 and 1.32 before any of these changes; whatever causes
it is pre-existing and we have not investigated it.

XenServer has fixes for all of the above issues which are undergoing
further testing before release to customers. We can't send these to
qemu-devel as we have made extensive use of Claude Opus 5 in all
aspects of this work from defining the test environment, the test
strategy and automation to implement it, post-run analysis and gap
detection, and finally implementation of the solution fixes. All text
in this message is mine with the exception of the data tables.

We did feel though that it was only right that we made you aware of
the issues we have found so that you can make your own decisions about
them and whether to address them.

I'm happy to have any further conversation on this as you deem
necessary.

Regards,

Mark
XenServer Storage Engineering.

Reply via email to