Stefano, Anthony, Paul, Hi guys, hope this finds you all well. We, in XenServer, have been doing some back to back Xen VM PV storage backend performance comparisons between tapdisk + VHD and QEMU + QCOW2. The image storage format appears to be largely irrelevant to the rest of this conversation. In doing so we have found a few significant issues in QEMU, mainly around the xen-block dataplane and adjacent code.
## Testing Environment Debian 12 VM, 8 vCPUS, 4G Ram on XenServer 9. Three virtual block devices, one for the OS on host local storage and two, unformatted, 20G test devices, one backed by tapdisk and one backed by QEMU. These disks themselves are then backed by a BRD based target hosted on a native installed Rocky 9.8 host. The two hosts are connected over a 25G, 9000 MTU, switched storage network whose switch is rated to run all ports at wire speed without contention. In the guest `xen_blkfront` is configured with `max_ring_page_order=3`, which is honoured by both tapdisk and QEMU. QEMU supports 4 but tapdisk caps at 3 so for consistency 3 was used. Performance data was collected by `fio` v3.33 from the default Debian 12 repo. Note: we did find during this that QEMU has a separate problem with feature publication when hot plugging a device - https://lists.gnu.org/archive/html/qemu-devel/2026-09/msg03492.html This was worked around by always cold booting the VM which then negotiates correctly. ## Results The following table is the full grid from the testing performed. Points are expressed as ratios of QEMU/tapdisk so any points < 1.0 show QEMU having a performance deficit. Each figure is the median of the per-repeat ratio over 12 repeats, the first discarded as unstable on a newly booted VM. The tapdisk arm is measured seconds away from the QEMU arm in every repeat, so that drift in the rig cancels rather than being attributed to either datapath. | block size | queue depth | read | write | |:-----------|------------:|-----:|------:| | 4k | 1 | 0.73 | 0.73 | | 64k | 1 | 0.81 | 0.83 | | 256k | 1 | 0.67 | 0.92 | | 1m | 1 | 0.67 | 0.76 | | 4m | 1 | 0.81 | 1.01 | | 4k | 32 | 1.90 | 1.32 | | 64k | 32 | 0.99 | 1.00 | | 256k | 32 | 1.14 | 0.84 | | 1m | 32 | 0.86 | 0.88 | | 4m | 32 | 0.87 | 0.86 | Median across the 20 cells is 0.86, with 5 of 20 at or above parity. The deficit is concentrated at queue depth 1, which is what pointed at per-request latency rather than throughput. Note that 1m QD1 read is the noisiest cell in the grid at 24% coefficient of variation across repeats, so treat 0.67 as "much worse" rather than as a precise figure; the 4k and 64k QD1 cells sit at 3-4% and are solid. 1m QD1 is noisy on both backends and is probably an artifact of the environment. ## Issues found 1. **Large (4m) blocksize read**. Analysis determined that this was impacted by QEMU being overly keen in notifying the frontend of completion responses being placed into the ring. In turn this results in a Dom0 to DomU context switch and all the associated hypervisor flushes and mitigations. This can also be observed in the `ctx` count tracker in `fio` itself. Tapdisk mitigates this by only sending the notification on the `final` completion, managed by a `bool` flag. For QEMU we can use built in primitives such as `qemu_bh_schedule` so that the notify is sent at the tail of processing the `iothread` events. 2. **Small (4k) Queue Depth 1 read/writes**. Analysis pointed at dataplane wakeup latency of ~57us being the culprit. Tapdisk mitigates this by implementing polling, for a defined time period, after receiving a request from the ring. QEMU contains built in support for equivalent polling which can be set a property on the `iothread`. However there is no plumbing for this in `xen_device_set_event_channel_context` as `io_poll_ready` is passed as NULL to `aio_set_fd_handler` so no polling occurs regardless of the value set for `poll-max-ns`. Adding an implementation for this goes most of the way to addressing this issue. Matching tapdisk we used 8,000,000, i.e. 8ms. This ensures that the poll window exceeds the backend service time. There should be no `evtchn` unmask in the `io_poll_end` as this will introduce a CPU spin in the poller. 3. **Overall performance deficit**. After addressing the above two issues, overall comparative performance was still found to be deficient relative to tapdisk. Analysis found that QEMU was making many more, much smaller, I/O requests to the block layer than tapdisk was for equivalent guest I/O patterns. This I/O pattern difference was traced to tapdisk performing merges of adjacent, same operation, I/O requests at the blkif level before submitting to the lower level format drivers. Equivalent behaviour can be achieved in QEMU using the `qemu_iovec` function family and issuing `blk_aio_pwritev`/`blk_aio_preadv` as necessary. Suitable completion fan-out is then needed to respond appropriately to the associated request IDs. ## Outcomes After addressing the above items the resulting comparison table is as follows, measured the same way on the same rig as the table above, again 12 repeats, first discarded. | block size | queue depth | read | write | |:-----------|------------:|-----:|------:| | 4k | 1 | 1.07 | 1.07 | | 64k | 1 | 1.02 | 1.03 | | 256k | 1 | 0.93 | 1.04 | | 1m | 1 | 0.80 | 1.00 | | 4m | 1 | 1.03 | 1.01 | | 4k | 32 | 2.24 | 1.50 | | 64k | 32 | 1.05 | 1.00 | | 256k | 32 | 1.32 | 0.99 | | 1m | 32 | 1.17 | 0.99 | | 4m | 32 | 1.06 | 0.98 | Median across the 20 cells moves from 0.86 to 1.03, and the number of cells at or above parity from 5 of 20 to 13 of 20. The cells still well below parity are 1m QD1 read at 0.80, 256k QD1 read at 0.93, the other deficient cells are 1-2% down. 4k QD1 operations benefited most from the fixes for issue 2, voluntary context switches per I/O went from 1.94 to 0.00, corresponding CPU usage went from 26.5% to 96.7% which is a significant increase but it was mostly doing useful work. We believe that the final 3 rows of QD32 operations are probably I/O bound on the iSCSI transport. One cell deserves comment so it is not mistaken for one of our fixes. 4k QD32 sits well above parity at 2.24 read and 1.50 write, but it was already at 1.90 and 1.32 before any of these changes; whatever causes it is pre-existing and we have not investigated it. XenServer has fixes for all of the above issues which are undergoing further testing before release to customers. We can't send these to qemu-devel as we have made extensive use of Claude Opus 5 in all aspects of this work from defining the test environment, the test strategy and automation to implement it, post-run analysis and gap detection, and finally implementation of the solution fixes. All text in this message is mine with the exception of the data tables. We did feel though that it was only right that we made you aware of the issues we have found so that you can make your own decisions about them and whether to address them. I'm happy to have any further conversation on this as you deem necessary. Regards, Mark XenServer Storage Engineering.
