The write granularity that a frontend reports is derived from the logical and the physical block size, which are properties of the device model, while the granularity that the medium requires and the write pointers that it holds belong to the backend. Nothing ties them together, and where they disagree a guest is told about a device it cannot use.
A backend can require a coarser granularity than the block sizes express. A 512e host managed disk has a logical block size of 512 and requires writes to a sequential zone to be a multiple of 4096, so attaching it with a finer physical block size reports a granularity that the disk will not accept, and the guest is given a plain I/O error for a request that it was told was valid. The write pointers can disagree in the same way. They outlive any particular use of the device, while the block sizes are chosen afresh every time it is attached, so a zone written with logical_block_size=512 leaves a write pointer that logical_block_size=4096 cannot address. A zone report expresses it in 512 byte sectors, so the guest is told about a position that does not fall on a logical block boundary, which it can neither read nor write, and the zone can then only be recovered by resetting it. The reverse is harmless: a pointer laid down with a larger logical block size is still a multiple of a smaller one. Check both at realize time, along with the zone size and the zone capacity, and refuse to start otherwise. The write pointers are already held in memory by the driver, so this costs no I/O. The check uses blkconf_zone_write_granularity(), the same value that a frontend reports to its guest and validates requests against, so the three cannot disagree. Signed-off-by: Niklas Cassel <[email protected]> --- hw/block/block.c | 63 ++++++++++++++++++++++++++++++++++++++++ hw/block/virtio-blk.c | 4 +++ include/hw/block/block.h | 1 + 3 files changed, 68 insertions(+) diff --git a/hw/block/block.c b/hw/block/block.c index 79b0a67cea..3bb8121e9e 100644 --- a/hw/block/block.c +++ b/hw/block/block.c @@ -206,6 +206,69 @@ uint32_t blkconf_zone_write_granularity(BlockConf *conf) return MAX(conf->logical_block_size, conf->physical_block_size); } +bool blkconf_check_zoned_geometry(BlockConf *conf, Error **errp) +{ + BlockDriverState *bs = blk_bs(conf->blk); + uint32_t wg; + + if (bs->bl.zoned == BLK_Z_NONE) { + return true; + } + + wg = blkconf_zone_write_granularity(conf); + + /* + * A backend that requires a coarser granularity than the one that is about + * to be reported would reject writes that the guest has been told are + * valid, leaving it with a plain I/O error. + */ + if (bs->bl.write_granularity > wg) { + error_setg(errp, "the backend requires writes to sequential zones to " + "be a multiple of %" PRIu32 " bytes, which is coarser than " + "the zone write granularity %" PRIu32 " that would be " + "reported", bs->bl.write_granularity, wg); + error_append_hint(errp, "Set physical_block_size=%" PRIu32 ".\n", + bs->bl.write_granularity); + return false; + } + + if (!QEMU_IS_ALIGNED(bs->bl.zone_size, wg)) { + error_setg(errp, "zone size %" PRIu64 " is not a multiple of the zone " + "write granularity %" PRIu32, bs->bl.zone_size, wg); + return false; + } + + /* + * A write pointer that is not a multiple of the write granularity does not + * fall on a logical block boundary, so the guest can neither read nor write + * at it and the zone can only be recovered by resetting it. A backend that + * records its write pointers, rather than reading them back from a device, + * can hand us such a pointer when the zones were written while the device + * was configured with a smaller logical block size. + */ + for (uint32_t i = 0; i < bs->bl.nr_zones; i++) { + uint64_t wp; + + if (bdrv_zone_is_conv(bs, i)) { + continue; + } + + wp = bs->wps->wp[i]; + if (!QEMU_IS_ALIGNED(wp, wg)) { + error_setg(errp, "write pointer 0x%" PRIx64 " of zone %" PRIu32 + " is not a multiple of the zone write granularity %" + PRIu32, wp, i, wg); + error_append_hint(errp, "The zones were written with a finer zone " + "write granularity. Reset them, or attach the " + "device with the block sizes that they were " + "written with.\n"); + return false; + } + } + + return true; +} + bool blkconf_apply_backend_options(BlockConf *conf, bool readonly, bool resizable, Error **errp) { diff --git a/hw/block/virtio-blk.c b/hw/block/virtio-blk.c index 449de307b4..74bd08b0ec 100644 --- a/hw/block/virtio-blk.c +++ b/hw/block/virtio-blk.c @@ -1840,6 +1840,10 @@ static void virtio_blk_device_realize(DeviceState *dev, Error **errp) return; } + if (!blkconf_check_zoned_geometry(&conf->conf, errp)) { + return; + } + bs = blk_bs(conf->conf.blk); if (bs->bl.zoned != BLK_Z_NONE) { virtio_add_feature(&s->host_features, VIRTIO_BLK_F_ZONED); diff --git a/include/hw/block/block.h b/include/hw/block/block.h index c769b1297d..c0a8b84df1 100644 --- a/include/hw/block/block.h +++ b/include/hw/block/block.h @@ -127,6 +127,7 @@ bool blkconf_blocksizes(BlockConf *conf, Error **errp); * physical block size, so this is the physical block size in practice. */ uint32_t blkconf_zone_write_granularity(BlockConf *conf); +bool blkconf_check_zoned_geometry(BlockConf *conf, Error **errp); bool blkconf_apply_backend_options(BlockConf *conf, bool readonly, bool resizable, Error **errp); -- 2.55.0
