The write pointers of a zoned device outlive any particular use of it,
whether the device keeps them itself or a backend records them, while the
logical block size is a property of the frontend and is chosen afresh
every time the device is attached. Nothing ties the two together. A zone
written while the device was configured with logical_block_size=512
leaves a write pointer that is a multiple of 512, and attaching the same
device with logical_block_size=4096 makes that pointer unaddressable.

Such a pointer is not merely misaligned. The guest addresses the device
in logical blocks, and a zone report expresses the write pointer in 512
byte sectors, so the guest is told about a position that does not fall on
a logical block boundary. It can neither read nor write there, and the
zone can only be recovered by resetting it. The reverse direction is
harmless: a pointer laid down with a larger logical block size is still a
multiple of a smaller one.

On a zoned null_blk device with a logical block size of 512, a 512 byte
append to a sequential zone leaves the write pointer half a logical block
into it:

  $ qemu-io --image-opts -n driver=host_device,filename=/dev/nullb0 \
        -c "zap -p 0x20000000 0x200" -c "zrp 0x20000000 1"
  start: 0x100000, len 0x80000, cap 0x80000, wptr 0x100001, zcond:2

Attaching that disk with logical_block_size=4096 handed the guest a zone
it could not write to.

Check at realize time that the zone size and every write pointer of a
sequential zone are multiples of the write granularity that the device is
about to report, and refuse to start otherwise. The write pointers are
already held in memory by the driver, so this costs no I/O.

The check uses blkconf_zone_write_granularity(), the same value that a
frontend reports to its guest and validates requests against, so the
three cannot disagree.

Signed-off-by: Niklas Cassel <[email protected]>
---
 hw/block/block.c         | 46 ++++++++++++++++++++++++++++++++++++++++
 hw/block/virtio-blk.c    |  4 ++++
 include/hw/block/block.h |  1 +
 3 files changed, 51 insertions(+)

diff --git a/hw/block/block.c b/hw/block/block.c
index 1c3135843d..b908e0ef64 100644
--- a/hw/block/block.c
+++ b/hw/block/block.c
@@ -208,6 +208,52 @@ uint32_t blkconf_zone_write_granularity(BlockConf *conf)
     return MAX(bs->bl.write_granularity, conf->logical_block_size);
 }
 
+bool blkconf_zoned(BlockConf *conf, Error **errp)
+{
+    BlockDriverState *bs = blk_bs(conf->blk);
+    uint32_t wg;
+
+    if (bs->bl.zoned == BLK_Z_NONE) {
+        return true;
+    }
+
+    wg = blkconf_zone_write_granularity(conf);
+
+    if (!QEMU_IS_ALIGNED(bs->bl.zone_size, wg)) {
+        error_setg(errp, "zone size %" PRIu64 " is not a multiple of the zone "
+                   "write granularity %" PRIu32, bs->bl.zone_size, wg);
+        return false;
+    }
+
+    /*
+     * A write pointer that is not a multiple of the write granularity does not
+     * fall on a logical block boundary, so the guest can neither read nor 
write
+     * at it and the zone can only be recovered by resetting it. A backend that
+     * records its write pointers, rather than reading them back from a device,
+     * can hand us such a pointer when the zones were written while the device
+     * was configured with a smaller logical block size.
+     */
+    for (uint32_t i = 0; i < bs->bl.nr_zones; i++) {
+        uint64_t wp = bs->wps->wp[i];
+
+        if (BDRV_ZT_IS_CONV(wp)) {
+            continue;
+        }
+
+        if (!QEMU_IS_ALIGNED(wp, wg)) {
+            error_setg(errp, "write pointer 0x%" PRIx64 " of zone %" PRIu32
+                       " is not a multiple of the zone write granularity %"
+                       PRIu32, wp, i, wg);
+            error_append_hint(errp, "The zones were written with a smaller "
+                              "logical_block_size. Reset them, or keep using "
+                              "the smaller size.\n");
+            return false;
+        }
+    }
+
+    return true;
+}
+
 bool blkconf_apply_backend_options(BlockConf *conf, bool readonly,
                                    bool resizable, Error **errp)
 {
diff --git a/hw/block/virtio-blk.c b/hw/block/virtio-blk.c
index 7977f4abe5..61f7b3cdc1 100644
--- a/hw/block/virtio-blk.c
+++ b/hw/block/virtio-blk.c
@@ -1821,6 +1821,10 @@ static void virtio_blk_device_realize(DeviceState *dev, 
Error **errp)
         return;
     }
 
+    if (!blkconf_zoned(&conf->conf, errp)) {
+        return;
+    }
+
     bs = blk_bs(conf->conf.blk);
     if (bs->bl.zoned != BLK_Z_NONE) {
         virtio_add_feature(&s->host_features, VIRTIO_BLK_F_ZONED);
diff --git a/include/hw/block/block.h b/include/hw/block/block.h
index f98525c01a..3c96de3222 100644
--- a/include/hw/block/block.h
+++ b/include/hw/block/block.h
@@ -122,6 +122,7 @@ bool blkconf_blocksizes(BlockConf *conf, Error **errp);
  * value to its guest and validate requests against it.
  */
 uint32_t blkconf_zone_write_granularity(BlockConf *conf);
+bool blkconf_zoned(BlockConf *conf, Error **errp);
 bool blkconf_apply_backend_options(BlockConf *conf, bool readonly,
                                    bool resizable, Error **errp);
 
-- 
2.55.0


Reply via email to