All of lore.kernel.org
 help / color / mirror / Atom feed
From: Niklas Cassel <cassel@kernel.org>
To: Damien Le Moal <dlemoal@kernel.org>
Cc: Stefan Hajnoczi <stefanha@redhat.com>,
	Kevin Wolf <kwolf@redhat.com>,
	"Michael S. Tsirkin" <mst@redhat.com>,
	Hanna Reitz <hreitz@redhat.com>, Sam Li <faithilikerun@gmail.com>,
	qemu-block@nongnu.org, qemu-devel@nongnu.org
Subject: Re: [PATCH v2 04/11] virtio-blk: check the write granularity of writes to sequential zones
Date: Fri, 4 Sep 2026 18:05:06 +0200	[thread overview]
Message-ID: <aprsMv0_DgfmMf9M@ryzen> (raw)
In-Reply-To: <d4b0f1a3-02b5-417e-928a-808b065e7cc2@kernel.org>

On Thu, Sep 03, 2026 at 09:51:01AM +0900, Damien Le Moal wrote:
> On 9/3/26 04:44, Niklas Cassel wrote:
> > All VIRTIO_BLK_T_OUT requests issued to sequential zones and all
> > VIRTIO_BLK_T_ZONE_APPEND requests must have an offset and a data size
> > that are multiples of the write granularity reported by the device
> > (virtio 1.4, 5.2.6.1), and a violation is reported as
> > VIRTIO_BLK_S_ZONE_UNALIGNED_WP (virtio 1.4, 5.2.6).
> > 
> > Neither request type was fully checked. Zone appends validated only the
> > offset, while writes were not checked at all.
> > 
> > Check the size of the appended data, and both the offset and the size of
> > a write, against blkconf_zone_write_granularity(), so that every request the
> > device accepts is one that the guest driver was told is valid. Writes to
> > conventional zones keep no alignment constraint beyond the logical block
> > size. The write path performs the check after virtio_blk_sect_range_ok()
> > so that the zone index derived from the guest supplied sector is known to
> > be in range.
> > 
> > Signed-off-by: Niklas Cassel <cassel@kernel.org>
> > ---
> >  hw/block/virtio-blk.c | 37 ++++++++++++++++++++++++++++++++++++-
> >  1 file changed, 36 insertions(+), 1 deletion(-)
> > 
> > diff --git a/hw/block/virtio-blk.c b/hw/block/virtio-blk.c
> > index b22aed04a1..247af53ae2 100644
> > --- a/hw/block/virtio-blk.c
> > +++ b/hw/block/virtio-blk.c
> > @@ -398,6 +398,32 @@ static bool virtio_blk_sect_range_ok(VirtIOBlock *dev,
> >      return true;
> >  }
> >  
> > +/*
> > + * Both the offset and the size of a write to a sequential zone must be a
> > + * multiple of the write granularity that the device reports. Conventional
> > + * zones are not constrained. The zone index is derived from a guest supplied
> > + * sector, so this must only be called once virtio_blk_sect_range_ok() has
> > + * bounded it.
> > + */
> > +static bool virtio_blk_zone_write_granularity_ok(VirtIOBlock *dev,
> > +                                                 uint64_t sector, size_t size)
> > +{
> > +    BlockDriverState *bs = blk_bs(dev->blk);
> > +    uint64_t offset = sector << BDRV_SECTOR_BITS;
> > +    uint32_t wg_mask;
> > +
> > +    if (bs->bl.zoned == BLK_Z_NONE) {
> > +        return true;
> > +    }
> > +
> > +    wg_mask = blkconf_zone_write_granularity(&dev->conf.conf) - 1;
> > +    if (!(offset & wg_mask) && !(size & wg_mask)) {
> > +        return true;
> > +    }
> > +
> > +    return BDRV_ZT_IS_CONV(bs->wps->wp[offset / bs->bl.zone_size]);
> > +}
> 
> Even though in practice I do not think it would ever happen, technically,
> conventional zones are not bound by the physical sector size/write granularity
> and can accept LBA aligned writes even when LBA size < physical block size
> (write granularity).

Well, that is how the code works already :)

And the function comment does also already say that conventional
zones are not constrained.

Yes, I guess swapping the order does make the code slightly easier to read,
so let me do that.


> 
> SO I think this should be:
> 
> +static bool virtio_blk_zone_write_granularity_ok(VirtIOBlock *dev,
> +                                                 uint64_t sector, size_t size)
> +{
> +    BlockDriverState *bs = blk_bs(dev->blk);
> +    uint64_t offset = sector << BDRV_SECTOR_BITS;
> +    uint32_t wg_mask;
> +
> +    if (bs->bl.zoned == BLK_Z_NONE) {
> +        return true;
> +    }
> +
> +    if (BDRV_ZT_IS_CONV(bs->wps->wp[offset / bs->bl.zone_size]) {
> +        return true;
> +    }
> +
> +    wg_mask = blkconf_zone_write_granularity(&dev->conf.conf) - 1;
> +    return !(offset & wg_mask) && !(size & wg_mask);
> +}
> 
> Side note not for this patch: it really would be nice to have a helper that
> gives a zone number from an offset, using a bit shift (bs->bl.zone_size_shift)
> instead of a costly division.

Will add a helper as a new patch.


Kind regards,
Niklas


  reply	other threads:[~2026-09-04 16:06 UTC|newest]

Thread overview: 16+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-09-02 19:44 [PATCH v2 00/11] block: fix the zone write granularity and the zone append limit Niklas Cassel
2026-09-02 19:44 ` [PATCH v2 01/11] block: widen BlockLimits.zone_size to uint64_t Niklas Cassel
2026-09-02 19:44 ` [PATCH v2 02/11] virtio-blk: do not merge requests across a zone boundary Niklas Cassel
2026-09-02 19:44 ` [PATCH v2 03/11] virtio-blk: report the effective zone write granularity Niklas Cassel
2026-09-02 19:44 ` [PATCH v2 04/11] virtio-blk: check the write granularity of writes to sequential zones Niklas Cassel
2026-09-03  0:51   ` Damien Le Moal
2026-09-04 16:05     ` Niklas Cassel [this message]
2026-09-02 19:44 ` [PATCH v2 05/11] hw/block: reject a zoned device whose write pointers are unaddressable Niklas Cassel
2026-09-03  0:52   ` Damien Le Moal
2026-09-02 19:44 ` [PATCH v2 06/11] block: reject zone appends that are not a multiple of the sector size Niklas Cassel
2026-09-02 19:44 ` [PATCH v2 07/11] file-posix: remove the zone append write granularity check Niklas Cassel
2026-09-02 19:44 ` [PATCH v2 08/11] file-posix: base the zone append limit on the transfer limit Niklas Cassel
2026-09-02 19:44 ` [PATCH v2 09/11] virtio-blk: derive the maximum zone append size Niklas Cassel
2026-09-02 19:44 ` [PATCH v2 10/11] file-posix: reject a zone append past the device capacity Niklas Cassel
2026-09-02 19:44 ` [PATCH v2 11/11] file-posix: reject a zone append to a full or conventional zone Niklas Cassel
2026-09-04 15:06   ` Niklas Cassel

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=aprsMv0_DgfmMf9M@ryzen \
    --to=cassel@kernel.org \
    --cc=dlemoal@kernel.org \
    --cc=faithilikerun@gmail.com \
    --cc=hreitz@redhat.com \
    --cc=kwolf@redhat.com \
    --cc=mst@redhat.com \
    --cc=qemu-block@nongnu.org \
    --cc=qemu-devel@nongnu.org \
    --cc=stefanha@redhat.com \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is an external index of several public inboxes,
see mirroring instructions on how to clone and mirror
all data and code used by this external index.