All of lore.kernel.org
 help / color / mirror / Atom feed
* [PATCH v2 00/11] block: fix the zone write granularity and the zone append limit
@ 2026-09-02 19:44 Niklas Cassel
  2026-09-02 19:44 ` [PATCH v2 01/11] block: widen BlockLimits.zone_size to uint64_t Niklas Cassel
                   ` (10 more replies)
  0 siblings, 11 replies; 16+ messages in thread
From: Niklas Cassel @ 2026-09-02 19:44 UTC (permalink / raw)
  To: Stefan Hajnoczi, Kevin Wolf, Hanna Reitz, Fam Zheng, John Snow,
	Denis V. Lunev, Michael S. Tsirkin
  Cc: Sam Li, Damien Le Moal, Niklas Cassel, qemu-block, qemu-devel

Hello Stefan, Kevin, and everyone else,

This series fixes how QEMU reports and enforces the two constraints a zoned
device places on a write to a sequential zone: the write granularity, and
the largest zone append it accepts. It also fixes two bugs in the zone
append emulation in file-posix.

Many of these patches are in preparation for Sam Li's zoned qcow2 series.
The first two patches in the series are taken directly from there, as they
are unrelated to qcow2.

Patches 3 to 5 concern the write granularity. virtio-blk reported the
logical block size while the driver enforced the backend value, so on a
512e SMR disk a guest could be told that a request was valid and get an
I/O error for it. Writes to sequential zones were not checked against the
granularity at all, and zone appends had only their offset checked, not
their length. Patch 5 then refuses at realize a device whose write
pointers the configured logical block size cannot address, which needs no
emulated backend to provoke: write 512 bytes to a zone of a null_blk
device and attach it with logical_block_size=4096.

Patches 6 to 9 concern the append limit. The sector invariant moves to
bdrv_co_zone_append(), since a write pointer is tracked in sectors and
cannot represent anything finer, and file-posix drops its own check which
conflated that with the coarser granularity of the medium. file-posix then
stops reporting zone_append_max_bytes, which bounds REQ_OP_ZONE_APPEND, an
operation it never issues: it appends with an ordinary pwritev(), so
max_hw_transfer is the limit that applies. Finally virtio-blk derives what
it advertises rather than passing BlockLimits.max_append_sectors through,
which made an unset field mean "zone append unsupported" rather than "no
limit of its own", and Linux refuses to attach a zoned device that reports
zero.

Patches 10 and 11 fix the write pointer that raw_co_prw() substitutes for
the offset of an append. An offset that is never bounded against the
device derives an out of range zone index and reads past the write pointer
array, which qemu-io can reach. An append to a full zone uses a pointer
recorded at the end of the zone, so the data is written into the next zone
and success is returned; a guest can reach that one, because nothing in
virtio-blk checks whether a zone is full.


Changes since v1:
-Dropped patch 03/12 from v1, as it has landed in the QEMU master branch.
-Picked up Reviewed-by tags from Damien (thanks!)
-Addressed review comments from Damien.
-Improved commit message for patch 02/11.


Niklas Cassel (9):
  virtio-blk: report the effective zone write granularity
  virtio-blk: check the write granularity of writes to sequential zones
  hw/block: reject a zoned device whose write pointers are unaddressable
  block: reject zone appends that are not a multiple of the sector size
  file-posix: remove the zone append write granularity check
  file-posix: base the zone append limit on the transfer limit
  virtio-blk: derive the maximum zone append size
  file-posix: reject a zone append past the device capacity
  file-posix: reject a zone append to a full or conventional zone

Sam Li (2):
  block: widen BlockLimits.zone_size to uint64_t
  virtio-blk: do not merge requests across a zone boundary

 block/block-backend.c             |  11 ++++
 block/file-posix.c                |  65 +++++++++++-------
 block/io.c                        |  19 ++++++
 hw/block/block.c                  |  53 +++++++++++++++
 hw/block/virtio-blk.c             | 105 ++++++++++++++++++++++++++----
 include/block/block-io.h          |   6 ++
 include/block/block_int-common.h  |   2 +-
 include/hw/block/block.h          |   9 +++
 include/system/block-backend-io.h |   1 +
 9 files changed, 234 insertions(+), 37 deletions(-)

-- 
2.55.0



^ permalink raw reply	[flat|nested] 16+ messages in thread

end of thread, other threads:[~2026-09-04 16:06 UTC | newest]

Thread overview: 16+ messages (download: mbox.gz follow: Atom feed
-- links below jump to the message on this page --
2026-09-02 19:44 [PATCH v2 00/11] block: fix the zone write granularity and the zone append limit Niklas Cassel
2026-09-02 19:44 ` [PATCH v2 01/11] block: widen BlockLimits.zone_size to uint64_t Niklas Cassel
2026-09-02 19:44 ` [PATCH v2 02/11] virtio-blk: do not merge requests across a zone boundary Niklas Cassel
2026-09-02 19:44 ` [PATCH v2 03/11] virtio-blk: report the effective zone write granularity Niklas Cassel
2026-09-02 19:44 ` [PATCH v2 04/11] virtio-blk: check the write granularity of writes to sequential zones Niklas Cassel
2026-09-03  0:51   ` Damien Le Moal
2026-09-04 16:05     ` Niklas Cassel
2026-09-02 19:44 ` [PATCH v2 05/11] hw/block: reject a zoned device whose write pointers are unaddressable Niklas Cassel
2026-09-03  0:52   ` Damien Le Moal
2026-09-02 19:44 ` [PATCH v2 06/11] block: reject zone appends that are not a multiple of the sector size Niklas Cassel
2026-09-02 19:44 ` [PATCH v2 07/11] file-posix: remove the zone append write granularity check Niklas Cassel
2026-09-02 19:44 ` [PATCH v2 08/11] file-posix: base the zone append limit on the transfer limit Niklas Cassel
2026-09-02 19:44 ` [PATCH v2 09/11] virtio-blk: derive the maximum zone append size Niklas Cassel
2026-09-02 19:44 ` [PATCH v2 10/11] file-posix: reject a zone append past the device capacity Niklas Cassel
2026-09-02 19:44 ` [PATCH v2 11/11] file-posix: reject a zone append to a full or conventional zone Niklas Cassel
2026-09-04 15:06   ` Niklas Cassel

This is an external index of several public inboxes,
see mirroring instructions on how to clone and mirror
all data and code used by this external index.