From: Niklas Cassel <cassel@kernel.org>
To: Stefan Hajnoczi <stefanha@redhat.com>,
Kevin Wolf <kwolf@redhat.com>, Hanna Reitz <hreitz@redhat.com>
Cc: Sam Li <faithilikerun@gmail.com>,
Damien Le Moal <dlemoal@kernel.org>,
Niklas Cassel <cassel@kernel.org>,
qemu-block@nongnu.org, qemu-devel@nongnu.org
Subject: [PATCH 09/12] file-posix: base the zone append limit on the transfer limit
Date: Tue, 25 Aug 2026 22:57:44 +0200 [thread overview]
Message-ID: <20260825205748.679968-10-cassel@kernel.org> (raw)
In-Reply-To: <20260825205748.679968-1-cassel@kernel.org>
The zone_append_max_bytes queue attribute is the largest
REQ_OP_ZONE_APPEND that the device accepts, and Linux has no interface
for issuing one from userspace: include/uapi has no such operation, and
every submitter of REQ_OP_ZONE_APPEND is in the kernel. Userspace writes
to a sequential zone with an ordinary write at the write pointer.
That is what this driver does. raw_co_zone_append() substitutes the write
pointer of the zone for the offset and hands the request to raw_co_prw(),
which reaches handle_aiocb_rw_vector() and issues a plain pwritev(). The
kernel never sees a zone append, so the attribute describes a limit on an
operation that is never issued.
Nor is it a limit that this driver runs into. The kernel splits a write
that exceeds the transfer limit rather than refusing it, so a larger
append succeeds: on a null_blk device whose zone_append_max_bytes is
130560, a 16 MiB append completes and advances the write pointer by
16 MiB. Reporting the attribute only understates what the driver can do,
because Linux derives it as a minimum that already includes max_sectors
and chunk_sectors.
Report max_hw_transfer instead, the limit that governs the write the
driver actually issues. There is no use in telling a guest that it may
append more than the device carries in one command, and a frontend then
does not have to reason about how this driver implements an append in
order to bound the value it advertises.
On a null_blk device with max_hw_sectors_kb of 127 and
zone_append_max_bytes of 130560, virtio-blk reports 255 sectors before
and after.
Signed-off-by: Niklas Cassel <cassel@kernel.org>
---
block/file-posix.c | 13 +++++++++----
1 file changed, 9 insertions(+), 4 deletions(-)
diff --git a/block/file-posix.c b/block/file-posix.c
index f267513a4e..e019cc3cd8 100644
--- a/block/file-posix.c
+++ b/block/file-posix.c
@@ -1488,10 +1488,15 @@ static void raw_refresh_zoned_limits(BlockDriverState *bs, struct stat *st,
}
bs->bl.nr_zones = ret;
- ret = get_sysfs_long_val(st, "zone_append_max_bytes");
- if (ret > 0) {
- bs->bl.max_append_sectors = ret >> BDRV_SECTOR_BITS;
- }
+ /*
+ * raw_co_zone_append() carries out an append as an ordinary write at the
+ * write pointer, so the zone_append_max_bytes attribute, which bounds an
+ * operation that this driver never issues, does not apply. The kernel
+ * splits a write that is larger than the transfer limit rather than
+ * refusing it, but there is no use in telling a guest that it may append
+ * more than the device carries in one command.
+ */
+ bs->bl.max_append_sectors = bs->bl.max_hw_transfer >> BDRV_SECTOR_BITS;
ret = get_sysfs_long_val(st, "zone_write_granularity");
if (ret >= 0) {
--
2.55.0
next prev parent reply other threads:[~2026-08-25 20:59 UTC|newest]
Thread overview: 25+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-08-25 20:57 [PATCH 00/12] block: fix the zone write granularity and the zone append limit Niklas Cassel
2026-08-25 20:57 ` [PATCH 01/12] block: widen BlockLimits.zone_size to uint64_t Niklas Cassel
2026-08-28 5:58 ` Damien Le Moal
2026-08-25 20:57 ` [PATCH 02/12] virtio-blk: do not merge writes across a zone boundary Niklas Cassel
2026-08-28 6:02 ` Damien Le Moal
2026-08-25 20:57 ` [PATCH 03/12] file-posix: fix zone write granularity assignment for zoned block devices Niklas Cassel
2026-08-25 20:57 ` [PATCH 04/12] virtio-blk: report the effective zone write granularity Niklas Cassel
2026-08-28 6:03 ` Damien Le Moal
2026-08-25 20:57 ` [PATCH 05/12] virtio-blk: check the write granularity of writes to sequential zones Niklas Cassel
2026-08-28 6:07 ` Damien Le Moal
2026-08-28 6:09 ` Damien Le Moal
2026-08-25 20:57 ` [PATCH 06/12] hw/block: reject a zoned device whose write pointers are unaddressable Niklas Cassel
2026-08-28 6:12 ` Damien Le Moal
2026-08-25 20:57 ` [PATCH 07/12] block: reject zone appends that are not a multiple of the sector size Niklas Cassel
2026-08-28 6:13 ` Damien Le Moal
2026-08-25 20:57 ` [PATCH 08/12] file-posix: remove the zone append write granularity check Niklas Cassel
2026-08-28 6:14 ` Damien Le Moal
2026-08-25 20:57 ` Niklas Cassel [this message]
2026-08-28 6:17 ` [PATCH 09/12] file-posix: base the zone append limit on the transfer limit Damien Le Moal
2026-08-25 20:57 ` [PATCH 10/12] virtio-blk: derive the maximum zone append size Niklas Cassel
2026-08-28 6:18 ` Damien Le Moal
2026-08-25 20:57 ` [PATCH 11/12] file-posix: reject a zone append past the device capacity Niklas Cassel
2026-08-28 6:18 ` Damien Le Moal
2026-08-25 20:57 ` [PATCH 12/12] file-posix: reject a zone append to a full or conventional zone Niklas Cassel
2026-08-28 6:19 ` Damien Le Moal
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=20260825205748.679968-10-cassel@kernel.org \
--to=cassel@kernel.org \
--cc=dlemoal@kernel.org \
--cc=faithilikerun@gmail.com \
--cc=hreitz@redhat.com \
--cc=kwolf@redhat.com \
--cc=qemu-block@nongnu.org \
--cc=qemu-devel@nongnu.org \
--cc=stefanha@redhat.com \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is an external index of several public inboxes,
see mirroring instructions on how to clone and mirror
all data and code used by this external index.