From: Niklas Cassel <cassel@kernel.org>
To: Stefan Hajnoczi <stefanha@redhat.com>,
Kevin Wolf <kwolf@redhat.com>, Hanna Reitz <hreitz@redhat.com>,
Fam Zheng <fam@euphon.net>
Cc: Sam Li <faithilikerun@gmail.com>,
Damien Le Moal <dlemoal@kernel.org>,
Niklas Cassel <cassel@kernel.org>,
qemu-block@nongnu.org, qemu-devel@nongnu.org
Subject: [PATCH v2 11/11] file-posix: reject a zone append to a full or conventional zone
Date: Wed, 2 Sep 2026 21:44:22 +0200 [thread overview]
Message-ID: <20260902194423.759355-12-cassel@kernel.org> (raw)
In-Reply-To: <20260902194423.759355-1-cassel@kernel.org>
raw_co_prw() replaces the offset of a zone append with the write pointer
of the addressed zone, which assumes that the stored value names a
position inside that zone. It does not in two cases.
A full zone has its write pointer recorded at the end of the zone, since
get_zones_wp() stores start + len for BLK_ZONE_COND_FULL. That is the
first sector of the following zone, so the append is submitted there. The
kernel accepts it whenever that zone is empty, because it is a legal
write at its write pointer, and the completion path advances the wrong
zone because it recomputes the zone index from the replaced offset. The
data is written to a zone that was never addressed and success is
returned:
zone 2 finished, then a 4 KiB append to zone 2:
After zap done, the append sector is 0x180000 <- zone 3
zone 2: wptr 0x180000, zcond:14 (full)
zone 3: wptr 0x180008 <- advanced
A conventional zone has no write pointer at all, and its array entry
carries only the type marker in the top bit, so the offset becomes
negative and the write fails with EINVAL. That is harmless but it reports
nothing about the actual mistake.
Reject both while the write pointer lock is held, since the state has to
be read and acted on atomically. check_zoned_request() in virtio-blk
refuses an append to a conventional zone, so that case needs a caller
that goes to the driver directly, but nothing there examines whether a
zone is full, so a guest can reach the misdirected write.
Fixes: 4751d09adcc3 ("block: introduce zone append write for zoned devices")
Signed-off-by: Niklas Cassel <cassel@kernel.org>
---
block/file-posix.c | 27 +++++++++++++++++++++++++--
block/io.c | 9 +++++++++
include/block/block-io.h | 6 ++++++
3 files changed, 40 insertions(+), 2 deletions(-)
diff --git a/block/file-posix.c b/block/file-posix.c
index 85d735c079..57ba8dde64 100644
--- a/block/file-posix.c
+++ b/block/file-posix.c
@@ -2565,8 +2565,31 @@ raw_co_prw(BlockDriverState *bs, int64_t *offset_ptr, uint64_t bytes,
bs->bl.zoned != BLK_Z_NONE) {
qemu_co_mutex_lock(&bs->wps->colock);
if (type & QEMU_AIO_ZONE_APPEND) {
- int index = offset / bs->bl.zone_size;
- offset = bs->wps->wp[index];
+ uint32_t index = offset / bs->bl.zone_size;
+ uint64_t wp = bs->wps->wp[index];
+
+ /*
+ * The write pointer of the addressed zone becomes the offset of
+ * the write, so it has to name a position inside that zone. It
+ * does not for a conventional zone, which has no write pointer and
+ * stores a type marker in the top bit instead, and it does not for
+ * a full zone, whose write pointer is reported at the zone end.
+ * Either would send the data to a zone that was never addressed.
+ */
+ if (BDRV_ZT_IS_CONV(wp)) {
+ error_report("zone append at offset 0x%" PRIx64 " addresses a "
+ "conventional zone", offset);
+ qemu_co_mutex_unlock(&bs->wps->colock);
+ return -EINVAL;
+ }
+ if (bdrv_zone_is_full(bs, index)) {
+ error_report("zone append at offset 0x%" PRIx64 " addresses a "
+ "full zone", offset);
+ qemu_co_mutex_unlock(&bs->wps->colock);
+ return -ENOSPC;
+ }
+
+ offset = wp;
}
}
#endif
diff --git a/block/io.c b/block/io.c
index 705dc73d76..3574e15a55 100644
--- a/block/io.c
+++ b/block/io.c
@@ -3371,6 +3371,15 @@ out:
return co.ret;
}
+bool bdrv_zone_is_full(BlockDriverState *bs, uint32_t index)
+{
+ uint64_t zone_end = MIN((uint64_t)(index + 1) * bs->bl.zone_size,
+ (uint64_t)bs->total_sectors << BDRV_SECTOR_BITS);
+ IO_CODE();
+
+ return bs->wps->wp[index] >= zone_end;
+}
+
void *qemu_blockalign(BlockDriverState *bs, size_t size)
{
IO_CODE();
diff --git a/include/block/block-io.h b/include/block/block-io.h
index d34d846bb2..2a9505312a 100644
--- a/include/block/block-io.h
+++ b/include/block/block-io.h
@@ -126,6 +126,12 @@ int coroutine_fn GRAPH_RDLOCK bdrv_co_zone_append(BlockDriverState *bs,
int64_t *offset,
QEMUIOVector *qiov,
BdrvRequestFlags flags);
+/*
+ * True when the write pointer of a zone has reached the end of the writable
+ * part of that zone, so that nothing more can be written to it until it is
+ * reset. The write pointer lock must be held when called.
+ */
+bool bdrv_zone_is_full(BlockDriverState *bs, uint32_t index);
bool bdrv_can_write_zeroes_with_unmap(BlockDriverState *bs);
--
2.55.0
next prev parent reply other threads:[~2026-09-02 19:46 UTC|newest]
Thread overview: 16+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-09-02 19:44 [PATCH v2 00/11] block: fix the zone write granularity and the zone append limit Niklas Cassel
2026-09-02 19:44 ` [PATCH v2 01/11] block: widen BlockLimits.zone_size to uint64_t Niklas Cassel
2026-09-02 19:44 ` [PATCH v2 02/11] virtio-blk: do not merge requests across a zone boundary Niklas Cassel
2026-09-02 19:44 ` [PATCH v2 03/11] virtio-blk: report the effective zone write granularity Niklas Cassel
2026-09-02 19:44 ` [PATCH v2 04/11] virtio-blk: check the write granularity of writes to sequential zones Niklas Cassel
2026-09-03 0:51 ` Damien Le Moal
2026-09-04 16:05 ` Niklas Cassel
2026-09-02 19:44 ` [PATCH v2 05/11] hw/block: reject a zoned device whose write pointers are unaddressable Niklas Cassel
2026-09-03 0:52 ` Damien Le Moal
2026-09-02 19:44 ` [PATCH v2 06/11] block: reject zone appends that are not a multiple of the sector size Niklas Cassel
2026-09-02 19:44 ` [PATCH v2 07/11] file-posix: remove the zone append write granularity check Niklas Cassel
2026-09-02 19:44 ` [PATCH v2 08/11] file-posix: base the zone append limit on the transfer limit Niklas Cassel
2026-09-02 19:44 ` [PATCH v2 09/11] virtio-blk: derive the maximum zone append size Niklas Cassel
2026-09-02 19:44 ` [PATCH v2 10/11] file-posix: reject a zone append past the device capacity Niklas Cassel
2026-09-02 19:44 ` Niklas Cassel [this message]
2026-09-04 15:06 ` [PATCH v2 11/11] file-posix: reject a zone append to a full or conventional zone Niklas Cassel
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=20260902194423.759355-12-cassel@kernel.org \
--to=cassel@kernel.org \
--cc=dlemoal@kernel.org \
--cc=faithilikerun@gmail.com \
--cc=fam@euphon.net \
--cc=hreitz@redhat.com \
--cc=kwolf@redhat.com \
--cc=qemu-block@nongnu.org \
--cc=qemu-devel@nongnu.org \
--cc=stefanha@redhat.com \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is an external index of several public inboxes,
see mirroring instructions on how to clone and mirror
all data and code used by this external index.