From: "Denis V. Lunev" <den@openvz.org>
To: qemu-devel@nongnu.org
Cc: qemu-block@nongnu.org, "Denis V. Lunev" <den@openvz.org>,
Vladimir Sementsov-Ogievskiy <vsementsov@yandex-team.ru>,
John Snow <jsnow@redhat.com>,
Andrey Drobyshev <andrey.drobyshev@virtuozzo.com>
Subject: [PATCH v3 9/9] block/block-copy: coalesce write-zeroes tasks
Date: Tue, 29 Sep 2026 17:51:25 +0200 [thread overview]
Message-ID: <20260929155125.3151111-10-den@openvz.org> (raw)
In-Reply-To: <20260929155125.3151111-1-den@openvz.org>
From: Denis V. Lunev <den@openvz.org>
Task size was capped to block_copy_chunk_size() regardless of method,
which sizes a task by the copy buffer. A write-zeroes task carries no
buffer, so that cap only splits one long known-zero run into a crowd of
small COPY_WRITE_ZEROES tasks, each with its own request against the
target.
block_copy_task_create() already decides from zero_bitmap how far a task
may run, so add block_copy_widen_zero_area() to that decision: for a
write-zeroes task, re-search the dirty area under BDRV_REQUEST_MAX_BYTES
instead of the buffer chunk size, and let the existing clamp cut it back
to where the zero run ends. That bound is INT_MAX rounded down to a
sector, so nothing has aligned it to cluster_size and it is aligned down
at the use site. Widening is safe: an overlapping caller waits on the
task's BlockReq via reqlist_wait_one() rather than observing it
mid-flight.
A request that large is only cheap where the target zeroes by metadata.
BDRV_REQ_NO_FALLBACK in supported_zero_flags rules out the targets which
cannot, qcow2 v2 and iscsi among them, but it is optimistic for the rest:
file-posix advertises it at open and only learns from a failing
fallocate. So a write-zeroes request asks for it while the target is
believed to oblige, which costs nothing when it does. A refusal ends the
widening, and the range is written in buffer-sized chunks, as is any
task widened before the refusal came. -ENOTSUP says the fast path is
unavailable, not that the target wrote nothing, so the range is
rewritten whole: zeroes over zeroes change nothing.
Signed-off-by: Denis V. Lunev <den@openvz.org>
CC: Vladimir Sementsov-Ogievskiy <vsementsov@yandex-team.ru>
CC: John Snow <jsnow@redhat.com>
CC: Andrey Drobyshev <andrey.drobyshev@virtuozzo.com>
---
block/block-copy.c | 94 ++++++++++++++++++++++++++++++++++++++++++----
1 file changed, 87 insertions(+), 7 deletions(-)
diff --git a/block/block-copy.c b/block/block-copy.c
index c0a3359938..94ca2a10ed 100644
--- a/block/block-copy.c
+++ b/block/block-copy.c
@@ -159,6 +159,8 @@ typedef struct BlockCopyState {
bool skip_unallocated; /* atomic */
/* State fields that use a thread-safe API */
BdrvDirtyBitmap *copy_bitmap;
+ /* Whether a write-zeroes task may widen; see block_copy_write_zeroes(). */
+ bool zero_widen; /* atomic */
/* Clusters reading as zero; allocated on demand, frozen once valid. */
HBitmap *zero_bitmap;
/* Published only after the scan, with skip_unallocated already false. */
@@ -187,6 +189,34 @@ static int64_t block_copy_chunk_size(BlockCopyState *s)
}
}
+/*
+ * A write-zeroes task carries no buffer, so it may cover far more than
+ * block_copy_chunk_size(). Return how far it may run; the caller clamps it
+ * to where the zero run ends.
+ */
+static int64_t block_copy_widen_zero_area(BlockCopyState *s,
+ BlockCopyCallState *call_state,
+ int64_t offset, int64_t search_end,
+ int64_t bytes)
+{
+ int64_t aligned = QEMU_ALIGN_DOWN(BDRV_REQUEST_MAX_BYTES,
+ s->cluster_size);
+ int64_t zero_chunk = MIN_NON_ZERO(MAX(aligned, s->cluster_size),
+ call_state->max_chunk);
+ int64_t wide_offset, wide_bytes;
+
+ if (!bdrv_dirty_bitmap_next_dirty_area(s->copy_bitmap, offset, search_end,
+ zero_chunk, &wide_offset,
+ &wide_bytes)) {
+ return bytes;
+ }
+
+ /* @offset is dirty, so the search cannot have moved past it. */
+ assert(wide_offset == offset);
+
+ return wide_bytes;
+}
+
/*
* Search for the first dirty area in offset/bytes range and create task at
* the beginning of it.
@@ -198,6 +228,7 @@ block_copy_task_create(BlockCopyState *s, BlockCopyCallState *call_state,
BlockCopyTask *task;
BlockCopyMethod method;
int64_t max_chunk;
+ int64_t search_end = offset + bytes;
QEMU_LOCK_GUARD(&s->lock);
max_chunk = MIN_NON_ZERO(block_copy_chunk_size(s), call_state->max_chunk);
@@ -219,6 +250,10 @@ block_copy_task_create(BlockCopyState *s, BlockCopyCallState *call_state,
if (hbitmap_get(s->zero_bitmap, offset)) {
method = COPY_WRITE_ZEROES;
+ if (qatomic_read(&s->zero_widen)) {
+ bytes = block_copy_widen_zero_area(s, call_state, offset,
+ search_end, bytes);
+ }
boundary = hbitmap_next_zero(s->zero_bitmap, offset, bytes);
} else {
boundary = hbitmap_next_dirty(s->zero_bitmap, offset, bytes);
@@ -465,6 +500,7 @@ BlockCopyState *block_copy_state_new(BdrvChild *source, BdrvChild *target,
.max_transfer = QEMU_ALIGN_DOWN(
block_copy_max_transfer(source, target),
cluster_size),
+ .zero_widen = target->bs->supported_zero_flags & BDRV_REQ_NO_FALLBACK,
};
s->discard_source = discard_source;
@@ -521,6 +557,56 @@ static coroutine_fn int block_copy_task_run(AioTaskPool *pool,
return 0;
}
+/*
+ * Widening a write-zeroes request only pays off where the target zeroes by
+ * metadata, so while the target is believed to oblige, a request asks to fail
+ * instead of falling back to writing the zeroes out. That costs nothing when
+ * it does oblige, and a refusal ends the widening for the rest of the run.
+ */
+static int coroutine_fn GRAPH_RDLOCK
+block_copy_write_zeroes(BlockCopyState *s, int64_t offset, int64_t bytes,
+ bool *error_is_read)
+{
+ BdrvRequestFlags flags = s->write_flags & ~BDRV_REQ_WRITE_COMPRESSED;
+ int64_t chunk;
+ int ret = 0;
+
+ if (qatomic_read(&s->zero_widen)) {
+ ret = bdrv_co_pwrite_zeroes(s->target, offset, bytes,
+ flags | BDRV_REQ_NO_FALLBACK);
+ if (ret != -ENOTSUP) {
+ goto out;
+ }
+
+ /* Redoing the range is safe: zeroes over zeroes change nothing. */
+ qatomic_set(&s->zero_widen, false);
+ }
+
+ WITH_QEMU_LOCK_GUARD(&s->lock) {
+ chunk = block_copy_chunk_size(s);
+ }
+
+ while (bytes) {
+ int64_t n = MIN(bytes, chunk);
+
+ ret = bdrv_co_pwrite_zeroes(s->target, offset, n, flags);
+ if (ret < 0) {
+ break;
+ }
+
+ offset += n;
+ bytes -= n;
+ }
+
+out:
+ if (ret < 0) {
+ trace_block_copy_write_zeroes_fail(s, offset, ret);
+ *error_is_read = false;
+ }
+
+ return ret;
+}
+
/*
* block_copy_do_copy
*
@@ -552,13 +638,7 @@ block_copy_do_copy(BlockCopyState *s, int64_t offset, int64_t bytes,
switch (*method) {
case COPY_WRITE_ZEROES:
- ret = bdrv_co_pwrite_zeroes(s->target, offset, nbytes, s->write_flags &
- ~BDRV_REQ_WRITE_COMPRESSED);
- if (ret < 0) {
- trace_block_copy_write_zeroes_fail(s, offset, ret);
- *error_is_read = false;
- }
- return ret;
+ return block_copy_write_zeroes(s, offset, nbytes, error_is_read);
case COPY_RANGE_SMALL:
case COPY_RANGE_FULL:
--
2.53.0
next prev parent reply other threads:[~2026-09-29 15:53 UTC|newest]
Thread overview: 14+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-09-29 15:51 [PATCH v3 0/9] block: cheaper zero handling in backup and commit Denis V. Lunev
2026-09-29 15:51 ` [PATCH v3 1/9] block/commit: pass BDRV_WANT_PRECISE to block-status Denis V. Lunev
2026-10-05 21:48 ` Eric Blake
2026-09-29 15:51 ` [PATCH v3 2/9] iotests/040: cover large and fragmented commit runs Denis V. Lunev
2026-09-29 15:51 ` [PATCH v3 3/9] block/commit: batch block-status queries Denis V. Lunev
2026-09-29 15:51 ` [PATCH v3 4/9] iotests/124: cover backup of zero clusters and holes Denis V. Lunev
2026-09-29 15:51 ` [PATCH v3 5/9] block/block-copy: don't reserve memory for zero tasks Denis V. Lunev
2026-09-29 15:51 ` [PATCH v3 6/9] block/block-copy: extract block_copy_set_task_method() Denis V. Lunev
2026-09-29 15:51 ` [PATCH v3 7/9] block/block-copy: track known-zero source clusters Denis V. Lunev
2026-09-29 15:51 ` [PATCH v3 8/9] block/backup: pre-fill zero_bitmap for full/bitmap Denis V. Lunev
2026-09-29 15:51 ` Denis V. Lunev [this message]
2026-09-30 7:43 ` [PATCH v3 9/9] block/block-copy: coalesce write-zeroes tasks Andrey Drobyshev
2026-10-05 8:59 ` [PATCH v3 0/9] block: cheaper zero handling in backup and commit Denis V. Lunev
2026-10-05 22:04 ` Eric Blake
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=20260929155125.3151111-10-den@openvz.org \
--to=den@openvz.org \
--cc=andrey.drobyshev@virtuozzo.com \
--cc=jsnow@redhat.com \
--cc=qemu-block@nongnu.org \
--cc=qemu-devel@nongnu.org \
--cc=vsementsov@yandex-team.ru \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is an external index of several public inboxes,
see mirroring instructions on how to clone and mirror
all data and code used by this external index.