All of lore.kernel.org
 help / color / mirror / Atom feed
From: Zhang Yi <yizhang089@gmail.com>
To: Ojaswin Mujoo <ojaswin@linux.ibm.com>,
	Zhang Yi <yi.zhang@huaweicloud.com>
Cc: linux-ext4@vger.kernel.org, linux-fsdevel@vger.kernel.org,
	linux-kernel@vger.kernel.org, tytso@mit.edu,
	adilger.kernel@dilger.ca, libaokun@linux.alibaba.com,
	jack@suse.cz, ritesh.list@gmail.com, djwong@kernel.org,
	hch@infradead.org, yi.zhang@huawei.com, chengzhihao1@huawei.com,
	yangerkun@huawei.com, wangkefeng.wang@huawei.com,
	yukuai@fnnas.com
Subject: Re: [PATCH v6 22/31] ext4: submit and wait for pending disksize-grow I/O on writeback
Date: Thu, 8 Oct 2026 16:53:24 +0800	[thread overview]
Message-ID: <08dfc1dc-adc6-41ed-aaa5-7af7836f71c7@gmail.com> (raw)
In-Reply-To: <asNnFY8N3NJwucWe@li-dc0c254c-257c-11b2-a85c-98b6c1322444.ibm.com>

On 10/5/2026 5:00 PM, Ojaswin Mujoo wrote:
> On Thu, Sep 03, 2026 at 08:35:34PM +0800, Zhang Yi wrote:
>> From: Zhang Yi <yi.zhang@huawei.com>
>>
>> When the current writeback pass begins beyond the disksize-grow-pending
>> zeroed EOF block, the ioend worker would otherwise have to wait for the
>> pending EOF block to complete before it can advance i_disksize.
>> Otherwise the old EOF block could be exposed as stale data once
>> i_disksize advances past it.
>>
>> Therefore, introduce the ioend mechanism for the pending range, tag
>> ioends that cover the pending zeroed EOF block which straddles
>> i_disksize with EXT4_IOMAP_IOEND_DISKSIZE_GROW_IO in
>> ext4_iomap_writeback_submit(), and clear the bit and wake up all waiters
>> in ext4_iomap_end_bio() when such an ioend completes.
>>
>> Clearing EXT4_IOMAP_IOEND_DISKSIZE_GROW_IO does not depend on whether
>> the disksize grow I/O succeeds. That is, even if the I/O fails, we still
>> allow subsequent writes in the range to update i_disksize. This is
>> consistent with the previous behavior, and we rely on data_err=abort to
>> prevent metadata updates when data write failures occur.
>>
>> In order to avoid the ioend that passes the pending range waiting for a
>> long time, proactively submit the pending range first in
>> ext4_iomap_writepages() so it completes in parallel with the rest of the
>> writeback.
>>
>> Note that the handling of discarding the zeroed EOF folio will be
>> processed later, otherwise the bit will be set forever.
>> EXT4_STATE_DISKSIZE_GROW_PENDING will be set after everthing is done.
>>
>> Signed-off-by: Zhang Yi <yi.zhang@huawei.com>
>> ---
>>   fs/ext4/ext4.h    |  6 ++++++
>>   fs/ext4/inode.c   | 53 ++++++++++++++++++++++++++++++++++++++++++++++-
>>   fs/ext4/page-io.c | 40 +++++++++++++++++++++++++++++++++++
>>   3 files changed, 98 insertions(+), 1 deletion(-)
>>
>> diff --git a/fs/ext4/ext4.h b/fs/ext4/ext4.h
>> index 1c3d736fb700..089dbd39c5c2 100644
>> --- a/fs/ext4/ext4.h
>> +++ b/fs/ext4/ext4.h
>> @@ -3986,6 +3986,12 @@ extern int ext4_move_extents(struct file *o_filp, struct file *d_filp,
>>   			     __u64 len, __u64 *moved_len);
>>   
>>   /* page-io.c */
>> +/*
>> + * The I/O range covers the zeroed EOF block that straddles i_disksize
>> + * and will advance it upon completion.
>> + */
>> +#define EXT4_IOMAP_IOEND_DISKSIZE_GROW_IO	1UL
>> +
>>   extern int __init ext4_init_pageio(void);
>>   extern void ext4_exit_pageio(void);
>>   extern ext4_io_end_t *ext4_init_io_end(struct inode *inode, gfp_t flags);
>> diff --git a/fs/ext4/inode.c b/fs/ext4/inode.c
>> index 05dd4ee805fb..4239be5a769f 100644
>> --- a/fs/ext4/inode.c
>> +++ b/fs/ext4/inode.c
>> @@ -4366,7 +4366,10 @@ static int ext4_iomap_writeback_submit(struct iomap_writepage_ctx *wpc,
>>   				       int error)
>>   {
>>   	struct iomap_ioend *ioend = wpc->wb_ctx;
>> -	struct ext4_inode_info *ei = EXT4_I(ioend->io_inode);
>> +	struct inode *inode = ioend->io_inode;
>> +	struct ext4_inode_info *ei = EXT4_I(inode);
>> +	unsigned int blocksize = i_blocksize(inode);
>> +	loff_t pstart, plen;
>>   
>>   	/*
>>   	 * After I/O completion, a worker needs to be scheduled when:
>> @@ -4379,6 +4382,21 @@ static int ext4_iomap_writeback_submit(struct iomap_writepage_ctx *wpc,
>>   	    test_opt(ioend->io_inode->i_sb, DATA_ERR_ABORT))
>>   		ioend->io_bio.bi_end_io = ext4_iomap_end_bio;
>>   
>> +	/*
>> +	 * Mark the I/O as DISKSIZE_GROW_IO by setting io_private to
>> +	 * EXT4_IOMAP_IOEND_DISKSIZE_GROW_IO if it covers the pending range.
>> +	 * Such I/O will allow or trigger i_disksize advancement in the
>> +	 * ioend worker.
>> +	 */
>> +	plen = ext4_iomap_get_disksize_pending_range(inode, &pstart);
>> +	if (plen &&
>> +	    round_down(ioend->io_offset, blocksize) <= pstart &&
>> +	    round_up(ioend->io_offset + ioend->io_size, blocksize) >=
>> +			pstart + plen) {
>> +		ioend->io_bio.bi_end_io = ext4_iomap_end_bio;
>> +		ioend->io_private = (void *)EXT4_IOMAP_IOEND_DISKSIZE_GROW_IO;
>> +	}
>> +
>>   	/*
>>   	 * ext4_iomap_end_bio() always defers endio processing, disable
>>   	 * generic BIO in task to avoid double deferral since we will use
>> @@ -4398,6 +4416,33 @@ static const struct iomap_writeback_ops ext4_writeback_ops = {
>>   	.writeback_submit = ext4_iomap_writeback_submit,
>>   };
>>   
>> +/*
>> + * If the current writeback range begins after the pending zeroed EOF
>> + * block range which straddles i_disksize, issue a separate writeback to
>> + * flush it first, so as to avoid prolonged waiting.
>> + */
>> +static void ext4_iomap_wb_submit_zeroed_eof(struct inode *inode,
>> +					    struct writeback_control *wbc)
>> +{
>> +	struct address_space *mapping = inode->i_mapping;
>> +	loff_t pstart, plen, range_start;
>> +
>> +	if (wbc->range_cyclic)
>> +		range_start = (loff_t)mapping->writeback_index << PAGE_SHIFT;
>> +	else
>> +		range_start = wbc->range_start;
>> +
>> +	plen = ext4_iomap_get_disksize_pending_range(inode, &pstart);
>> +	if (!plen || range_start < pstart + plen)
>> +		return;
> Hi Zhang,
> 
> Maybe we can check here if pstart lies withing the same folio as the
> writeback range, then we don't need an explicit flush as we will anyways
> flush out the whole folio. We can check against the
> mapping_min_folio_nrbytes perhaps.
> 

Hi Ojaswin,

Thank you for the suggestion. IIUC, do you mean to change it somewhat
like the following:

       if (!plen ||
           round_down(range_start, mapping_min_folio_nrbytes(mapping)) <
                      pstart + plen)
               return;

The benefit is that it covers the case of blocksize < PAGE_SIZE, where
the pending range occupies only part of the folio and range_start lands
in the middle or later part of that same folio. It does not help when
the pending range happens to be in a large folio, though, because
mapping_min_folio_nrbytes() is only a lower bound on the folio size.
That said, the pending range is the old EOF of the file and it is
unaligned, so that folio can more likely to be a single page than a
large one.

Or do you want it to be very precise, able to know exactly the size of
the folio containing the pending range and its positional relationship
with the writeback range? If the latter, that would require an exact
folio lookup, which seems a bit expensive and perhaps over-optimization
to me.

Thanks,
Yi.


  reply	other threads:[~2026-10-08  8:53 UTC|newest]

Thread overview: 95+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-09-03 12:35 [PATCH v6 00/31] ext4: use iomap for regular file's buffered I/O path Zhang Yi
2026-09-03 12:35 ` [PATCH v6 01/31] ext4: simplify size updating in ext4_setattr() Zhang Yi
2026-09-03 12:56   ` sashiko-bot
2026-09-03 12:35 ` [PATCH v6 02/31] ext4: factor out ext4_truncate_[up|down]() Zhang Yi
2026-09-03 13:09   ` sashiko-bot
2026-09-03 12:35 ` [PATCH v6 03/31] ext4: skip ordered I/O wait when zeroing beyond i_disksize block Zhang Yi
2026-09-03 13:14   ` sashiko-bot
2026-09-08 11:31     ` Zhang Yi
2026-09-24 11:40       ` Ojaswin Mujoo
2026-09-03 12:35 ` [PATCH v6 04/31] ext4: set EXT4_MAP_NEW flag for delayed allocated blocks Zhang Yi
2026-09-03 12:54   ` sashiko-bot
2026-09-24 11:42   ` Ojaswin Mujoo
2026-09-03 12:35 ` [PATCH v6 05/31] ext4: recheck extent status tree before block allocation Zhang Yi
2026-09-03 13:02   ` sashiko-bot
2026-09-09  9:27     ` Zhang Yi
2026-09-24 13:22   ` Ojaswin Mujoo
2026-09-28  7:15     ` Zhang Yi
2026-09-03 12:35 ` [PATCH v6 06/31] ext4: fix orig_mlen initialization in ext4_map_blocks() Zhang Yi
2026-09-03 12:55   ` sashiko-bot
2026-09-24 13:24   ` Ojaswin Mujoo
2026-09-03 12:35 ` [PATCH v6 07/31] ext4: allow ext4_map_blocks() to start its own transaction handle Zhang Yi
2026-09-03 13:02   ` sashiko-bot
2026-09-28 10:19   ` Ojaswin Mujoo
2026-09-03 12:35 ` [PATCH v6 08/31] ext4: avoid unnecessary transaction in ext4_map_blocks() for unwritten extents Zhang Yi
2026-09-03 13:06   ` sashiko-bot
2026-09-28 10:21   ` Ojaswin Mujoo
2026-09-28 11:57     ` Zhang Yi
2026-09-03 12:35 ` [PATCH v6 09/31] ext4: skip block allocation for holes in the data submission path Zhang Yi
2026-09-03 13:46   ` sashiko-bot
2026-09-09  7:13     ` Zhang Yi
2026-09-28 10:36   ` Ojaswin Mujoo
2026-09-28 12:10     ` Zhang Yi
2026-09-03 12:35 ` [PATCH v6 10/31] ext4: add iomap address space operations for buffered I/O Zhang Yi
2026-09-03 12:57   ` sashiko-bot
2026-09-03 12:35 ` [PATCH v6 11/31] ext4: implement buffered read path using iomap Zhang Yi
2026-09-03 13:12   ` sashiko-bot
2026-10-05  9:23   ` Theodore Tso
2026-10-05 14:01     ` Theodore Tso
2026-09-03 12:35 ` [PATCH v6 12/31] ext4: pass out extent seq counter when mapping da blocks Zhang Yi
2026-09-03 13:05   ` sashiko-bot
2026-09-03 12:35 ` [PATCH v6 13/31] ext4: do not use data=ordered mode for inodes using buffered iomap path Zhang Yi
2026-09-03 13:13   ` sashiko-bot
2026-09-29 14:20   ` Ojaswin Mujoo
2026-09-03 12:35 ` [PATCH v6 14/31] ext4: implement buffered write path using iomap Zhang Yi
2026-09-03 13:23   ` sashiko-bot
2026-09-10 12:14     ` Zhang Yi
2026-09-03 12:35 ` [PATCH v6 15/31] ext4: implement writeback " Zhang Yi
2026-09-03 13:36   ` sashiko-bot
2026-09-12  8:28     ` Zhang Yi
2026-09-30  8:51       ` Ojaswin Mujoo
2026-09-30  9:26         ` Zhang Yi
2026-09-30 10:36           ` Ojaswin Mujoo
2026-09-03 12:35 ` [PATCH v6 16/31] ext4: implement mmap " Zhang Yi
2026-09-03 13:29   ` sashiko-bot
2026-09-03 12:35 ` [PATCH v6 17/31] ext4: implement partial block zero range " Zhang Yi
2026-09-03 13:29   ` sashiko-bot
2026-09-30 12:24   ` Ojaswin Mujoo
2026-09-03 12:35 ` [PATCH v6 18/31] ext4: drain writeback before removing extents on the iomap path Zhang Yi
2026-09-03 13:19   ` sashiko-bot
2026-10-01 11:05   ` Ojaswin Mujoo
2026-09-03 12:35 ` [PATCH v6 19/31] ext4: add block mapping tracepoints for iomap buffered I/O path Zhang Yi
2026-09-03 13:19   ` sashiko-bot
2026-09-03 12:35 ` [PATCH v6 20/31] ext4: disable online defrag when inode using " Zhang Yi
2026-09-03 13:25   ` sashiko-bot
2026-09-03 12:35 ` [PATCH v6 21/31] ext4: add EXT4_STATE_DISKSIZE_GROW_PENDING state bit and helpers Zhang Yi
2026-09-03 13:19   ` sashiko-bot
2026-09-03 12:35 ` [PATCH v6 22/31] ext4: submit and wait for pending disksize-grow I/O on writeback Zhang Yi
2026-09-03 13:40   ` sashiko-bot
2026-09-14  7:27     ` Zhang Yi
2026-10-05  9:00   ` Ojaswin Mujoo
2026-10-08  8:53     ` Zhang Yi [this message]
2026-10-08  9:15       ` Ojaswin Mujoo
2026-10-08 12:22         ` Zhang Yi
2026-09-03 12:35 ` [PATCH v6 23/31] ext4: advance i_disksize to i_size upon disksize-grow I/O completion Zhang Yi
2026-09-03 13:31   ` sashiko-bot
2026-09-03 12:35 ` [PATCH v6 24/31] ext4: defer i_disksize update while DISKSIZE_GROW_PENDING is set Zhang Yi
2026-09-03 13:38   ` sashiko-bot
2026-09-03 12:35 ` [PATCH v6 25/31] ext4: submit and wait for disksize-grow I/O in fallocate paths Zhang Yi
2026-09-03 13:42   ` sashiko-bot
2026-09-03 12:35 ` [PATCH v6 26/31] ext4: clear DISKSIZE_GROW_PENDING on truncate or error Zhang Yi
2026-09-03 13:46   ` sashiko-bot
2026-09-14  7:37     ` Zhang Yi
2026-09-03 12:35 ` [PATCH v6 27/31] ext4: set DISKSIZE_GROW_PENDING after zeroing unaligned EOF block Zhang Yi
2026-09-03 13:36   ` sashiko-bot
2026-10-05  8:55   ` Ojaswin Mujoo
2026-10-08 12:20     ` Zhang Yi
2026-09-03 12:35 ` [PATCH v6 28/31] ext4: add tracepoints for DISKSIZE_GROW_PENDING set, clear, and wait Zhang Yi
2026-09-03 13:34   ` sashiko-bot
2026-09-03 12:40 ` [PATCH v6 29/31] ext4: add tracepoints for EOF block zeroing and disksize-grow I/O Zhang Yi
2026-09-03 13:32   ` sashiko-bot
2026-09-03 12:40 ` [PATCH v6 30/31] ext4: partially enable iomap for the buffered I/O path of regular files Zhang Yi
2026-09-03 14:04   ` sashiko-bot
2026-09-14  8:38     ` Zhang Yi
2026-09-03 12:40 ` [PATCH v6 31/31] ext4: introduce a mount option for iomap buffered I/O path Zhang Yi
2026-09-03 13:46   ` sashiko-bot

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=08dfc1dc-adc6-41ed-aaa5-7af7836f71c7@gmail.com \
    --to=yizhang089@gmail.com \
    --cc=adilger.kernel@dilger.ca \
    --cc=chengzhihao1@huawei.com \
    --cc=djwong@kernel.org \
    --cc=hch@infradead.org \
    --cc=jack@suse.cz \
    --cc=libaokun@linux.alibaba.com \
    --cc=linux-ext4@vger.kernel.org \
    --cc=linux-fsdevel@vger.kernel.org \
    --cc=linux-kernel@vger.kernel.org \
    --cc=ojaswin@linux.ibm.com \
    --cc=ritesh.list@gmail.com \
    --cc=tytso@mit.edu \
    --cc=wangkefeng.wang@huawei.com \
    --cc=yangerkun@huawei.com \
    --cc=yi.zhang@huawei.com \
    --cc=yi.zhang@huaweicloud.com \
    --cc=yukuai@fnnas.com \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is an external index of several public inboxes,
see mirroring instructions on how to clone and mirror
all data and code used by this external index.