From: Zhang Yi <yizhang089@gmail.com>
To: sashiko-reviews@lists.linux.dev, Zhang Yi <yi.zhang@huaweicloud.com>
Cc: linux-ext4@vger.kernel.org, tytso@mit.edu
Subject: Re: [PATCH v6 26/31] ext4: clear DISKSIZE_GROW_PENDING on truncate or error
Date: Mon, 14 Sep 2026 15:37:13 +0800 [thread overview]
Message-ID: <9f221a5b-281c-45d4-8cc7-e63b66a2c6cf@gmail.com> (raw)
In-Reply-To: <20260903134656.C9FDE1F00ACA@smtp.kernel.org>
On 9/3/2026 9:46 PM, sashiko-bot@kernel.org wrote:
> Thank you for your contribution! Sashiko AI review found 2 potential issue(s) to consider:
> - [High] Unconditional advancement of i_disksize in ext4_truncate_down exposes unwritten data and corrupts the disksize pending state.
> - [Medium] EXT4_STATE_DISKSIZE_GROW_PENDING is not cleared when evicting an inode with i_nlink > 0 or a bad inode, leading to a spurious WARN_ON_ONCE in ext4_destroy_inode().
> --
>
> commit ce48fe2572983c4095fed45d85327a228bcb48f0
> Author: Zhang Yi <yi.zhang@huawei.com>
>
> ext4: clear DISKSIZE_GROW_PENDING on truncate or error
>
> This commit clears the DISKSIZE_GROW_PENDING state on error paths, such
> as folio discard and unlinked inode eviction, as well as during truncation.
> It aims to prevent writeback from blocking indefinitely when a zeroed EOF
> block is discarded, and adds a warning if an inode is destroyed with the
> bit still set.
>
>> diff --git a/fs/ext4/inode.c b/fs/ext4/inode.c
>> index 173ac69e9b953..f38437fa91064 100644
>> --- a/fs/ext4/inode.c
>> +++ b/fs/ext4/inode.c
>
> [ ... ]
>
>> @@ -273,6 +273,8 @@ void ext4_evict_inode(struct inode *inode)
>>
>> if (ext4_should_order_data(inode))
>> ext4_begin_ordered_truncate(inode, 0);
>> + if (ext4_inode_buffered_iomap(inode))
>> + ext4_iomap_clear_disksize_pending(inode);
>> truncate_inode_pages_final(&inode->i_data);
>>
>
> [Severity: Medium]
> Does this miss clearing the pending bit for linked inodes (i_nlink > 0)
> or bad inodes?
>
> If writeback fails for an inode with a pending zeroed EOF block, and the
> filesystem is unmounted, evict_inodes() unconditionally evicts all inodes.
> Earlier in ext4_evict_inode(), the code truncates pages for linked inodes
> and then jumps to the no_delete label:
>
> if (inode->i_nlink) {
> ...
> truncate_inode_pages_final(&inode->i_data);
> ...
> goto no_delete;
> }
>
> This jump bypasses the new ext4_iomap_clear_disksize_pending() cleanup here.
>
> Could this result in the pending bit remaining set and triggering the new
> WARN_ON_ONCE in ext4_destroy_inode() when the inode is finally destroyed?
>
This is false positive, the pending bit is always cleared before an
inode can be destroyed:
- ext4_iomap_end_bio() clears it when the pending I/O completes,
regardless of whether that I/O succeeded, and a failure before
submission goes through ext4_iomap_discard_folio(), which clears it
as well.
- An inode with i_nlink > 0 is only evicted after its dirty pages have
been written out.
- The only case where the pending block is not written back is an I/O
error or a forced shutdown, and both set EXT4_ERROR_FS / the
emergency state, which the WARN_ON_ONCE() already skips.
A bad inode can never have the bit set since the bit is set only from
ext4_block_zero_eof() on the buffered I/O path, while make_bad_inode()
is only reached from __ext4_iget() and from the mount-time journal
inode load, both before the inode can be used for any I/O.
> [ ... ]
>
>> @@ -6765,7 +6776,18 @@ static int ext4_truncate_down(struct inode *inode, loff_t oldsize,
>> start_lblk = newsize > 0 ? (newsize - 1) >> inode->i_blkbits : 0;
>> ext4_fc_track_range(handle, inode, start_lblk, EXT_MAX_BLOCKS - 1);
>>
>> - ext4_set_inode_size(inode, newsize);
>> + down_write(&EXT4_I(inode)->i_data_sem);
>> + /*
>> + * Truncate the zeroed EOF block invalidates the pending disksize
>> + * update, so clear the disksize-grow-pending state.
>> + */
>> + if (ext4_test_inode_state(inode, EXT4_STATE_DISKSIZE_GROW_PENDING) &&
>> + (newsize <= EXT4_I(inode)->i_disksize))
>> + ext4_iomap_clear_disksize_pending(inode);
>> +
>> + i_size_write(inode, newsize);
>> + __ext4_set_i_disksize(inode, newsize);
>> + up_write(&EXT4_I(inode)->i_data_sem);
>>
>> ret = ext4_mark_inode_dirty(handle, inode);
>> ext4_journal_stop(handle);
>
> [Severity: High]
> Could unconditionally advancing i_disksize here expose unwritten data or leak
> the pending bit?
>
> If ftruncate() is called with a newsize that is smaller than i_size but
> strictly greater than i_disksize, the condition
> (newsize <= EXT4_I(inode)->i_disksize) evaluates to false. The pending bit is
> not cleared.
>
> However, __ext4_set_i_disksize(inode, newsize) then unconditionally advances
> i_disksize past the safely written data.
>
> If a crash occurs before the dirty pages up to newsize are written back, will
> this expose stale or uninitialized data on disk since data ordering is not
> guaranteed for the iomap path?
>
> Additionally, if newsize is block-aligned, would the writeback thread later
> hitting ext4_iomap_get_disksize_pending_range() encounter a WARN_ON_ONCE
> because i_disksize was updated to a block-aligned value, preventing the ioend
> from ever clearing the pending bit?
>
This is also false positive, __ext4_set_i_disksize() does not advance
i_disksize while the GROW_PENDING bit is set. So with
newsize > i_disksize, i_disksize is left untouched and stays unaligned,
and the WARN_ON_ONCE() in ext4_iomap_get_disksize_pending_range()
cannot trigger.
Keeping the bit in that case is intentional. The new EOF is still inside
the pending zeroed block, so the zeroed data must reach the disk before
i_disksize may be advanced past it.
Thanks,
Yi.
next prev parent reply other threads:[~2026-09-14 7:37 UTC|newest]
Thread overview: 96+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-09-03 12:35 [PATCH v6 00/31] ext4: use iomap for regular file's buffered I/O path Zhang Yi
2026-09-03 12:35 ` [PATCH v6 01/31] ext4: simplify size updating in ext4_setattr() Zhang Yi
2026-09-03 12:56 ` sashiko-bot
2026-09-03 12:35 ` [PATCH v6 02/31] ext4: factor out ext4_truncate_[up|down]() Zhang Yi
2026-09-03 13:09 ` sashiko-bot
2026-09-03 12:35 ` [PATCH v6 03/31] ext4: skip ordered I/O wait when zeroing beyond i_disksize block Zhang Yi
2026-09-03 13:14 ` sashiko-bot
2026-09-08 11:31 ` Zhang Yi
2026-09-24 11:40 ` Ojaswin Mujoo
2026-09-03 12:35 ` [PATCH v6 04/31] ext4: set EXT4_MAP_NEW flag for delayed allocated blocks Zhang Yi
2026-09-03 12:54 ` sashiko-bot
2026-09-24 11:42 ` Ojaswin Mujoo
2026-09-03 12:35 ` [PATCH v6 05/31] ext4: recheck extent status tree before block allocation Zhang Yi
2026-09-03 13:02 ` sashiko-bot
2026-09-09 9:27 ` Zhang Yi
2026-09-24 13:22 ` Ojaswin Mujoo
2026-09-28 7:15 ` Zhang Yi
2026-09-03 12:35 ` [PATCH v6 06/31] ext4: fix orig_mlen initialization in ext4_map_blocks() Zhang Yi
2026-09-03 12:55 ` sashiko-bot
2026-09-24 13:24 ` Ojaswin Mujoo
2026-09-03 12:35 ` [PATCH v6 07/31] ext4: allow ext4_map_blocks() to start its own transaction handle Zhang Yi
2026-09-03 13:02 ` sashiko-bot
2026-09-28 10:19 ` Ojaswin Mujoo
2026-09-03 12:35 ` [PATCH v6 08/31] ext4: avoid unnecessary transaction in ext4_map_blocks() for unwritten extents Zhang Yi
2026-09-03 13:06 ` sashiko-bot
2026-09-28 10:21 ` Ojaswin Mujoo
2026-09-28 11:57 ` Zhang Yi
2026-09-03 12:35 ` [PATCH v6 09/31] ext4: skip block allocation for holes in the data submission path Zhang Yi
2026-09-03 13:46 ` sashiko-bot
2026-09-09 7:13 ` Zhang Yi
2026-09-28 10:36 ` Ojaswin Mujoo
2026-09-28 12:10 ` Zhang Yi
2026-09-03 12:35 ` [PATCH v6 10/31] ext4: add iomap address space operations for buffered I/O Zhang Yi
2026-09-03 12:57 ` sashiko-bot
2026-09-03 12:35 ` [PATCH v6 11/31] ext4: implement buffered read path using iomap Zhang Yi
2026-09-03 13:12 ` sashiko-bot
2026-10-05 9:23 ` Theodore Tso
2026-10-05 14:01 ` Theodore Tso
2026-09-03 12:35 ` [PATCH v6 12/31] ext4: pass out extent seq counter when mapping da blocks Zhang Yi
2026-09-03 13:05 ` sashiko-bot
2026-09-03 12:35 ` [PATCH v6 13/31] ext4: do not use data=ordered mode for inodes using buffered iomap path Zhang Yi
2026-09-03 13:13 ` sashiko-bot
2026-09-29 14:20 ` Ojaswin Mujoo
2026-09-03 12:35 ` [PATCH v6 14/31] ext4: implement buffered write path using iomap Zhang Yi
2026-09-03 13:23 ` sashiko-bot
2026-09-10 12:14 ` Zhang Yi
2026-09-03 12:35 ` [PATCH v6 15/31] ext4: implement writeback " Zhang Yi
2026-09-03 13:36 ` sashiko-bot
2026-09-12 8:28 ` Zhang Yi
2026-09-30 8:51 ` Ojaswin Mujoo
2026-09-30 9:26 ` Zhang Yi
2026-09-30 10:36 ` Ojaswin Mujoo
2026-09-03 12:35 ` [PATCH v6 16/31] ext4: implement mmap " Zhang Yi
2026-09-03 13:29 ` sashiko-bot
2026-09-03 12:35 ` [PATCH v6 17/31] ext4: implement partial block zero range " Zhang Yi
2026-09-03 13:29 ` sashiko-bot
2026-09-30 12:24 ` Ojaswin Mujoo
2026-09-03 12:35 ` [PATCH v6 18/31] ext4: drain writeback before removing extents on the iomap path Zhang Yi
2026-09-03 13:19 ` sashiko-bot
2026-10-01 11:05 ` Ojaswin Mujoo
2026-09-03 12:35 ` [PATCH v6 19/31] ext4: add block mapping tracepoints for iomap buffered I/O path Zhang Yi
2026-09-03 13:19 ` sashiko-bot
2026-09-03 12:35 ` [PATCH v6 20/31] ext4: disable online defrag when inode using " Zhang Yi
2026-09-03 13:25 ` sashiko-bot
2026-09-03 12:35 ` [PATCH v6 21/31] ext4: add EXT4_STATE_DISKSIZE_GROW_PENDING state bit and helpers Zhang Yi
2026-09-03 13:19 ` sashiko-bot
2026-09-03 12:35 ` [PATCH v6 22/31] ext4: submit and wait for pending disksize-grow I/O on writeback Zhang Yi
2026-09-03 13:40 ` sashiko-bot
2026-09-14 7:27 ` Zhang Yi
2026-10-05 9:00 ` Ojaswin Mujoo
2026-10-08 8:53 ` Zhang Yi
2026-10-08 9:15 ` Ojaswin Mujoo
2026-10-08 12:22 ` Zhang Yi
2026-09-03 12:35 ` [PATCH v6 23/31] ext4: advance i_disksize to i_size upon disksize-grow I/O completion Zhang Yi
2026-09-03 13:31 ` sashiko-bot
2026-09-03 12:35 ` [PATCH v6 24/31] ext4: defer i_disksize update while DISKSIZE_GROW_PENDING is set Zhang Yi
2026-09-03 13:38 ` sashiko-bot
2026-09-03 12:35 ` [PATCH v6 25/31] ext4: submit and wait for disksize-grow I/O in fallocate paths Zhang Yi
2026-09-03 13:42 ` sashiko-bot
2026-09-03 12:35 ` [PATCH v6 26/31] ext4: clear DISKSIZE_GROW_PENDING on truncate or error Zhang Yi
2026-09-03 13:46 ` sashiko-bot
2026-09-14 7:37 ` Zhang Yi [this message]
2026-09-03 12:35 ` [PATCH v6 27/31] ext4: set DISKSIZE_GROW_PENDING after zeroing unaligned EOF block Zhang Yi
2026-09-03 13:36 ` sashiko-bot
2026-10-05 8:55 ` Ojaswin Mujoo
2026-10-08 12:20 ` Zhang Yi
2026-09-03 12:35 ` [PATCH v6 28/31] ext4: add tracepoints for DISKSIZE_GROW_PENDING set, clear, and wait Zhang Yi
2026-09-03 13:34 ` sashiko-bot
2026-09-03 12:40 ` [PATCH v6 29/31] ext4: add tracepoints for EOF block zeroing and disksize-grow I/O Zhang Yi
2026-09-03 13:32 ` sashiko-bot
2026-09-03 12:40 ` [PATCH v6 30/31] ext4: partially enable iomap for the buffered I/O path of regular files Zhang Yi
2026-09-03 14:04 ` sashiko-bot
2026-09-14 8:38 ` Zhang Yi
2026-09-03 12:40 ` [PATCH v6 31/31] ext4: introduce a mount option for iomap buffered I/O path Zhang Yi
2026-09-03 13:46 ` sashiko-bot
2026-10-09 8:56 ` [PATCH v6 00/31] ext4: use iomap for regular file's " Theodore Ts'o
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=9f221a5b-281c-45d4-8cc7-e63b66a2c6cf@gmail.com \
--to=yizhang089@gmail.com \
--cc=linux-ext4@vger.kernel.org \
--cc=sashiko-reviews@lists.linux.dev \
--cc=tytso@mit.edu \
--cc=yi.zhang@huaweicloud.com \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox