From: Chao Yu <chao@kernel.org>
To: yonggil.song@samsung.com, "jaegeuk@kernel.org" <jaegeuk@kernel.org>
Cc: chao@kernel.org,
"linux-f2fs-devel@lists.sourceforge.net"
<linux-f2fs-devel@lists.sourceforge.net>,
"linux-kernel@vger.kernel.org" <linux-kernel@vger.kernel.org>,
Dongjin Kim <dongjin_.kim@samsung.com>,
Daejun Park <daejun7.park@samsung.com>
Subject: Re: (2) [PATCH] f2fs: issue multi-device flushes in parallel
Date: Thu, 6 Aug 2026 10:46:53 +0800 [thread overview]
Message-ID: <259edaa7-9747-45dd-9b0c-08dd8ecfabe0@kernel.org> (raw)
In-Reply-To: <20260805224648epcms2p8359e0c2a6b0728dcc3ee3d540c195189@epcms2p8>
On 8/6/26 06:46, Yonggil Song wrote:
> Hi Chao,
>
> Thanks for the review.
>
> On 8/3/26 14:59, Chao Yu wrote:
>> On 7/13/26 13:57, Yonggil Song wrote:
>>> On a multi-device setup, submit_flush_wait() walked the dirty devices
>>> in order and aborted the whole loop on the first device whose flush
>>> failed, leaving the remaining dirty devices un-flushed. Each device
>>> still needs its own data made durable, so a failure on one device must
>>> not skip the others. It also waited for one device's flush to complete
>>> before issuing the next, even though the devices have independent
>>> flush queues and could be flushed concurrently.
>>>
>>> Flush every dirty device best-effort and in parallel instead: build
>>> one PREFLUSH bio per dirty device, submit them all, then wait for
>>> every completion, returning the first error seen (0 if all succeed).
>>> This bounds the flush window by the slowest device rather than the sum
>>> of all of them. No caller depends on the previous early-abort
>>> behaviour -- fsync only checks whether the return value is zero
>>> (fs/f2fs/file.c). The checkpoint path (f2fs_flush_device_cache) is
>>> unaffected; this only touches the fsync flush path.
>>>
>>> Signed-off-by: Yonggil Song <yonggil.song@samsung.com>
>>> ---
>>> fs/f2fs/segment.c | 84 +++++++++++++++++++++++++++++++++++++++++++++++++++----
>>> 1 file changed, 78 insertions(+), 6 deletions(-)
>>>
>>> diff --git a/fs/f2fs/segment.c b/fs/f2fs/segment.c
>>> index d71ddb3ee918..51d7f76e3d1d 100644
>>> --- a/fs/f2fs/segment.c
>>> +++ b/fs/f2fs/segment.c
>>> @@ -566,24 +566,96 @@ static int __submit_flush_wait(struct f2fs_sb_info *sbi,
>>> return ret;
>>> }
>>>
>>> -static int submit_flush_wait(struct f2fs_sb_info *sbi, nid_t ino)
>>> +static void f2fs_flush_end_io(struct bio *bio)
>>> +{
>>> + complete(bio->bi_private);
>>> +}
>>> +
>>> +struct f2fs_flush_bio {
>>> + struct bio bio;
>>> + struct completion wait;
>>> +};
>>> +
>>> +/*
>>> + * Flush every dirty device best-effort: a failure on one device must not
>>> + * skip the flush on the remaining dirty devices, since each device still
>>> + * needs its own data made durable. Report the first error.
>>> + */
>>> +static int submit_flush_wait_serial(struct f2fs_sb_info *sbi, nid_t ino)
>>> {
>>> int ret = 0;
>>> int i;
>>>
>>> - if (!f2fs_is_multi_device(sbi))
>>> - return __submit_flush_wait(sbi, sbi->sb->s_bdev);
>>> + for (i = 0; i < sbi->s_ndevs; i++) {
>>> + int err;
>>> +
>>> + if (!f2fs_is_dirty_device(sbi, ino, i, FLUSH_INO))
>>> + continue;
>>> + err = __submit_flush_wait(sbi, FDEV(i).bdev);
>>> + if (err && !ret)
>>> + ret = err;
>>> + }
>>> + return ret;
>>> +}
>>> +
>>> +/*
>>> + * Same best-effort/first-error contract as submit_flush_wait_serial(), but
>>> + * issue every dirty device's flush before waiting for any of them, so the
>>> + * per-device flush latencies overlap instead of adding up in series. Fall
>>> + * back to the serial path if the bio array cannot be allocated.
>>> + */
>>> +static int submit_flush_wait_parallel(struct f2fs_sb_info *sbi, nid_t ino)
>>> +{
>>> + struct f2fs_flush_bio *flush_bio;
>>> + unsigned long devices = 0;
>>> + int ret = 0;
>>> + int i;
>>> +
>>> + flush_bio = kcalloc(sbi->s_ndevs, sizeof(*flush_bio), GFP_NOFS);
>>
>> Can we allocate flush_bio in local stack? so that we don't need to fallback
>> to submit_flush_wait_serial() for low memory case?
>
> I tried the stack first, but the array is sized by MAX_DEVICES (8) and
> each entry is a struct bio (136 bytes) plus a struct completion (88
> bytes), so the frame grows to 1832 bytes and trips
> -Wframe-larger-than=1024:
>
> fs/f2fs/segment.c: In function 'submit_flush_wait':
> fs/f2fs/segment.c:661:1: warning: the frame size of 1832 bytes is
> larger than 1024 bytes [-Wframe-larger-than=]
Ah, alright.
>
> Since the allocation is small and bounded, how about keeping it on the
> heap but allocating it with GFP_NOFS | __GFP_NOFAIL, which cannot fail?
> That still lets us drop the serial fallback path entirely:
>
> flush_bio = kmalloc(array_size(sbi->s_ndevs, sizeof(*flush_bio)),
> GFP_NOFS | __GFP_NOFAIL);
>
> v2 would also add REQ_SYNC to the flush bios to match what
> submit_bio_wait() sets in the single-device blkdev_issue_flush() path.
Looks fine to me.
>>
>>> + if (!flush_bio)
>>> + return submit_flush_wait_serial(sbi, ino);
>>>
>>> for (i = 0; i < sbi->s_ndevs; i++) {
>>> if (!f2fs_is_dirty_device(sbi, ino, i, FLUSH_INO))
>>> continue;
>>> - ret = __submit_flush_wait(sbi, FDEV(i).bdev);
>>> - if (ret)
>>> - break;
>>> +
>>> + bio_init(&flush_bio[i].bio, FDEV(i).bdev, NULL, 0,
>>> + REQ_OP_WRITE | REQ_PREFLUSH);
>>> + init_completion(&flush_bio[i].wait);
>>> + flush_bio[i].bio.bi_private = &flush_bio[i].wait;
>>> + flush_bio[i].bio.bi_end_io = f2fs_flush_end_io;
>>> + submit_bio(&flush_bio[i].bio);
>>> + devices |= BIT(i);
>>> + }
>>> +
>>> + for (i = 0; i < sbi->s_ndevs; i++) {
>>> + int err;
>>> +
>>> + if (!(devices & BIT(i)))
>>> + continue;
>>> +
>>> + wait_for_completion(&flush_bio[i].wait);
>>> + err = blk_status_to_errno(flush_bio[i].bio.bi_status);
>>> + trace_f2fs_issue_flush(FDEV(i).bdev, test_opt(sbi, NOBARRIER),
>>> + test_opt(sbi, FLUSH_MERGE), err);
>>> + if (!err)
>>> + f2fs_update_iostat(sbi, NULL, FS_FLUSH_IO, 0);
>>> + else if (!ret)
>>> + ret = err;
>>> + bio_uninit(&flush_bio[i].bio);
>>
>> It may leak bio reference previously? For local bio variable, there is no such issue.
>
> I don't think this leaks: the bios are initialized with bio_init() inside
> the kcalloc'ed array (not taken from a bioset), every submitted bio is
> waited for right here, each one gets bio_uninit() after its completion
> is reaped, and the array is kfree'd — so there is no reference left to
> put. The point would be moot in v2 anyway since the fallback path is
> gone.
Oh, you're right, we don't need to put bio due to bio is not allocated from
bio_alloc().
>
> If the __GFP_NOFAIL alternative looks good to you, I'll send v2 that
> way.
Yes, please go ahead.
Thanks,
>
> Thanks,
> Yonggil
>>
>> Thanks,
>>
>>> }
>>> + kfree(flush_bio);
>>> return ret;
>>> }
>>>
>>> +static int submit_flush_wait(struct f2fs_sb_info *sbi, nid_t ino)
>>> +{
>>> + if (!f2fs_is_multi_device(sbi))
>>> + return __submit_flush_wait(sbi, sbi->sb->s_bdev);
>>> +
>>> + return submit_flush_wait_parallel(sbi, ino);
>>> +}
>>> +
>>> static int issue_flush_thread(void *data)
>>> {
>>> struct f2fs_sb_info *sbi = data;
>>
>>
>
prev parent reply other threads:[~2026-08-06 2:46 UTC|newest]
Thread overview: 4+ messages / expand[flat|nested] mbox.gz Atom feed top
[not found] <CGME20260713055711epcms2p712e7add62211e42e995c54a92fd4ff0c@epcms2p7>
2026-07-13 5:57 ` [PATCH] f2fs: issue multi-device flushes in parallel Yonggil Song
2026-08-03 6:59 ` Chao Yu
2026-08-05 22:46 ` Yonggil Song
2026-08-06 2:46 ` Chao Yu [this message]
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=259edaa7-9747-45dd-9b0c-08dd8ecfabe0@kernel.org \
--to=chao@kernel.org \
--cc=daejun7.park@samsung.com \
--cc=dongjin_.kim@samsung.com \
--cc=jaegeuk@kernel.org \
--cc=linux-f2fs-devel@lists.sourceforge.net \
--cc=linux-kernel@vger.kernel.org \
--cc=yonggil.song@samsung.com \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox