From: Chao Yu via Linux-f2fs-devel <linux-f2fs-devel@lists.sourceforge.net>
To: yonggil.song@samsung.com, "jaegeuk@kernel.org" <jaegeuk@kernel.org>
Cc: "linux-kernel@vger.kernel.org" <linux-kernel@vger.kernel.org>,
Dongjin Kim <dongjin_.kim@samsung.com>,
"linux-f2fs-devel@lists.sourceforge.net"
<linux-f2fs-devel@lists.sourceforge.net>
Subject: Re: [f2fs-dev] [PATCH] f2fs: issue multi-device flushes in parallel
Date: Mon, 3 Aug 2026 14:59:31 +0800 [thread overview]
Message-ID: <9fc81e1e-ec13-4cc0-9aed-8270bf1a4157@kernel.org> (raw)
In-Reply-To: <20260713055711epcms2p712e7add62211e42e995c54a92fd4ff0c@epcms2p7>
On 7/13/26 13:57, Yonggil Song wrote:
> On a multi-device setup, submit_flush_wait() walked the dirty devices
> in order and aborted the whole loop on the first device whose flush
> failed, leaving the remaining dirty devices un-flushed. Each device
> still needs its own data made durable, so a failure on one device must
> not skip the others. It also waited for one device's flush to complete
> before issuing the next, even though the devices have independent
> flush queues and could be flushed concurrently.
>
> Flush every dirty device best-effort and in parallel instead: build
> one PREFLUSH bio per dirty device, submit them all, then wait for
> every completion, returning the first error seen (0 if all succeed).
> This bounds the flush window by the slowest device rather than the sum
> of all of them. No caller depends on the previous early-abort
> behaviour -- fsync only checks whether the return value is zero
> (fs/f2fs/file.c). The checkpoint path (f2fs_flush_device_cache) is
> unaffected; this only touches the fsync flush path.
>
> Signed-off-by: Yonggil Song <yonggil.song@samsung.com>
> ---
> fs/f2fs/segment.c | 84 +++++++++++++++++++++++++++++++++++++++++++++++++++----
> 1 file changed, 78 insertions(+), 6 deletions(-)
>
> diff --git a/fs/f2fs/segment.c b/fs/f2fs/segment.c
> index d71ddb3ee918..51d7f76e3d1d 100644
> --- a/fs/f2fs/segment.c
> +++ b/fs/f2fs/segment.c
> @@ -566,24 +566,96 @@ static int __submit_flush_wait(struct f2fs_sb_info *sbi,
> return ret;
> }
>
> -static int submit_flush_wait(struct f2fs_sb_info *sbi, nid_t ino)
> +static void f2fs_flush_end_io(struct bio *bio)
> +{
> + complete(bio->bi_private);
> +}
> +
> +struct f2fs_flush_bio {
> + struct bio bio;
> + struct completion wait;
> +};
> +
> +/*
> + * Flush every dirty device best-effort: a failure on one device must not
> + * skip the flush on the remaining dirty devices, since each device still
> + * needs its own data made durable. Report the first error.
> + */
> +static int submit_flush_wait_serial(struct f2fs_sb_info *sbi, nid_t ino)
> {
> int ret = 0;
> int i;
>
> - if (!f2fs_is_multi_device(sbi))
> - return __submit_flush_wait(sbi, sbi->sb->s_bdev);
> + for (i = 0; i < sbi->s_ndevs; i++) {
> + int err;
> +
> + if (!f2fs_is_dirty_device(sbi, ino, i, FLUSH_INO))
> + continue;
> + err = __submit_flush_wait(sbi, FDEV(i).bdev);
> + if (err && !ret)
> + ret = err;
> + }
> + return ret;
> +}
> +
> +/*
> + * Same best-effort/first-error contract as submit_flush_wait_serial(), but
> + * issue every dirty device's flush before waiting for any of them, so the
> + * per-device flush latencies overlap instead of adding up in series. Fall
> + * back to the serial path if the bio array cannot be allocated.
> + */
> +static int submit_flush_wait_parallel(struct f2fs_sb_info *sbi, nid_t ino)
> +{
> + struct f2fs_flush_bio *flush_bio;
> + unsigned long devices = 0;
> + int ret = 0;
> + int i;
> +
> + flush_bio = kcalloc(sbi->s_ndevs, sizeof(*flush_bio), GFP_NOFS);
Can we allocate flush_bio in local stack? so that we don't need to fallback
to submit_flush_wait_serial() for low memory case?
> + if (!flush_bio)
> + return submit_flush_wait_serial(sbi, ino);
>
> for (i = 0; i < sbi->s_ndevs; i++) {
> if (!f2fs_is_dirty_device(sbi, ino, i, FLUSH_INO))
> continue;
> - ret = __submit_flush_wait(sbi, FDEV(i).bdev);
> - if (ret)
> - break;
> +
> + bio_init(&flush_bio[i].bio, FDEV(i).bdev, NULL, 0,
> + REQ_OP_WRITE | REQ_PREFLUSH);
> + init_completion(&flush_bio[i].wait);
> + flush_bio[i].bio.bi_private = &flush_bio[i].wait;
> + flush_bio[i].bio.bi_end_io = f2fs_flush_end_io;
> + submit_bio(&flush_bio[i].bio);
> + devices |= BIT(i);
> + }
> +
> + for (i = 0; i < sbi->s_ndevs; i++) {
> + int err;
> +
> + if (!(devices & BIT(i)))
> + continue;
> +
> + wait_for_completion(&flush_bio[i].wait);
> + err = blk_status_to_errno(flush_bio[i].bio.bi_status);
> + trace_f2fs_issue_flush(FDEV(i).bdev, test_opt(sbi, NOBARRIER),
> + test_opt(sbi, FLUSH_MERGE), err);
> + if (!err)
> + f2fs_update_iostat(sbi, NULL, FS_FLUSH_IO, 0);
> + else if (!ret)
> + ret = err;
> + bio_uninit(&flush_bio[i].bio);
It may leak bio reference previously? For local bio variable, there is no such issue.
Thanks,
> }
> + kfree(flush_bio);
> return ret;
> }
>
> +static int submit_flush_wait(struct f2fs_sb_info *sbi, nid_t ino)
> +{
> + if (!f2fs_is_multi_device(sbi))
> + return __submit_flush_wait(sbi, sbi->sb->s_bdev);
> +
> + return submit_flush_wait_parallel(sbi, ino);
> +}
> +
> static int issue_flush_thread(void *data)
> {
> struct f2fs_sb_info *sbi = data;
_______________________________________________
Linux-f2fs-devel mailing list
Linux-f2fs-devel@lists.sourceforge.net
https://lists.sourceforge.net/lists/listinfo/linux-f2fs-devel
WARNING: multiple messages have this Message-ID (diff)
From: Chao Yu <chao@kernel.org>
To: yonggil.song@samsung.com, "jaegeuk@kernel.org" <jaegeuk@kernel.org>
Cc: chao@kernel.org,
"linux-f2fs-devel@lists.sourceforge.net"
<linux-f2fs-devel@lists.sourceforge.net>,
"linux-kernel@vger.kernel.org" <linux-kernel@vger.kernel.org>,
Dongjin Kim <dongjin_.kim@samsung.com>,
Daejun Park <daejun7.park@samsung.com>
Subject: Re: [PATCH] f2fs: issue multi-device flushes in parallel
Date: Mon, 3 Aug 2026 14:59:31 +0800 [thread overview]
Message-ID: <9fc81e1e-ec13-4cc0-9aed-8270bf1a4157@kernel.org> (raw)
In-Reply-To: <20260713055711epcms2p712e7add62211e42e995c54a92fd4ff0c@epcms2p7>
On 7/13/26 13:57, Yonggil Song wrote:
> On a multi-device setup, submit_flush_wait() walked the dirty devices
> in order and aborted the whole loop on the first device whose flush
> failed, leaving the remaining dirty devices un-flushed. Each device
> still needs its own data made durable, so a failure on one device must
> not skip the others. It also waited for one device's flush to complete
> before issuing the next, even though the devices have independent
> flush queues and could be flushed concurrently.
>
> Flush every dirty device best-effort and in parallel instead: build
> one PREFLUSH bio per dirty device, submit them all, then wait for
> every completion, returning the first error seen (0 if all succeed).
> This bounds the flush window by the slowest device rather than the sum
> of all of them. No caller depends on the previous early-abort
> behaviour -- fsync only checks whether the return value is zero
> (fs/f2fs/file.c). The checkpoint path (f2fs_flush_device_cache) is
> unaffected; this only touches the fsync flush path.
>
> Signed-off-by: Yonggil Song <yonggil.song@samsung.com>
> ---
> fs/f2fs/segment.c | 84 +++++++++++++++++++++++++++++++++++++++++++++++++++----
> 1 file changed, 78 insertions(+), 6 deletions(-)
>
> diff --git a/fs/f2fs/segment.c b/fs/f2fs/segment.c
> index d71ddb3ee918..51d7f76e3d1d 100644
> --- a/fs/f2fs/segment.c
> +++ b/fs/f2fs/segment.c
> @@ -566,24 +566,96 @@ static int __submit_flush_wait(struct f2fs_sb_info *sbi,
> return ret;
> }
>
> -static int submit_flush_wait(struct f2fs_sb_info *sbi, nid_t ino)
> +static void f2fs_flush_end_io(struct bio *bio)
> +{
> + complete(bio->bi_private);
> +}
> +
> +struct f2fs_flush_bio {
> + struct bio bio;
> + struct completion wait;
> +};
> +
> +/*
> + * Flush every dirty device best-effort: a failure on one device must not
> + * skip the flush on the remaining dirty devices, since each device still
> + * needs its own data made durable. Report the first error.
> + */
> +static int submit_flush_wait_serial(struct f2fs_sb_info *sbi, nid_t ino)
> {
> int ret = 0;
> int i;
>
> - if (!f2fs_is_multi_device(sbi))
> - return __submit_flush_wait(sbi, sbi->sb->s_bdev);
> + for (i = 0; i < sbi->s_ndevs; i++) {
> + int err;
> +
> + if (!f2fs_is_dirty_device(sbi, ino, i, FLUSH_INO))
> + continue;
> + err = __submit_flush_wait(sbi, FDEV(i).bdev);
> + if (err && !ret)
> + ret = err;
> + }
> + return ret;
> +}
> +
> +/*
> + * Same best-effort/first-error contract as submit_flush_wait_serial(), but
> + * issue every dirty device's flush before waiting for any of them, so the
> + * per-device flush latencies overlap instead of adding up in series. Fall
> + * back to the serial path if the bio array cannot be allocated.
> + */
> +static int submit_flush_wait_parallel(struct f2fs_sb_info *sbi, nid_t ino)
> +{
> + struct f2fs_flush_bio *flush_bio;
> + unsigned long devices = 0;
> + int ret = 0;
> + int i;
> +
> + flush_bio = kcalloc(sbi->s_ndevs, sizeof(*flush_bio), GFP_NOFS);
Can we allocate flush_bio in local stack? so that we don't need to fallback
to submit_flush_wait_serial() for low memory case?
> + if (!flush_bio)
> + return submit_flush_wait_serial(sbi, ino);
>
> for (i = 0; i < sbi->s_ndevs; i++) {
> if (!f2fs_is_dirty_device(sbi, ino, i, FLUSH_INO))
> continue;
> - ret = __submit_flush_wait(sbi, FDEV(i).bdev);
> - if (ret)
> - break;
> +
> + bio_init(&flush_bio[i].bio, FDEV(i).bdev, NULL, 0,
> + REQ_OP_WRITE | REQ_PREFLUSH);
> + init_completion(&flush_bio[i].wait);
> + flush_bio[i].bio.bi_private = &flush_bio[i].wait;
> + flush_bio[i].bio.bi_end_io = f2fs_flush_end_io;
> + submit_bio(&flush_bio[i].bio);
> + devices |= BIT(i);
> + }
> +
> + for (i = 0; i < sbi->s_ndevs; i++) {
> + int err;
> +
> + if (!(devices & BIT(i)))
> + continue;
> +
> + wait_for_completion(&flush_bio[i].wait);
> + err = blk_status_to_errno(flush_bio[i].bio.bi_status);
> + trace_f2fs_issue_flush(FDEV(i).bdev, test_opt(sbi, NOBARRIER),
> + test_opt(sbi, FLUSH_MERGE), err);
> + if (!err)
> + f2fs_update_iostat(sbi, NULL, FS_FLUSH_IO, 0);
> + else if (!ret)
> + ret = err;
> + bio_uninit(&flush_bio[i].bio);
It may leak bio reference previously? For local bio variable, there is no such issue.
Thanks,
> }
> + kfree(flush_bio);
> return ret;
> }
>
> +static int submit_flush_wait(struct f2fs_sb_info *sbi, nid_t ino)
> +{
> + if (!f2fs_is_multi_device(sbi))
> + return __submit_flush_wait(sbi, sbi->sb->s_bdev);
> +
> + return submit_flush_wait_parallel(sbi, ino);
> +}
> +
> static int issue_flush_thread(void *data)
> {
> struct f2fs_sb_info *sbi = data;
next prev parent reply other threads:[~2026-08-03 6:59 UTC|newest]
Thread overview: 8+ messages / expand[flat|nested] mbox.gz Atom feed top
[not found] <CGME20260713055711epcms2p712e7add62211e42e995c54a92fd4ff0c@epcms2p7>
2026-07-13 5:57 ` [f2fs-dev] [PATCH] f2fs: issue multi-device flushes in parallel Yonggil Song
2026-07-13 5:57 ` Yonggil Song
2026-08-03 6:59 ` Chao Yu via Linux-f2fs-devel [this message]
2026-08-03 6:59 ` Chao Yu
2026-08-05 22:46 ` [f2fs-dev] (2) " Yonggil Song
2026-08-05 22:46 ` Yonggil Song
2026-08-06 2:46 ` [f2fs-dev] (2) " Chao Yu via Linux-f2fs-devel
2026-08-06 2:46 ` Chao Yu
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=9fc81e1e-ec13-4cc0-9aed-8270bf1a4157@kernel.org \
--to=linux-f2fs-devel@lists.sourceforge.net \
--cc=chao@kernel.org \
--cc=dongjin_.kim@samsung.com \
--cc=jaegeuk@kernel.org \
--cc=linux-kernel@vger.kernel.org \
--cc=yonggil.song@samsung.com \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is an external index of several public inboxes,
see mirroring instructions on how to clone and mirror
all data and code used by this external index.