From: Yu Kuai <yukuai1@huaweicloud.com>
To: Yu Kuai <yukuai1@huaweicloud.com>, xni@redhat.com, song@kernel.org
Cc: linux-raid@vger.kernel.org, linux-kernel@vger.kernel.org,
yi.zhang@huawei.com, yangerkun@huawei.com,
"yukuai (C)" <yukuai3@huawei.com>
Subject: Re: [PATCH -next v2 3/7] md: delay choosing sync direction to md_start_sync()
Date: Tue, 15 Aug 2023 14:00:28 +0800 [thread overview]
Message-ID: <bb11d6ca-978a-8e1d-e721-d9d84c9dc5e3@huaweicloud.com> (raw)
In-Reply-To: <20230815030957.509535-4-yukuai1@huaweicloud.com>
Hi,
在 2023/08/15 11:09, Yu Kuai 写道:
> From: Yu Kuai <yukuai3@huawei.com>
>
> Before this patch, for read-write array:
>
> 1) md_check_recover() found that something need to be done, and it'll
> try to grab 'reconfig_mutex'. The case that md_check_recover() need
> to do something:
> - array is not suspend;
> - super_block need to be updated;
> - 'MD_RECOVERY_NEEDED' or ''MD_RECOVERY_DONE' is set;
> - unusual case related to safemode;
>
> 2) if 'MD_RECOVERY_RUNNING' is not set, and 'MD_RECOVERY_NEEDED' is set,
> md_check_recover() will try to choose a sync direction, and then
> queue a work md_start_sync().
>
> 3) md_start_sync() register sync_thread;
>
> After this patch,
>
> 1) is the same;
> 2) if 'MD_RECOVERY_RUNNING' is not set, and 'MD_RECOVERY_NEEDED' is set,
> queue a work md_start_sync() directly;
> 3) md_start_sync() will try to choose a sync direction, and then
> register sync_thread();
>
> Because 'MD_RECOVERY_RUNNING' is cleared when sync_thread is done, 2)
> and 3) is always ran in serial and they can never concurrent, this
> change should not introduce any behavior change for now.
>
> Also fix a problem that md_start_sync() can clear 'MD_RECOVERY_RUNNING'
> without protection in error path, which might affect the logical in
> md_check_recovery().
>
> The advantage to change this is that array reconfiguration is
> independent from daemon now, and it'll be much easier to synchronize it
> with io, consider that io may rely on daemon thread to be done.
>
> Signed-off-by: Yu Kuai <yukuai3@huawei.com>
> ---
> drivers/md/md.c | 70 ++++++++++++++++++++++++++-----------------------
> 1 file changed, 37 insertions(+), 33 deletions(-)
>
> diff --git a/drivers/md/md.c b/drivers/md/md.c
> index 4846ff6d25b0..03615b0e9fe1 100644
> --- a/drivers/md/md.c
> +++ b/drivers/md/md.c
> @@ -9291,6 +9291,22 @@ static bool md_choose_sync_direction(struct mddev *mddev, int *spares)
> static void md_start_sync(struct work_struct *ws)
> {
> struct mddev *mddev = container_of(ws, struct mddev, sync_work);
> + int spares = 0;
> +
> + mddev_lock_nointr(mddev);
> +
> + if (!md_choose_sync_direction(mddev, &spares))
> + goto not_running;
> +
> + if (!mddev->pers->sync_request)
> + goto not_running;
> +
> + /*
> + * We are adding a device or devices to an array which has the bitmap
> + * stored on all devices. So make sure all bitmap pages get written.
> + */
> + if (spares)
> + md_bitmap_write_all(mddev->bitmap);
>
> rcu_assign_pointer(mddev->sync_thread,
> md_register_thread(md_do_sync, mddev, "resync"));
> @@ -9298,20 +9314,27 @@ static void md_start_sync(struct work_struct *ws)
> pr_warn("%s: could not start resync thread...\n",
> mdname(mddev));
> /* leave the spares where they are, it shouldn't hurt */
> - clear_bit(MD_RECOVERY_SYNC, &mddev->recovery);
> - clear_bit(MD_RECOVERY_RESHAPE, &mddev->recovery);
> - clear_bit(MD_RECOVERY_REQUESTED, &mddev->recovery);
> - clear_bit(MD_RECOVERY_CHECK, &mddev->recovery);
> - clear_bit(MD_RECOVERY_RUNNING, &mddev->recovery);
> - wake_up(&resync_wait);
> - if (test_and_clear_bit(MD_RECOVERY_RECOVER,
> - &mddev->recovery))
> - if (mddev->sysfs_action)
> - sysfs_notify_dirent_safe(mddev->sysfs_action);
> - } else
> - md_wakeup_thread(mddev->sync_thread);
> + goto not_running;
> + }
> +
> + mddev_unlock(mddev);
> + md_wakeup_thread(mddev->sync_thread);
> sysfs_notify_dirent_safe(mddev->sysfs_action);
> md_new_event();
> + return;
> +
> +not_running:
> + clear_bit(MD_RECOVERY_SYNC, &mddev->recovery);
> + clear_bit(MD_RECOVERY_RESHAPE, &mddev->recovery);
> + clear_bit(MD_RECOVERY_REQUESTED, &mddev->recovery);
> + clear_bit(MD_RECOVERY_CHECK, &mddev->recovery);
> + clear_bit(MD_RECOVERY_RUNNING, &mddev->recovery);
> + mddev_unlock(mddev);
> +
> + wake_up(&resync_wait);
> + if (test_and_clear_bit(MD_RECOVERY_RECOVER, &mddev->recovery) &&
> + mddev->sysfs_action)
> + sysfs_notify_dirent_safe(mddev->sysfs_action);
> }
>
> /*
> @@ -9379,7 +9402,6 @@ void md_check_recovery(struct mddev *mddev)
> return;
>
> if (mddev_trylock(mddev)) {
> - int spares = 0;
> bool try_set_sync = mddev->safemode != 0;
>
> if (!mddev->external && mddev->safemode == 1)
> @@ -9467,29 +9489,11 @@ void md_check_recovery(struct mddev *mddev)
> clear_bit(MD_RECOVERY_DONE, &mddev->recovery);
>
> if (!test_and_clear_bit(MD_RECOVERY_NEEDED, &mddev->recovery) ||
> - test_bit(MD_RECOVERY_FROZEN, &mddev->recovery))
> - goto not_running;
> - if (!md_choose_sync_direction(mddev, &spares))
> - goto not_running;
> - if (mddev->pers->sync_request) {
> - if (spares) {
> - /* We are adding a device or devices to an array
> - * which has the bitmap stored on all devices.
> - * So make sure all bitmap pages get written
> - */
> - md_bitmap_write_all(mddev->bitmap);
> - }
> + test_bit(MD_RECOVERY_FROZEN, &mddev->recovery)) {
Sorry that I made a mistake here while rebasing v2, here should be
!test_bit(MD_RECOVERY_FROZEN, &mddev->recovery)
With this fixed, there are no new regression for mdadm tests using loop
devicein my VM.
Thanks,
Kuai
> queue_work(md_misc_wq, &mddev->sync_work);
> - goto unlock;
> - }
> - not_running:
> - if (!mddev->sync_thread) {
> + } else {
> clear_bit(MD_RECOVERY_RUNNING, &mddev->recovery);
> wake_up(&resync_wait);
> - if (test_and_clear_bit(MD_RECOVERY_RECOVER,
> - &mddev->recovery))
> - if (mddev->sysfs_action)
> - sysfs_notify_dirent_safe(mddev->sysfs_action);
> }
> unlock:
> wake_up(&mddev->sb_wait);
>
next prev parent reply other threads:[~2023-08-15 6:01 UTC|newest]
Thread overview: 20+ messages / expand[flat|nested] mbox.gz Atom feed top
2023-08-15 3:09 [PATCH -next v2 0/7] md: make rdev addition and removal independent from daemon thread Yu Kuai
2023-08-15 3:09 ` [PATCH -next v2 1/7] md: use separate work_struct for md_start_sync() Yu Kuai
2023-08-15 3:09 ` [PATCH -next v2 2/7] md: factor out a helper to choose sync direction from md_check_recovery() Yu Kuai
2023-08-17 7:58 ` Mariusz Tkaczyk
2023-08-17 21:49 ` Song Liu
2023-08-20 1:44 ` Yu Kuai
2023-08-15 3:09 ` [PATCH -next v2 3/7] md: delay choosing sync direction to md_start_sync() Yu Kuai
2023-08-15 6:00 ` Yu Kuai [this message]
2023-08-15 15:54 ` Song Liu
2023-08-16 1:07 ` Yu Kuai
2023-08-17 21:53 ` Song Liu
2023-08-20 1:45 ` Yu Kuai
2023-08-16 6:38 ` Xiao Ni
2023-08-20 2:04 ` Yu Kuai
2023-08-15 3:09 ` [PATCH -next v2 4/7] md: factor out a helper rdev_removeable() from remove_and_add_spares() Yu Kuai
2023-08-15 3:09 ` [PATCH -next v2 5/7] md: factor out a helper rdev_is_spare() " Yu Kuai
2023-08-15 3:09 ` [PATCH -next v2 6/7] md: factor out a helper rdev_addable() " Yu Kuai
2023-08-15 3:09 ` [PATCH -next v2 7/7] md: delay remove_and_add_spares() for read only array to md_start_sync() Yu Kuai
2023-08-16 7:18 ` Xiao Ni
2023-08-20 2:19 ` Yu Kuai
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=bb11d6ca-978a-8e1d-e721-d9d84c9dc5e3@huaweicloud.com \
--to=yukuai1@huaweicloud.com \
--cc=linux-kernel@vger.kernel.org \
--cc=linux-raid@vger.kernel.org \
--cc=song@kernel.org \
--cc=xni@redhat.com \
--cc=yangerkun@huawei.com \
--cc=yi.zhang@huawei.com \
--cc=yukuai3@huawei.com \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox;
as well as URLs for NNTP newsgroup(s).