All of lore.kernel.org
 help / color / mirror / Atom feed
From: John Garry <john.g.garry@oracle.com>
To: sashiko-reviews@lists.linux.dev
Cc: linux-scsi@vger.kernel.org
Subject: Re: [PATCH v6 06/27] libmultipath: Add delayed removal support
Date: Wed, 29 Jul 2026 13:15:03 +0100	[thread overview]
Message-ID: <b9e1fe6b-3d89-4e4c-99d2-bfda982243f3@oracle.com> (raw)
In-Reply-To: <20260729120812.555561F000E9@smtp.kernel.org>

On 29/07/2026 13:08, sashiko-bot@kernel.org wrote:
> Thank you for your contribution! Sashiko AI review found 6 potential issue(s) to consider:
> - [High] mpath_remove_head_work() unconditionally removes the multipath head without verifying if paths have been concurrently added, leading to destruction of an active device.
> - [High] Missing cancellation of remove_work during multipath head teardown leads to a use-after-free when the delayed work expires.
> - [High] mpath_head->drv_module is never initialized, leading to silent failures or crashes on module unload.
> - [High] Disabling delayed removal via sysfs leaves queued I/Os stuck indefinitely.
> - [Medium] Module reference count leak when mpath_can_remove_head() is called multiple times.
> - [Medium] Large delayed removal values cause immediate removal instead of a long delay.
> --
> 
> --- Patch [6]: [PATCH v6 06/27] libmultipath: Add delayed removal support ---
> commit 87d6b13fc148098f06b11013c5f12523eb0d477f
> Author: John Garry <john.garry@linux.dev>
> 
> libmultipath: Add delayed removal support
> 
> Add support for delayed removal, same as exists for NVMe, to keep the
> multipath disk and cdev present for intermittent periods of no available
> path.
> 
>> diff --git a/lib/multipath.c b/lib/multipath.c
>> index 6b8f6e846987a..fada92bf9754c 100644
>> --- a/lib/multipath.c
>> +++ b/lib/multipath.c
>> @@ -66,6 +66,9 @@ int mpath_add_device(struct mpath_device *mpath_device, struct gendisk *disk,
>>   	list_add_tail_rcu(&mpath_device->siblings, &mpath_head->dev_list);
>>   	mutex_unlock(&mpath_head->lock);
>>   
>> +	if (cancel_delayed_work(&mpath_head->remove_work))
>> +		module_put(mpath_head->drv_module);
>> +
>>   	return 0;
>>   }
> 
> [Severity: High]
> Can this race with the removal work?
> 
> If the delayed work is already executing, cancel_delayed_work() returns
> false, but mpath_add_device() still returns success. Meanwhile, could
> mpath_remove_head_work() execute without taking mpath_head->lock or
> verifying if the device list is empty, unintentionally destroying the
> multipath head while the driver believes a new path was successfully
> added?
> 

the removal callback should take care of this


> [ ... ]
> 
>> @@ -481,6 +493,39 @@ static void mpath_requeue_work(struct work_struct *work)
>>   	}
>>   }
>>   
>> +bool mpath_can_remove_head(struct mpath_head *mpath_head)
>> +{
>> +	unsigned long delay;
>> +	bool remove = false;
>> +
>> +	mutex_lock(&mpath_head->lock);
>> +	/*
>> +	 * Ensure that no one could remove this module while the head
>> +	 * remove work is pending.
>> +	 */
>> +	if (mpath_head_queue_if_no_path(mpath_head) &&
>> +	    !check_mul_overflow(mpath_head->delayed_removal_secs, HZ, &delay) &&
> 
> [Severity: Medium]
> What happens if a user configures a very large delayed removal value?
> 
> If check_mul_overflow() detects an overflow (e.g., on 32-bit systems where
> a large sysfs input overflows unsigned long), it fails the condition and
> falls through to the else block, setting remove = true. Does this cause an
> immediate teardown instead of clamping the delay to a maximum value?

yes, so don't use a silly big value

> 
>> +	    try_module_get(mpath_head->drv_module)) {
>> +		mod_delayed_work(mpath_wq, &mpath_head->remove_work, delay);
> 
> [Severity: Medium]
> Could this lead to a module reference leak?
> 
> If mod_delayed_work() returns true (meaning the work was already pending
> and its timer was just modified), it doesn't queue a new work item.
> Since try_module_get() was called unconditionally beforehand, and the
> work callback only drops one module reference, do subsequent calls to
> mpath_can_remove_head() leak a module reference permanently?

it should not happen - that is the job of the caller to ensure

> 
>> +	} else {
>> +		remove = true;
>> +	}
>> +
>> +	mutex_unlock(&mpath_head->lock);
>> +	return remove;
>> +}
> 
> [ ... ]
> 
>> @@ -529,6 +574,9 @@ int mpath_alloc_head_disk(struct mpath_head *mpath_head,
>>   	mpath_head->disk->private_data = mpath_head;
>>   	mpath_head->disk->fops = &mpath_ops;
>>   
>> +	INIT_DELAYED_WORK(&mpath_head->remove_work, mpath_remove_head_work);
>> +	mpath_head->delayed_removal_secs = 0;
>> +
> 
> [Severity: High]
> Is drv_module properly initialized anywhere?

yes

> 
> It seems drv_module is added to struct mpath_head but never assigned in
> mpath_alloc_head_disk() or elsewhere in this patch. If it defaults to
> NULL, try_module_get(NULL) succeeds but pins no module. Could this allow
> the driver to unload while the delayed work is still pending, leading to
> a kernel panic when the timer fires?
> 
> [Severity: High]
> Is the remove_work guaranteed to be cancelled during object teardown?

it must be assigned

> 
> The standard teardown sequence calls mpath_put_disk() and
> mpath_head_uninit(). If cancel_delayed_work_sync(&mpath_head->remove_work)
> is omitted during teardown, could the timer remain active on the
> workqueue and execute on the freed mpath_head structure later?
> 
>>   	set_bit(GD_SUPPRESS_PART_SCAN, &mpath_head->disk->state);
>>   
>>   	return 0;
> 
> [ ... ]
> 
>> @@ -572,6 +620,47 @@ void mpath_device_set_live(struct mpath_device *mpath_device)
>>   }
>>   EXPORT_SYMBOL_GPL(mpath_device_set_live);
>>   
>> +ssize_t mpath_delayed_removal_secs_show(struct mpath_head *mpath_head,
>> +					char *buf)
>> +{
>> +	int ret;
>> +
>> +	mutex_lock(&mpath_head->lock);
>> +	ret = sysfs_emit(buf, "%u\n", mpath_head->delayed_removal_secs);
>> +	mutex_unlock(&mpath_head->lock);
>> +
>> +	return ret;
>> +}
>> +EXPORT_SYMBOL_GPL(mpath_delayed_removal_secs_show);
>> +
>> +ssize_t mpath_delayed_removal_secs_store(struct mpath_head *mpath_head,
>> +			const char *buf, size_t count)
>> +{
>> +	unsigned int sec;
>> +	ssize_t ret;
>> +
>> +	ret = kstrtouint(buf, 0, &sec);
>> +	if (ret < 0)
>> +		return ret;
>> +
>> +	mutex_lock(&mpath_head->lock);
>> +	mpath_head->delayed_removal_secs = sec;
>> +	if (sec)
>> +		set_bit(MPATH_HEAD_QUEUE_IF_NO_PATH, &mpath_head->flags);
>> +	else
>> +		clear_bit(MPATH_HEAD_QUEUE_IF_NO_PATH, &mpath_head->flags);
>> +	mutex_unlock(&mpath_head->lock);
>> +
>> +	/*
>> +	 * Ensure that update to MPATH_HEAD_QUEUE_IF_NO_PATH is seen
>> +	 * by its reader.
>> +	 */
>> +	mpath_synchronize(mpath_head);
>> +
>> +	return count;
>> +}
> 
> [Severity: High]
> If a user disables delayed removal by setting delayed_removal_secs to 0 via
> sysfs while paths are missing, how do pending I/Os get handled?
> 
> We clear the flag and call mpath_synchronize(), but shouldn't we also
> call mpath_schedule_requeue_work() here to flush and fail the I/Os that
> are already parked in mpath_head->requeue_list?

I don't see how this is related to requeue work

> 


  reply	other threads:[~2026-07-29 12:15 UTC|newest]

Thread overview: 63+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-07-29 10:50 [PATCH v6 00/27] Native SCSI Multipath support John Garry
2026-07-29 10:50 ` [PATCH v6 01/27] libmultipath: Add initial framework John Garry
2026-07-29 10:50 ` [PATCH v6 02/27] libmultipath: Add basic gendisk support John Garry
2026-07-29 11:20   ` sashiko-bot
2026-07-29 11:36     ` John Garry
2026-07-29 10:50 ` [PATCH v6 03/27] libmultipath: Add path selection support John Garry
2026-07-29 11:33   ` sashiko-bot
2026-07-29 11:41     ` John Garry
2026-07-29 10:50 ` [PATCH v6 04/27] libmultipath: Add bio handling John Garry
2026-07-29 10:50 ` [PATCH v6 05/27] libmultipath: Add support for mpath_device management John Garry
2026-07-29 11:57   ` sashiko-bot
2026-07-29 12:11     ` John Garry
2026-07-29 10:50 ` [PATCH v6 06/27] libmultipath: Add delayed removal support John Garry
2026-07-29 12:08   ` sashiko-bot
2026-07-29 12:15     ` John Garry [this message]
2026-07-29 10:50 ` [PATCH v6 07/27] libmultipath: Add sysfs helpers John Garry
2026-07-29 12:28   ` sashiko-bot
2026-07-29 12:51     ` John Garry
2026-07-29 10:50 ` [PATCH v6 08/27] libmultipath: Add support for block device IOCTL John Garry
2026-07-29 12:39   ` sashiko-bot
2026-07-29 12:53     ` John Garry
2026-07-29 10:50 ` [PATCH v6 09/27] libmultipath: Add mpath_bdev_getgeo() John Garry
2026-07-29 10:50 ` [PATCH v6 10/27] libmultipath: Add mpath_bdev_get_unique_id() John Garry
2026-07-29 10:50 ` [PATCH v6 11/27] scsi-multipath: introduce basic SCSI device support John Garry
2026-07-29 10:50 ` [PATCH v6 12/27] scsi-multipath: introduce scsi_device head structure John Garry
2026-07-29 13:46   ` sashiko-bot
2026-07-29 14:06     ` John Garry
2026-07-29 10:50 ` [PATCH v6 13/27] scsi-multipath: provide sysfs link from to scsi_device John Garry
2026-07-29 10:50 ` [PATCH v6 14/27] scsi-multipath: support iopolicy John Garry
2026-07-29 14:07   ` sashiko-bot
2026-07-29 14:11     ` John Garry
2026-07-29 10:50 ` [PATCH v6 15/27] scsi-multipath: clone each bio John Garry
2026-07-29 14:23   ` sashiko-bot
2026-07-29 14:25     ` John Garry
2026-07-29 10:50 ` [PATCH v6 16/27] scsi-multipath: clear path when device is blocked John Garry
2026-07-29 14:42   ` sashiko-bot
2026-07-29 14:51     ` John Garry
2026-07-29 10:50 ` [PATCH v6 17/27] scsi-multipath: revalidate paths upon device unblock John Garry
2026-07-29 14:54   ` sashiko-bot
2026-07-29 15:27     ` John Garry
2026-07-29 10:50 ` [PATCH v6 18/27] scsi-multipath: failover handling John Garry
2026-07-29 15:14   ` sashiko-bot
2026-07-29 15:29     ` John Garry
2026-07-29 10:50 ` [PATCH v6 19/27] scsi-multipath: provide callbacks for path state John Garry
2026-07-29 15:43   ` sashiko-bot
2026-07-29 16:54     ` John Garry
2026-07-29 10:51 ` [PATCH v6 20/27] scsi-multipath: add scsi_mpath_{start,end}_request() John Garry
2026-07-29 16:12   ` sashiko-bot
2026-07-29 16:55     ` John Garry
2026-07-29 10:51 ` [PATCH v6 21/27] scsi-multipath: add delayed disk removal support John Garry
2026-07-29 16:26   ` sashiko-bot
2026-07-29 16:56     ` John Garry
2026-07-29 10:51 ` [PATCH v6 22/27] scsi: sd: add multipath disk class John Garry
2026-07-29 16:33   ` sashiko-bot
2026-07-29 16:57     ` John Garry
2026-07-29 10:51 ` [PATCH v6 23/27] scsi: sd: add multipath disk attr groups John Garry
2026-07-29 16:48   ` sashiko-bot
2026-07-29 17:05     ` John Garry
2026-07-29 10:51 ` [PATCH v6 24/27] scsi: sd: support multipath disk John Garry
2026-07-29 17:01   ` sashiko-bot
2026-07-29 10:51 ` [PATCH v6 25/27] scsi: sd: add mpath_dev file John Garry
2026-07-29 10:51 ` [PATCH v6 26/27] scsi: sd: add mpath_numa_nodes dev attribute John Garry
2026-07-29 10:51 ` [PATCH v6 27/27] scsi: sd: add mpath_queue_depth " John Garry

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=b9e1fe6b-3d89-4e4c-99d2-bfda982243f3@oracle.com \
    --to=john.g.garry@oracle.com \
    --cc=linux-scsi@vger.kernel.org \
    --cc=sashiko-reviews@lists.linux.dev \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is an external index of several public inboxes,
see mirroring instructions on how to clone and mirror
all data and code used by this external index.