Linux SCSI subsystem development
 help / color / mirror / Atom feed
From: John Garry <john.g.garry@oracle.com>
To: sashiko-reviews@lists.linux.dev
Cc: linux-scsi@vger.kernel.org
Subject: Re: [PATCH v5 03/27] libmultipath: Add path selection support
Date: Mon, 27 Jul 2026 16:20:57 +0100	[thread overview]
Message-ID: <362edeb3-20be-415d-8719-163cfcee99fb@oracle.com> (raw)
In-Reply-To: <20260727150304.A28121F00A3E@smtp.kernel.org>

On 27/07/2026 16:03, sashiko-bot@kernel.org wrote:
> Thank you for your contribution! Sashiko AI review found 4 potential issue(s) to consider:
> - [High] Round-robin path selection can enter an infinite loop under the SRCU read lock if the current path is concurrently removed.
> - [High] If `blk_alloc_disk()` fails, an `ERR_PTR` is left in `mpath_head->disk`, which will cause a kernel panic during cleanup.
> - [High] Missing requeue list flushes in `mpath_device_set_live` and `mpath_remove_disk` cause IO stalls and bio leaks.
> - [High] `mpath_round_robin_path` incorrectly returns a deleted path when the path list becomes singular.
> --
> 
> --- Patch [3]: [PATCH v5 03/27] libmultipath: Add path selection support ---
> commit 88605e74a8a44a0f658b820940be9c7cdbbebc25
> Author: John Garry <john.garry@linux.dev>
> 
> libmultipath: Add path selection support
> 
> This commit adds path selection support for NVMe ANA, abstracting it into
> enum mpath_access_state. It introduces round-robin, NUMA, and queue-depth
> selection modes along with iopolicy configuration functions.
> 
>> diff --git a/lib/multipath.c b/lib/multipath.c
>> index d2e334259f3c1..2b22c74760868 100644
>> --- a/lib/multipath.c
>> +++ b/lib/multipath.c
> 
> [ ... ]
> 
>> +static struct mpath_device *mpath_round_robin_path(
>> +				struct mpath_head *mpath_head)
>> +{
>> +	struct mpath_device *mpath_device, *found = NULL;
>> +	int node = numa_node_id();
>> +	enum mpath_access_state access_state_old;
>> +	struct mpath_device *old =
>> +			srcu_dereference(mpath_head->current_path[node],
>> +				&mpath_head->srcu);
>> +
>> +	if (unlikely(!old))
>> +		return __mpath_find_path(mpath_head, node);
>> +
>> +	if (list_is_singular(&mpath_head->dev_list)) {
>> +		if (mpath_path_is_disabled(mpath_head, old))
>> +			return NULL;
>> +		return old;
>> +	}
> 
> [Severity: High]
> If the cached path (old) was concurrently unlinked from the list, is it
> possible that list_is_singular() evaluates to true because there is exactly
> one other valid path remaining? If so, does this mistakenly return the
> unlinked old path (or NULL) instead of evaluating the single remaining valid
> path?
> 
>> +
>> +	for (mpath_device = mpath_next_dev(mpath_head, old);
>> +	    mpath_device && mpath_device != old;
>> +	    mpath_device = mpath_next_dev(mpath_head, mpath_device)) {
> 
> [Severity: High]
> If the currently cached path (old) is concurrently removed via
> list_del_rcu(), it is unlinked from dev_list. Since mpath_next_dev()
> iterates through the list, will this loop fail to encounter the unlinked
> old pointer? If all remaining paths are non-optimized or disabled, could
> this result in an infinite loop inside an SRCU read-side critical section?

locking is added later, which is being ignored

> 
>> +
>> +		if (mpath_path_is_disabled(mpath_head, mpath_device))
>> +			continue;
> 
> [ ... ]
> 
>> @@ -75,6 +312,7 @@ void mpath_remove_disk(struct mpath_head *mpath_head)
>>   	if (test_and_clear_bit(MPATH_HEAD_DISK_LIVE, &mpath_head->flags)) {
>>   		struct gendisk *disk = mpath_head->disk;
>>   
>> +		mpath_synchronize(mpath_head);
>>   		del_gendisk(disk);
>>   	}
> 
> [Severity: High]
> When the disk is removed, are there potentially stranded bios left in the
> mpath_head->requeue_list? Without a call to mpath_schedule_requeue_work()
> here to flush the list, do these queued bios leak?

that is done elsewhere

> 
>>   }
> 
> [ ... ]
> 
>> @@ -96,7 +334,7 @@ int mpath_alloc_head_disk(struct mpath_head *mpath_head,
>>   			struct queue_limits *lim, int numa_node)
>>   {
>>   	if (!mpath_head->disk_groups || !mpath_head->parent ||
>> -	    mpath_head->disk)
>> +	     !mpath_head->iopolicy || mpath_head->disk)
>>   		return -EINVAL;
>>   
>>   	mpath_head->disk = blk_alloc_disk(lim, numa_node);
> 
> [Severity: High]
> If blk_alloc_disk() fails, does it leave an ERR_PTR in mpath_head->disk?
> If the error pointer is not cleared here, would the cleanup path in
> mpath_put_disk() mistakenly evaluate "if (!mpath_head->disk)" as false
> and attempt to call put_disk() on the error pointer?

already commented on

> 
> [ ... ]
> 
>> @@ -130,6 +368,19 @@ void mpath_device_set_live(struct mpath_device *mpath_device)
>>   		}
>>   		queue_work(mpath_wq, &mpath_head->partition_scan_work);
>>   	}
>> +
>> +	mutex_lock(&mpath_head->lock);
>> +	if (mpath_path_is_optimized(mpath_head, mpath_device)) {
>> +		int node, srcu_idx;
>> +
>> +		srcu_idx = srcu_read_lock(&mpath_head->srcu);
>> +		for_each_online_node(node)
>> +			__mpath_find_path(mpath_head, node);
>> +		srcu_read_unlock(&mpath_head->srcu, srcu_idx);
>> +	}
>> +	mutex_unlock(&mpath_head->lock);
>> +
>> +	mpath_synchronize(mpath_head);
>>   }
> 
> [Severity: High]
> When a new path becomes available and transitions to live, does this function
> need to call mpath_schedule_requeue_work()? If bios were held in the
> requeue_list while no paths were available, could omitting this flush lead
> to permanent I/O stalls?

I don't think so

> 


  reply	other threads:[~2026-07-27 15:21 UTC|newest]

Thread overview: 71+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-07-27 14:42 [PATCH v5 00/27] Native SCSI Multipath support John Garry
2026-07-27 14:42 ` [PATCH v5 01/27] libmultipath: Add initial framework John Garry
2026-07-27 15:05   ` sashiko-bot
2026-07-27 15:12     ` John Garry
2026-07-27 14:42 ` [PATCH v5 02/27] libmultipath: Add basic gendisk support John Garry
2026-07-27 15:04   ` sashiko-bot
2026-07-27 15:15     ` John Garry
2026-07-27 14:42 ` [PATCH v5 03/27] libmultipath: Add path selection support John Garry
2026-07-27 15:03   ` sashiko-bot
2026-07-27 15:20     ` John Garry [this message]
2026-07-27 14:42 ` [PATCH v5 04/27] libmultipath: Add bio handling John Garry
2026-07-27 14:42 ` [PATCH v5 05/27] libmultipath: Add support for mpath_device management John Garry
2026-07-27 15:03   ` sashiko-bot
2026-07-27 15:23     ` John Garry
2026-07-27 14:42 ` [PATCH v5 06/27] libmultipath: Add delayed removal support John Garry
2026-07-27 15:00   ` sashiko-bot
2026-07-27 15:25     ` John Garry
2026-07-27 14:42 ` [PATCH v5 07/27] libmultipath: Add sysfs helpers John Garry
2026-07-27 15:03   ` sashiko-bot
2026-07-27 15:33     ` John Garry
2026-07-27 14:42 ` [PATCH v5 08/27] libmultipath: Add support for block device IOCTL John Garry
2026-07-27 15:08   ` sashiko-bot
2026-07-27 15:31     ` John Garry
2026-07-27 14:42 ` [PATCH v5 09/27] libmultipath: Add mpath_bdev_getgeo() John Garry
2026-07-27 14:42 ` [PATCH v5 10/27] libmultipath: Add mpath_bdev_get_unique_id() John Garry
2026-07-27 14:42 ` [PATCH v5 11/27] scsi-multipath: introduce basic SCSI device support John Garry
2026-07-27 18:59   ` sashiko-bot
2026-07-27 14:42 ` [PATCH v5 12/27] scsi-multipath: introduce scsi_device head structure John Garry
2026-07-27 15:15   ` sashiko-bot
2026-07-27 15:37     ` John Garry
2026-07-27 14:42 ` [PATCH v5 13/27] scsi-multipath: provide sysfs link from to scsi_device John Garry
2026-07-27 15:07   ` sashiko-bot
2026-07-27 16:21     ` John Garry
2026-07-27 14:42 ` [PATCH v5 14/27] scsi-multipath: support iopolicy John Garry
2026-07-27 15:06   ` sashiko-bot
2026-07-27 15:39     ` John Garry
2026-07-27 14:42 ` [PATCH v5 15/27] scsi-multipath: clone each bio John Garry
2026-07-27 15:21   ` sashiko-bot
2026-07-27 15:40     ` John Garry
2026-07-27 14:42 ` [PATCH v5 16/27] scsi-multipath: clear path when device is blocked John Garry
2026-07-27 15:14   ` sashiko-bot
2026-07-27 15:44     ` John Garry
2026-07-27 14:42 ` [PATCH v5 17/27] scsi-multipath: revalidate paths upon device unblock John Garry
2026-07-27 15:17   ` sashiko-bot
2026-07-27 16:05     ` John Garry
2026-07-27 14:42 ` [PATCH v5 18/27] scsi-multipath: failover handling John Garry
2026-07-27 14:42 ` [PATCH v5 19/27] scsi-multipath: provide callbacks for path state John Garry
2026-07-27 15:24   ` sashiko-bot
2026-07-27 16:07     ` John Garry
2026-07-27 14:42 ` [PATCH v5 20/27] scsi-multipath: add scsi_mpath_{start,end}_request() John Garry
2026-07-27 15:25   ` sashiko-bot
2026-07-27 16:18     ` John Garry
2026-07-27 14:42 ` [PATCH v5 21/27] scsi-multipath: add delayed disk removal support John Garry
2026-07-27 15:23   ` sashiko-bot
2026-07-27 16:20     ` John Garry
2026-07-27 14:42 ` [PATCH v5 22/27] scsi: sd: add multipath disk class John Garry
2026-07-27 15:14   ` sashiko-bot
2026-07-27 16:21     ` John Garry
2026-07-27 14:42 ` [PATCH v5 23/27] scsi: sd: add multipath disk attr groups John Garry
2026-07-27 15:20   ` sashiko-bot
2026-07-27 16:22     ` John Garry
2026-07-27 14:42 ` [PATCH v5 24/27] scsi: sd: support multipath disk John Garry
2026-07-27 15:20   ` sashiko-bot
2026-07-27 16:33     ` John Garry
2026-07-27 14:42 ` [PATCH v5 25/27] scsi: sd: add mpath_dev file John Garry
2026-07-27 15:33   ` sashiko-bot
2026-07-27 16:24     ` John Garry
2026-07-27 14:42 ` [PATCH v5 26/27] scsi: sd: add mpath_numa_nodes dev attribute John Garry
2026-07-27 14:42 ` [PATCH v5 27/27] scsi: sd: add mpath_queue_depth " John Garry
2026-07-27 15:28   ` sashiko-bot
2026-07-27 15:29     ` John Garry

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=362edeb3-20be-415d-8719-163cfcee99fb@oracle.com \
    --to=john.g.garry@oracle.com \
    --cc=linux-scsi@vger.kernel.org \
    --cc=sashiko-reviews@lists.linux.dev \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox