All of lore.kernel.org
 help / color / mirror / Atom feed
From: sashiko-bot@kernel.org
To: "John Garry" <john.g.garry@oracle.com>
Cc: linux-scsi@vger.kernel.org
Subject: Re: [PATCH v5 03/27] libmultipath: Add path selection support
Date: Mon, 27 Jul 2026 15:03:03 +0000	[thread overview]
Message-ID: <20260727150304.A28121F00A3E@smtp.kernel.org> (raw)
In-Reply-To: <20260727144234.3802948-4-john.g.garry@oracle.com>

Thank you for your contribution! Sashiko AI review found 4 potential issue(s) to consider:
- [High] Round-robin path selection can enter an infinite loop under the SRCU read lock if the current path is concurrently removed.
- [High] If `blk_alloc_disk()` fails, an `ERR_PTR` is left in `mpath_head->disk`, which will cause a kernel panic during cleanup.
- [High] Missing requeue list flushes in `mpath_device_set_live` and `mpath_remove_disk` cause IO stalls and bio leaks.
- [High] `mpath_round_robin_path` incorrectly returns a deleted path when the path list becomes singular.
--

--- Patch [3]: [PATCH v5 03/27] libmultipath: Add path selection support ---
commit 88605e74a8a44a0f658b820940be9c7cdbbebc25
Author: John Garry <john.garry@linux.dev>

libmultipath: Add path selection support

This commit adds path selection support for NVMe ANA, abstracting it into
enum mpath_access_state. It introduces round-robin, NUMA, and queue-depth
selection modes along with iopolicy configuration functions.

> diff --git a/lib/multipath.c b/lib/multipath.c
> index d2e334259f3c1..2b22c74760868 100644
> --- a/lib/multipath.c
> +++ b/lib/multipath.c

[ ... ]

> +static struct mpath_device *mpath_round_robin_path(
> +				struct mpath_head *mpath_head)
> +{
> +	struct mpath_device *mpath_device, *found = NULL;
> +	int node = numa_node_id();
> +	enum mpath_access_state access_state_old;
> +	struct mpath_device *old =
> +			srcu_dereference(mpath_head->current_path[node],
> +				&mpath_head->srcu);
> +
> +	if (unlikely(!old))
> +		return __mpath_find_path(mpath_head, node);
> +
> +	if (list_is_singular(&mpath_head->dev_list)) {
> +		if (mpath_path_is_disabled(mpath_head, old))
> +			return NULL;
> +		return old;
> +	}

[Severity: High]
If the cached path (old) was concurrently unlinked from the list, is it
possible that list_is_singular() evaluates to true because there is exactly
one other valid path remaining? If so, does this mistakenly return the
unlinked old path (or NULL) instead of evaluating the single remaining valid
path?

> +
> +	for (mpath_device = mpath_next_dev(mpath_head, old);
> +	    mpath_device && mpath_device != old;
> +	    mpath_device = mpath_next_dev(mpath_head, mpath_device)) {

[Severity: High]
If the currently cached path (old) is concurrently removed via
list_del_rcu(), it is unlinked from dev_list. Since mpath_next_dev()
iterates through the list, will this loop fail to encounter the unlinked
old pointer? If all remaining paths are non-optimized or disabled, could
this result in an infinite loop inside an SRCU read-side critical section?

> +
> +		if (mpath_path_is_disabled(mpath_head, mpath_device))
> +			continue;

[ ... ]

> @@ -75,6 +312,7 @@ void mpath_remove_disk(struct mpath_head *mpath_head)
>  	if (test_and_clear_bit(MPATH_HEAD_DISK_LIVE, &mpath_head->flags)) {
>  		struct gendisk *disk = mpath_head->disk;
>  
> +		mpath_synchronize(mpath_head);
>  		del_gendisk(disk);
>  	}

[Severity: High]
When the disk is removed, are there potentially stranded bios left in the
mpath_head->requeue_list? Without a call to mpath_schedule_requeue_work()
here to flush the list, do these queued bios leak?

>  }

[ ... ]

> @@ -96,7 +334,7 @@ int mpath_alloc_head_disk(struct mpath_head *mpath_head,
>  			struct queue_limits *lim, int numa_node)
>  {
>  	if (!mpath_head->disk_groups || !mpath_head->parent ||
> -	    mpath_head->disk)
> +	     !mpath_head->iopolicy || mpath_head->disk)
>  		return -EINVAL;
>  
>  	mpath_head->disk = blk_alloc_disk(lim, numa_node);

[Severity: High]
If blk_alloc_disk() fails, does it leave an ERR_PTR in mpath_head->disk?
If the error pointer is not cleared here, would the cleanup path in
mpath_put_disk() mistakenly evaluate "if (!mpath_head->disk)" as false
and attempt to call put_disk() on the error pointer?

[ ... ]

> @@ -130,6 +368,19 @@ void mpath_device_set_live(struct mpath_device *mpath_device)
>  		}
>  		queue_work(mpath_wq, &mpath_head->partition_scan_work);
>  	}
> +
> +	mutex_lock(&mpath_head->lock);
> +	if (mpath_path_is_optimized(mpath_head, mpath_device)) {
> +		int node, srcu_idx;
> +
> +		srcu_idx = srcu_read_lock(&mpath_head->srcu);
> +		for_each_online_node(node)
> +			__mpath_find_path(mpath_head, node);
> +		srcu_read_unlock(&mpath_head->srcu, srcu_idx);
> +	}
> +	mutex_unlock(&mpath_head->lock);
> +
> +	mpath_synchronize(mpath_head);
>  }

[Severity: High]
When a new path becomes available and transitions to live, does this function
need to call mpath_schedule_requeue_work()? If bios were held in the
requeue_list while no paths were available, could omitting this flush lead
to permanent I/O stalls?

-- 
Sashiko AI review · https://sashiko.dev/#/patchset/20260727144234.3802948-1-john.g.garry@oracle.com?part=3

  reply	other threads:[~2026-07-27 15:03 UTC|newest]

Thread overview: 71+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-07-27 14:42 [PATCH v5 00/27] Native SCSI Multipath support John Garry
2026-07-27 14:42 ` [PATCH v5 01/27] libmultipath: Add initial framework John Garry
2026-07-27 15:05   ` sashiko-bot
2026-07-27 15:12     ` John Garry
2026-07-27 14:42 ` [PATCH v5 02/27] libmultipath: Add basic gendisk support John Garry
2026-07-27 15:04   ` sashiko-bot
2026-07-27 15:15     ` John Garry
2026-07-27 14:42 ` [PATCH v5 03/27] libmultipath: Add path selection support John Garry
2026-07-27 15:03   ` sashiko-bot [this message]
2026-07-27 15:20     ` John Garry
2026-07-27 14:42 ` [PATCH v5 04/27] libmultipath: Add bio handling John Garry
2026-07-27 14:42 ` [PATCH v5 05/27] libmultipath: Add support for mpath_device management John Garry
2026-07-27 15:03   ` sashiko-bot
2026-07-27 15:23     ` John Garry
2026-07-27 14:42 ` [PATCH v5 06/27] libmultipath: Add delayed removal support John Garry
2026-07-27 15:00   ` sashiko-bot
2026-07-27 15:25     ` John Garry
2026-07-27 14:42 ` [PATCH v5 07/27] libmultipath: Add sysfs helpers John Garry
2026-07-27 15:03   ` sashiko-bot
2026-07-27 15:33     ` John Garry
2026-07-27 14:42 ` [PATCH v5 08/27] libmultipath: Add support for block device IOCTL John Garry
2026-07-27 15:08   ` sashiko-bot
2026-07-27 15:31     ` John Garry
2026-07-27 14:42 ` [PATCH v5 09/27] libmultipath: Add mpath_bdev_getgeo() John Garry
2026-07-27 14:42 ` [PATCH v5 10/27] libmultipath: Add mpath_bdev_get_unique_id() John Garry
2026-07-27 14:42 ` [PATCH v5 11/27] scsi-multipath: introduce basic SCSI device support John Garry
2026-07-27 18:59   ` sashiko-bot
2026-07-27 14:42 ` [PATCH v5 12/27] scsi-multipath: introduce scsi_device head structure John Garry
2026-07-27 15:15   ` sashiko-bot
2026-07-27 15:37     ` John Garry
2026-07-27 14:42 ` [PATCH v5 13/27] scsi-multipath: provide sysfs link from to scsi_device John Garry
2026-07-27 15:07   ` sashiko-bot
2026-07-27 16:21     ` John Garry
2026-07-27 14:42 ` [PATCH v5 14/27] scsi-multipath: support iopolicy John Garry
2026-07-27 15:06   ` sashiko-bot
2026-07-27 15:39     ` John Garry
2026-07-27 14:42 ` [PATCH v5 15/27] scsi-multipath: clone each bio John Garry
2026-07-27 15:21   ` sashiko-bot
2026-07-27 15:40     ` John Garry
2026-07-27 14:42 ` [PATCH v5 16/27] scsi-multipath: clear path when device is blocked John Garry
2026-07-27 15:14   ` sashiko-bot
2026-07-27 15:44     ` John Garry
2026-07-27 14:42 ` [PATCH v5 17/27] scsi-multipath: revalidate paths upon device unblock John Garry
2026-07-27 15:17   ` sashiko-bot
2026-07-27 16:05     ` John Garry
2026-07-27 14:42 ` [PATCH v5 18/27] scsi-multipath: failover handling John Garry
2026-07-27 14:42 ` [PATCH v5 19/27] scsi-multipath: provide callbacks for path state John Garry
2026-07-27 15:24   ` sashiko-bot
2026-07-27 16:07     ` John Garry
2026-07-27 14:42 ` [PATCH v5 20/27] scsi-multipath: add scsi_mpath_{start,end}_request() John Garry
2026-07-27 15:25   ` sashiko-bot
2026-07-27 16:18     ` John Garry
2026-07-27 14:42 ` [PATCH v5 21/27] scsi-multipath: add delayed disk removal support John Garry
2026-07-27 15:23   ` sashiko-bot
2026-07-27 16:20     ` John Garry
2026-07-27 14:42 ` [PATCH v5 22/27] scsi: sd: add multipath disk class John Garry
2026-07-27 15:14   ` sashiko-bot
2026-07-27 16:21     ` John Garry
2026-07-27 14:42 ` [PATCH v5 23/27] scsi: sd: add multipath disk attr groups John Garry
2026-07-27 15:20   ` sashiko-bot
2026-07-27 16:22     ` John Garry
2026-07-27 14:42 ` [PATCH v5 24/27] scsi: sd: support multipath disk John Garry
2026-07-27 15:20   ` sashiko-bot
2026-07-27 16:33     ` John Garry
2026-07-27 14:42 ` [PATCH v5 25/27] scsi: sd: add mpath_dev file John Garry
2026-07-27 15:33   ` sashiko-bot
2026-07-27 16:24     ` John Garry
2026-07-27 14:42 ` [PATCH v5 26/27] scsi: sd: add mpath_numa_nodes dev attribute John Garry
2026-07-27 14:42 ` [PATCH v5 27/27] scsi: sd: add mpath_queue_depth " John Garry
2026-07-27 15:28   ` sashiko-bot
2026-07-27 15:29     ` John Garry

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=20260727150304.A28121F00A3E@smtp.kernel.org \
    --to=sashiko-bot@kernel.org \
    --cc=john.g.garry@oracle.com \
    --cc=linux-scsi@vger.kernel.org \
    --cc=sashiko-reviews@lists.linux.dev \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is an external index of several public inboxes,
see mirroring instructions on how to clone and mirror
all data and code used by this external index.