All of lore.kernel.org
 help / color / mirror / Atom feed
From: sashiko-bot@kernel.org
To: "Mykola Marzhan" <mykola@meshstor.io>
Cc: linux-raid@vger.kernel.org, linux-pci@vger.kernel.org, yukuai@fygo.io
Subject: Re: [PATCH v6 10/12] nvme-rdma: return BLK_STS_P2PDMA for unsupported P2P transfers
Date: Sun, 26 Jul 2026 12:10:13 +0000	[thread overview]
Message-ID: <20260726121014.4337F1F000E9@smtp.kernel.org> (raw)
In-Reply-To: <20260726113229.2154269-11-mykola@meshstor.io>

Thank you for your contribution! Sashiko AI review found 2 potential issue(s) to consider:

New issues:
- [High] DMA API violation causing stale PI metadata (Data Corruption).

Pre-existing issues:
- [High] Incomplete fix: multipath I/O accounting leak on nvme_rdma_post_send() failure.
--

--- Patch [10]: [PATCH v6 10/12] nvme-rdma: return BLK_STS_P2PDMA for unsupported P2P transfers ---
commit d3d8febf3662843e41333a6adc3c3ea1d88ff373
Author: Mykola Marzhan <mykola@meshstor.io>

nvme-rdma: return BLK_STS_P2PDMA for unsupported P2P transfers

This commit propagates -EREMOTEIO from ib_dma_map_sgtable_attrs() as BLK_STS_P2PDMA
to prevent infinite multipath requeues. It also delays nvme_start_request() until
after data mapping succeeds to avoid multipath accounting leaks on mapping errors.

Cc: stable@vger.kernel.org # v7.1: requires BLK_STS_P2PDMA and the ib_dma_map_sgtable_attrs() conversion

> diff --git a/drivers/nvme/host/rdma.c b/drivers/nvme/host/rdma.c
> index 63830334c73e2..c38d949502b7c 100644
> --- a/drivers/nvme/host/rdma.c
> +++ b/drivers/nvme/host/rdma.c
> @@ -2033,8 +2029,6 @@ static blk_status_t nvme_rdma_queue_rq(struct blk_mq_hw_ctx *hctx,
>  	if (ret)
>  		goto unmap_qe;
>  
> -	nvme_start_request(rq);
> -
>  	if (IS_ENABLED(CONFIG_BLK_DEV_INTEGRITY) &&
>  	    queue->pi_support &&
>  	    (c->common.opcode == nvme_cmd_write ||
> @@ -2051,6 +2045,8 @@ static blk_status_t nvme_rdma_queue_rq(struct blk_mq_hw_ctx *hctx,
>  		goto err;
>  	}
>  
> +	nvme_start_request(rq);
> +

[Severity: High]
Does moving nvme_start_request() after nvme_rdma_map_data() violate the DMA API?

nvme_start_request() calls blk_mq_start_request(), which invokes
blk_integrity_prepare(). This modifies the PI metadata buffers in CPU memory
to update reference tags.

Because nvme_rdma_map_data() has already mapped these buffers for DMA
(DMA_TO_DEVICE), could CPU writes happening after the map operation remain
unsynced on non-coherent architectures or when bounce buffers are used?

This might cause the device to read stale PI metadata, leading to PI
verification failures or data corruption.

>  	sqe->cqe.done = nvme_rdma_send_done;
>  
>  	ib_dma_sync_single_for_device(dev, sqe->dma,

[ ... ]

> @@ -2070,6 +2066,9 @@ static blk_status_t nvme_rdma_queue_rq(struct blk_mq_hw_ctx *hctx,
>  		ret = nvme_host_path_error(rq);
>  	else if (err == -ENOMEM || err == -EAGAIN)
>  		ret = BLK_STS_RESOURCE;
> +	/* Peer memory unreachable from this device: don't retry. */
> +	else if (err == -EREMOTEIO)
> +		ret = BLK_STS_P2PDMA;
>  	else
>  		ret = BLK_STS_IOERR;

[Severity: High]
This is a pre-existing issue, but I noticed a potential multipath I/O
accounting leak if nvme_rdma_post_send() fails.

The commit message notes that starting the request before a failure causes
multipath accounting leaks when blk-mq hard-completes the request directly.

If nvme_rdma_post_send() fails (e.g., returns -EINVAL), execution jumps to
err_unmap and falls through to the err: block, returning a hard error status
like BLK_STS_IOERR:

	err = nvme_rdma_post_send(queue, sqe, ...);
	if (unlikely(err))
		goto err_unmap;
	...
err:
	...
	else
		ret = BLK_STS_IOERR;

Since nvme_start_request() was already called, will blk-mq hard-complete the
request directly, bypassing the driver's nvme_mpath_end_request() and leaking
multipath accounting like bd_stats->in_flight and nr_active?

This could permanently block block device teardown.

>  	nvme_cleanup_cmd(rq);

-- 
Sashiko AI review · https://sashiko.dev/#/patchset/20260726113229.2154269-1-mykola@meshstor.io?part=10

  reply	other threads:[~2026-07-26 12:10 UTC|newest]

Thread overview: 25+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-07-26 11:32 [PATCH v6 00/12] block,md,nvme: correct handling of unsupported P2PDMA transfers Mykola Marzhan
2026-07-26 11:32 ` [PATCH v6 01/12] block: add BLK_STS_P2PDMA for unsupported peer-to-peer transfers Mykola Marzhan
2026-07-26 12:06   ` sashiko-bot
2026-07-26 11:32 ` [PATCH v6 02/12] md: ensure REQ_NOMERGE is set on P2PDMA bios Mykola Marzhan
2026-07-26 11:59   ` sashiko-bot
2026-07-26 11:32 ` [PATCH v6 03/12] md/raid1: serialize non-write-behind writes on CollisionCheck rdevs Mykola Marzhan
2026-07-26 12:10   ` sashiko-bot
2026-07-26 11:32 ` [PATCH v6 04/12] md/raid1: don't use write-behind for P2PDMA bios Mykola Marzhan
2026-07-26 12:01   ` sashiko-bot
2026-07-26 11:32 ` [PATCH v6 05/12] md/raid1,raid10: factor out raid1_write_error() helper Mykola Marzhan
2026-07-26 11:57   ` sashiko-bot
2026-07-26 11:32 ` [PATCH v6 06/12] md/raid1,raid10: keep REQ_NOMERGE on narrow_write_error() retry clones Mykola Marzhan
2026-07-26 12:09   ` sashiko-bot
2026-07-26 11:32 ` [PATCH v6 07/12] md/raid1,raid10: skip futile retries on P2PDMA mapping failures Mykola Marzhan
2026-07-26 12:03   ` sashiko-bot
2026-07-26 11:32 ` [PATCH v6 08/12] md/raid1,raid10: set IO_BLOCKED in case of BLK_STS_P2PDMA Mykola Marzhan
2026-07-26 12:06   ` sashiko-bot
2026-07-26 11:32 ` [PATCH v6 09/12] nvme-rdma: use ib_dma_map_sgtable_attrs() Mykola Marzhan
2026-07-26 12:01   ` sashiko-bot
2026-07-26 11:32 ` [PATCH v6 10/12] nvme-rdma: return BLK_STS_P2PDMA for unsupported P2P transfers Mykola Marzhan
2026-07-26 12:10   ` sashiko-bot [this message]
2026-07-26 11:32 ` [PATCH v6 11/12] nvme-rdma: ratelimit the map-failure error message Mykola Marzhan
2026-07-26 11:53   ` sashiko-bot
2026-07-26 11:32 ` [PATCH v6 12/12] nvme-rdma: factor out the scatterlist DMA mapping helper Mykola Marzhan
2026-07-26 12:02   ` sashiko-bot

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=20260726121014.4337F1F000E9@smtp.kernel.org \
    --to=sashiko-bot@kernel.org \
    --cc=linux-pci@vger.kernel.org \
    --cc=linux-raid@vger.kernel.org \
    --cc=mykola@meshstor.io \
    --cc=sashiko-reviews@lists.linux.dev \
    --cc=yukuai@fygo.io \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is an external index of several public inboxes,
see mirroring instructions on how to clone and mirror
all data and code used by this external index.