linux-rdma.vger.kernel.org archive mirror
 help / color / mirror / Atom feed
From: sashiko-bot@kernel.org
To: "Konstantin Taranov" <kotaranov@linux.microsoft.com>
Cc: linux-rdma@vger.kernel.org
Subject: Re: [PATCH rdma-next v3 07/10] RDMA/mana_ib: Poll RC completions using PSN and FSN progress
Date: Wed, 07 Oct 2026 15:48:38 +0000	[thread overview]
Message-ID: <sashiko-outbox-163163@kernel.org> (raw)
In-Reply-To: <20261007153330.2979799-8-kotaranov@linux.microsoft.com>

Thank you for your contribution! Sashiko AI review found 2 potential issue(s) to consider:
- [High] rdma/mana_ib: RC SQ queue->tail stalls due to missing wqe_size_in_bu initialization
- [High] rdma/mana_ib: mana_ib_arm_cq returns 0 when has_pending_cqe is true, missing events
--

commit e4f126a1330a4d8df81cbe90bf16cac4e71247c6
Author: Konstantin Taranov <kotaranov@microsoft.com>

RDMA/mana_ib: Poll RC completions using PSN and FSN progress

This commit decodes RC receive completions, armed-PSN notifications, and MMQ
completions. It retires requester and MMQ shadow entries using their
sequence numbers, honors cross-queue fence dependencies, and suppresses
successful WCs for unsignaled work requests.

> diff --git a/drivers/infiniband/hw/mana/cq.c b/drivers/infiniband/hw/mana/cq.c
> --- a/drivers/infiniband/hw/mana/cq.c
> +++ b/drivers/infiniband/hw/mana/cq.c

[ ... ]

> @@ -324,8 +330,19 @@ static void mana_complete_send(struct mana_ib_qp *qp,
>  	if (vendor_error || !(wqe->flags & MANA_WQ_NO_SIGNAL_WC))
>  		mana_fill_wc(qp, poll, wqe, wqe->send_opcode, vendor_error);
>  
> -	queue = mana_qp_get_sq(qp)->kmem;
> -	queue->tail += wqe->wqe_size_in_bu;
> +	if (mmq) {
> +		queue = qp->rc_qp.queues[MANA_RC_SEND_QUEUE_MMQ].kmem;
> +		queue->tail += wqe->wqe_size_in_bu;
> +	} else {
> +		queue = mana_qp_get_sq(qp)->kmem;
> +		if (wqe->send_opcode == IB_WC_RDMA_READ) {
> +			qp->rc_qp.queues[MANA_RC_RECV_QUEUE_REQUESTER].kmem->tail +=
> +				wqe->wqe_size_in_bu;
> +			queue->tail += qp->rc_qp.wqe_size_in_bu;
> +		} else {
> +			queue->tail += wqe->wqe_size_in_bu;

[Severity: High]
For RC Send and Write shadow WQEs, wqe_size_in_bu is left as 0 during
posting. When adding it to the queue tail here in mana_complete_send(),
could this stall space accounting because queue->tail won't advance?

This could eventually cause mana_gd_wq_avail_space() to calculate incorrect
available space, permanently hanging the RC send queue under standard
SEND/WRITE traffic.

> +		}
> +	}
>  	shadow_queue_advance_consumer(shadow);
>  }

[ ... ]

> @@ -390,17 +422,105 @@ static void handle_rq_cqe(struct mana_ib_qp *qp, struct gdma_comp *cqe,

[ ... ]

> +static bool mana_advance_send(struct mana_ib_qp *qp, bool mmq, u32 seq,
> +			      struct mana_cq_poll *poll)
> +{
> +	struct shadow_queue *shadow = mmq ? &qp->shadow_mmq : &qp->shadow_sq;
> +	struct shadow_queue *other = mmq ? &qp->shadow_sq : &qp->shadow_mmq;
> +	struct shadow_wqe_header *shadow_wqe;
> +
> +	while ((shadow_wqe = shadow_queue_get_next_to_consume(shadow))) {
> +		if (PSN_LT(seq, mmq ? shadow_wqe->fsn : shadow_wqe->wqe_offset_or_psn))
> +			break;
> +		if (poll->produced == poll->budget)
> +			return false;

[Severity: High]
Returning false here when poll->produced == poll->budget causes a hardware
CQE to be cached in software by setting cq->has_pending_cqe to true.

If an application subsequently calls ib_req_notify_cq() with
IB_CQ_REPORT_MISSED_EVENTS, wouldn't the driver return 0 because it misses
checking this software-cached CQE?

This could cause standard RDMA ULPs to hang permanently waiting for an
interrupt, as the CQ is falsely reported as empty.

> +		if (shadow_wqe->flags & MANA_WQ_FENCE_WC) {
> +			u32 dependency = mmq ? shadow_wqe->wqe_offset_or_psn : shadow_wqe->fsn;
> +			struct shadow_wqe_header *other_wqe;

-- 
Sashiko AI review · https://sashiko.dev/#/patchset/20261007153330.2979799-1-kotaranov@linux.microsoft.com?part=7

  reply	other threads:[~2026-10-07 15:48 UTC|newest]

Thread overview: 21+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-10-07 15:33 [PATCH rdma-next v3 00/10] RDMA/mana_ib: Add kernel RC and fast registration support Konstantin Taranov
2026-10-07 15:33 ` [PATCH rdma-next v3 01/10] RDMA/mana_ib: Allocate and map fast-registration MRs Konstantin Taranov
2026-10-07 15:46   ` sashiko-bot
2026-10-07 15:33 ` [PATCH rdma-next v3 02/10] RDMA/mana: Create and destroy kernel RC QPs Konstantin Taranov
2026-10-07 15:50   ` sashiko-bot
2026-10-07 15:33 ` [PATCH rdma-next v3 03/10] net/mana: Extend GDMA encoding for new RDMA WQEs Konstantin Taranov
2026-10-07 15:47   ` sashiko-bot
2026-10-07 15:33 ` [PATCH rdma-next v3 04/10] RDMA/mana_ib: Maintain kernel RC QP state Konstantin Taranov
2026-10-07 15:46   ` sashiko-bot
2026-10-07 15:33 ` [PATCH rdma-next v3 05/10] RDMA/mana_ib: Post receive WRs on kernel RC QPs Konstantin Taranov
2026-10-07 15:43   ` sashiko-bot
2026-10-07 15:33 ` [PATCH rdma-next v3 06/10] RDMA/mana_ib: Post send and memory-management WRs on " Konstantin Taranov
2026-10-07 15:53   ` sashiko-bot
2026-10-07 15:33 ` [PATCH rdma-next v3 07/10] RDMA/mana_ib: Poll RC completions using PSN and FSN progress Konstantin Taranov
2026-10-07 15:48   ` sashiko-bot [this message]
2026-10-07 15:33 ` [PATCH rdma-next v3 08/10] RDMA/mana_ib: Flush and notify CQs when kernel QPs enter ERR Konstantin Taranov
2026-10-07 15:46   ` sashiko-bot
2026-10-07 15:33 ` [PATCH rdma-next v3 09/10] RDMA/mana_ib: Handle error CQEs for RC QPs Konstantin Taranov
2026-10-07 15:51   ` sashiko-bot
2026-10-07 15:33 ` [PATCH rdma-next v3 10/10] RDMA/mana_ib: Drain kernel receive and send queues Konstantin Taranov
2026-10-07 15:51   ` sashiko-bot

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=sashiko-outbox-163163@kernel.org \
    --to=sashiko-bot@kernel.org \
    --cc=kotaranov@linux.microsoft.com \
    --cc=linux-rdma@vger.kernel.org \
    --cc=sashiko-reviews@lists.linux.dev \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox;
as well as URLs for NNTP newsgroup(s).