From: sashiko-bot@kernel.org
To: "Konstantin Taranov" <kotaranov@linux.microsoft.com>
Cc: linux-rdma@vger.kernel.org
Subject: Re: [PATCH rdma-next v2 5/5] RDMA/mana_ib: Poll UD completions and flush software error QPs
Date: Fri, 18 Sep 2026 14:20:38 +0000 [thread overview]
Message-ID: <20260918142039.5964D1F00898@smtp.kernel.org> (raw)
In-Reply-To: <20260918140407.2922819-6-kotaranov@linux.microsoft.com>
Thank you for your contribution! Sashiko AI review found 1 potential issue(s) to consider:
- [High] RDMA/mana_ib: Send Queue deadlock when handling unsignaled Work Requests
--
commit d64ff9be78f7ee54ec40f144bfa6a3c96bf9df22
Author: Konstantin Taranov <kotaranov@microsoft.com>
RDMA/mana_ib: Poll UD completions and flush software error QPs
This commit updates the driver to build UD and GSI work completions directly
from each hardware CQE and shadow entry. It introduces a WC-budgeted poll
context and replaces the all-QP CQ membership list with persistent send and
receive error-QP lists.
> diff --git a/drivers/infiniband/hw/mana/wr.c b/drivers/infiniband/hw/mana/wr.c
> index eb671b45b2976..15faca66179d1 100644
> --- a/drivers/infiniband/hw/mana/wr.c
> +++ b/drivers/infiniband/hw/mana/wr.c
[ ... ]
> @@ -144,6 +145,7 @@ static int mana_ib_post_send_ud_one(struct mana_ib_qp *qp, const struct ib_ud_wr
> memset(shadow_wqe, 0, sizeof(*shadow_wqe));
> shadow_wqe->wr_id = wr->wr.wr_id;
> shadow_wqe->wqe_size_in_bu = wqe_info.wqe_size_in_bu;
> + shadow_wqe->wqe_offset_or_psn = wqe_info.wqe_offset;
> shadow_queue_advance_producer(&qp->shadow_sq);
>
> return 0;
[Severity: High]
Does this code lose track of which Work Requests were posted as unsignaled?
In mana_ib_post_send_ud_one(), shadow_wqe is zeroed out, but the
MANA_WQ_NO_SIGNAL_WC flag is never conditionally set in shadow_wqe->flags based
on the unsignaled status of the WR.
> diff --git a/drivers/infiniband/hw/mana/cq.c b/drivers/infiniband/hw/mana/cq.c
> index 6764e2de0c9f9..55d7a7d1cd69b 100644
> --- a/drivers/infiniband/hw/mana/cq.c
> +++ b/drivers/infiniband/hw/mana/cq.c
[ ... ]
> @@ -236,126 +288,156 @@ int mana_ib_arm_cq(struct ib_cq *ibcq, enum ib_cq_notify_flags flags)
[ ... ]
> +static void handle_ud_sq_cqe(struct mana_ib_qp *qp, struct mana_rdma_cqe *rdma_cqe,
> + struct mana_cq_poll *poll)
> {
> + struct shadow_wqe_header *wqe = shadow_queue_get_next_to_consume(&qp->shadow_sq);
> + u32 offset = rdma_cqe->ud_send.tx_wqe_offset & MANA_WQE_OFFSET_MASK;
>
> + if (!wqe || wqe->wqe_offset_or_psn != offset)
> return;
>
> + mana_complete_send(qp, poll, rdma_cqe->ud_send.vendor_error);
> }
[Severity: High]
Can this strict offset check cause the Send Queue to become permanently
deadlocked if the user posts an unsignaled Work Request?
When an unsignaled WR is posted and succeeds, the hardware will skip
generating a CQE for it. The next CQE generated will belong to a subsequent
signaled WR.
Because there is no loop here to advance the shadow queue past unsignaled WRs,
wqe will continue to point to the older unsignaled WR, while offset belongs
to the newer signaled WR. This mismatch causes handle_ud_sq_cqe() to return
immediately. The CQE is discarded and the shadow queue is never advanced,
permanently stalling the queue.
--
Sashiko AI review · https://sashiko.dev/#/patchset/20260918140407.2922819-1-kotaranov@linux.microsoft.com?part=5
prev parent reply other threads:[~2026-09-18 14:20 UTC|newest]
Thread overview: 11+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-09-18 14:04 [PATCH rdma-next v2 0/5] RDMA/mana_ib: Streamline kernel UD/GSI posting and completion handling Konstantin Taranov
2026-09-18 14:04 ` [PATCH rdma-next v2 1/5] RDMA/mana_ib: Optimize shadow queue bookkeeping Konstantin Taranov
2026-09-18 14:17 ` sashiko-bot
2026-09-18 14:04 ` [PATCH rdma-next v2 2/5] RDMA/mana_ib: Revise UD send posting and WQE definitions Konstantin Taranov
2026-09-18 14:17 ` sashiko-bot
2026-09-18 14:04 ` [PATCH rdma-next v2 3/5] RDMA/mana_ib: Revise UD receive posting with GDMA_WR_IB_SGL Konstantin Taranov
2026-09-18 14:13 ` sashiko-bot
2026-09-18 14:04 ` [PATCH rdma-next v2 4/5] RDMA/mana_ib: Make kernel CQ arming robust Konstantin Taranov
2026-09-18 14:18 ` sashiko-bot
2026-09-18 14:04 ` [PATCH rdma-next v2 5/5] RDMA/mana_ib: Poll UD completions and flush software error QPs Konstantin Taranov
2026-09-18 14:20 ` sashiko-bot [this message]
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=20260918142039.5964D1F00898@smtp.kernel.org \
--to=sashiko-bot@kernel.org \
--cc=kotaranov@linux.microsoft.com \
--cc=linux-rdma@vger.kernel.org \
--cc=sashiko-reviews@lists.linux.dev \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox