From: sashiko-bot@kernel.org
To: "Dongliang Qin" <cccccccccccc777777@gmail.com>
Cc: linux-rdma@vger.kernel.org
Subject: Re: [PATCH 3/4] RDMA/rxe: Invalidate MWs on QP destroy
Date: Mon, 28 Sep 2026 16:16:35 +0000 [thread overview]
Message-ID: <20260928161636.021E81F000FF@smtp.kernel.org> (raw)
In-Reply-To: <20260928155351.3222978-4-cccccccccccc777777@gmail.com>
Thank you for your contribution! Sashiko AI review found 4 potential issue(s) to consider:
- [High] rdma/rxe: type mismatch in xa_for_each causes memory corruption
- [High] rdma/rxe: Use-After-Free during MW pool iteration under RCU
- [High] rdma/rxe: concurrent MW bind during QP destruction causes hang and UAF
- [Medium] rdma/rxe: soft lockup during MW pool traversal in rxe_invalidate_mws
--
commit 80671c88acbdcad0b751f787a4de4bc92545ab54
Author: Dongliang Qin <cccccccccccc777777@gmail.com>
RDMA/rxe: Invalidate MWs on QP destroy
A type-2 MW holds a reference to the QP that bound it. If the QP is
destroyed while an MW is still bound, the MW keeps the QP alive and the
responder can retain a stale QP association. This commit invalidates MWs
bound to the QP before removing the QP from its pool.
Cc: stable@vger.kernel.org
> diff --git a/drivers/infiniband/sw/rxe/rxe_mw.c b/drivers/infiniband/sw/rxe/rxe_mw.c
> index 82e9fef89b6c3..02d20ede1a789 100644
> --- a/drivers/infiniband/sw/rxe/rxe_mw.c
> +++ b/drivers/infiniband/sw/rxe/rxe_mw.c
[ ... ]
> +void rxe_invalidate_mws(struct rxe_qp *qp)
> +{
> + struct rxe_dev *rxe = to_rdev(qp->ibqp.device);
> + struct rxe_mw *mw;
> + unsigned long index;
> +
> + rcu_read_lock();
> + xa_for_each(&rxe->mw_pool.xa, index, mw) {
[Severity: High]
Does this loop incorrectly assume the XArray returns struct rxe_mw * directly?
The XArray mw_pool.xa stores struct rxe_pool_elem * elements, but
xa_for_each() assigns the returned pointer directly to mw (which is declared
as struct rxe_mw *).
Since elem is not at offset 0 in struct rxe_mw (it is preceded by struct
ib_mw), would any subsequent access like rxe_get(mw) evaluate to &mw->elem
with a non-zero offset, leading to out-of-bounds memory access?
[Severity: High]
Could this iteration race with concurrent MW deallocation?
The loop relies on rcu_read_lock() to protect object memory during lookup,
but the RDMA core synchronously frees the Memory Window without an RCU grace
period:
drivers/infiniband/core/uverbs_main.c:uverbs_dealloc_mw() {
...
atomic_dec(&pd->usecnt);
kfree(mw);
return ret;
}
If a user destroys a QP while concurrently deallocating an MW, could
xa_for_each() return a pointer to an MW that is simultaneously kfreed,
causing rxe_get(mw) to access freed memory?
[Severity: Medium]
Could iterating this pool under RCU cause a soft lockup?
If a user creates a massive number of MWs (up to RXE_MAX_MW, which is
524,288) and then destroys a QP, this loop will iterate the entire
populated pool.
Since this executes strictly under rcu_read_lock() and takes spin_lock_bh()
for each present MW, preemption and bottom-halves are disabled for the
entire duration of this potentially massive loop without yielding. Will
this lead to watchdogs firing?
> + if (!rxe_get(mw))
> + continue;
> +
> + spin_lock_bh(&mw->lock);
> + if (mw->qp == qp)
> + rxe_do_invalidate_mw(mw);
> + spin_unlock_bh(&mw->lock);
> +
> + rxe_put(mw);
> + }
> + rcu_read_unlock();
> +}
[ ... ]
> diff --git a/drivers/infiniband/sw/rxe/rxe_verbs.c b/drivers/infiniband/sw/rxe/rxe_verbs.c
> index 8553c8402c619..c0d77cad4906b 100644
> --- a/drivers/infiniband/sw/rxe/rxe_verbs.c
> +++ b/drivers/infiniband/sw/rxe/rxe_verbs.c
[ ... ]
> @@ -658,6 +658,8 @@ static int rxe_destroy_qp(struct ib_qp *ibqp, struct ib_udata *udata)
> goto err_out;
> }
>
> + rxe_invalidate_mws(qp);
> +
> err = rxe_cleanup(qp);
[Severity: High]
Is there a race condition between invalidating MWs and stopping the QP
tasklets?
The invalidation happens before the QP tasklets are stopped in rxe_cleanup(),
leaving a race window. The requester tasklet can concurrently process a bind
request:
drivers/infiniband/sw/rxe/rxe_req.c:rxe_requester() {
...
case IB_WR_BIND_MW:
ret = rxe_bind_mw(qp, wqe);
...
}
Could this concurrent bind request elevate the QP refcount after the
invalidation pass has already finished, causing the subsequent rxe_cleanup()
to hang indefinitely (or UAF if it times out) while waiting for the refcount
to drop?
--
Sashiko AI review · https://sashiko.dev/#/patchset/20260928155351.3222978-1-cccccccccccc777777@gmail.com?part=3
next prev parent reply other threads:[~2026-09-28 16:16 UTC|newest]
Thread overview: 9+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-09-28 15:53 [PATCH 0/4] RDMA/rxe: Fix MW/MR lifetime races Dongliang Qin
2026-09-28 15:53 ` [PATCH 1/4] RDMA/rxe: Take MR reference under MW lock Dongliang Qin
2026-09-28 16:03 ` sashiko-bot
2026-09-28 15:53 ` [PATCH 2/4] RDMA/rxe: Reserve MR state during MW binding Dongliang Qin
2026-09-28 16:03 ` sashiko-bot
2026-09-28 15:53 ` [PATCH 3/4] RDMA/rxe: Invalidate MWs on QP destroy Dongliang Qin
2026-09-28 16:16 ` sashiko-bot [this message]
2026-09-28 15:53 ` [PATCH 4/4] RDMA/rxe: Do not force cleanup on pool timeout Dongliang Qin
2026-09-28 16:06 ` sashiko-bot
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=20260928161636.021E81F000FF@smtp.kernel.org \
--to=sashiko-bot@kernel.org \
--cc=cccccccccccc777777@gmail.com \
--cc=linux-rdma@vger.kernel.org \
--cc=sashiko-reviews@lists.linux.dev \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is an external index of several public inboxes,
see mirroring instructions on how to clone and mirror
all data and code used by this external index.