* [PATCH] RDMA/rxe: refuse to destroy a QP with type 2 MWs still bound
@ 2026-09-30 18:22 Youngsung Ahn
2026-09-30 18:41 ` sashiko-bot
0 siblings, 1 reply; 2+ messages in thread
From: Youngsung Ahn @ 2026-09-30 18:22 UTC (permalink / raw)
To: zyjzyj2000; +Cc: linux-rdma, security
From: Youngsung Ahn <ays511.kr@gmail.com>
Binding a type 2 memory window takes a reference on the QP:
rxe_do_bind_mw() does rxe_get(qp) and stores it in mw->qp. That
reference is released only when the window is invalidated
(rxe_do_invalidate_mw() on IB_WR_LOCAL_INV) or the window itself is
destroyed (rxe_mw_cleanup()).
rxe_qp_chk_destroy() does not account for those references -- it only
refuses when qp->mcg_num is non-zero. So an unprivileged user can bind
a type 2 MW to a QP and then destroy the QP while the binding is still
live. __rxe_cleanup() cannot drop the QP refcount to zero (the MW still
holds a reference), hits its -ETIMEDOUT path, and frees the QP anyway;
mw->qp is left dangling. When the MW is later invalidated or
deallocated, rxe_mw_cleanup()/rxe_do_invalidate_mw() do rxe_put(mw->qp)
on the freed struct rxe_qp -- a use-after-free write (refcount_dec, and
complete()/list work if it reaches zero) reachable from userspace.
IB_WR_BIND_MW is a local operation, so no peer is required.
Track the number of type 2 MWs bound to a QP and refuse to destroy the
QP while that count is non-zero, returning -EBUSY exactly as the
existing multicast-attachment check does. The window must be
invalidated or deallocated, which drops the reference and the count,
before the QP can be destroyed.
Fixes: 8700e3e7c485 ("Soft RoCE driver")
Signed-off-by: Youngsung Ahn <ays511.kr@gmail.com>
---
Notes (not part of the commit):
Reproduced on a KASAN x86-64 build of 7.3-rc4 as uid 1000: deterministic
KASAN slab-use-after-free write via rxe_mw_cleanup() -> rxe_put() on the
freed struct rxe_qp (bind a type 2 MW to an RTS QP, destroy the QP, then
deallocate the MW). Present unchanged in mainline 551c722f4080
(2026-09-29) and rdma for-next. This patch is compile-tested
(KASAN+RDMA_RXE); the counter-based fix itself has not been runtime-tested
in this form. Found through manual review; per security-bugs.rst this is
public and a reproducer can be shared on request. The Fixes: tag points at
the base driver and should be refined to the type 2 MW bind support commit
when preparing for merge.
drivers/infiniband/sw/rxe/rxe_mw.c | 3 +++
drivers/infiniband/sw/rxe/rxe_qp.c | 10 ++++++++++
drivers/infiniband/sw/rxe/rxe_verbs.h | 1 +
3 files changed, 14 insertions(+)
diff --git a/drivers/infiniband/sw/rxe/rxe_mw.c b/drivers/infiniband/sw/rxe/rxe_mw.c
index bddb7a25..a2e3a6ce 100644
--- a/drivers/infiniband/sw/rxe/rxe_mw.c
+++ b/drivers/infiniband/sw/rxe/rxe_mw.c
@@ -161,6 +161,7 @@ static void rxe_do_bind_mw(struct rxe_qp *qp, struct rxe_send_wqe *wqe,
if (mw->ibmw.type == IB_MW_TYPE_2) {
rxe_get(qp);
+ atomic_inc(&qp->mw_bind_num);
mw->qp = qp;
}
}
@@ -245,6 +246,7 @@ static void rxe_do_invalidate_mw(struct rxe_mw *mw)
/* valid type 2 MW will always have a QP pointer */
qp = mw->qp;
mw->qp = NULL;
+ atomic_dec(&qp->mw_bind_num);
rxe_put(qp);
/* valid type 2 MW will always have an MR pointer */
@@ -332,6 +334,7 @@ void rxe_mw_cleanup(struct rxe_pool_elem *elem)
struct rxe_qp *qp = mw->qp;
mw->qp = NULL;
+ atomic_dec(&qp->mw_bind_num);
rxe_put(qp);
}
diff --git a/drivers/infiniband/sw/rxe/rxe_qp.c b/drivers/infiniband/sw/rxe/rxe_qp.c
index 311f285d..fea55bc3 100644
--- a/drivers/infiniband/sw/rxe/rxe_qp.c
+++ b/drivers/infiniband/sw/rxe/rxe_qp.c
@@ -858,6 +858,16 @@ int rxe_qp_chk_destroy(struct rxe_qp *qp)
return -EBUSY;
}
+ /* An attempt to destroy a QP while it still holds references for
+ * bound type 2 memory windows will fail immediately. Otherwise the
+ * QP is freed with those references outstanding and the windows are
+ * left pointing at freed memory (use-after-free).
+ */
+ if (atomic_read(&qp->mw_bind_num)) {
+ rxe_dbg_qp(qp, "Attempt to destroy while type 2 MWs are bound\n");
+ return -EBUSY;
+ }
+
return 0;
}
diff --git a/drivers/infiniband/sw/rxe/rxe_verbs.h b/drivers/infiniband/sw/rxe/rxe_verbs.h
index 0f5ffd94..80df376e 100644
--- a/drivers/infiniband/sw/rxe/rxe_verbs.h
+++ b/drivers/infiniband/sw/rxe/rxe_verbs.h
@@ -261,6 +261,7 @@ struct rxe_qp {
struct rxe_av alt_av;
atomic_t mcg_num;
+ atomic_t mw_bind_num;
struct sk_buff_head req_pkts;
struct sk_buff_head resp_pkts;
--
2.43.0
^ permalink raw reply related [flat|nested] 2+ messages in thread* Re: [PATCH] RDMA/rxe: refuse to destroy a QP with type 2 MWs still bound
2026-09-30 18:22 [PATCH] RDMA/rxe: refuse to destroy a QP with type 2 MWs still bound Youngsung Ahn
@ 2026-09-30 18:41 ` sashiko-bot
0 siblings, 0 replies; 2+ messages in thread
From: sashiko-bot @ 2026-09-30 18:41 UTC (permalink / raw)
To: Youngsung Ahn; +Cc: linux-rdma
Thank you for your contribution! Sashiko AI review found 1 potential issue(s) to consider:
- [Critical] RDMA/rxe: TOCTOU race in rxe_qp_chk_destroy() allows UAF
--
commit 94e08f1cc2290d06c7fcd154d031cde9d5a0054c
Author: Youngsung Ahn <ays511.kr@gmail.com>
RDMA/rxe: refuse to destroy a QP with type 2 MWs still bound
This commit adds tracking for type 2 memory windows bound to a QP via an
atomic counter. It attempts to prevent the destruction of a QP if there are
outstanding type 2 MW bindings by returning -EBUSY.
> diff --git a/drivers/infiniband/sw/rxe/rxe_mw.c b/drivers/infiniband/sw/rxe/rxe_mw.c
> index bddb7a2578313..a2e3a6ce09463 100644
> --- a/drivers/infiniband/sw/rxe/rxe_mw.c
> +++ b/drivers/infiniband/sw/rxe/rxe_mw.c
[ ... ]
> @@ -161,6 +161,7 @@ static void rxe_do_bind_mw(struct rxe_qp *qp, struct rxe_send_wqe *wqe,
>
> if (mw->ibmw.type == IB_MW_TYPE_2) {
> rxe_get(qp);
> + atomic_inc(&qp->mw_bind_num);
> mw->qp = qp;
> }
> }
[Severity: Critical]
Is it possible for the worker task to execute this concurrently during QP
destruction?
Since qp->valid is not cleared until rxe_qp_do_cleanup() runs, the requester
can freely process pending BIND_MW work requests during the cleanup delay:
drivers/infiniband/sw/rxe/rxe_req.c:rxe_requester() {
...
if (unlikely(!qp->valid)) {
...
}
If this runs concurrently, couldn't the counter be incremented too late,
holding the reference that triggers the 50-second timeout during cleanup?
> diff --git a/drivers/infiniband/sw/rxe/rxe_qp.c b/drivers/infiniband/sw/rxe/rxe_qp.c
> index 311f285d78a6b..fea55bc3b0cb9 100644
> --- a/drivers/infiniband/sw/rxe/rxe_qp.c
> +++ b/drivers/infiniband/sw/rxe/rxe_qp.c
[ ... ]
> @@ -858,6 +858,16 @@ int rxe_qp_chk_destroy(struct rxe_qp *qp)
> return -EBUSY;
> }
>
> + /* An attempt to destroy a QP while it still holds references for
> + * bound type 2 memory windows will fail immediately. Otherwise the
> + * QP is freed with those references outstanding and the windows are
> + * left pointing at freed memory (use-after-free).
> + */
> + if (atomic_read(&qp->mw_bind_num)) {
> + rxe_dbg_qp(qp, "Attempt to destroy while type 2 MWs are bound\n");
> + return -EBUSY;
> + }
> +
> return 0;
> }
[Severity: Critical]
Does this lockless check fully prevent the regression?
Because this check happens before the QP is drained or the worker task is
stopped, a racing userspace application could post an IB_WR_BIND_MW work
request that is processed immediately after this check passes.
When that happens, the delayed bind will cause rxe_cleanup() to hit its
50-second timeout waiting for the reference to drop. However, the timeout
error from rxe_cleanup() is ignored:
drivers/infiniband/sw/rxe/rxe_verbs.c:rxe_destroy_qp() {
...
err = rxe_cleanup(qp);
if (err)
rxe_err_qp(qp, "cleanup failed, err = %d\n", err);
return 0;
}
Since it unconditionally returns 0, won't the core still free the QP memory
while the memory window retains a dangling pointer?
--
Sashiko AI review · https://sashiko.dev/#/patchset/20260930182223.2511145-1-ays511.kr@gmail.com?part=1
^ permalink raw reply [flat|nested] 2+ messages in thread
end of thread, other threads:[~2026-09-30 18:41 UTC | newest]
Thread overview: 2+ messages (download: mbox.gz follow: Atom feed
-- links below jump to the message on this page --
2026-09-30 18:22 [PATCH] RDMA/rxe: refuse to destroy a QP with type 2 MWs still bound Youngsung Ahn
2026-09-30 18:41 ` sashiko-bot
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox