Linux RDMA and InfiniBand development
 help / color / mirror / Atom feed
* [PATCH] RDMA/rxe: refuse to destroy a QP with type 2 MWs still bound
@ 2026-09-30 18:22 Youngsung Ahn
  2026-09-30 18:41 ` sashiko-bot
  0 siblings, 1 reply; 2+ messages in thread
From: Youngsung Ahn @ 2026-09-30 18:22 UTC (permalink / raw)
  To: zyjzyj2000; +Cc: linux-rdma, security

From: Youngsung Ahn <ays511.kr@gmail.com>

Binding a type 2 memory window takes a reference on the QP:
rxe_do_bind_mw() does rxe_get(qp) and stores it in mw->qp. That
reference is released only when the window is invalidated
(rxe_do_invalidate_mw() on IB_WR_LOCAL_INV) or the window itself is
destroyed (rxe_mw_cleanup()).

rxe_qp_chk_destroy() does not account for those references -- it only
refuses when qp->mcg_num is non-zero. So an unprivileged user can bind
a type 2 MW to a QP and then destroy the QP while the binding is still
live. __rxe_cleanup() cannot drop the QP refcount to zero (the MW still
holds a reference), hits its -ETIMEDOUT path, and frees the QP anyway;
mw->qp is left dangling. When the MW is later invalidated or
deallocated, rxe_mw_cleanup()/rxe_do_invalidate_mw() do rxe_put(mw->qp)
on the freed struct rxe_qp -- a use-after-free write (refcount_dec, and
complete()/list work if it reaches zero) reachable from userspace.
IB_WR_BIND_MW is a local operation, so no peer is required.

Track the number of type 2 MWs bound to a QP and refuse to destroy the
QP while that count is non-zero, returning -EBUSY exactly as the
existing multicast-attachment check does. The window must be
invalidated or deallocated, which drops the reference and the count,
before the QP can be destroyed.

Fixes: 8700e3e7c485 ("Soft RoCE driver")
Signed-off-by: Youngsung Ahn <ays511.kr@gmail.com>
---
Notes (not part of the commit):
Reproduced on a KASAN x86-64 build of 7.3-rc4 as uid 1000: deterministic
KASAN slab-use-after-free write via rxe_mw_cleanup() -> rxe_put() on the
freed struct rxe_qp (bind a type 2 MW to an RTS QP, destroy the QP, then
deallocate the MW). Present unchanged in mainline 551c722f4080
(2026-09-29) and rdma for-next. This patch is compile-tested
(KASAN+RDMA_RXE); the counter-based fix itself has not been runtime-tested
in this form. Found through manual review; per security-bugs.rst this is
public and a reproducer can be shared on request. The Fixes: tag points at
the base driver and should be refined to the type 2 MW bind support commit
when preparing for merge.

 drivers/infiniband/sw/rxe/rxe_mw.c    |  3 +++
 drivers/infiniband/sw/rxe/rxe_qp.c    | 10 ++++++++++
 drivers/infiniband/sw/rxe/rxe_verbs.h |  1 +
 3 files changed, 14 insertions(+)

diff --git a/drivers/infiniband/sw/rxe/rxe_mw.c b/drivers/infiniband/sw/rxe/rxe_mw.c
index bddb7a25..a2e3a6ce 100644
--- a/drivers/infiniband/sw/rxe/rxe_mw.c
+++ b/drivers/infiniband/sw/rxe/rxe_mw.c
@@ -161,6 +161,7 @@ static void rxe_do_bind_mw(struct rxe_qp *qp, struct rxe_send_wqe *wqe,
 
 	if (mw->ibmw.type == IB_MW_TYPE_2) {
 		rxe_get(qp);
+		atomic_inc(&qp->mw_bind_num);
 		mw->qp = qp;
 	}
 }
@@ -245,6 +246,7 @@ static void rxe_do_invalidate_mw(struct rxe_mw *mw)
 	/* valid type 2 MW will always have a QP pointer */
 	qp = mw->qp;
 	mw->qp = NULL;
+	atomic_dec(&qp->mw_bind_num);
 	rxe_put(qp);
 
 	/* valid type 2 MW will always have an MR pointer */
@@ -332,6 +334,7 @@ void rxe_mw_cleanup(struct rxe_pool_elem *elem)
 		struct rxe_qp *qp = mw->qp;
 
 		mw->qp = NULL;
+		atomic_dec(&qp->mw_bind_num);
 		rxe_put(qp);
 	}
 
diff --git a/drivers/infiniband/sw/rxe/rxe_qp.c b/drivers/infiniband/sw/rxe/rxe_qp.c
index 311f285d..fea55bc3 100644
--- a/drivers/infiniband/sw/rxe/rxe_qp.c
+++ b/drivers/infiniband/sw/rxe/rxe_qp.c
@@ -858,6 +858,16 @@ int rxe_qp_chk_destroy(struct rxe_qp *qp)
 		return -EBUSY;
 	}
 
+	/* An attempt to destroy a QP while it still holds references for
+	 * bound type 2 memory windows will fail immediately.  Otherwise the
+	 * QP is freed with those references outstanding and the windows are
+	 * left pointing at freed memory (use-after-free).
+	 */
+	if (atomic_read(&qp->mw_bind_num)) {
+		rxe_dbg_qp(qp, "Attempt to destroy while type 2 MWs are bound\n");
+		return -EBUSY;
+	}
+
 	return 0;
 }
 
diff --git a/drivers/infiniband/sw/rxe/rxe_verbs.h b/drivers/infiniband/sw/rxe/rxe_verbs.h
index 0f5ffd94..80df376e 100644
--- a/drivers/infiniband/sw/rxe/rxe_verbs.h
+++ b/drivers/infiniband/sw/rxe/rxe_verbs.h
@@ -261,6 +261,7 @@ struct rxe_qp {
 	struct rxe_av		alt_av;
 
 	atomic_t		mcg_num;
+	atomic_t		mw_bind_num;
 
 	struct sk_buff_head	req_pkts;
 	struct sk_buff_head	resp_pkts;
--
2.43.0

^ permalink raw reply related	[flat|nested] 2+ messages in thread

* Re: [PATCH] RDMA/rxe: refuse to destroy a QP with type 2 MWs still bound
  2026-09-30 18:22 [PATCH] RDMA/rxe: refuse to destroy a QP with type 2 MWs still bound Youngsung Ahn
@ 2026-09-30 18:41 ` sashiko-bot
  0 siblings, 0 replies; 2+ messages in thread
From: sashiko-bot @ 2026-09-30 18:41 UTC (permalink / raw)
  To: Youngsung Ahn; +Cc: linux-rdma

Thank you for your contribution! Sashiko AI review found 1 potential issue(s) to consider:
- [Critical] RDMA/rxe: TOCTOU race in rxe_qp_chk_destroy() allows UAF
--

commit 94e08f1cc2290d06c7fcd154d031cde9d5a0054c
Author: Youngsung Ahn <ays511.kr@gmail.com>

    RDMA/rxe: refuse to destroy a QP with type 2 MWs still bound

This commit adds tracking for type 2 memory windows bound to a QP via an
atomic counter. It attempts to prevent the destruction of a QP if there are
outstanding type 2 MW bindings by returning -EBUSY.

> diff --git a/drivers/infiniband/sw/rxe/rxe_mw.c b/drivers/infiniband/sw/rxe/rxe_mw.c
> index bddb7a2578313..a2e3a6ce09463 100644
> --- a/drivers/infiniband/sw/rxe/rxe_mw.c
> +++ b/drivers/infiniband/sw/rxe/rxe_mw.c
[ ... ]
> @@ -161,6 +161,7 @@ static void rxe_do_bind_mw(struct rxe_qp *qp, struct rxe_send_wqe *wqe,
>  
>  	if (mw->ibmw.type == IB_MW_TYPE_2) {
>  		rxe_get(qp);
> +		atomic_inc(&qp->mw_bind_num);
>  		mw->qp = qp;
>  	}
>  }

[Severity: Critical]
Is it possible for the worker task to execute this concurrently during QP
destruction?

Since qp->valid is not cleared until rxe_qp_do_cleanup() runs, the requester
can freely process pending BIND_MW work requests during the cleanup delay:

drivers/infiniband/sw/rxe/rxe_req.c:rxe_requester() {
    ...
    if (unlikely(!qp->valid)) {
    ...
}

If this runs concurrently, couldn't the counter be incremented too late,
holding the reference that triggers the 50-second timeout during cleanup?

> diff --git a/drivers/infiniband/sw/rxe/rxe_qp.c b/drivers/infiniband/sw/rxe/rxe_qp.c
> index 311f285d78a6b..fea55bc3b0cb9 100644
> --- a/drivers/infiniband/sw/rxe/rxe_qp.c
> +++ b/drivers/infiniband/sw/rxe/rxe_qp.c
[ ... ]
> @@ -858,6 +858,16 @@ int rxe_qp_chk_destroy(struct rxe_qp *qp)
>  		return -EBUSY;
>  	}
>  
> +	/* An attempt to destroy a QP while it still holds references for
> +	 * bound type 2 memory windows will fail immediately.  Otherwise the
> +	 * QP is freed with those references outstanding and the windows are
> +	 * left pointing at freed memory (use-after-free).
> +	 */
> +	if (atomic_read(&qp->mw_bind_num)) {
> +		rxe_dbg_qp(qp, "Attempt to destroy while type 2 MWs are bound\n");
> +		return -EBUSY;
> +	}
> +
>  	return 0;
>  }

[Severity: Critical]
Does this lockless check fully prevent the regression?

Because this check happens before the QP is drained or the worker task is
stopped, a racing userspace application could post an IB_WR_BIND_MW work
request that is processed immediately after this check passes.

When that happens, the delayed bind will cause rxe_cleanup() to hit its
50-second timeout waiting for the reference to drop. However, the timeout
error from rxe_cleanup() is ignored:

drivers/infiniband/sw/rxe/rxe_verbs.c:rxe_destroy_qp() {
    ...
    err = rxe_cleanup(qp);
    if (err)
        rxe_err_qp(qp, "cleanup failed, err = %d\n", err);

    return 0;
}

Since it unconditionally returns 0, won't the core still free the QP memory
while the memory window retains a dangling pointer?

-- 
Sashiko AI review · https://sashiko.dev/#/patchset/20260930182223.2511145-1-ays511.kr@gmail.com?part=1

^ permalink raw reply	[flat|nested] 2+ messages in thread

end of thread, other threads:[~2026-09-30 18:41 UTC | newest]

Thread overview: 2+ messages (download: mbox.gz follow: Atom feed
-- links below jump to the message on this page --
2026-09-30 18:22 [PATCH] RDMA/rxe: refuse to destroy a QP with type 2 MWs still bound Youngsung Ahn
2026-09-30 18:41 ` sashiko-bot

This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox