Linux RDMA and InfiniBand development
 help / color / mirror / Atom feed
* [PATCH v3] RDMA/rxe: validate cur_sge on the receive path before indexing dma->sge[]
@ 2026-10-04 15:33 Youngsung Ahn
  2026-10-04 15:40 ` sashiko-bot
  0 siblings, 1 reply; 2+ messages in thread
From: Youngsung Ahn @ 2026-10-04 15:33 UTC (permalink / raw)
  To: Zhu Yanjun, Leon Romanovsky; +Cc: linux-rdma

A user QP's receive queue is an mmap'd ring. The user provider writes
the WQEs into it and post_recv never touches them, so every field of
struct rxe_recv_wqe comes from user space. When the responder picks up a
WQE it only checks dma.num_sge (in rxe_get_recv_wqe() and
get_srq_wqe()). dma.cur_sge, dma.sge_offset and dma.resid are copied
into qp->resp.srq_wqe without any check.

copy_data() then uses cur_sge as an index right away:

	struct rxe_sge *sge = &dma->sge[dma->cur_sge];

It dereferences that pointer (sge->length) before the in-loop
"dma->cur_sge >= dma->num_sge" test, which only runs after the first
sge++. cur_sge is a u32, so the pointer can land far outside
dma->sge[], which has RXE_MAX_SGE entries and sits inside the
kmalloc-2k struct rxe_qp. That gives an information leak, because the
out-of-bounds sge->length can be recovered from the work-completion
status, and it also gives a DoS. The send queue already got a
num_sge/cur_sge check in commit 126c757e4cd4 ("RDMA/rxe: Validate
num_sge/cur_sge before indexing wqe->dma.sge[]"). The receive path did
not.

Reject an out-of-range cur_sge before the first dereference. Read the
field once with READ_ONCE() into a local, then check and index that
same local. This matters because copy_data() also runs on the send path
(finish_packet() -> copy_data(&wqe->dma, ...)), and there dma points
straight at the mmap'd ring instead of the kernel copy the responder
makes. So the value that is checked must be the value that is used. The
kernel ULP path always sets cur_sge = 0, so it is not affected.

This is not a complete bound for the send path, and I would rather say
so than imply otherwise. num_sge is read from the ring as well, and the
retransmit path reaches dma->sge[] through advance_dma_data(), which
still has the same unchecked index. I will send those separately.

Fixes: 8700e3e7c485 ("Soft RoCE driver")
Signed-off-by: Youngsung Ahn <ays511.kr@gmail.com>
Assisted-by: LLM
---
v3: read dma->cur_sge once with READ_ONCE() into a local, then check and
index that same local, so the value that is checked is the value that is
used. Automated review pointed out that v2 read the field twice. I am
also sending this as a standalone mail and I dropped the Cc: stable
trailer.

As the commit message says, this is not a complete bound for the send
path. num_sge also comes from the ring, so a concurrent write can still
widen the window, and the retransmit path reaches dma->sge[] through
advance_dma_data() instead of copy_data(). A complete fix would pass the
queue's max_sge into copy_data(), instead of comparing two fields that
both come from the ring. I did not want to change the function signature
in a fix patch without asking first. Please tell me if you prefer that.

Test results. The kernel is an unpatched KASAN x86-64 build of 7.3-rc4
(93f51579e7df). The reproducer runs as a normal user (uid 1000), against
an rxe link on lo. It writes a recv WQE straight into the mmap'd RQ ring
with dma.cur_sge = 0x30, then posts a 64-byte SEND to the same QP:

  BUG: KASAN: slab-out-of-bounds in copy_data+0xa5/0x2d0
  Read of size 4 at addr ffff88810164c858 by task kworker/u8:1/32
  CPU: 1 UID: 0 PID: 32 Comm: kworker/u8:1 Not tainted 7.3.0-rc4 #3
  Workqueue: rxe_wq do_work
  Call Trace:
   kasan_report+0xce/0x100
   copy_data+0xa5/0x2d0
   rxe_receiver+0x20a6/0x3cb0
   do_work+0xb6/0x250
   process_one_work+0x3d1/0x790
   worker_thread+0x296/0x500
   kthread+0x194/0x1e0
   ret_from_fork+0x2ac/0x3c0
  The buggy address belongs to the object at ffff88810164c000
   which belongs to the cache kmalloc-2k of size 2048
  The buggy address is located 96 bytes to the right of
   allocated 2040-byte region [ffff88810164c000, ffff88810164c7f8)

The 2040-byte kmalloc-2k region is the struct rxe_qp from the commit
message. A second report follows at copy_data+0xda, which is the
sge->length read inside the loop. The report names the rxe_wq kworker
because that is where the responder consumes the WQE.

With a larger cur_sge (0x40000) the same read lands outside any live
object, and KASAN reports use-after-free instead. So the index is not
just off by a few entries:

  BUG: KASAN: use-after-free in copy_data+0xa5/0x2d0

The code is unchanged in mainline 551c722f4080 (2026-09-29). This patch
itself is only compile-tested (KASAN + CONFIG_RDMA_RXE). I have not run
the reproducer again on a patched kernel.

The reproducer is a small C program. It uses the legacy uverbs write ABI,
so it does not need libibverbs. The fix is not merged yet, so I am not
posting it on the list. I can send it to you off-list, as I did for the
wr_opcode patch.

 drivers/infiniband/sw/rxe/rxe_mr.c | 13 ++++++++++++-
 1 file changed, 12 insertions(+), 1 deletion(-)

diff --git a/drivers/infiniband/sw/rxe/rxe_mr.c b/drivers/infiniband/sw/rxe/rxe_mr.c
index 71d9ea477289..6c7d6919e5d5 100644
--- a/drivers/infiniband/sw/rxe/rxe_mr.c
+++ b/drivers/infiniband/sw/rxe/rxe_mr.c
@@ -422,7 +422,8 @@ int copy_data(
 	enum rxe_mr_copy_dir	dir)
 {
 	int			bytes;
-	struct rxe_sge		*sge	= &dma->sge[dma->cur_sge];
+	struct rxe_sge		*sge;
+	u32			cur_sge	= READ_ONCE(dma->cur_sge);
 	int			offset	= dma->sge_offset;
 	int			resid	= dma->resid;
 	struct rxe_mr		*mr	= NULL;
@@ -437,6 +438,16 @@ int copy_data(
 		goto err2;
 	}
 
+	/*
+	 * dma may live in a user-mapped queue, so read cur_sge once and
+	 * validate the value that is actually used to index sge[].
+	 */
+	if (unlikely(cur_sge >= dma->num_sge)) {
+		err = -EINVAL;
+		goto err2;
+	}
+	sge = &dma->sge[cur_sge];
+
 	if (sge->length && (offset < sge->length)) {
 		mr = lookup_mr(pd, access, sge->lkey, RXE_LOOKUP_LOCAL);
 		if (!mr) {
-- 
2.43.0


^ permalink raw reply related	[flat|nested] 2+ messages in thread

* Re: [PATCH v3] RDMA/rxe: validate cur_sge on the receive path before indexing dma->sge[]
  2026-10-04 15:33 [PATCH v3] RDMA/rxe: validate cur_sge on the receive path before indexing dma->sge[] Youngsung Ahn
@ 2026-10-04 15:40 ` sashiko-bot
  0 siblings, 0 replies; 2+ messages in thread
From: sashiko-bot @ 2026-10-04 15:40 UTC (permalink / raw)
  To: Youngsung Ahn; +Cc: linux-rdma

> A user QP's receive queue is an mmap'd ring. The user provider writes
> the WQEs into it and post_recv never touches them, so every field of
> struct rxe_recv_wqe comes from user space. When the responder picks up a
> WQE it only checks dma.num_sge (in rxe_get_recv_wqe() and
> get_srq_wqe()). dma.cur_sge, dma.sge_offset and dma.resid are copied
> into qp->resp.srq_wqe without any check.
> 
> copy_data() then uses cur_sge as an index right away:
> 
> 	struct rxe_sge *sge = &dma->sge[dma->cur_sge];
> 
> It dereferences that pointer (sge->length) before the in-loop
> "dma->cur_sge >= dma->num_sge" test, which only runs after the first
> sge++. cur_sge is a u32, so the pointer can land far outside
> dma->sge[], which has RXE_MAX_SGE entries and sits inside the
> [ ... ]
> retransmit path reaches dma->sge[] through advance_dma_data(), which
> still has the same unchecked index. I will send those separately.
> 
> Fixes: 8700e3e7c485 ("Soft RoCE driver")
> Signed-off-by: Youngsung Ahn <ays511.kr@gmail.com>

Sashiko has reviewed this patch and found no issues. It looks great!

-- 
Sashiko AI review · https://sashiko.dev/#/patchset/20261004153325.2952372-1-ays511.kr@gmail.com?part=1


^ permalink raw reply	[flat|nested] 2+ messages in thread

end of thread, other threads:[~2026-10-04 15:40 UTC | newest]

Thread overview: 2+ messages (download: mbox.gz follow: Atom feed
-- links below jump to the message on this page --
2026-10-04 15:33 [PATCH v3] RDMA/rxe: validate cur_sge on the receive path before indexing dma->sge[] Youngsung Ahn
2026-10-04 15:40 ` sashiko-bot

This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox