Linux RDMA and InfiniBand development
 help / color / mirror / Atom feed
From: Youngsung Ahn <ays511.kr@gmail.com>
To: Zhu Yanjun <yanjun.zhu@linux.dev>, Leon Romanovsky <leon@kernel.org>
Cc: linux-rdma@vger.kernel.org
Subject: [PATCH v3] RDMA/rxe: bound the ODP page index against the umem in rxe_check_pagefault()
Date: Mon,  5 Oct 2026 00:32:46 +0900	[thread overview]
Message-ID: <20261004153246.2951263-1-ays511.kr@gmail.com> (raw)

An ODP MR has a page array, map.pfn_list[]. Its size is computed from
the umem start address. But rxe_check_pagefault() builds the index into
that array from an iova that is relative to the same start, while the
access itself was range-checked against the MR's iova (ibmr.iova).

Those two bases do not have to match. reg_mr only requires that start
and hca_va have the same page offset, so a user can register an MR with
hca_va = start + N*PAGE. The difference then becomes an index skew of
(hca_va - start) >> PAGE_SHIFT. Nothing compares the result with the
size of pfn_list[], so the read goes past the end of the array.

An unprivileged local user can reach this with a single RDMA operation
on a self-connected rxe QP. KASAN reports a slab-out-of-bounds read. If
the out-of-bounds qword happens to look like a writable HMM pfn, rxe
then copies the payload into a page the user never registered.

Reject an index that falls outside the umem and take the fault path
instead. That path rejects the out-of-range access in
ib_umem_odp_map_dma_and_lock(), which does check the umem range. The
other users of pfn_list[], __rxe_odp_mr_copy() and the atomic helpers,
run only after a successful map, so this single check covers them too.

Fixes: 2fae67ab63db ("RDMA/rxe: Add support for Send/Recv/Write/Read with ODP")
Signed-off-by: Youngsung Ahn <ays511.kr@gmail.com>
Assisted-by: LLM
---
v3: no code change. I am sending it as a standalone mail this time,
because v2 was sent as a reply to v1. I also dropped the Cc: stable
trailer. Both were asked for on the list.

Test results. The kernel is an unpatched KASAN x86-64 build of 7.3-rc4
(93f51579e7df). The reproducer runs as uid 1000, against an rxe link on
lo. It registers an ODP MR with length 0x8000 and
hca_va = start + 0x8000, that is 8 pages of skew. So map.pfn_list[] has
8 entries and the valid indices are 0..7. The reproducer then posts a
recv WQE in which every field is legal (num_sge=1, cur_sge=0,
sge_offset=0, sge.addr == ibmr.iova) and sends 64 bytes to its own QPN.
The skewed index lands on entry 8:

  BUG: KASAN: slab-out-of-bounds in rxe_odp_map_range_and_lock+0x13b/0x280
  Read of size 8 at addr ffff888102bf3f40 by task kworker/u8:1/32
  CPU: 0 UID: 0 PID: 32 Comm: kworker/u8:1 Not tainted 7.3.0-rc4 #3
  Workqueue: rxe_wq do_work
  Call Trace:
   kasan_report+0xce/0x100
   rxe_odp_map_range_and_lock+0x13b/0x280
   rxe_odp_mr_copy+0x7f/0x1e0
   copy_data+0x124/0x2d0
   rxe_receiver+0x20a6/0x3cb0
   do_work+0xb6/0x250
   process_one_work+0x3d1/0x790
   worker_thread+0x296/0x500
   kthread+0x194/0x1e0
   ret_from_fork+0x2ac/0x3c0
  Allocated by task 86:
   rxe_odp_mr_init_user+0x52/0x180
  The buggy address belongs to the object at ffff888102bf3f00
   which belongs to the cache kmalloc-64 of size 64
  The buggy address is located 0 bytes to the right of
   allocated 64-byte region [ffff888102bf3f00, ffff888102bf3f40)

The 64-byte region is map.pfn_list[] itself, which is 8 entries of 8
bytes. The "Allocated by" stack names rxe_odp_mr_init_user(), so the
object really is the pfn_list of this MR. The read is exactly one entry
past the end. That is the (hca_va - start) >> PAGE_SHIFT skew described
above. rxe_check_pagefault() is inlined, so the report names
rxe_odp_map_range_and_lock() instead.

The code is unchanged in mainline 551c722f4080 (2026-09-29). This patch
itself is only compile-tested (KASAN + CONFIG_RDMA_RXE). I have not run
the reproducer again on a patched kernel.

You may prefer to index relative to ibmr.iova instead, the way the
non-ODP rxe_mr_iova_to_index() does, so that MRs with a legitimate skew
keep working. This patch takes the smaller and safer route. Please tell
me if you want the other approach.

The reproducer needs CONFIG_INFINIBAND_ON_DEMAND_PAGING. The fix is not
merged yet, so I am not posting it on the list. I can send it to you
off-list, as I did for the wr_opcode patch.

 drivers/infiniband/sw/rxe/rxe_odp.c | 5 +++++
 1 file changed, 5 insertions(+)

diff --git a/drivers/infiniband/sw/rxe/rxe_odp.c b/drivers/infiniband/sw/rxe/rxe_odp.c
index ab21b620e94c..07a30c60fc62 100644
--- a/drivers/infiniband/sw/rxe/rxe_odp.c
+++ b/drivers/infiniband/sw/rxe/rxe_odp.c
@@ -136,6 +136,11 @@ static inline bool rxe_check_pagefault(struct ib_umem_odp *umem_odp, u64 iova,
 	while (addr < iova + length) {
 		idx = (addr - ib_umem_start(umem_odp)) >> umem_odp->page_shift;
 
+		if (idx >= ib_umem_odp_num_pages(umem_odp)) {
+			need_fault = true;
+			break;
+		}
+
 		if ((umem_odp->map.pfn_list[idx] & access) != access) {
 			need_fault = true;
 			break;
-- 
2.43.0


             reply	other threads:[~2026-10-04 15:33 UTC|newest]

Thread overview: 2+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-10-04 15:32 Youngsung Ahn [this message]
2026-10-04 15:50 ` [PATCH v3] RDMA/rxe: bound the ODP page index against the umem in rxe_check_pagefault() sashiko-bot

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=20261004153246.2951263-1-ays511.kr@gmail.com \
    --to=ays511.kr@gmail.com \
    --cc=leon@kernel.org \
    --cc=linux-rdma@vger.kernel.org \
    --cc=yanjun.zhu@linux.dev \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox