* [PATCH v3] RDMA/rxe: bound the ODP page index against the umem in rxe_check_pagefault()
@ 2026-10-04 15:32 Youngsung Ahn
2026-10-04 15:50 ` sashiko-bot
0 siblings, 1 reply; 2+ messages in thread
From: Youngsung Ahn @ 2026-10-04 15:32 UTC (permalink / raw)
To: Zhu Yanjun, Leon Romanovsky; +Cc: linux-rdma
An ODP MR has a page array, map.pfn_list[]. Its size is computed from
the umem start address. But rxe_check_pagefault() builds the index into
that array from an iova that is relative to the same start, while the
access itself was range-checked against the MR's iova (ibmr.iova).
Those two bases do not have to match. reg_mr only requires that start
and hca_va have the same page offset, so a user can register an MR with
hca_va = start + N*PAGE. The difference then becomes an index skew of
(hca_va - start) >> PAGE_SHIFT. Nothing compares the result with the
size of pfn_list[], so the read goes past the end of the array.
An unprivileged local user can reach this with a single RDMA operation
on a self-connected rxe QP. KASAN reports a slab-out-of-bounds read. If
the out-of-bounds qword happens to look like a writable HMM pfn, rxe
then copies the payload into a page the user never registered.
Reject an index that falls outside the umem and take the fault path
instead. That path rejects the out-of-range access in
ib_umem_odp_map_dma_and_lock(), which does check the umem range. The
other users of pfn_list[], __rxe_odp_mr_copy() and the atomic helpers,
run only after a successful map, so this single check covers them too.
Fixes: 2fae67ab63db ("RDMA/rxe: Add support for Send/Recv/Write/Read with ODP")
Signed-off-by: Youngsung Ahn <ays511.kr@gmail.com>
Assisted-by: LLM
---
v3: no code change. I am sending it as a standalone mail this time,
because v2 was sent as a reply to v1. I also dropped the Cc: stable
trailer. Both were asked for on the list.
Test results. The kernel is an unpatched KASAN x86-64 build of 7.3-rc4
(93f51579e7df). The reproducer runs as uid 1000, against an rxe link on
lo. It registers an ODP MR with length 0x8000 and
hca_va = start + 0x8000, that is 8 pages of skew. So map.pfn_list[] has
8 entries and the valid indices are 0..7. The reproducer then posts a
recv WQE in which every field is legal (num_sge=1, cur_sge=0,
sge_offset=0, sge.addr == ibmr.iova) and sends 64 bytes to its own QPN.
The skewed index lands on entry 8:
BUG: KASAN: slab-out-of-bounds in rxe_odp_map_range_and_lock+0x13b/0x280
Read of size 8 at addr ffff888102bf3f40 by task kworker/u8:1/32
CPU: 0 UID: 0 PID: 32 Comm: kworker/u8:1 Not tainted 7.3.0-rc4 #3
Workqueue: rxe_wq do_work
Call Trace:
kasan_report+0xce/0x100
rxe_odp_map_range_and_lock+0x13b/0x280
rxe_odp_mr_copy+0x7f/0x1e0
copy_data+0x124/0x2d0
rxe_receiver+0x20a6/0x3cb0
do_work+0xb6/0x250
process_one_work+0x3d1/0x790
worker_thread+0x296/0x500
kthread+0x194/0x1e0
ret_from_fork+0x2ac/0x3c0
Allocated by task 86:
rxe_odp_mr_init_user+0x52/0x180
The buggy address belongs to the object at ffff888102bf3f00
which belongs to the cache kmalloc-64 of size 64
The buggy address is located 0 bytes to the right of
allocated 64-byte region [ffff888102bf3f00, ffff888102bf3f40)
The 64-byte region is map.pfn_list[] itself, which is 8 entries of 8
bytes. The "Allocated by" stack names rxe_odp_mr_init_user(), so the
object really is the pfn_list of this MR. The read is exactly one entry
past the end. That is the (hca_va - start) >> PAGE_SHIFT skew described
above. rxe_check_pagefault() is inlined, so the report names
rxe_odp_map_range_and_lock() instead.
The code is unchanged in mainline 551c722f4080 (2026-09-29). This patch
itself is only compile-tested (KASAN + CONFIG_RDMA_RXE). I have not run
the reproducer again on a patched kernel.
You may prefer to index relative to ibmr.iova instead, the way the
non-ODP rxe_mr_iova_to_index() does, so that MRs with a legitimate skew
keep working. This patch takes the smaller and safer route. Please tell
me if you want the other approach.
The reproducer needs CONFIG_INFINIBAND_ON_DEMAND_PAGING. The fix is not
merged yet, so I am not posting it on the list. I can send it to you
off-list, as I did for the wr_opcode patch.
drivers/infiniband/sw/rxe/rxe_odp.c | 5 +++++
1 file changed, 5 insertions(+)
diff --git a/drivers/infiniband/sw/rxe/rxe_odp.c b/drivers/infiniband/sw/rxe/rxe_odp.c
index ab21b620e94c..07a30c60fc62 100644
--- a/drivers/infiniband/sw/rxe/rxe_odp.c
+++ b/drivers/infiniband/sw/rxe/rxe_odp.c
@@ -136,6 +136,11 @@ static inline bool rxe_check_pagefault(struct ib_umem_odp *umem_odp, u64 iova,
while (addr < iova + length) {
idx = (addr - ib_umem_start(umem_odp)) >> umem_odp->page_shift;
+ if (idx >= ib_umem_odp_num_pages(umem_odp)) {
+ need_fault = true;
+ break;
+ }
+
if ((umem_odp->map.pfn_list[idx] & access) != access) {
need_fault = true;
break;
--
2.43.0
^ permalink raw reply related [flat|nested] 2+ messages in thread
* Re: [PATCH v3] RDMA/rxe: bound the ODP page index against the umem in rxe_check_pagefault()
2026-10-04 15:32 [PATCH v3] RDMA/rxe: bound the ODP page index against the umem in rxe_check_pagefault() Youngsung Ahn
@ 2026-10-04 15:50 ` sashiko-bot
0 siblings, 0 replies; 2+ messages in thread
From: sashiko-bot @ 2026-10-04 15:50 UTC (permalink / raw)
To: Youngsung Ahn; +Cc: linux-rdma
> An ODP MR has a page array, map.pfn_list[]. Its size is computed from
> the umem start address. But rxe_check_pagefault() builds the index into
> that array from an iova that is relative to the same start, while the
> access itself was range-checked against the MR's iova (ibmr.iova).
>
> Those two bases do not have to match. reg_mr only requires that start
> and hca_va have the same page offset, so a user can register an MR with
> hca_va = start + N*PAGE. The difference then becomes an index skew of
> (hca_va - start) >> PAGE_SHIFT. Nothing compares the result with the
> size of pfn_list[], so the read goes past the end of the array.
>
> An unprivileged local user can reach this with a single RDMA operation
> on a self-connected rxe QP. KASAN reports a slab-out-of-bounds read. If
> the out-of-bounds qword happens to look like a writable HMM pfn, rxe
> then copies the payload into a page the user never registered.
>
> Reject an index that falls outside the umem and take the fault path
> instead. That path rejects the out-of-range access in
> ib_umem_odp_map_dma_and_lock(), which does check the umem range. The
> other users of pfn_list[], __rxe_odp_mr_copy() and the atomic helpers,
> run only after a successful map, so this single check covers them too.
>
> Fixes: 2fae67ab63db ("RDMA/rxe: Add support for Send/Recv/Write/Read with ODP")
> Signed-off-by: Youngsung Ahn <ays511.kr@gmail.com>
Sashiko has reviewed this patch and found no issues. It looks great!
--
Sashiko AI review · https://sashiko.dev/#/patchset/20261004153246.2951263-1-ays511.kr@gmail.com?part=1
^ permalink raw reply [flat|nested] 2+ messages in thread
end of thread, other threads:[~2026-10-04 15:50 UTC | newest]
Thread overview: 2+ messages (download: mbox.gz follow: Atom feed
-- links below jump to the message on this page --
2026-10-04 15:32 [PATCH v3] RDMA/rxe: bound the ODP page index against the umem in rxe_check_pagefault() Youngsung Ahn
2026-10-04 15:50 ` sashiko-bot
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox