Linux RDMA and InfiniBand development
 help / color / mirror / Atom feed
From: Zhu Yanjun <yanjun.zhu@linux.dev>
To: Tymbark7372 <tymbark7372@proton.me>,
	"yanjun.zhu@linux.dev" <yanjun.zhu@linux.dev>
Cc: linux-rdma@vger.kernel.org, zyjzyj2000@gmail.com, jgg@nvidia.com,
	leonro@nvidia.com, stable@vger.kernel.org
Subject: Re: [PATCH 0/4] RDMA/rxe: Fix u64 iova-overflow family in MR/ODP/RESP/MW paths
Date: Mon, 25 May 2026 13:54:36 -0700	[thread overview]
Message-ID: <2cb75dc7-c1d5-4a7c-94a9-1c8a09f122d7@linux.dev> (raw)
In-Reply-To: <QuzKB10MggFmGWmfyd7fR6r_cI8exwmdYrXemdK6h2MFg-lN9-iRPtdKuL7aGu5GGEODHzOvsj4yFKWL7trcCiIxoT8-FVbKaZyolnO8tOA=@proton.me>


在 2026/5/25 4:15, Tymbark7372 写道:
> Hi Zhu,
>
> Thanks for the reply.  The iova on USER MRs is fully attacker-controlled
> at three entry points; cited file:line so it's verifiable:
>
> 1. Registration via /dev/infiniband/uverbs0:
>     drivers/infiniband/core/uverbs_cmd.c:746 does
>     `mr->iova = cmd.hca_va`.  `cmd.hca_va` comes straight from the
>     user's `struct ib_uverbs_reg_mr` over the UAPI.  The only check
>     (line 714) is that the low PAGE_SHIFT bits of cmd.hca_va match
>     cmd.start -- high bits are not constrained.  Line 731 passes
>     `cmd.hca_va` as the `iova` argument to
>     `pd->device->ops.reg_user_mr()`, which in rxe is
>     `rxe_reg_user_mr()`.  No subsystem rewriting on this path.
>
> 2. Work-request posting via ibv_post_send():
>     drivers/infiniband/sw/rxe/rxe_verbs.c:822 copies the user's
>     `ibwr->sg_list` directly into the WQE; `sge->addr` is the iova
>     that subsequently flows to `mr_check_range()` on SEND/RDMA-WRITE.
>     User chooses any u64.
>
> 3. Network path (RoCEv2):
>     drivers/infiniband/sw/rxe/rxe_resp.c:409 reads
>     `qp->resp.va = reth_va(pkt)` directly from the RETH wire header
>     on inbound RDMA-WRITE/READ/ATOMIC.  Line 1357 does the same in
>     duplicate_request.  A peer on UDP/4791 with a known QPN/rkey
>     controls iova directly.
>
> PoC is a small libibverbs program: register a destination MR with
> hca_va = 0xFFFFFFFFFFFFFC00, then post a single ibv_post_send()
> with IBV_WR_RDMA_WRITE pointing at that MR's rkey, sge.length = 0x400.
> On v7.1.0-rc3 + KASAN this hits:
>
>    WARNING ... at rxe_mr_iova_to_index+0x135/0x180
>    ...
>    BUG: unable to handle page fault for address: ffff887b3ba27250
>    RIP: 0010:rxe_mr_copy+0x20d/0x5e0
>
> Reproduced as unprivileged uid with /dev/infiniband/uverbs0 open.
> Happy to send the four PoC sources (one per sibling) inline if you'd
> like to reproduce.

Thank you very much. If you could share a PoC source, I would really

appreciate it, as it would help us reproduce the issue locally.

If sharing the PoC is not convenient, it would also be very helpful if

you could post your reproduction evidence or logs publicly on the

community thread, so others can independently verify and confirm

the issue as well.

Thanks a lot.

Zhu Yanjun

>
> Tymbark7372
>
>
>
> On Friday, May 22nd, 2026 at 4:54 AM, Zhu Yanjun <yanjun.zhu@linux.dev> wrote:
>
>> 在 2026/5/21 12:44, Tymbark7372 写道:
>>> This patchset fixes a family of u64 overflow bugs in the rxe Soft-RoCE
>>> driver.  All four sites share one root cause: addition of an
>>> attacker-influenced iova/addr (u64) with an attacker-influenced
>>> length/resid (size_t/u32/int promoted to u64), without overflow
>>> check, leading to an OOB read/write primitive in the rxe responder
>>> workqueue.
>> The core premise of these commits is that a user-space program can
>> arbitrarily set the IOVA via /dev/infiniband/uverbs0. I am still
>> skeptical about this, as my understanding was that the IOVA is managed
>> by the subsystem and difficult for a user to modify. If it is indeed
>> possible for a user to control or change this IOVA, then I am completely
>> fine with this patchset.
>>
>> Thanks a lot.
>> Zhu Yanjun
>>
>>> I originally reported these to security@kernel.org.  Jason Gunthorpe
>>> confirmed that rxe and siw are development-only drivers without
>>> embargo handling and asked me to send patches publicly, so I'm
>>> posting here per his direction.  security@kernel.org is intentionally
>>> not in Cc per Jason's instruction.
>>>
>>> This is a resend of the patches I sent earlier today as attachments.
>>> Zhu Yanjun pointed out attachments aren't the convention and asked
>>> for inline format via git send-email.
>>>
>>> Patches:
>>>
>>>     1/4: rxe_mr.c mr_check_range
>>>          The USER/MEM_REG case computes iova + length and compares to
>>>          mr->ibmr.iova + mr->ibmr.length.  Both additions wrap in u64.
>>>          Use check_add_overflow() for both ends.
>>>
>>>     2/4: rxe_odp.c rxe_check_pagefault
>>>          Loop condition addr < iova + length wraps when iova is near
>>>          U64_MAX and length is positive.  Compute iova_end with
>>>          check_add_overflow() once and use it in the loop condition.
>>>
>>>     3/4: rxe_resp.c duplicate_request
>>>          Third clause iova + resid > res->read.va_org + res->read.length
>>>          has u64 wrap on both sides.  Use check_add_overflow() for both
>>>          ends.  (Site A in check_rkey, also in rxe_resp.c, calls into
>>>          mr_check_range and is closed by patch 1.)
>>>
>>>     4/4: rxe_mw.c rxe_check_bind_mw
>>>          Same wrap class as patch 1.  Found by sibling-site grep; not on
>>>          the OOB-write path of the three primary bugs but a
>>>          structurally-identical u64 wrap that would let an attacker bind
>>>          a memory window outside its parent MR's range.
>>>
>>> Verification:
>>>
>>> Each of the three primary sibling triggers (patches 1, 2, 3) has been
>>> exercised on v7.1.0-rc3 + KASAN in QEMU as the OOB-write case.
>>> Patches 1 and 3 produce a single-page-fault Oops in rxe_mr_copy after
>>> the wrap.  Patch 2 produces a single-page-fault Oops in
>>> rxe_odp_mr_copy.  All three are triggered by a single ibv_post_send
>>> from an unprivileged local user with /dev/infiniband/uverbs0 open.
>>> A working LPE exploit demonstrated end-to-end privilege escalation
>>> via the rxe_odp path under the verification config (KASAN dev-build,
>>> selinux=0, nokaslr).  Full PoC and writeup were attached to the
>>> original security@kernel.org submission.
>>>
>>> After applying all four patches, the same triggers no longer fire;
>>> the wrap checks correctly reject the attacker iova.  Re-tested in the
>>> same QEMU+KASAN configuration.
>>>
>>> The trigger PoCs are simple libibverbs programs (one per sibling)
>>> that I am happy to provide on request.
>>>
>>> Fixes / stable:
>>>
>>>     1/4: Fixes 8700e3e7c485 ("Soft RoCE driver"), v4.8+
>>>     2/4: Fixes 2fae67ab63db ("RDMA/rxe: Add support for Send/Recv/Write/Read with ODP"), v6.15+
>>>     3/4: Fixes 8700e3e7c485 ("Soft RoCE driver"), v4.8+
>>>     4/4: Fixes 8700e3e7c485 ("Soft RoCE driver"), v4.8+
>>>
>>> Pre-f04d5b3d916c LTS branches carry the older wrap form
>>>     iova > mr->ibmr.iova + mr->ibmr.length - length
>>> instead of the current `iova + length > ...` shape.  Patches 1, 3, 4
>>> will need a backport variant for those branches; I can provide on
>>> request.
>>>
>>> Tymbark7372 (4):
>>>     RDMA/rxe: Fix u64 iova+length overflow in mr_check_range
>>>     RDMA/rxe: Fix u64 iova+length overflow in rxe_check_pagefault
>>>     RDMA/rxe: Fix u64 iova+resid overflow in duplicate_request
>>>     RDMA/rxe: Fix u64 addr+length overflow in rxe_check_bind_mw
>>>
>>>    drivers/infiniband/sw/rxe/rxe_mr.c   | 12 +++++++++---
>>>    drivers/infiniband/sw/rxe/rxe_odp.c  | 10 ++++++++--
>>>    drivers/infiniband/sw/rxe/rxe_resp.c | 11 ++++++++---
>>>    drivers/infiniband/sw/rxe/rxe_mw.c   | 11 ++++++++---
>>>    4 files changed, 33 insertions(+), 11 deletions(-)
>>>
>>> --
>>> 2.43.0
>>>
>>
-- 
Best Regards,
Yanjun.Zhu


  reply	other threads:[~2026-05-25 20:54 UTC|newest]

Thread overview: 11+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-05-21 19:44 [PATCH 0/4] RDMA/rxe: Fix u64 iova-overflow family in MR/ODP/RESP/MW paths Tymbark7372
2026-05-21 19:44 ` [PATCH 1/4] RDMA/rxe: Fix u64 iova+length overflow in mr_check_range Tymbark7372
2026-05-22  4:44   ` Greg KH
2026-05-21 19:44 ` [PATCH 2/4] RDMA/rxe: Fix u64 iova+length overflow in rxe_check_pagefault Tymbark7372
2026-05-21 19:44 ` [PATCH 3/4] RDMA/rxe: Fix u64 iova+resid overflow in duplicate_request Tymbark7372
2026-05-21 19:44 ` [PATCH 4/4] RDMA/rxe: Fix u64 addr+length overflow in rxe_check_bind_mw Tymbark7372
2026-05-22  2:54 ` [PATCH 0/4] RDMA/rxe: Fix u64 iova-overflow family in MR/ODP/RESP/MW paths Zhu Yanjun
2026-05-25 11:15   ` Tymbark7372
2026-05-25 20:54     ` Zhu Yanjun [this message]
  -- strict thread matches above, loose matches on Subject: below --
2026-05-21 15:39 Tymbark7372
2026-05-21 15:46 ` Zhu Yanjun

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=2cb75dc7-c1d5-4a7c-94a9-1c8a09f122d7@linux.dev \
    --to=yanjun.zhu@linux.dev \
    --cc=jgg@nvidia.com \
    --cc=leonro@nvidia.com \
    --cc=linux-rdma@vger.kernel.org \
    --cc=stable@vger.kernel.org \
    --cc=tymbark7372@proton.me \
    --cc=zyjzyj2000@gmail.com \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox