Linux RDMA and InfiniBand development
 help / color / mirror / Atom feed
* [PATCH v3] RDMA/rxe: bound the WQE opcode before indexing rxe_wr_opcode_info[]
@ 2026-10-02 16:45 Youngsung Ahn
  2026-10-03  1:33 ` sashiko-bot
  2026-10-04  3:40 ` Zhu Yanjun
  0 siblings, 2 replies; 3+ messages in thread
From: Youngsung Ahn @ 2026-10-02 16:45 UTC (permalink / raw)
  To: Zhu Yanjun, Leon Romanovsky; +Cc: linux-rdma

wr_opcode_mask() indexes the global rxe_wr_opcode_info[] array with
a caller-supplied opcode and no bounds check. For a user QP the WQE
comes from an mmap'd SQ ring, so wqe->wr.opcode is an
attacker-controlled __u32 read back by the requester
(req_next_wqe() -> rxe_requester()); rxe_post_send() takes the
qp->is_user branch and never runs validate_send_wr(), whose only
opcode check is itself !wr_opcode_mask() (it indexes before it
checks). rxe_wr_opcode_info[] has entries only up to IB_WR_REG_MR,
so an out-of-range opcode reads out of bounds and the result is used
as a mask and dereferenced; observed as a KASAN global-out-of-bounds
"Read of size 4" and a wild-pointer oops.

Return a zero mask for an out-of-range opcode, matching how
validate_send_wr() already treats a zero mask (an invalid WR), so no
path indexes the array out of bounds.

Fixes: 8700e3e7c485 ("Soft RoCE driver")
Signed-off-by: Youngsung Ahn <ays511.kr@gmail.com>
Assisted-by: LLM
---
Notes (not part of the commit):
v3: no code change. Sent as a standalone message rather than as a reply to
the v2 posting, and dropped the Cc: stable@vger.kernel.org trailer, both
as asked on-list.

On the question about req_retry(): req_retry() does not reject the WQE
itself, it only re-posts it. A zero mask makes every "mask & WR_*" test
there false, so an out-of-range opcode yields iova = 0, the dma cursor
reset, and neither retry_first_write_send() nor the read fix-up; the WQE
is left in wqe_state_posted. Once the requester dispatches it, wqe->mask
stays 0 so WR_LOCAL_OP_MASK is not taken, and next_opcode() matches no
case for an out-of-range wqe->wr.opcode and returns -EINVAL, so
rxe_requester() sets wqe->status = IB_WC_LOC_QP_OP_ERR and takes the err
path, which moves the QP to IB_QPS_ERR and lets the completer post the
completion.

wr_opcode_mask() is the single choke point all three callers
(validate_send_wr(), req_next_wqe(), req_retry()) funnel through, and only
validate_send_wr() rejects the WQE on a zero mask. In req_retry()
wr_opcode_mask() is called at the top of the loop body, before the
wqe_state_posted / wqe_state_done checks, so the out-of-bounds read
happens even for WQEs the loop then skips.

ARRAY_SIZE() is not usable here: rxe_opcode.h only declares "extern
struct rxe_wr_opcode_info rxe_wr_opcode_info[]". IB_WR_REG_MR (32) is the
highest index initialised in rxe_opcode.c, and on x86-64 the object is
1320 bytes with a 40-byte stride, i.e. 33 elements, so valid indices are
0..32.

The bug was reproduced on an unpatched KASAN x86-64 build of 7.3-rc4, with
the reproducer run as uid 1000 (the splat itself is attributed to the
rxe_wq kworker, since that is where the WQE is consumed): KASAN
global-out-of-bounds read of size 4 at rxe_wr_opcode_info, and an Oops or
GPF for larger opcodes. Present unchanged in mainline 551c722f4080
(2026-09-29). This patch itself is only compile-tested (KASAN +
CONFIG_RDMA_RXE). I am posting the reproducer on-list in reply to the v2
discussion of this patch.

 drivers/infiniband/sw/rxe/rxe_loc.h | 3 +++
 1 file changed, 3 insertions(+)

diff --git a/drivers/infiniband/sw/rxe/rxe_loc.h b/drivers/infiniband/sw/rxe/rxe_loc.h
index 64d636bf80fd..dcc0788db90a 100644
--- a/drivers/infiniband/sw/rxe/rxe_loc.h
+++ b/drivers/infiniband/sw/rxe/rxe_loc.h
@@ -184,6 +184,9 @@ void rxe_comp_queue_pkt(struct rxe_qp *qp, struct sk_buff *skb);
 
 static inline unsigned int wr_opcode_mask(int opcode, struct rxe_qp *qp)
 {
+	if (unlikely(opcode < 0 || opcode > IB_WR_REG_MR))
+		return 0;
+
 	return rxe_wr_opcode_info[opcode].mask[qp->ibqp.qp_type];
 }
 
-- 
2.43.0


^ permalink raw reply related	[flat|nested] 3+ messages in thread

* Re: [PATCH v3] RDMA/rxe: bound the WQE opcode before indexing rxe_wr_opcode_info[]
  2026-10-02 16:45 [PATCH v3] RDMA/rxe: bound the WQE opcode before indexing rxe_wr_opcode_info[] Youngsung Ahn
@ 2026-10-03  1:33 ` sashiko-bot
  2026-10-04  3:40 ` Zhu Yanjun
  1 sibling, 0 replies; 3+ messages in thread
From: sashiko-bot @ 2026-10-03  1:33 UTC (permalink / raw)
  To: Youngsung Ahn; +Cc: linux-rdma

> wr_opcode_mask() indexes the global rxe_wr_opcode_info[] array with
> a caller-supplied opcode and no bounds check. For a user QP the WQE
> comes from an mmap'd SQ ring, so wqe->wr.opcode is an
> attacker-controlled __u32 read back by the requester
> (req_next_wqe() -> rxe_requester()); rxe_post_send() takes the
> qp->is_user branch and never runs validate_send_wr(), whose only
> opcode check is itself !wr_opcode_mask() (it indexes before it
> checks). rxe_wr_opcode_info[] has entries only up to IB_WR_REG_MR,
> so an out-of-range opcode reads out of bounds and the result is used
> as a mask and dereferenced; observed as a KASAN global-out-of-bounds
> "Read of size 4" and a wild-pointer oops.
> 
> Return a zero mask for an out-of-range opcode, matching how
> validate_send_wr() already treats a zero mask (an invalid WR), so no
> path indexes the array out of bounds.
> 
> Fixes: 8700e3e7c485 ("Soft RoCE driver")
> Signed-off-by: Youngsung Ahn <ays511.kr@gmail.com>

Sashiko has reviewed this patch and found no issues. It looks great!

-- 
Sashiko AI review · https://sashiko.dev/#/patchset/20261002164552.294749-1-ays511.kr@gmail.com?part=1


^ permalink raw reply	[flat|nested] 3+ messages in thread

* Re: [PATCH v3] RDMA/rxe: bound the WQE opcode before indexing rxe_wr_opcode_info[]
  2026-10-02 16:45 [PATCH v3] RDMA/rxe: bound the WQE opcode before indexing rxe_wr_opcode_info[] Youngsung Ahn
  2026-10-03  1:33 ` sashiko-bot
@ 2026-10-04  3:40 ` Zhu Yanjun
  1 sibling, 0 replies; 3+ messages in thread
From: Zhu Yanjun @ 2026-10-04  3:40 UTC (permalink / raw)
  To: Youngsung Ahn, Leon Romanovsky, yanjun.zhu@linux.dev; +Cc: linux-rdma


在 2026/10/2 9:45, Youngsung Ahn 写道:
> wr_opcode_mask() indexes the global rxe_wr_opcode_info[] array with
> a caller-supplied opcode and no bounds check. For a user QP the WQE
> comes from an mmap'd SQ ring, so wqe->wr.opcode is an
> attacker-controlled __u32 read back by the requester
> (req_next_wqe() -> rxe_requester()); rxe_post_send() takes the
> qp->is_user branch and never runs validate_send_wr(), whose only
> opcode check is itself !wr_opcode_mask() (it indexes before it
> checks). rxe_wr_opcode_info[] has entries only up to IB_WR_REG_MR,
> so an out-of-range opcode reads out of bounds and the result is used
> as a mask and dereferenced; observed as a KASAN global-out-of-bounds
> "Read of size 4" and a wild-pointer oops.
>
> Return a zero mask for an out-of-range opcode, matching how
> validate_send_wr() already treats a zero mask (an invalid WR), so no
> path indexes the array out of bounds.
>
> Fixes: 8700e3e7c485 ("Soft RoCE driver")
> Signed-off-by: Youngsung Ahn <ays511.kr@gmail.com>

I checked this code snippet. The wr_opcode_mask() function is called in 
three places,

and this change does not introduce any problems in those call sites.

The logs also indicate that this change fixes the reported problem.

Reviewed-by: Zhu Yanjun yanjun.zhu@linux.dev

Zhu Yanjun

> Assisted-by: LLM
> ---
> Notes (not part of the commit):
> v3: no code change. Sent as a standalone message rather than as a reply to
> the v2 posting, and dropped the Cc: stable@vger.kernel.org trailer, both
> as asked on-list.
>
> On the question about req_retry(): req_retry() does not reject the WQE
> itself, it only re-posts it. A zero mask makes every "mask & WR_*" test
> there false, so an out-of-range opcode yields iova = 0, the dma cursor
> reset, and neither retry_first_write_send() nor the read fix-up; the WQE
> is left in wqe_state_posted. Once the requester dispatches it, wqe->mask
> stays 0 so WR_LOCAL_OP_MASK is not taken, and next_opcode() matches no
> case for an out-of-range wqe->wr.opcode and returns -EINVAL, so
> rxe_requester() sets wqe->status = IB_WC_LOC_QP_OP_ERR and takes the err
> path, which moves the QP to IB_QPS_ERR and lets the completer post the
> completion.
>
> wr_opcode_mask() is the single choke point all three callers
> (validate_send_wr(), req_next_wqe(), req_retry()) funnel through, and only
> validate_send_wr() rejects the WQE on a zero mask. In req_retry()
> wr_opcode_mask() is called at the top of the loop body, before the
> wqe_state_posted / wqe_state_done checks, so the out-of-bounds read
> happens even for WQEs the loop then skips.
>
> ARRAY_SIZE() is not usable here: rxe_opcode.h only declares "extern
> struct rxe_wr_opcode_info rxe_wr_opcode_info[]". IB_WR_REG_MR (32) is the
> highest index initialised in rxe_opcode.c, and on x86-64 the object is
> 1320 bytes with a 40-byte stride, i.e. 33 elements, so valid indices are
> 0..32.
>
> The bug was reproduced on an unpatched KASAN x86-64 build of 7.3-rc4, with
> the reproducer run as uid 1000 (the splat itself is attributed to the
> rxe_wq kworker, since that is where the WQE is consumed): KASAN
> global-out-of-bounds read of size 4 at rxe_wr_opcode_info, and an Oops or
> GPF for larger opcodes. Present unchanged in mainline 551c722f4080
> (2026-09-29). This patch itself is only compile-tested (KASAN +
> CONFIG_RDMA_RXE). I am posting the reproducer on-list in reply to the v2
> discussion of this patch.
>
>   drivers/infiniband/sw/rxe/rxe_loc.h | 3 +++
>   1 file changed, 3 insertions(+)
>
> diff --git a/drivers/infiniband/sw/rxe/rxe_loc.h b/drivers/infiniband/sw/rxe/rxe_loc.h
> index 64d636bf80fd..dcc0788db90a 100644
> --- a/drivers/infiniband/sw/rxe/rxe_loc.h
> +++ b/drivers/infiniband/sw/rxe/rxe_loc.h
> @@ -184,6 +184,9 @@ void rxe_comp_queue_pkt(struct rxe_qp *qp, struct sk_buff *skb);
>   
>   static inline unsigned int wr_opcode_mask(int opcode, struct rxe_qp *qp)
>   {
> +	if (unlikely(opcode < 0 || opcode > IB_WR_REG_MR))
> +		return 0;
> +
>   	return rxe_wr_opcode_info[opcode].mask[qp->ibqp.qp_type];
>   }
>   

-- 
Best Regards,
Yanjun.Zhu


^ permalink raw reply	[flat|nested] 3+ messages in thread

end of thread, other threads:[~2026-10-04  3:40 UTC | newest]

Thread overview: 3+ messages (download: mbox.gz follow: Atom feed
-- links below jump to the message on this page --
2026-10-02 16:45 [PATCH v3] RDMA/rxe: bound the WQE opcode before indexing rxe_wr_opcode_info[] Youngsung Ahn
2026-10-03  1:33 ` sashiko-bot
2026-10-04  3:40 ` Zhu Yanjun

This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox