* [PATCH v3] RDMA/rxe: bound the WQE opcode before indexing rxe_wr_opcode_info[]
@ 2026-10-02 16:45 Youngsung Ahn
2026-10-03 1:33 ` sashiko-bot
2026-10-04 3:40 ` Zhu Yanjun
0 siblings, 2 replies; 3+ messages in thread
From: Youngsung Ahn @ 2026-10-02 16:45 UTC (permalink / raw)
To: Zhu Yanjun, Leon Romanovsky; +Cc: linux-rdma
wr_opcode_mask() indexes the global rxe_wr_opcode_info[] array with
a caller-supplied opcode and no bounds check. For a user QP the WQE
comes from an mmap'd SQ ring, so wqe->wr.opcode is an
attacker-controlled __u32 read back by the requester
(req_next_wqe() -> rxe_requester()); rxe_post_send() takes the
qp->is_user branch and never runs validate_send_wr(), whose only
opcode check is itself !wr_opcode_mask() (it indexes before it
checks). rxe_wr_opcode_info[] has entries only up to IB_WR_REG_MR,
so an out-of-range opcode reads out of bounds and the result is used
as a mask and dereferenced; observed as a KASAN global-out-of-bounds
"Read of size 4" and a wild-pointer oops.
Return a zero mask for an out-of-range opcode, matching how
validate_send_wr() already treats a zero mask (an invalid WR), so no
path indexes the array out of bounds.
Fixes: 8700e3e7c485 ("Soft RoCE driver")
Signed-off-by: Youngsung Ahn <ays511.kr@gmail.com>
Assisted-by: LLM
---
Notes (not part of the commit):
v3: no code change. Sent as a standalone message rather than as a reply to
the v2 posting, and dropped the Cc: stable@vger.kernel.org trailer, both
as asked on-list.
On the question about req_retry(): req_retry() does not reject the WQE
itself, it only re-posts it. A zero mask makes every "mask & WR_*" test
there false, so an out-of-range opcode yields iova = 0, the dma cursor
reset, and neither retry_first_write_send() nor the read fix-up; the WQE
is left in wqe_state_posted. Once the requester dispatches it, wqe->mask
stays 0 so WR_LOCAL_OP_MASK is not taken, and next_opcode() matches no
case for an out-of-range wqe->wr.opcode and returns -EINVAL, so
rxe_requester() sets wqe->status = IB_WC_LOC_QP_OP_ERR and takes the err
path, which moves the QP to IB_QPS_ERR and lets the completer post the
completion.
wr_opcode_mask() is the single choke point all three callers
(validate_send_wr(), req_next_wqe(), req_retry()) funnel through, and only
validate_send_wr() rejects the WQE on a zero mask. In req_retry()
wr_opcode_mask() is called at the top of the loop body, before the
wqe_state_posted / wqe_state_done checks, so the out-of-bounds read
happens even for WQEs the loop then skips.
ARRAY_SIZE() is not usable here: rxe_opcode.h only declares "extern
struct rxe_wr_opcode_info rxe_wr_opcode_info[]". IB_WR_REG_MR (32) is the
highest index initialised in rxe_opcode.c, and on x86-64 the object is
1320 bytes with a 40-byte stride, i.e. 33 elements, so valid indices are
0..32.
The bug was reproduced on an unpatched KASAN x86-64 build of 7.3-rc4, with
the reproducer run as uid 1000 (the splat itself is attributed to the
rxe_wq kworker, since that is where the WQE is consumed): KASAN
global-out-of-bounds read of size 4 at rxe_wr_opcode_info, and an Oops or
GPF for larger opcodes. Present unchanged in mainline 551c722f4080
(2026-09-29). This patch itself is only compile-tested (KASAN +
CONFIG_RDMA_RXE). I am posting the reproducer on-list in reply to the v2
discussion of this patch.
drivers/infiniband/sw/rxe/rxe_loc.h | 3 +++
1 file changed, 3 insertions(+)
diff --git a/drivers/infiniband/sw/rxe/rxe_loc.h b/drivers/infiniband/sw/rxe/rxe_loc.h
index 64d636bf80fd..dcc0788db90a 100644
--- a/drivers/infiniband/sw/rxe/rxe_loc.h
+++ b/drivers/infiniband/sw/rxe/rxe_loc.h
@@ -184,6 +184,9 @@ void rxe_comp_queue_pkt(struct rxe_qp *qp, struct sk_buff *skb);
static inline unsigned int wr_opcode_mask(int opcode, struct rxe_qp *qp)
{
+ if (unlikely(opcode < 0 || opcode > IB_WR_REG_MR))
+ return 0;
+
return rxe_wr_opcode_info[opcode].mask[qp->ibqp.qp_type];
}
--
2.43.0
^ permalink raw reply related [flat|nested] 3+ messages in thread* Re: [PATCH v3] RDMA/rxe: bound the WQE opcode before indexing rxe_wr_opcode_info[]
2026-10-02 16:45 [PATCH v3] RDMA/rxe: bound the WQE opcode before indexing rxe_wr_opcode_info[] Youngsung Ahn
@ 2026-10-03 1:33 ` sashiko-bot
2026-10-04 3:40 ` Zhu Yanjun
1 sibling, 0 replies; 3+ messages in thread
From: sashiko-bot @ 2026-10-03 1:33 UTC (permalink / raw)
To: Youngsung Ahn; +Cc: linux-rdma
> wr_opcode_mask() indexes the global rxe_wr_opcode_info[] array with
> a caller-supplied opcode and no bounds check. For a user QP the WQE
> comes from an mmap'd SQ ring, so wqe->wr.opcode is an
> attacker-controlled __u32 read back by the requester
> (req_next_wqe() -> rxe_requester()); rxe_post_send() takes the
> qp->is_user branch and never runs validate_send_wr(), whose only
> opcode check is itself !wr_opcode_mask() (it indexes before it
> checks). rxe_wr_opcode_info[] has entries only up to IB_WR_REG_MR,
> so an out-of-range opcode reads out of bounds and the result is used
> as a mask and dereferenced; observed as a KASAN global-out-of-bounds
> "Read of size 4" and a wild-pointer oops.
>
> Return a zero mask for an out-of-range opcode, matching how
> validate_send_wr() already treats a zero mask (an invalid WR), so no
> path indexes the array out of bounds.
>
> Fixes: 8700e3e7c485 ("Soft RoCE driver")
> Signed-off-by: Youngsung Ahn <ays511.kr@gmail.com>
Sashiko has reviewed this patch and found no issues. It looks great!
--
Sashiko AI review · https://sashiko.dev/#/patchset/20261002164552.294749-1-ays511.kr@gmail.com?part=1
^ permalink raw reply [flat|nested] 3+ messages in thread* Re: [PATCH v3] RDMA/rxe: bound the WQE opcode before indexing rxe_wr_opcode_info[]
2026-10-02 16:45 [PATCH v3] RDMA/rxe: bound the WQE opcode before indexing rxe_wr_opcode_info[] Youngsung Ahn
2026-10-03 1:33 ` sashiko-bot
@ 2026-10-04 3:40 ` Zhu Yanjun
1 sibling, 0 replies; 3+ messages in thread
From: Zhu Yanjun @ 2026-10-04 3:40 UTC (permalink / raw)
To: Youngsung Ahn, Leon Romanovsky, yanjun.zhu@linux.dev; +Cc: linux-rdma
在 2026/10/2 9:45, Youngsung Ahn 写道:
> wr_opcode_mask() indexes the global rxe_wr_opcode_info[] array with
> a caller-supplied opcode and no bounds check. For a user QP the WQE
> comes from an mmap'd SQ ring, so wqe->wr.opcode is an
> attacker-controlled __u32 read back by the requester
> (req_next_wqe() -> rxe_requester()); rxe_post_send() takes the
> qp->is_user branch and never runs validate_send_wr(), whose only
> opcode check is itself !wr_opcode_mask() (it indexes before it
> checks). rxe_wr_opcode_info[] has entries only up to IB_WR_REG_MR,
> so an out-of-range opcode reads out of bounds and the result is used
> as a mask and dereferenced; observed as a KASAN global-out-of-bounds
> "Read of size 4" and a wild-pointer oops.
>
> Return a zero mask for an out-of-range opcode, matching how
> validate_send_wr() already treats a zero mask (an invalid WR), so no
> path indexes the array out of bounds.
>
> Fixes: 8700e3e7c485 ("Soft RoCE driver")
> Signed-off-by: Youngsung Ahn <ays511.kr@gmail.com>
I checked this code snippet. The wr_opcode_mask() function is called in
three places,
and this change does not introduce any problems in those call sites.
The logs also indicate that this change fixes the reported problem.
Reviewed-by: Zhu Yanjun yanjun.zhu@linux.dev
Zhu Yanjun
> Assisted-by: LLM
> ---
> Notes (not part of the commit):
> v3: no code change. Sent as a standalone message rather than as a reply to
> the v2 posting, and dropped the Cc: stable@vger.kernel.org trailer, both
> as asked on-list.
>
> On the question about req_retry(): req_retry() does not reject the WQE
> itself, it only re-posts it. A zero mask makes every "mask & WR_*" test
> there false, so an out-of-range opcode yields iova = 0, the dma cursor
> reset, and neither retry_first_write_send() nor the read fix-up; the WQE
> is left in wqe_state_posted. Once the requester dispatches it, wqe->mask
> stays 0 so WR_LOCAL_OP_MASK is not taken, and next_opcode() matches no
> case for an out-of-range wqe->wr.opcode and returns -EINVAL, so
> rxe_requester() sets wqe->status = IB_WC_LOC_QP_OP_ERR and takes the err
> path, which moves the QP to IB_QPS_ERR and lets the completer post the
> completion.
>
> wr_opcode_mask() is the single choke point all three callers
> (validate_send_wr(), req_next_wqe(), req_retry()) funnel through, and only
> validate_send_wr() rejects the WQE on a zero mask. In req_retry()
> wr_opcode_mask() is called at the top of the loop body, before the
> wqe_state_posted / wqe_state_done checks, so the out-of-bounds read
> happens even for WQEs the loop then skips.
>
> ARRAY_SIZE() is not usable here: rxe_opcode.h only declares "extern
> struct rxe_wr_opcode_info rxe_wr_opcode_info[]". IB_WR_REG_MR (32) is the
> highest index initialised in rxe_opcode.c, and on x86-64 the object is
> 1320 bytes with a 40-byte stride, i.e. 33 elements, so valid indices are
> 0..32.
>
> The bug was reproduced on an unpatched KASAN x86-64 build of 7.3-rc4, with
> the reproducer run as uid 1000 (the splat itself is attributed to the
> rxe_wq kworker, since that is where the WQE is consumed): KASAN
> global-out-of-bounds read of size 4 at rxe_wr_opcode_info, and an Oops or
> GPF for larger opcodes. Present unchanged in mainline 551c722f4080
> (2026-09-29). This patch itself is only compile-tested (KASAN +
> CONFIG_RDMA_RXE). I am posting the reproducer on-list in reply to the v2
> discussion of this patch.
>
> drivers/infiniband/sw/rxe/rxe_loc.h | 3 +++
> 1 file changed, 3 insertions(+)
>
> diff --git a/drivers/infiniband/sw/rxe/rxe_loc.h b/drivers/infiniband/sw/rxe/rxe_loc.h
> index 64d636bf80fd..dcc0788db90a 100644
> --- a/drivers/infiniband/sw/rxe/rxe_loc.h
> +++ b/drivers/infiniband/sw/rxe/rxe_loc.h
> @@ -184,6 +184,9 @@ void rxe_comp_queue_pkt(struct rxe_qp *qp, struct sk_buff *skb);
>
> static inline unsigned int wr_opcode_mask(int opcode, struct rxe_qp *qp)
> {
> + if (unlikely(opcode < 0 || opcode > IB_WR_REG_MR))
> + return 0;
> +
> return rxe_wr_opcode_info[opcode].mask[qp->ibqp.qp_type];
> }
>
--
Best Regards,
Yanjun.Zhu
^ permalink raw reply [flat|nested] 3+ messages in thread
end of thread, other threads:[~2026-10-04 3:40 UTC | newest]
Thread overview: 3+ messages (download: mbox.gz follow: Atom feed
-- links below jump to the message on this page --
2026-10-02 16:45 [PATCH v3] RDMA/rxe: bound the WQE opcode before indexing rxe_wr_opcode_info[] Youngsung Ahn
2026-10-03 1:33 ` sashiko-bot
2026-10-04 3:40 ` Zhu Yanjun
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox