Linux RDMA and InfiniBand development
 help / color / mirror / Atom feed
* [PATCH] rds: fix out-of-bounds conn_path walk in peer gen update
@ 2026-10-08 16:40 Henry Martin
  2026-10-09 16:41 ` sashiko-bot
  0 siblings, 1 reply; 2+ messages in thread
From: Henry Martin @ 2026-10-08 16:40 UTC (permalink / raw)
  To: Allison Henderson, David S . Miller, Eric Dumazet, Jakub Kicinski,
	Paolo Abeni, Simon Horman
  Cc: netdev, linux-rdma, linux-kernel, Henry Martin, stable

A connection on the IB transport (t_mp_capable unset) is allocated
with exactly one rds_conn_path slot, yet the peer's genetic number
update routine is hardcoded to walk RDS_MPATH_WORKERS entries whenever
the peer announces a non-zero gen number that differs from the one we
cached.  Everything past the first entry lands on whatever happens to
follow the single-slot array in the same slab slot group, and the loop
then reads and modifies struct fields there (lock acquired,
cp_next_tx_seq cleared, retransmit list inspected).  A malicious peer
can therefore force consecutive out-of-bounds writes into neighbouring
kmalloc-512 objects (KASAN: "Write of size 4 in
_raw_spin_lock_irqsave").

The allocation and the walk now agree on the same expression
(trans->t_mp_capable ? RDS_MPATH_WORKERS : 1), so the loop can never
outlive the array regardless of what the peer reports.

This vulnerability was discovered by Tencent CodeBuddy Security.

Cc: stable@vger.kernel.org
Fixes: ba3d1f480c7a ("net/rds: size a connection's path set by the transport it ends up with")
Signed-off-by: Henry Martin <bsdhenrymartin@gmail.com>
---
 net/rds/recv.c | 3 ++-
 1 file changed, 2 insertions(+), 1 deletion(-)

--- a/net/rds/recv.c
+++ b/net/rds/recv.c
@@ -144,7 +144,8 @@
 	if (peer_gen_num != 0) {
 		if (conn->c_peer_gen_num != 0 &&
 		    peer_gen_num != conn->c_peer_gen_num) {
-			for (i = 0; i < RDS_MPATH_WORKERS; i++) {
+			for (i = 0; i < (conn->c_trans->t_mp_capable ?
+			     RDS_MPATH_WORKERS : 1); i++) {
 				struct rds_conn_path *cp;

 				cp = &conn->c_path[i];
--
2.43.7

^ permalink raw reply	[flat|nested] 2+ messages in thread

* Re: [PATCH] rds: fix out-of-bounds conn_path walk in peer gen update
  2026-10-08 16:40 [PATCH] rds: fix out-of-bounds conn_path walk in peer gen update Henry Martin
@ 2026-10-09 16:41 ` sashiko-bot
  0 siblings, 0 replies; 2+ messages in thread
From: sashiko-bot @ 2026-10-09 16:41 UTC (permalink / raw)
  To: Henry Martin; +Cc: linux-rdma

Thank you for your contribution! Sashiko AI review found 1 potential issue(s) to consider:
- [High] net/rds: race condition in rds_conn_peer_gen_update causes sequence number reset
--

commit 717052475aa8589026dc0030beaf8862f5697d69
Author: Henry Martin <bsdhenrymartin@gmail.com>

rds: fix out-of-bounds conn_path walk in peer gen update

This commit modifies rds_conn_peer_gen_update() to bound the loop over
c_path entries. It checks whether the transport is multipath capable and
walks either RDS_MPATH_WORKERS entries or just 1, preventing out-of-bounds
memory accesses on IB transports.

> diff --git a/net/rds/recv.c b/net/rds/recv.c
> index 6204e577a90ae..afc810a2e3cf9 100644
> --- a/net/rds/recv.c
> +++ b/net/rds/recv.c
> @@ -144,7 +144,8 @@ static void rds_conn_peer_gen_update(struct rds_connection *conn,
>  	if (peer_gen_num != 0) {
>  		if (conn->c_peer_gen_num != 0 &&
>  		    peer_gen_num != conn->c_peer_gen_num) {
> -			for (i = 0; i < RDS_MPATH_WORKERS; i++) {
> +			for (i = 0; i < (conn->c_trans->t_mp_capable ?
> +			     RDS_MPATH_WORKERS : 1); i++) {
>  				struct rds_conn_path *cp;
>  
>  				cp = &conn->c_path[i];

[Severity: High]
Does this code have a race condition that could cause sequence numbers to be
reset during active transmission?

When a peer reboots and establishes multiple multipath TCP connections
concurrently, handshake pings arrive on these paths simultaneously.

Because rds_conn_peer_gen_update() lacks connection-level serialization
across paths, multiple threads can evaluate the generation number
simultaneously:

	if (conn->c_peer_gen_num != 0 &&
	    peer_gen_num != conn->c_peer_gen_num) {

If Thread A and Thread B both evaluate this condition as true and enter the
block, Thread A might complete the reset loop and the path could resume
active transmission.

If Thread B is delayed or preempted, could it then execute the loop and
reset cp_next_tx_seq and cp_next_rx_seq after Thread A has finished?

This would break sequence monotonicity and fragment reassembly, leading
to packet drops and connection stalling.

Furthermore, is the update to the connection state protected against
this race?

net/rds/recv.c:rds_conn_peer_gen_update() {
    ...
		}
		conn->c_peer_gen_num = peer_gen_num;
	}
}

This unprotected write to the shared connection state doesn't seem to
prevent concurrent threads from observing the old value and incorrectly
assuming they must execute the reset block.

-- 
Sashiko AI review · https://sashiko.dev/#/patchset/20261008164041.961961-1-bsdhenrymartin@gmail.com?part=1

^ permalink raw reply	[flat|nested] 2+ messages in thread

end of thread, other threads:[~2026-10-09 16:41 UTC | newest]

Thread overview: 2+ messages (download: mbox.gz follow: Atom feed
-- links below jump to the message on this page --
2026-10-08 16:40 [PATCH] rds: fix out-of-bounds conn_path walk in peer gen update Henry Martin
2026-10-09 16:41 ` sashiko-bot

This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox