Linux RDMA and InfiniBand development
 help / color / mirror / Atom feed
* [PATCH net v3] net/rds: ib: zero the unwritten tail of a short receive fragment
@ 2026-10-08 12:49 Shubham Antil
  2026-10-08 12:53 ` netdev-bot+sinfo
  2026-10-09 12:49 ` sashiko-bot
  0 siblings, 2 replies; 3+ messages in thread
From: Shubham Antil @ 2026-10-08 12:49 UTC (permalink / raw)
  To: netdev, linux-rdma
  Cc: Allison Henderson, David S . Miller, Eric Dumazet, Jakub Kicinski,
	Paolo Abeni, Simon Horman, linux-kernel, Giovanni Vignone,
	sungbyeongchan

rds_ib_process_recv() accepts an RDS/IB fragment once the receive
completion reports at least an RDS header, and then trusts the
header-declared message length h_len: rds_ib_inc_copy_to_user() later
copies up to h_len bytes from the fragment pages to userspace on
recvmsg().

When a fragment carries fewer payload bytes than h_len accounts for, the
bytes between the end of the received payload and h_len are read from the
unwritten tail of the fragment page.  That page comes from the per-CPU
receive cache and is not zeroed, so recvmsg() returns stale page contents
to userspace.

Zero the unwritten tail of a short fragment page on receipt, so the copy
can only ever return received payload or zeros.  Fragments that fill the
page are unaffected.

Confirmed under KMSAN: without this change recvmsg() returns stale
frag-page bytes and KMSAN flags an uninitialised copy in
rds_ib_inc_copy_to_user(); with it the same over-read returns zeros and
KMSAN is clean.

Fixes: 1e23b3ee0e94 ("RDS/IB: Receive datagrams via IB")
Reported-by: sungbyeongchan <tjdqudcks0424@naver.com>
Closes: https://lore.kernel.org/netdev/20261006205204.1322102-1-tjdqudcks0424@naver.com/
Reported-by: Shubham Antil <shubham@octane.security>
Reported-by: Giovanni Vignone <gio@octane.security>
Reported-by: Robert van Eijk <robert@octane.security>
Reported-by: Paolo Gentry <paolo@octane.security>
Assisted-by: Claude:claude-opus-4-8
Signed-off-by: Shubham Antil <shubham@octane.security>
---
v3:
 - Zero the unwritten fragment tail instead of dropping the connection.
   A zero-copy send with a small iovec segment can legitimately produce a
   non-final fragment shorter than RDS_FRAG_SIZE, so the v2 length check
   could force a reconnect on well-formed traffic; zeroing the tail avoids
   that while still keeping uninitialised bytes out of recvmsg().
   (Raised by the Sashiko review.)
 - Drop Reviewed-by, since the approach changed.
v2:
 - Add Reported-by (sungbyeongchan, first to report, and the Octane
   Security reporters).
 net/rds/ib_recv.c | 14 ++++++++++++++
 1 file changed, 14 insertions(+)

diff --git a/net/rds/ib_recv.c b/net/rds/ib_recv.c
index bd6cb3ffa..9695c7f24 100644
--- a/net/rds/ib_recv.c
+++ b/net/rds/ib_recv.c
@@ -35,6 +35,7 @@
 #include <linux/slab.h>
 #include <linux/pci.h>
 #include <linux/dma-mapping.h>
+#include <linux/highmem.h>
 #include <rdma/rdma_cm.h>
 
 #include "rds_single_path.h"
@@ -949,6 +950,19 @@ static void rds_ib_process_recv(struct rds_connection *conn,
 		}
 	}
 
+	/* The device wrote only data_len bytes into this recycled fragment
+	 * page.  Zero the unwritten tail so a later rds_ib_inc_copy_to_user(),
+	 * which copies up to the peer-declared h_len, cannot hand stale page
+	 * contents to userspace.
+	 */
+	if (data_len < RDS_FRAG_SIZE) {
+		void *frag = kmap_local_page(sg_page(&recv->r_frag->f_sg));
+
+		memset(frag + recv->r_frag->f_sg.offset + data_len, 0,
+		       RDS_FRAG_SIZE - data_len);
+		kunmap_local(frag);
+	}
+
 	list_add_tail(&recv->r_frag->f_item, &ibinc->ii_frags);
 	recv->r_frag = NULL;
 
-- 
2.43.0


^ permalink raw reply related	[flat|nested] 3+ messages in thread

end of thread, other threads:[~2026-10-09 12:49 UTC | newest]

Thread overview: 3+ messages (download: mbox.gz follow: Atom feed
-- links below jump to the message on this page --
2026-10-08 12:49 [PATCH net v3] net/rds: ib: zero the unwritten tail of a short receive fragment Shubham Antil
2026-10-08 12:53 ` netdev-bot+sinfo
2026-10-09 12:49 ` sashiko-bot

This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox