* [PATCH net v4 1/2] net/rds: keep the connected peer's scope id apart from the bound one
2026-10-08 3:14 [PATCH net v4 0/2] net/rds: scope-aware sendmsg connection cache Allison Henderson
@ 2026-10-08 3:14 ` Allison Henderson
2026-10-09 11:57 ` Simon Horman
2026-10-08 3:14 ` [PATCH net v4 2/2] net/rds: include the scope id in the sendmsg connection cache check Allison Henderson
1 sibling, 1 reply; 5+ messages in thread
From: Allison Henderson @ 2026-10-08 3:14 UTC (permalink / raw)
To: netdev, linux-rdma, pabeni, edumazet, kuba, horms; +Cc: achender
rds_connect() has nowhere to keep the scope of a link-local peer, so
it stores it in rs_bound_scope_id, the scope of the socket's own
bound address, for rds_bind() to check a later link-local bind
against. That field is then also what a send without a destination
uses as the request's scope, what getpeername() and recvmsg() report
as the peer's scope, and what a later bind() overwrites:
rds_add_bound() stores the bound address's scope unconditionally,
which for a non-link-local address is 0.
So after connect(fe80::x%ifA) followed by bind(global), a send() with
no destination asks for fe80::x with scope 0. rds_conn_lookup() keys
on the interface, so that finds or creates a connection with c_dev_if
0, which the TCP transport then tries to connect through
sin6_scope_id 0 and tcp_v6_connect() rejects for a link-local peer:
the connected socket's data is queued on a connection that can never
come up. In the other order, bind(global) then connect(fe80::x%ifB),
the connect leaves a global-bound socket with rs_bound_scope_id ifB,
so sends to other link-local peers are refused as off-link and sends
to global peers inherit a meaningless interface.
Give the connected peer its own rs_conn_scope_id. rds_connect()
records it there and leaves the bound scope alone, rds_bind()'s
connected-socket check compares against it, and the destination-less
send takes its scope from it - falling back to the bound scope for a
non-link-local peer, so that send() and sendto() to the connected
peer compute the same scope, and so the same connection.
getpeername() reports the connected peer with its own scope, and
recvmsg() reports a sender with the connected peer's scope when there
is one and the bound scope otherwise, as before.
Found by inspection while reworking the sendmsg connection cache; not
observed in the field.
Fixes: 1e2b44e78eea ("rds: Enable RDS IPv6 support")
Assisted-by: Claude-Code:claude-fable-5
Signed-off-by: Allison Henderson <achender@kernel.org>
---
net/rds/af_rds.c | 15 +++++++++------
net/rds/bind.c | 4 ++--
net/rds/rds.h | 2 ++
net/rds/recv.c | 3 ++-
net/rds/send.c | 6 +++++-
5 files changed, 20 insertions(+), 10 deletions(-)
diff --git a/net/rds/af_rds.c b/net/rds/af_rds.c
index d5defe9172e3..e84d00bbe9a1 100644
--- a/net/rds/af_rds.c
+++ b/net/rds/af_rds.c
@@ -136,8 +136,7 @@ static int rds_getname(struct socket *sock, struct sockaddr *uaddr,
sin6->sin6_port = rs->rs_conn_port;
sin6->sin6_addr = rs->rs_conn_addr;
sin6->sin6_flowinfo = 0;
- /* scope_id is the same as in the bound address. */
- sin6->sin6_scope_id = rs->rs_bound_scope_id;
+ sin6->sin6_scope_id = rs->rs_conn_scope_id;
uaddr_len = sizeof(*sin6);
}
} else {
@@ -572,6 +571,7 @@ static int rds_connect(struct socket *sock, struct sockaddr_unsized *uaddr,
}
ipv6_addr_set_v4mapped(sin->sin_addr.s_addr, &rs->rs_conn_addr);
rs->rs_conn_port = sin->sin_port;
+ rs->rs_conn_scope_id = 0;
break;
#if IS_ENABLED(CONFIG_IPV6)
@@ -616,11 +616,14 @@ static int rds_connect(struct socket *sock, struct sockaddr_unsized *uaddr,
ret = -EINVAL;
break;
}
- /* Remember the connected address scope ID. It will
- * be checked against the binding local address when
- * the socket is bound.
+ /* Remember the connected address scope ID. It is
+ * checked against the binding local address when
+ * the socket is bound, and gives a send without a
+ * destination its scope.
*/
- rs->rs_bound_scope_id = sin6->sin6_scope_id;
+ rs->rs_conn_scope_id = sin6->sin6_scope_id;
+ } else {
+ rs->rs_conn_scope_id = 0;
}
rs->rs_conn_addr = sin6->sin6_addr;
rs->rs_conn_port = sin6->sin6_port;
diff --git a/net/rds/bind.c b/net/rds/bind.c
index f800d920d969..3ac59cd512a2 100644
--- a/net/rds/bind.c
+++ b/net/rds/bind.c
@@ -233,8 +233,8 @@ int rds_bind(struct socket *sock, struct sockaddr_unsized *uaddr, int addr_len)
* non-link local address (scope_id is 0).
*/
if (!ipv6_addr_any(&rs->rs_conn_addr) && scope_id &&
- rs->rs_bound_scope_id &&
- scope_id != rs->rs_bound_scope_id) {
+ rs->rs_conn_scope_id &&
+ scope_id != rs->rs_conn_scope_id) {
ret = -EINVAL;
goto out;
}
diff --git a/net/rds/rds.h b/net/rds/rds.h
index 2db49573dacd..9b1ffc49c0a6 100644
--- a/net/rds/rds.h
+++ b/net/rds/rds.h
@@ -646,6 +646,8 @@ struct rds_sock {
struct in6_addr rs_conn_addr;
#define rs_conn_addr_v4 rs_conn_addr.s6_addr32[3]
__be16 rs_conn_port;
+ /* scope of rs_conn_addr when it is link-local, 0 otherwise */
+ __u32 rs_conn_scope_id;
struct rds_transport *rs_transport;
/*
diff --git a/net/rds/recv.c b/net/rds/recv.c
index 6204e577a90a..9d81c76320ea 100644
--- a/net/rds/recv.c
+++ b/net/rds/recv.c
@@ -789,7 +789,8 @@ int rds_recvmsg(struct socket *sock, struct msghdr *msg, size_t size,
sin6->sin6_port = inc->i_hdr.h_sport;
sin6->sin6_addr = inc->i_saddr;
sin6->sin6_flowinfo = 0;
- sin6->sin6_scope_id = rs->rs_bound_scope_id;
+ sin6->sin6_scope_id = rs->rs_conn_scope_id ?:
+ rs->rs_bound_scope_id;
msg->msg_namelen = sizeof(*sin6);
}
}
diff --git a/net/rds/send.c b/net/rds/send.c
index 1afa981e5c06..7b525a6f7eac 100644
--- a/net/rds/send.c
+++ b/net/rds/send.c
@@ -1256,7 +1256,11 @@ int rds_sendmsg(struct socket *sock, struct msghdr *msg, size_t payload_len)
lock_sock(sk);
daddr = rs->rs_conn_addr;
dport = rs->rs_conn_port;
- scope_id = rs->rs_bound_scope_id;
+ /* The connected peer's scope for a link-local peer, else
+ * the bound scope, as an explicit send to that peer would
+ * compute below.
+ */
+ scope_id = rs->rs_conn_scope_id ?: rs->rs_bound_scope_id;
release_sock(sk);
}
--
2.25.1
^ permalink raw reply related [flat|nested] 5+ messages in thread* [PATCH net v4 2/2] net/rds: include the scope id in the sendmsg connection cache check
2026-10-08 3:14 [PATCH net v4 0/2] net/rds: scope-aware sendmsg connection cache Allison Henderson
2026-10-08 3:14 ` [PATCH net v4 1/2] net/rds: keep the connected peer's scope id apart from the bound one Allison Henderson
@ 2026-10-08 3:14 ` Allison Henderson
2026-10-09 11:57 ` Simon Horman
1 sibling, 1 reply; 5+ messages in thread
From: Allison Henderson @ 2026-10-08 3:14 UTC (permalink / raw)
To: netdev, linux-rdma, pabeni, edumazet, kuba, horms; +Cc: achender
rds_sendmsg() reuses the connection cached in rs->rs_conn when its
peer address and ToS match the request. The interface index is part
of a connection's identity as well: rds_conn_create_outgoing() passes
the request's scope_id down as dev_if, and rds_conn_lookup() compares
c_dev_if, so sends to the same link-local address through two
interfaces are two different connections. The cache-hit test never
looked at it.
A socket bound to a non-link-local address has rs_bound_scope_id 0,
and the scope check at the top of rds_sendmsg() accepts any non-zero
destination scope for such a socket. So after a send to fe80::x%ifA,
a send to fe80::x%ifB hits the cached ifA connection and the datagram
leaves through ifA, to whichever peer answers to that address there.
Compare c_dev_if with the request's scope_id in the cache test, so
that such a send takes the lookup path and finds, or creates, the ifB
connection - the same key the lookup uses. With the previous patch a
send without a destination carries the connected peer's scope, or the
bound scope for a non-link-local peer, which is what an explicit send
to that peer computes, so the comparison is exact for it too.
Found by inspection while reworking this cache for the connection
lifetime series on net-next; not observed in the field.
Fixes: 1e2b44e78eea ("rds: Enable RDS IPv6 support")
Assisted-by: Claude-Code:claude-fable-5
Signed-off-by: Allison Henderson <achender@kernel.org>
---
net/rds/send.c | 3 ++-
1 file changed, 2 insertions(+), 1 deletion(-)
diff --git a/net/rds/send.c b/net/rds/send.c
index 7b525a6f7eac..f38967f65af1 100644
--- a/net/rds/send.c
+++ b/net/rds/send.c
@@ -1346,7 +1346,8 @@ int rds_sendmsg(struct socket *sock, struct msghdr *msg, size_t payload_len)
/* rds_conn_create has a spinlock that runs with IRQ off.
* Caching the conn in the socket helps a lot. */
if (rs->rs_conn && ipv6_addr_equal(&rs->rs_conn->c_faddr, &daddr) &&
- rs->rs_tos == rs->rs_conn->c_tos) {
+ rs->rs_tos == rs->rs_conn->c_tos &&
+ rs->rs_conn->c_dev_if == scope_id) {
conn = rs->rs_conn;
} else {
conn = rds_conn_create_outgoing(sock_net(sock->sk),
--
2.25.1
^ permalink raw reply related [flat|nested] 5+ messages in thread