linux-rdma.vger.kernel.org archive mirror
 help / color / mirror / Atom feed
From: Allison Henderson <achender@kernel.org>
To: netdev@vger.kernel.org, linux-rdma@vger.kernel.org,
	pabeni@redhat.com, edumazet@google.com, kuba@kernel.org,
	horms@kernel.org
Cc: achender@kernel.org
Subject: [PATCH net-next v7 11/12] net/rds: pin the connection across RDMA-CM event handling
Date: Sat, 26 Sep 2026 23:14:47 -0700	[thread overview]
Message-ID: <20260927061448.167862-12-achender@kernel.org> (raw)
In-Reply-To: <20260927061448.167862-1-achender@kernel.org>

rds_rdma_cm_event_handler_cmn() picks the connection up from
cm_id->context, which carries no reference, and holds c_cm_lock - a
mutex that lives in the connection's path array - across the transport
callbacks.  Before this series that was already a use-after-free
whenever a callback destroyed the connection, since rds_conn_destroy()
freed it synchronously and the handler's mutex_unlock() ran on freed
memory; the one such callback, rds_ib_cm_connect_complete() on a
protocol version below 3.1, has meanwhile been switched to
rds_conn_drop() by commit f97d8c7bab78 ("rds: ib: use rds_conn_drop()
on protocol version mismatch"), which also removed the deadlock that
destroy took on c_cm_lock.

Now that a connection is freed by its last reference, none of the
callbacks the handler dispatches drops a reference on the connection
it was handed: the version-mismatch path only drops the connection,
and rds_ib_cm_handle_connect() puts the reference rds_conn_create()
gave it, on a connection the listener's cm_id never pointed at.  What
can reach zero while an event is in flight are the holders outside
the handler - the destroy's initial reference, a socket's cache, a
parent's c_passive, an inc.  Today the shutdown pass destroys the
cm_id, and rdma_destroy_id() waits for a running handler, before the
initial reference is dropped, so every event is ordered ahead of the
free; the pin is defensive, keeping the handler correct without
leaning on that ordering.  Take a reference for the duration of the
handler, and ignore the event if the connection is already at zero
references rather than handle it.

rds_ib_cm_initiate_connect() has the same hole on the active side: a
connection whose destroy began while its address and route were
resolving reaches RDMA_CM_EVENT_ROUTE_RESOLVED and sets up a QP -
taking a device reference in rds_ib_add_conn() - after the destroy's
shutdown pass has run, or with that pass waiting on c_cm_lock behind
the handler.  Nothing would ever release the QP, the cm_id or the
device reference, and rds_ib_exit() would wait forever for the
device.  Return before the QP is set up when the destroy is pending;
the cm_id is still ic->i_cm_id, so the shutdown destroys it.

rds_ib_cm_handle_connect() has the mirror-image hole: a connection
whose destroy has already quiesced it sits in RDS_CONN_DOWN with no
cm_id, which is exactly the state the DOWN -> CONNECTING transition
claims.  A connect request arriving then would install a new cm_id and
QP on a connection that is only waiting for its last reference to go
away, and nothing would tear them down again.  Re-check
rds_destroy_pending() under c_cm_lock and reject the request instead.

Assisted-by: Claude-Code:claude-fable-5
Signed-off-by: Allison Henderson <achender@kernel.org>
---
 net/rds/ib_cm.c          | 22 ++++++++++++++++++++--
 net/rds/rdma_transport.c | 19 ++++++++++++++++++-
 2 files changed, 38 insertions(+), 3 deletions(-)

diff --git a/net/rds/ib_cm.c b/net/rds/ib_cm.c
index ee018dd230e9..948fbf4b6a85 100644
--- a/net/rds/ib_cm.c
+++ b/net/rds/ib_cm.c
@@ -874,6 +874,13 @@ int rds_ib_cm_handle_connect(struct rdma_cm_id *cm_id,
 	 * see the comment above rds_queue_reconnect()
 	 */
 	mutex_lock(&conn->c_cm_lock);
+	/* A destroy that has already quiesced this conn leaves it in
+	 * RDS_CONN_DOWN with no cm_id, exactly what the transition
+	 * below would happily claim; nothing would tear the new cm_id
+	 * and QP down again before the conn is freed.  Reject instead.
+	 */
+	if (rds_destroy_pending(conn))
+		goto out;
 	if (!rds_conn_transition(conn, RDS_CONN_DOWN, RDS_CONN_CONNECTING)) {
 		if (rds_conn_state(conn) == RDS_CONN_UP) {
 			rdsdebug("incoming connect while connecting\n");
@@ -928,8 +935,8 @@ int rds_ib_cm_handle_connect(struct rdma_cm_id *cm_id,
 		mutex_unlock(&conn->c_cm_lock);
 		/* Drop the reference rds_conn_create() handed us.  The
 		 * conn stays reachable through cm_id->context without a
-		 * reference of its own for now; the CM event handler is
-		 * given one of its own by a following patch.
+		 * reference of its own; rds_rdma_cm_event_handler_cmn()
+		 * takes one for the duration of each event it handles.
 		 */
 		rds_conn_put(conn);
 	}
@@ -948,6 +955,17 @@ int rds_ib_cm_initiate_connect(struct rdma_cm_id *cm_id, bool isv6)
 	union rds_ib_conn_priv dp;
 	int ret;
 
+	/* A destroy that began while the address and route were being
+	 * resolved has already quiesced this conn, or is waiting on
+	 * c_cm_lock to do so.  Setting up a QP now would leave it - and
+	 * the device reference rds_ib_add_conn() takes - with no
+	 * shutdown pass left to tear them down.  The id we were handed
+	 * is still ic->i_cm_id, so return success and let that shutdown
+	 * destroy it, rather than have the rdma_cm destroy it on error.
+	 */
+	if (rds_destroy_pending(conn))
+		return 0;
+
 	/* If the peer doesn't do protocol negotiation, we must
 	 * default to RDSv3.0 */
 	rds_ib_set_protocol(conn, RDS_PROTOCOL_4_1);
diff --git a/net/rds/rdma_transport.c b/net/rds/rdma_transport.c
index b15cf316b23a..3dda7cf76ebb 100644
--- a/net/rds/rdma_transport.c
+++ b/net/rds/rdma_transport.c
@@ -63,6 +63,21 @@ static int rds_rdma_cm_event_handler_cmn(struct rdma_cm_id *cm_id,
 	if (cm_id->device->node_type == RDMA_NODE_IB_CA)
 		trans = &rds_ib_transport;
 
+	/* cm_id->context carries no reference of its own.  Pin the
+	 * connection for the duration of the handler, since the mutex
+	 * released at out: lives in the connection's path array.  None
+	 * of the callbacks below drops a reference on this connection,
+	 * and the shutdown destroys the cm_id - waiting for a running
+	 * handler - before the last reference can go, so this is
+	 * defensive.  A connection already at zero references gets no
+	 * events handled.
+	 */
+	if (conn && !rds_conn_get_unless_zero(conn)) {
+		rdsdebug("conn %p id %p is being freed, ignoring event\n",
+			 conn, cm_id);
+		return 0;
+	}
+
 	/* Prevent shutdown from tearing down the connection
 	 * while we're executing. */
 	if (conn) {
@@ -171,8 +186,10 @@ static int rds_rdma_cm_event_handler_cmn(struct rdma_cm_id *cm_id,
 	}
 
 out:
-	if (conn)
+	if (conn) {
 		mutex_unlock(&conn->c_cm_lock);
+		rds_conn_put(conn);
+	}
 
 	rdsdebug("id %p event %u (%s) handling ret %d\n", cm_id, event->event,
 		 rdma_event_msg(event->event), ret);
-- 
2.25.1


  parent reply	other threads:[~2026-09-27  6:14 UTC|newest]

Thread overview: 37+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-09-27  6:14 [PATCH net-next v7 00/12] net/rds: make connection lifetime reference-counted Allison Henderson
2026-09-27  6:14 ` [PATCH net-next v7 01/12] net/rds: ib: don't enable interrupts in rds_ib_conn_free() Allison Henderson
2026-09-28  6:14   ` sashiko-bot
2026-09-27  6:14 ` [PATCH net-next v7 02/12] net/rds: undo conn_alloc() the same way on every __rds_conn_create() exit Allison Henderson
2026-09-28  6:15   ` sashiko-bot
2026-09-27  6:14 ` [PATCH net-next v7 03/12] net/rds: guard every work-requeueing site with rds_destroy_pending() Allison Henderson
2026-09-28  6:15   ` sashiko-bot
2026-10-01  6:16   ` netdev-bot+sashiko
2026-09-27  6:14 ` [PATCH net-next v7 04/12] net/rds: make rds_destroy_pending() report a connection's own destroy Allison Henderson
2026-09-28  6:15   ` sashiko-bot
2026-10-01  6:16   ` netdev-bot+sashiko
2026-09-27  6:14 ` [PATCH net-next v7 05/12] net/rds: split connection destroy into quiesce and kref-governed free Allison Henderson
2026-09-28  6:15   ` sashiko-bot
2026-10-01  6:16   ` netdev-bot+sashiko
2026-09-27  6:14 ` [PATCH net-next v7 06/12] net/rds: wait for connections to be freed on transport unload Allison Henderson
2026-09-28  6:15   ` sashiko-bot
2026-10-01  6:16   ` netdev-bot+sashiko
2026-09-27  6:14 ` [PATCH net-next v7 07/12] net/rds: unlink transport nodes before a possibly deferred connection free Allison Henderson
2026-09-28  6:15   ` sashiko-bot
2026-10-01  6:16   ` netdev-bot+sashiko
2026-09-27  6:14 ` [PATCH net-next v7 08/12] net/rds: hold a connection reference from struct rds_incoming Allison Henderson
2026-09-28  6:15   ` sashiko-bot
2026-10-01  6:16   ` netdev-bot+sashiko
2026-09-27  6:14 ` [PATCH net-next v7 09/12] net/rds: take cp_lock to purge cp_send_queue in the quiesce Allison Henderson
2026-09-28  6:15   ` sashiko-bot
2026-10-01  6:16   ` netdev-bot+sashiko
2026-09-27  6:14 ` [PATCH net-next v7 10/12] net/rds: hold connection references in lookup, sockets and c_passive Allison Henderson
2026-09-28  6:15   ` sashiko-bot
2026-10-01  6:16   ` netdev-bot+sashiko
2026-09-27  6:14 ` Allison Henderson [this message]
2026-09-28  6:15   ` [PATCH net-next v7 11/12] net/rds: pin the connection across RDMA-CM event handling sashiko-bot
2026-10-01  6:16   ` netdev-bot+sashiko
2026-09-27  6:14 ` [PATCH net-next v7 12/12] net/rds: drop rds_conn_count in favor of t_conn_count Allison Henderson
2026-09-28  6:15   ` sashiko-bot
2026-10-01  6:16   ` netdev-bot+sashiko
2026-10-02 19:37 ` [PATCH net-next v7 00/12] net/rds: make connection lifetime reference-counted Jakub Kicinski
2026-10-02 21:26   ` Allison Henderson

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=20260927061448.167862-12-achender@kernel.org \
    --to=achender@kernel.org \
    --cc=edumazet@google.com \
    --cc=horms@kernel.org \
    --cc=kuba@kernel.org \
    --cc=linux-rdma@vger.kernel.org \
    --cc=netdev@vger.kernel.org \
    --cc=pabeni@redhat.com \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox;
as well as URLs for NNTP newsgroup(s).