From mboxrd@z Thu Jan 1 00:00:00 1970 Received: from smtp.kernel.org (aws-us-west-2-korg-mail-alma10-1.taild15c8.ts.net [100.103.45.18]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 5B29B3BB124; Sun, 27 Sep 2026 06:14:54 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=100.103.45.18 ARC-Seal:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1790489695; cv=none; b=N5ZKToWfWNKoD1ViBmoEuBgVGHb6VbCDRc11akfhSyEGyVNftt+mo2ivc6fsmISxxp1L+tDreTGvTmlLvWYNn0Xm48PBR2hRrw+Fwg39986IETJKtx0V06XlcHb6UiiWm8gPnIhNXzCqlIH4YqYa6vTRbdB/WlPeakhfO68MyXY= ARC-Message-Signature:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1790489695; c=relaxed/simple; bh=/H+JLQOGtaVSGcSIeIn5R2Npqjts2eyk8UMFB8aXWao=; h=From:To:Cc:Subject:Date:Message-Id:In-Reply-To:References: MIME-Version; b=F7X83pKdITWPCjSBlZb0byPmVJQUf4gaMBShSPZQiOsAS5EqSq/buGwRjLGg6LTl9rtehTgWtKZZgQTBNL+NjDAnfCjsRE3KJ2H7l2+tL3VQrZni46h34D0hvqIKiqlV2mLcjoULd3tOkwQcS2JQzYW9SQiqz1og//JzFKxgVPY= ARC-Authentication-Results:i=1; smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b=SbKp8YTn; arc=none smtp.client-ip=100.103.45.18 Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b="SbKp8YTn" Received: by smtp.kernel.org (Postfix) with ESMTPSA id E232E1F0089A; Sun, 27 Sep 2026 06:14:53 +0000 (UTC) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=kernel.org; s=k20260515; t=1790489694; bh=evSiua2yG/GBu3Gejj20k+Hcvdq1/P1gzRjYZTx5P44=; h=From:To:Cc:Subject:Date:In-Reply-To:References; b=SbKp8YTn12OXMudYEn251W2taVzXSCXwHjuybCa9KjbpDVT9+3NqbW58cBBMzAmH+ GAF4F3z9gijauUivdZ7IM/SVuZ/QxoSQ2yRfb/GjdLgOBxM0iwqZdu7MuJZYNVmgSg X0pnqFRv3zLT+NnF2MbfmLJ9iepN8MoxEzfRm5ijKSLjEtJ6GMZyuiWzqDJT4i48PM g9Bd4++8gJY7Ms5SolaprYs/nbMXc6Dt96gvxkoqstIFTUiyjbmxWoU4OYvQ5K+KGK j6j+Y0nmcQCk4rlJUoW2zHAgoKUrwF+THtidY5ZK1QfBRsj+cz1zU3ulRJfG6IaK6N EJgg7k0gledEg== From: Allison Henderson To: netdev@vger.kernel.org, linux-rdma@vger.kernel.org, pabeni@redhat.com, edumazet@google.com, kuba@kernel.org, horms@kernel.org Cc: achender@kernel.org Subject: [PATCH net-next v7 09/12] net/rds: take cp_lock to purge cp_send_queue in the quiesce Date: Sat, 26 Sep 2026 23:14:45 -0700 Message-Id: <20260927061448.167862-10-achender@kernel.org> X-Mailer: git-send-email 2.25.1 In-Reply-To: <20260927061448.167862-1-achender@kernel.org> References: <20260927061448.167862-1-achender@kernel.org> Precedence: bulk X-Mailing-List: linux-rdma@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: 8bit rds_conn_path_quiesce() empties cp_send_queue by walking it with no lock held, while every path that adds to that queue - rds_send_queue_rm(), rds_send_probe(), the retransmit requeue in rds_send_path_reset() - does so under cp_lock. The unlocked walk was justified by the destroy running with nothing else alive: at every destroy trigger there is - netns teardown and module unload - no socket can still be sending on the connection, since a bound socket pins its transport module and a namespace closes its sockets before its RDS connections are torn down. That argument still holds, but it is an argument about the callers, not a property of the code. Splice the queue away under cp_lock and drop the message references outside it, so that the one walker of the list follows the same lock discipline as its adders. This does not by itself make a sender that is still running safe - nothing here tells such a sender that the queue is closed - and it does not need to, since no such sender exists at any destroy trigger. While at it, make the purge coherent with rds_send_drop_to(), the other path that removes messages from a connection queue. drop_to decides whether it owns the queue's reference by test_and_clear on RDS_MSG_ON_CONN; the purge left that bit set, so a message it had already put could be put a second time by a close() or RDS_CANCEL_SENT_TO racing it. Clear the bit under cp_lock as part of the splice. With that, a message that is still on its socket's send queue is no longer a fatal condition for the purge - the socket side keeps its own reference and retires it on close - so the BUG_ON(!list_empty(&rm->m_sock_item)) goes as well. Assisted-by: Claude-Code:claude-fable-5 Signed-off-by: Allison Henderson --- net/rds/connection.c | 24 +++++++++++++++++++----- 1 file changed, 19 insertions(+), 5 deletions(-) diff --git a/net/rds/connection.c b/net/rds/connection.c index 83e59fcaccea..1d7932cac035 100644 --- a/net/rds/connection.c +++ b/net/rds/connection.c @@ -540,6 +540,8 @@ void rds_conn_shutdown(struct rds_conn_path *cp) static void rds_conn_path_quiesce(struct rds_conn_path *cp) { struct rds_message *rm, *rtmp; + unsigned long flags; + LIST_HEAD(purge); if (!cp->cp_transport_data) return; @@ -551,12 +553,24 @@ static void rds_conn_path_quiesce(struct rds_conn_path *cp) rds_conn_path_drop(cp, true); flush_work(&cp->cp_down_w); - /* tear down queued messages */ - list_for_each_entry_safe(rm, rtmp, - &cp->cp_send_queue, - m_conn_item) { + /* Tear down queued messages. Every path that adds to + * cp_send_queue does so under cp_lock; take it here too. No + * sender can still be running at any destroy trigger, so this is + * lock discipline rather than a race fix: nothing here tells a + * sender that the queue is closed. + */ + spin_lock_irqsave(&cp->cp_lock, flags); + list_splice_init(&cp->cp_send_queue, &purge); + /* Give up the queue's claim on each message while still under + * the lock, so that rds_send_drop_to(), which decides ownership + * of the connection-queue reference by this bit, neither drops + * it a second time nor unlinks the message from our list. + */ + list_for_each_entry(rm, &purge, m_conn_item) + clear_bit(RDS_MSG_ON_CONN, &rm->m_flags); + spin_unlock_irqrestore(&cp->cp_lock, flags); + list_for_each_entry_safe(rm, rtmp, &purge, m_conn_item) { list_del_init(&rm->m_conn_item); - BUG_ON(!list_empty(&rm->m_sock_item)); rds_message_put(rm); } if (cp->cp_xmit_rm) -- 2.25.1