From mboxrd@z Thu Jan 1 00:00:00 1970 Received: from smtp.kernel.org (aws-us-west-2-korg-mail-alma10-1.taild15c8.ts.net [100.103.45.18]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 88A9B4DD3D1; Wed, 30 Sep 2026 15:42:36 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=100.103.45.18 ARC-Seal:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1790782959; cv=none; b=gz+mWROqY9d1WxMcRvt5A37QQPPYj7CwJHwv+YY+9rhkbyf566SaCkrhB4OJxDA1gLejAWHCvVchgQuN5U10T/WqUPfX8GcmIVgTz/+aQdvE//1hnSN8WefnIi6MOl+faUPgL8Kvnievw3mZ3GBLwo+CLn8ZkOTEA5p786H/d4c= ARC-Message-Signature:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1790782959; c=relaxed/simple; bh=8YTUsHkHLAA9TLrBYcLaRVwZk66GA2YyqfC2+cHgoOg=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=iybyeHDcR4aV/WAu2HVAPvGBZlstUmXS6P7oK0dnc3x0EwA0OueB8fLT42bxDZkVK3x+oNtQE9FqDw/laadAr9xWI/vB2agcYWk4NiEu3piTk9Ld31ILxja4wASuyLjMqLBeVg6Jn/uJ9k/Xk/aKOyK1pKdA+b1zb6i/gyY9soc= ARC-Authentication-Results:i=1; smtp.subspace.kernel.org; dkim=pass (1024-bit key) header.d=linuxfoundation.org header.i=@linuxfoundation.org header.b=gjtu5uc4; arc=none smtp.client-ip=100.103.45.18 Authentication-Results: smtp.subspace.kernel.org; dkim=pass (1024-bit key) header.d=linuxfoundation.org header.i=@linuxfoundation.org header.b="gjtu5uc4" Received: by smtp.kernel.org (Postfix) with ESMTPSA id 0961F1F00893; Wed, 30 Sep 2026 15:42:32 +0000 (UTC) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=linuxfoundation.org; s=korg; t=1790782953; bh=2dGTljO03PaIuKbTwQHlAEokPkVvamBtcWw8ZtnqYe0=; h=From:To:Cc:Subject:Date:In-Reply-To:References; b=gjtu5uc4sAGA6p3E4cWxVkbFrg5A2IvHb/KdboNlA1PlH18YI7h3tNQ5+DgtVanRu CHBy3vvCJ85HFa8LrsZVLEG107/HIDvCmuuMQ3R/Sya3Hhi6U6itM7nlFgp0anQpa8 KeKGO0Oec2mKAtOPEECotk5DfMnUem5F+eZ7w000= From: Greg Kroah-Hartman To: stable@vger.kernel.org Cc: Greg Kroah-Hartman , patches@lists.linux.dev, Allison Henderson , Jakub Kicinski , Sasha Levin Subject: [PATCH 5.10 228/595] net/rds: use wq_has_sleeper() in release_in_xmit() Date: Wed, 30 Sep 2026 17:22:01 +0200 Message-ID: <20260930152352.627992321@linuxfoundation.org> X-Mailer: git-send-email 2.55.0 In-Reply-To: <20260930152347.700140858@linuxfoundation.org> References: <20260930152347.700140858@linuxfoundation.org> User-Agent: quilt/0.69 X-stable: review X-Patchwork-Hint: ignore Precedence: bulk X-Mailing-List: patches@lists.linux.dev List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: 8bit 5.10-stable review patch. If anyone has any objections, please let me know. ------------------ From: Allison Henderson [ Upstream commit 6d0c8b7073913011459cf968cbbadd341e166bc3 ] release_in_xmit() clears RDS_IN_XMIT with clear_bit_unlock() and then checks waitqueue_active() to decide whether anyone needs waking. clear_bit_unlock() is only a release operation: it orders the critical section before the bit clear, but does not order the subsequent plain load of the wait queue head after it. The waiter side does the mirror image - it adds itself to the wait queue and then tests the bit. That is the classic store-buffering pattern: the releasing CPU can read the wait queue as empty while the waiting CPU still reads the bit as set, so the sleeper is never woken. The waiters are rds_conn_shutdown() and rds_tcp_reset_callbacks(), both in uninterruptible wait_event() with no timeout. A lost wake-up strands the shutdown worker on its single-threaded workqueue until some other sender releases the bit again - and on a connection that is being torn down precisely because it failed, there may never be another sender. The barrier used to be there: release_in_xmit() did clear_bit() followed by smp_mb__after_atomic() until commit 1422f28826d2 ("rds: introduce acquire/release ordering in acquire/release_in_xmit()") folded both into clear_bit_unlock(), which strengthened the lock hand-off but silently dropped the full barrier the wake-up check depends on. The refill counterpart, release_refill() in net/rds/ib_recv.c, still carries its smp_mb__after_atomic() for exactly this reason. Use wq_has_sleeper(), which is waitqueue_active() preceded by the required full barrier. Fixes: 1422f28826d2 ("rds: introduce acquire/release ordering in acquire/release_in_xmit()") Signed-off-by: Allison Henderson Link: https://patch.msgid.link/20260828223921.202913-2-achender@kernel.org Signed-off-by: Jakub Kicinski Signed-off-by: Sasha Levin --- net/rds/send.c | 7 ++++++- 1 file changed, 6 insertions(+), 1 deletion(-) diff --git a/net/rds/send.c b/net/rds/send.c index 6d49fd32a78c9..910013b844d89 100644 --- a/net/rds/send.c +++ b/net/rds/send.c @@ -114,8 +114,13 @@ static void release_in_xmit(struct rds_conn_path *cp) * hot path and finding waiters is very rare. We don't want to walk * the system-wide hashed waitqueue buckets in the fast path only to * almost never find waiters. + * + * wq_has_sleeper() supplies the full barrier that orders the wait + * queue read after the bit clear; clear_bit_unlock() alone is only + * a release and would let this check read a stale empty queue, + * losing the wake-up. */ - if (waitqueue_active(&cp->cp_waitq)) + if (wq_has_sleeper(&cp->cp_waitq)) wake_up_all(&cp->cp_waitq); } -- 2.53.0