From mboxrd@z Thu Jan 1 00:00:00 1970 Return-Path: X-Spam-Checker-Version: SpamAssassin 3.4.0 (2014-02-07) on aws-us-west-2-korg-lkml-1.web.codeaurora.org Received: from bombadil.infradead.org (bombadil.infradead.org [198.137.202.133]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.lore.kernel.org (Postfix) with ESMTPS id 3F226CA5FBE for ; Tue, 29 Sep 2026 08:23:35 +0000 (UTC) DKIM-Signature: v=1; a=rsa-sha256; q=dns/txt; c=relaxed/relaxed; d=lists.infradead.org; s=bombadil.20210309; h=Sender:List-Subscribe:List-Help :List-Post:List-Archive:List-Unsubscribe:List-Id:Content-Transfer-Encoding: MIME-Version:References:In-Reply-To:Message-ID:Date:Subject:Cc:To:From: Reply-To:Content-Type:Content-ID:Content-Description:Resent-Date:Resent-From: Resent-Sender:Resent-To:Resent-Cc:Resent-Message-ID:List-Owner; bh=YKI1nQjCRL2uE20uHP6f+tRLz/3aALJMUkPrXOsf6RA=; b=nxnd7EdTy7pO0ZajZxjY9lL0Qu cnYg+56JJcoJGg+y4fZZ3nnVtZpxtD7r5LniFMc+V9XrAuoonNKgf8QlZm16+1nwCX9HfjLPUhoHz GDfKfl91jDlCa6oLcDq5BSRuRNlfHYrIR3usVQpl3bBUXa7zNwnoUvIjvO0pKZF1LcondZ53Kb//B DhWsjA5R+thZ38C0T+vIOBCX/NoDfeDofM0U3BshGq03FJxu5mo5F4KumVyxM4ybvq30Hq+aITck3 BQEUyPvmQDTuvG5Rfx2JyIrqXydyaaXG4elaEFe38rKHB3ssZQIur4FFrvlVBSAKVRHpU/tfsdjrP E5Sc8E1Q==; Received: from localhost ([::1] helo=bombadil.infradead.org) by bombadil.infradead.org with esmtp (Exim 4.99.1 #2 (Red Hat Linux)) id 1xBT7e-00000002jRj-3JPQ; Tue, 29 Sep 2026 08:23:34 +0000 Received: from tor.source.kernel.org ([2600:3c04:e001:324:0:1991:8:25]) by bombadil.infradead.org with esmtps (Exim 4.99.1 #2 (Red Hat Linux)) id 1xBS6C-00000002a1f-12yx for linux-nvme@lists.infradead.org; Tue, 29 Sep 2026 07:18:00 +0000 Received: from smtp.kernel.org (quasi.space.kernel.org [100.103.45.18]) by tor.source.kernel.org (Postfix) with ESMTP id A1BDA60A59; Tue, 29 Sep 2026 07:17:59 +0000 (UTC) Received: by smtp.kernel.org (Postfix) with ESMTPSA id DB0A71F00893; Tue, 29 Sep 2026 07:17:55 +0000 (UTC) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=kernel.org; s=k20260515; t=1790666279; bh=YKI1nQjCRL2uE20uHP6f+tRLz/3aALJMUkPrXOsf6RA=; h=From:To:Cc:Subject:Date:In-Reply-To:References; b=UQr1/UcbrS2Udwr23E9XKOfZTKssIy6dd33xd2uXULdHlDV39fOaEV7oEEITgW0l5 V5HmfeOsvplx9ODkxUmOO4igyTryxb8pMckyIurSKiRT30Gpse4CWjfdJvznU1a2Ig KwGPpI+V/VXMw4iKFihnanfTbLTCzz0KtF1h1X9BQfOL+AC2gHF8hFs4vq1fyddljl Due4q/+KOceOLH/3jIJEf3PGP9xlb/lwWUxCeMdQ4SmMoqNnL/oaWe3d7AA/x0krXm sbgrRbGaxzY9obpAhGayjFdDQdu9StK8iFd7fQ2WuwvISH9w/CbXhesLWvTS3Q9ANc +nJ5IdNtB3GNg== From: Eric Dumazet To: "David S . Miller" , Jakub Kicinski , Paolo Abeni Cc: Simon Horman , Neal Cardwell , Kuniyuki Iwashima , edumazet@google.com, netdev@vger.kernel.org, Alexander Aring , David Teigland , gfs2@lists.linux.dev, John Fastabend , Jakub Sitnicki , Sabrina Dubroca , Jiayuan Chen , Matthieu Baerts , Mat Martineau , Geliang Tang , mptcp@lists.linux.dev, Wen Gu , Dust Li , "D. Wythe" , Chuck Lever , Jeff Layton , NeilBrown , Olga Kornievskaia , Dai Ngo , Tom Talpey , Trond Myklebust , Anna Schumaker , linux-nfs@vger.kernel.org, Allison Henderson , rds-devel@oss.oracle.com, Philipp Reisner , Lars Ellenberg , =?UTF-8?q?Christoph=20B=C3=B6hmwalder?= , Jens Axboe , drbd-dev@lists.linux.dev, Keith Busch , Christoph Hellwig , Sagi Grimberg , Chaitanya Kulkarni , linux-nvme@lists.infradead.org, Ilya Dryomov , Alex Markuze , Viacheslav Dubeyko , ceph-devel@vger.kernel.org, Eric Dumazet , stable@vger.kernel.org Subject: [PATCH v3 net-next 1/9] dlm: fix send buffer backpressure handling Date: Tue, 29 Sep 2026 07:17:35 +0000 Message-ID: <20260929071743.23624-2-edumazet@kernel.org> X-Mailer: git-send-email 2.56.0.rc1.315.gc6ed9934b7-goog In-Reply-To: <20260929071743.23624-1-edumazet@kernel.org> References: <20260929071743.23624-1-edumazet@kernel.org> MIME-Version: 1.0 Content-Transfer-Encoding: 8bit X-Mailman-Approved-At: Tue, 29 Sep 2026 01:23:30 -0700 X-BeenThere: linux-nvme@lists.infradead.org X-Mailman-Version: 2.1.34 Precedence: list List-Id: List-Unsubscribe: , List-Archive: List-Post: List-Help: List-Subscribe: , Sender: "Linux-nvme" Errors-To: linux-nvme-bounces+linux-nvme=archiver.kernel.org@lists.infradead.org lowcomms.c tests and clears SOCKWQ_ASYNC_NOSPACE in con->sock->flags, but this bit has not been stored there for ten years. Commit 9cd3e072b0be ("net: rename SOCK_ASYNC_NOSPACE and SOCK_ASYNC_WAITDATA") mechanically renamed the two dlm users, then commit ceb5d58b2170 ("net: fix sock_wake_async() rcu protection") moved the bit from socket->flags to the RCU protected socket_wq->flags, where it is reachable only through sk_set_bit() and sk_clear_bit(), and is only maintained for sockets having SOCK_FASYNC set. dlm uses kernel sockets, which never have SOCK_FASYNC set, and never sets the bit itself, so the test in send_to_sock() has been false ever since. The consequence is that when sock_sendmsg() returns -EAGAIN because the socket send buffer is full, dlm no longer sets CF_APP_LIMITED, does not increment sk_write_pending, and does not return DLM_IO_END to wait for lowcomms_write_space(). It returns DLM_IO_RESCHED instead, and process_send_sockets() immediately requeues the send work. A connection to a peer that is slow to drain thus keeps cycling through sock_sendmsg() and -EAGAIN, burning CPU, instead of sleeping until TCP reports that space is available again. Test SOCK_NOSPACE instead. This is the bit that lives in socket->flags, that TCP sets whenever sendmsg() returns -EAGAIN for lack of send buffer space (tcp_sendmsg_locked() and sk_stream_wait_memory()), and that lowcomms_write_space() already clears. This restores the semantics dlm had before the bit moved. Also remove the clear_bit() of SOCKWQ_ASYNC_NOSPACE from lowcomms_write_space(), for the same reason. SCTP connections are deliberately left as they are. SCTP does not set SOCK_NOSPACE, and it never calls sk->sk_write_space(): sctp_wfree() ends up in sctp_wake_up_waiters(), which calls sctp_write_space() directly. lowcomms_write_space() is thus never invoked for an SCTP connection, and the test added here stays false, so send_to_sock() keeps returning DLM_IO_RESCHED as it does today. This is the only safe behavior, as returning DLM_IO_END would wait for a callback that never comes. Fixes: ceb5d58b2170 ("net: fix sock_wake_async() rcu protection") Cc: stable@vger.kernel.org Cc: Alexander Aring Cc: David Teigland Acked-by: Alexander Aring Reviewed-by: Kuniyuki Iwashima Signed-off-by: Eric Dumazet --- fs/dlm/lowcomms.c | 6 ++---- 1 file changed, 2 insertions(+), 4 deletions(-) diff --git a/fs/dlm/lowcomms.c b/fs/dlm/lowcomms.c index 2aff1c7c17de49c41f9fd02a1fae00bcc9e6af5b..abe9ae4c643fae061d2e168c03219a4f4e1bc007 100644 --- a/fs/dlm/lowcomms.c +++ b/fs/dlm/lowcomms.c @@ -522,10 +522,8 @@ static void lowcomms_write_space(struct sock *sk) clear_bit(SOCK_NOSPACE, &con->sock->flags); spin_lock_bh(&con->writequeue_lock); - if (test_and_clear_bit(CF_APP_LIMITED, &con->flags)) { + if (test_and_clear_bit(CF_APP_LIMITED, &con->flags)) con->sock->sk->sk_write_pending--; - clear_bit(SOCKWQ_ASYNC_NOSPACE, &con->sock->flags); - } lowcomms_queue_swork(con); spin_unlock_bh(&con->writequeue_lock); @@ -1391,7 +1389,7 @@ static int send_to_sock(struct connection *con) if (ret == -EAGAIN || ret == 0) { lock_sock(con->sock->sk); spin_lock_bh(&con->writequeue_lock); - if (test_bit(SOCKWQ_ASYNC_NOSPACE, &con->sock->flags) && + if (test_bit(SOCK_NOSPACE, &con->sock->flags) && !test_and_set_bit(CF_APP_LIMITED, &con->flags)) { /* Notify TCP that we're limited by the * application window size. -- 2.56.0.rc1.315.gc6ed9934b7-goog