CEPH filesystem development
 help / color / mirror / Atom feed
From: Eric Dumazet <edumazet@kernel.org>
To: "David S . Miller" <davem@davemloft.net>,
	Jakub Kicinski <kuba@kernel.org>, Paolo Abeni <pabeni@redhat.com>
Cc: "Simon Horman" <horms@kernel.org>,
	"Neal Cardwell" <ncardwell@google.com>,
	"Kuniyuki Iwashima" <kuniyu@google.com>,
	edumazet@google.com, netdev@vger.kernel.org,
	"Alexander Aring" <aahringo@redhat.com>,
	"David Teigland" <teigland@redhat.com>,
	gfs2@lists.linux.dev, "John Fastabend" <john.fastabend@gmail.com>,
	"Jakub Sitnicki" <jakub@cloudflare.com>,
	"Sabrina Dubroca" <sd@queasysnail.net>,
	"Jiayuan Chen" <jiayuan.chen@linux.dev>,
	"Matthieu Baerts" <matttbe@kernel.org>,
	"Mat Martineau" <martineau@kernel.org>,
	"Geliang Tang" <geliang@kernel.org>,
	mptcp@lists.linux.dev, "Wen Gu" <guwen@linux.alibaba.com>,
	"Dust Li" <dust.li@linux.alibaba.com>,
	"D. Wythe" <alibuda@linux.alibaba.com>,
	"Chuck Lever" <cel@kernel.org>,
	"Jeff Layton" <jlayton@kernel.org>, NeilBrown <neil@brown.name>,
	"Olga Kornievskaia" <okorniev@redhat.com>,
	"Dai Ngo" <Dai.Ngo@oracle.com>, "Tom Talpey" <tom@talpey.com>,
	"Trond Myklebust" <trondmy@kernel.org>,
	"Anna Schumaker" <anna@kernel.org>,
	linux-nfs@vger.kernel.org,
	"Allison Henderson" <achender@kernel.org>,
	rds-devel@oss.oracle.com,
	"Philipp Reisner" <philipp.reisner@linbit.com>,
	"Lars Ellenberg" <lars.ellenberg@linbit.com>,
	"Christoph Böhmwalder" <christoph.boehmwalder@linbit.com>,
	"Jens Axboe" <axboe@kernel.dk>,
	drbd-dev@lists.linux.dev, "Keith Busch" <kbusch@kernel.org>,
	"Christoph Hellwig" <hch@lst.de>,
	"Sagi Grimberg" <sagi@grimberg.me>,
	"Chaitanya Kulkarni" <kch@nvidia.com>,
	linux-nvme@lists.infradead.org,
	"Ilya Dryomov" <idryomov@gmail.com>,
	"Alex Markuze" <amarkuze@redhat.com>,
	"Viacheslav Dubeyko" <slava@dubeyko.com>,
	ceph-devel@vger.kernel.org, "Eric Dumazet" <edumazet@kernel.org>,
	stable@vger.kernel.org
Subject: [PATCH v3 net-next 1/9] dlm: fix send buffer backpressure handling
Date: Tue, 29 Sep 2026 07:17:35 +0000	[thread overview]
Message-ID: <20260929071743.23624-2-edumazet@kernel.org> (raw)
In-Reply-To: <20260929071743.23624-1-edumazet@kernel.org>

lowcomms.c tests and clears SOCKWQ_ASYNC_NOSPACE in con->sock->flags,
but this bit has not been stored there for ten years.

Commit 9cd3e072b0be ("net: rename SOCK_ASYNC_NOSPACE and
SOCK_ASYNC_WAITDATA") mechanically renamed the two dlm users, then
commit ceb5d58b2170 ("net: fix sock_wake_async() rcu protection") moved
the bit from socket->flags to the RCU protected socket_wq->flags, where
it is reachable only through sk_set_bit() and sk_clear_bit(), and is
only maintained for sockets having SOCK_FASYNC set.  dlm uses kernel
sockets, which never have SOCK_FASYNC set, and never sets the bit
itself, so the test in send_to_sock() has been false ever since.

The consequence is that when sock_sendmsg() returns -EAGAIN because the
socket send buffer is full, dlm no longer sets CF_APP_LIMITED, does not
increment sk_write_pending, and does not return DLM_IO_END to wait for
lowcomms_write_space().  It returns DLM_IO_RESCHED instead, and
process_send_sockets() immediately requeues the send work.  A connection
to a peer that is slow to drain thus keeps cycling through
sock_sendmsg() and -EAGAIN, burning CPU, instead of sleeping until TCP
reports that space is available again.

Test SOCK_NOSPACE instead.  This is the bit that lives in socket->flags,
that TCP sets whenever sendmsg() returns -EAGAIN for lack of send buffer
space (tcp_sendmsg_locked() and sk_stream_wait_memory()), and that
lowcomms_write_space() already clears.  This restores the semantics dlm
had before the bit moved.

Also remove the clear_bit() of SOCKWQ_ASYNC_NOSPACE from
lowcomms_write_space(), for the same reason.

SCTP connections are deliberately left as they are.  SCTP does not set
SOCK_NOSPACE, and it never calls sk->sk_write_space(): sctp_wfree() ends
up in sctp_wake_up_waiters(), which calls sctp_write_space() directly.
lowcomms_write_space() is thus never invoked for an SCTP connection, and
the test added here stays false, so send_to_sock() keeps returning
DLM_IO_RESCHED as it does today.  This is the only safe behavior, as
returning DLM_IO_END would wait for a callback that never comes.

Fixes: ceb5d58b2170 ("net: fix sock_wake_async() rcu protection")
Cc: stable@vger.kernel.org
Cc: Alexander Aring <aahringo@redhat.com>
Cc: David Teigland <teigland@redhat.com>
Acked-by: Alexander Aring <aahringo@redhat.com>
Reviewed-by: Kuniyuki Iwashima <kuniyu@google.com>
Signed-off-by: Eric Dumazet <edumazet@kernel.org>
---
 fs/dlm/lowcomms.c | 6 ++----
 1 file changed, 2 insertions(+), 4 deletions(-)

diff --git a/fs/dlm/lowcomms.c b/fs/dlm/lowcomms.c
index 2aff1c7c17de49c41f9fd02a1fae00bcc9e6af5b..abe9ae4c643fae061d2e168c03219a4f4e1bc007 100644
--- a/fs/dlm/lowcomms.c
+++ b/fs/dlm/lowcomms.c
@@ -522,10 +522,8 @@ static void lowcomms_write_space(struct sock *sk)
 	clear_bit(SOCK_NOSPACE, &con->sock->flags);
 
 	spin_lock_bh(&con->writequeue_lock);
-	if (test_and_clear_bit(CF_APP_LIMITED, &con->flags)) {
+	if (test_and_clear_bit(CF_APP_LIMITED, &con->flags))
 		con->sock->sk->sk_write_pending--;
-		clear_bit(SOCKWQ_ASYNC_NOSPACE, &con->sock->flags);
-	}
 
 	lowcomms_queue_swork(con);
 	spin_unlock_bh(&con->writequeue_lock);
@@ -1391,7 +1389,7 @@ static int send_to_sock(struct connection *con)
 	if (ret == -EAGAIN || ret == 0) {
 		lock_sock(con->sock->sk);
 		spin_lock_bh(&con->writequeue_lock);
-		if (test_bit(SOCKWQ_ASYNC_NOSPACE, &con->sock->flags) &&
+		if (test_bit(SOCK_NOSPACE, &con->sock->flags) &&
 		    !test_and_set_bit(CF_APP_LIMITED, &con->flags)) {
 			/* Notify TCP that we're limited by the
 			 * application window size.
-- 
2.56.0.rc1.315.gc6ed9934b7-goog


  reply	other threads:[~2026-09-29  7:17 UTC|newest]

Thread overview: 21+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-09-29  7:17 [PATCH v3 net-next 0/9] tcp: avoid struct socket cache line miss in tcp_check_space() Eric Dumazet
2026-09-29  7:17 ` Eric Dumazet [this message]
2026-10-02  1:19   ` [PATCH v3 net-next 1/9] dlm: fix send buffer backpressure handling netdev-bot+sashiko
2026-10-02  8:33     ` Eric Dumazet
2026-10-02 14:07       ` Alexander Aring
2026-09-29  7:17 ` [PATCH v3 net-next 2/9] net: add sk_set_nospace() and sk_clear_nospace() Eric Dumazet
2026-10-02  1:19   ` netdev-bot+sashiko
2026-09-29  7:17 ` [PATCH v3 net-next 3/9] sunrpc: use " Eric Dumazet
2026-09-29 15:21   ` Chuck Lever
2026-09-29  7:17 ` [PATCH v3 net-next 4/9] rds: use sk_set_nospace() Eric Dumazet
2026-09-30  1:59   ` Allison Henderson
2026-09-29  7:17 ` [PATCH v3 net-next 5/9] dlm: use sk_set_nospace() and sk_clear_nospace() Eric Dumazet
2026-09-29  7:17 ` [PATCH v3 net-next 6/9] drbd: use sk_set_nospace() Eric Dumazet
2026-09-29 13:36   ` Christoph Böhmwalder
2026-09-29  7:17 ` [PATCH v3 net-next 7/9] nvme-tcp: use sk_clear_nospace() Eric Dumazet
2026-09-29  7:17 ` [PATCH v3 net-next 8/9] libceph: " Eric Dumazet
2026-09-29  7:17 ` [PATCH v3 net-next 9/9] tcp: add tp->tcp_nospace Eric Dumazet
2026-10-02  1:19   ` netdev-bot+sashiko
2026-09-29  7:24 ` [PATCH v3 net-next 0/9] tcp: avoid struct socket cache line miss in tcp_check_space() netdev-bot+sinfo
2026-09-29  7:30   ` Eric Dumazet
2026-10-05 23:30 ` patchwork-bot+netdevbpf

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=20260929071743.23624-2-edumazet@kernel.org \
    --to=edumazet@kernel.org \
    --cc=Dai.Ngo@oracle.com \
    --cc=aahringo@redhat.com \
    --cc=achender@kernel.org \
    --cc=alibuda@linux.alibaba.com \
    --cc=amarkuze@redhat.com \
    --cc=anna@kernel.org \
    --cc=axboe@kernel.dk \
    --cc=cel@kernel.org \
    --cc=ceph-devel@vger.kernel.org \
    --cc=christoph.boehmwalder@linbit.com \
    --cc=davem@davemloft.net \
    --cc=drbd-dev@lists.linux.dev \
    --cc=dust.li@linux.alibaba.com \
    --cc=edumazet@google.com \
    --cc=geliang@kernel.org \
    --cc=gfs2@lists.linux.dev \
    --cc=guwen@linux.alibaba.com \
    --cc=hch@lst.de \
    --cc=horms@kernel.org \
    --cc=idryomov@gmail.com \
    --cc=jakub@cloudflare.com \
    --cc=jiayuan.chen@linux.dev \
    --cc=jlayton@kernel.org \
    --cc=john.fastabend@gmail.com \
    --cc=kbusch@kernel.org \
    --cc=kch@nvidia.com \
    --cc=kuba@kernel.org \
    --cc=kuniyu@google.com \
    --cc=lars.ellenberg@linbit.com \
    --cc=linux-nfs@vger.kernel.org \
    --cc=linux-nvme@lists.infradead.org \
    --cc=martineau@kernel.org \
    --cc=matttbe@kernel.org \
    --cc=mptcp@lists.linux.dev \
    --cc=ncardwell@google.com \
    --cc=neil@brown.name \
    --cc=netdev@vger.kernel.org \
    --cc=okorniev@redhat.com \
    --cc=pabeni@redhat.com \
    --cc=philipp.reisner@linbit.com \
    --cc=rds-devel@oss.oracle.com \
    --cc=sagi@grimberg.me \
    --cc=sd@queasysnail.net \
    --cc=slava@dubeyko.com \
    --cc=stable@vger.kernel.org \
    --cc=teigland@redhat.com \
    --cc=tom@talpey.com \
    --cc=trondmy@kernel.org \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox