Netdev List
 help / color / mirror / Atom feed
* [PATCH v3 net-next 0/9] tcp: avoid struct socket cache line miss in tcp_check_space()
@ 2026-09-29  7:17 Eric Dumazet
  2026-09-29  7:17 ` [PATCH v3 net-next 1/9] dlm: fix send buffer backpressure handling Eric Dumazet
                   ` (10 more replies)
  0 siblings, 11 replies; 21+ messages in thread
From: Eric Dumazet @ 2026-09-29  7:17 UTC (permalink / raw)
  To: David S . Miller, Jakub Kicinski, Paolo Abeni
  Cc: Simon Horman, Neal Cardwell, Kuniyuki Iwashima, edumazet, netdev,
	Alexander Aring, David Teigland, gfs2, John Fastabend,
	Jakub Sitnicki, Sabrina Dubroca, Jiayuan Chen, Matthieu Baerts,
	Mat Martineau, Geliang Tang, mptcp, Wen Gu, Dust Li, D. Wythe,
	Chuck Lever, Jeff Layton, NeilBrown, Olga Kornievskaia, Dai Ngo,
	Tom Talpey, Trond Myklebust, Anna Schumaker, linux-nfs,
	Allison Henderson, rds-devel, Philipp Reisner, Lars Ellenberg,
	Christoph Böhmwalder, Jens Axboe, drbd-dev, Keith Busch,
	Christoph Hellwig, Sagi Grimberg, Chaitanya Kulkarni, linux-nvme,
	Ilya Dryomov, Alex Markuze, Viacheslav Dubeyko, ceph-devel,
	Eric Dumazet

tcp_check_space() is called on every incoming ACK and every transmitted
packet, and tests SOCK_NOSPACE in sk->sk_socket->flags.  Because struct
socket lives in its own cache line and the TCP fast paths do not touch
it for anything else, that test pulls in an extra cache line that is
cold when the working set of active sockets is large.

This series mirrors SOCK_NOSPACE into a new u8 field, tp->tcp_nospace,
placed right after tp->chrono_type in the tcp_sock_write_tx cache line
group (fitting in an existing 3-byte hole before chrono_start, so no
other field moves and sizeof(struct tcp_sock) is unchanged).  Both the
transmit and ACK fast paths already touch that cache line, making the
fast-path test in tcp_check_space() free of extra cache misses while
keeping __tcp_check_space() authoritative on SOCK_NOSPACE.

To maintain the invariant (SOCK_NOSPACE set => tp->tcp_nospace set) from
a single choke point:

 - Patch 1 fixes a long-standing bug in dlm where SOCKWQ_ASYNC_NOSPACE
   was tested and cleared on con->sock->flags instead of SOCK_NOSPACE.
 - Patch 2 introduces sk_set_nospace() and sk_clear_nospace() and
   converts the core networking setters and clearers.
 - Patches 3-8 convert the remaining in-kernel callers (sunrpc, rds,
   dlm, drbd, nvme-tcp, libceph) so that no open-coded set_bit() or
   clear_bit() of SOCK_NOSPACE remains in the tree.
 - Patch 9 adds tp->tcp_nospace, wires it into sk_set_nospace() and
   sk_clear_nospace(), and switches tcp_check_space() to test it.

v3:
 - Rebase and CC subsystem maintainers (Jakub)
 - Link to v2: https://lore.kernel.org/netdev/20260924134729.2047213-1-edumazet@google.com/

v2:
 - Order set_bit(SOCK_NOSPACE) before tp->tcp_nospace = 1 in
   sk_set_nospace() and tp->tcp_nospace = 0 before
   clear_bit(SOCK_NOSPACE) in sk_clear_nospace() so a concurrent
   lockless tcp_poll() cannot leave SOCK_NOSPACE set with
   tp->tcp_nospace cleared (Sashiko)
 - Move tp->tcp_nospace to the 3-byte hole after tp->chrono_type so no
   field in struct tcp_sock shifts after the AccECN bitfield additions
   (Sashiko)
 - Clarify comment and Patch 2 changelog wording (Sashiko)
 - Link to v1: https://lore.kernel.org/netdev/20260922122721.3568295-1-edumazet@google.com/

Eric Dumazet (9):
  dlm: fix send buffer backpressure handling
  net: add sk_set_nospace() and sk_clear_nospace()
  sunrpc: use sk_set_nospace() and sk_clear_nospace()
  rds: use sk_set_nospace()
  dlm: use sk_set_nospace() and sk_clear_nospace()
  drbd: use sk_set_nospace()
  nvme-tcp: use sk_clear_nospace()
  libceph: use sk_clear_nospace()
  tcp: add tp->tcp_nospace

 .../networking/net_cachelines/tcp_sock.rst    |  1 +
 drivers/block/drbd/drbd_worker.c              |  3 +-
 drivers/nvme/host/tcp.c                       |  2 +-
 drivers/nvme/target/tcp.c                     |  2 +-
 fs/dlm/lowcomms.c                             | 10 ++--
 include/linux/tcp.h                           |  3 ++
 include/net/sock.h                            |  2 +
 include/net/tcp.h                             | 30 +++++++++++-
 net/ceph/messenger.c                          |  2 +-
 net/core/sock.c                               | 46 ++++++++++++++++++-
 net/core/stream.c                             |  6 +--
 net/ipv4/tcp.c                                |  5 +-
 net/ipv4/tcp_bpf.c                            |  2 +-
 net/ipv4/tcp_input.c                          |  8 +++-
 net/kcm/kcmsock.c                             |  4 +-
 net/mptcp/protocol.c                          |  4 +-
 net/mptcp/subflow.c                           |  5 ++
 net/rds/tcp_send.c                            |  5 +-
 net/smc/af_smc.c                              |  2 +-
 net/smc/smc_tx.c                              |  6 +--
 net/sunrpc/svcsock.c                          |  4 +-
 net/sunrpc/xprtsock.c                         |  4 +-
 net/tls/tls_sw.c                              |  2 +-
 23 files changed, 121 insertions(+), 37 deletions(-)

-- 
2.56.0.rc1.315.gc6ed9934b7-goog


^ permalink raw reply	[flat|nested] 21+ messages in thread

* [PATCH v3 net-next 1/9] dlm: fix send buffer backpressure handling
  2026-09-29  7:17 [PATCH v3 net-next 0/9] tcp: avoid struct socket cache line miss in tcp_check_space() Eric Dumazet
@ 2026-09-29  7:17 ` Eric Dumazet
  2026-10-02  1:19   ` netdev-bot+sashiko
  2026-09-29  7:17 ` [PATCH v3 net-next 2/9] net: add sk_set_nospace() and sk_clear_nospace() Eric Dumazet
                   ` (9 subsequent siblings)
  10 siblings, 1 reply; 21+ messages in thread
From: Eric Dumazet @ 2026-09-29  7:17 UTC (permalink / raw)
  To: David S . Miller, Jakub Kicinski, Paolo Abeni
  Cc: Simon Horman, Neal Cardwell, Kuniyuki Iwashima, edumazet, netdev,
	Alexander Aring, David Teigland, gfs2, John Fastabend,
	Jakub Sitnicki, Sabrina Dubroca, Jiayuan Chen, Matthieu Baerts,
	Mat Martineau, Geliang Tang, mptcp, Wen Gu, Dust Li, D. Wythe,
	Chuck Lever, Jeff Layton, NeilBrown, Olga Kornievskaia, Dai Ngo,
	Tom Talpey, Trond Myklebust, Anna Schumaker, linux-nfs,
	Allison Henderson, rds-devel, Philipp Reisner, Lars Ellenberg,
	Christoph Böhmwalder, Jens Axboe, drbd-dev, Keith Busch,
	Christoph Hellwig, Sagi Grimberg, Chaitanya Kulkarni, linux-nvme,
	Ilya Dryomov, Alex Markuze, Viacheslav Dubeyko, ceph-devel,
	Eric Dumazet, stable

lowcomms.c tests and clears SOCKWQ_ASYNC_NOSPACE in con->sock->flags,
but this bit has not been stored there for ten years.

Commit 9cd3e072b0be ("net: rename SOCK_ASYNC_NOSPACE and
SOCK_ASYNC_WAITDATA") mechanically renamed the two dlm users, then
commit ceb5d58b2170 ("net: fix sock_wake_async() rcu protection") moved
the bit from socket->flags to the RCU protected socket_wq->flags, where
it is reachable only through sk_set_bit() and sk_clear_bit(), and is
only maintained for sockets having SOCK_FASYNC set.  dlm uses kernel
sockets, which never have SOCK_FASYNC set, and never sets the bit
itself, so the test in send_to_sock() has been false ever since.

The consequence is that when sock_sendmsg() returns -EAGAIN because the
socket send buffer is full, dlm no longer sets CF_APP_LIMITED, does not
increment sk_write_pending, and does not return DLM_IO_END to wait for
lowcomms_write_space().  It returns DLM_IO_RESCHED instead, and
process_send_sockets() immediately requeues the send work.  A connection
to a peer that is slow to drain thus keeps cycling through
sock_sendmsg() and -EAGAIN, burning CPU, instead of sleeping until TCP
reports that space is available again.

Test SOCK_NOSPACE instead.  This is the bit that lives in socket->flags,
that TCP sets whenever sendmsg() returns -EAGAIN for lack of send buffer
space (tcp_sendmsg_locked() and sk_stream_wait_memory()), and that
lowcomms_write_space() already clears.  This restores the semantics dlm
had before the bit moved.

Also remove the clear_bit() of SOCKWQ_ASYNC_NOSPACE from
lowcomms_write_space(), for the same reason.

SCTP connections are deliberately left as they are.  SCTP does not set
SOCK_NOSPACE, and it never calls sk->sk_write_space(): sctp_wfree() ends
up in sctp_wake_up_waiters(), which calls sctp_write_space() directly.
lowcomms_write_space() is thus never invoked for an SCTP connection, and
the test added here stays false, so send_to_sock() keeps returning
DLM_IO_RESCHED as it does today.  This is the only safe behavior, as
returning DLM_IO_END would wait for a callback that never comes.

Fixes: ceb5d58b2170 ("net: fix sock_wake_async() rcu protection")
Cc: stable@vger.kernel.org
Cc: Alexander Aring <aahringo@redhat.com>
Cc: David Teigland <teigland@redhat.com>
Acked-by: Alexander Aring <aahringo@redhat.com>
Reviewed-by: Kuniyuki Iwashima <kuniyu@google.com>
Signed-off-by: Eric Dumazet <edumazet@kernel.org>
---
 fs/dlm/lowcomms.c | 6 ++----
 1 file changed, 2 insertions(+), 4 deletions(-)

diff --git a/fs/dlm/lowcomms.c b/fs/dlm/lowcomms.c
index 2aff1c7c17de49c41f9fd02a1fae00bcc9e6af5b..abe9ae4c643fae061d2e168c03219a4f4e1bc007 100644
--- a/fs/dlm/lowcomms.c
+++ b/fs/dlm/lowcomms.c
@@ -522,10 +522,8 @@ static void lowcomms_write_space(struct sock *sk)
 	clear_bit(SOCK_NOSPACE, &con->sock->flags);
 
 	spin_lock_bh(&con->writequeue_lock);
-	if (test_and_clear_bit(CF_APP_LIMITED, &con->flags)) {
+	if (test_and_clear_bit(CF_APP_LIMITED, &con->flags))
 		con->sock->sk->sk_write_pending--;
-		clear_bit(SOCKWQ_ASYNC_NOSPACE, &con->sock->flags);
-	}
 
 	lowcomms_queue_swork(con);
 	spin_unlock_bh(&con->writequeue_lock);
@@ -1391,7 +1389,7 @@ static int send_to_sock(struct connection *con)
 	if (ret == -EAGAIN || ret == 0) {
 		lock_sock(con->sock->sk);
 		spin_lock_bh(&con->writequeue_lock);
-		if (test_bit(SOCKWQ_ASYNC_NOSPACE, &con->sock->flags) &&
+		if (test_bit(SOCK_NOSPACE, &con->sock->flags) &&
 		    !test_and_set_bit(CF_APP_LIMITED, &con->flags)) {
 			/* Notify TCP that we're limited by the
 			 * application window size.
-- 
2.56.0.rc1.315.gc6ed9934b7-goog


^ permalink raw reply related	[flat|nested] 21+ messages in thread

* [PATCH v3 net-next 2/9] net: add sk_set_nospace() and sk_clear_nospace()
  2026-09-29  7:17 [PATCH v3 net-next 0/9] tcp: avoid struct socket cache line miss in tcp_check_space() Eric Dumazet
  2026-09-29  7:17 ` [PATCH v3 net-next 1/9] dlm: fix send buffer backpressure handling Eric Dumazet
@ 2026-09-29  7:17 ` Eric Dumazet
  2026-10-02  1:19   ` netdev-bot+sashiko
  2026-09-29  7:17 ` [PATCH v3 net-next 3/9] sunrpc: use " Eric Dumazet
                   ` (8 subsequent siblings)
  10 siblings, 1 reply; 21+ messages in thread
From: Eric Dumazet @ 2026-09-29  7:17 UTC (permalink / raw)
  To: David S . Miller, Jakub Kicinski, Paolo Abeni
  Cc: Simon Horman, Neal Cardwell, Kuniyuki Iwashima, edumazet, netdev,
	Alexander Aring, David Teigland, gfs2, John Fastabend,
	Jakub Sitnicki, Sabrina Dubroca, Jiayuan Chen, Matthieu Baerts,
	Mat Martineau, Geliang Tang, mptcp, Wen Gu, Dust Li, D. Wythe,
	Chuck Lever, Jeff Layton, NeilBrown, Olga Kornievskaia, Dai Ngo,
	Tom Talpey, Trond Myklebust, Anna Schumaker, linux-nfs,
	Allison Henderson, rds-devel, Philipp Reisner, Lars Ellenberg,
	Christoph Böhmwalder, Jens Axboe, drbd-dev, Keith Busch,
	Christoph Hellwig, Sagi Grimberg, Chaitanya Kulkarni, linux-nvme,
	Ilya Dryomov, Alex Markuze, Viacheslav Dubeyko, ceph-devel,
	Eric Dumazet

SOCK_NOSPACE lives in sk->sk_socket->flags and is manipulated from
about thirty places in the tree, all of them open coding the
sk->sk_socket dereference, some with a NULL check, some without,
and some reaching the struct socket by yet another path.

Add sk_set_nospace() and sk_clear_nospace() helpers and convert the
core networking setters and clearers to them; the following patches
convert the remaining subsystem callers so that
"git grep _bit(SOCK_NOSPACE" only reports the two helpers and the
remaining test_bit() sites.

Callers that had no NULL check are all called from user context with a
socket attached, so folding the check into the helpers only makes them
more robust.

No functional change intended.

This is a preparation patch: a following one gives TCP a cheaper
private copy of this bit, and needs a single choke point to keep it
in sync.

Reviewed-by: Kuniyuki Iwashima <kuniyu@google.com>
Signed-off-by: Eric Dumazet <edumazet@kernel.org>
---
 include/net/sock.h   |  2 ++
 net/core/sock.c      | 37 +++++++++++++++++++++++++++++++++++--
 net/core/stream.c    |  6 +++---
 net/ipv4/tcp.c       |  4 ++--
 net/ipv4/tcp_bpf.c   |  2 +-
 net/kcm/kcmsock.c    |  4 ++--
 net/mptcp/protocol.c |  4 ++--
 net/smc/af_smc.c     |  2 +-
 net/smc/smc_tx.c     |  6 +++---
 net/tls/tls_sw.c     |  2 +-
 10 files changed, 52 insertions(+), 17 deletions(-)

diff --git a/include/net/sock.h b/include/net/sock.h
index 60ea55dc18854a9759f5df618cc8c904d2323e95..f4dd2e105171386f1b68f808175457c51aff183a 100644
--- a/include/net/sock.h
+++ b/include/net/sock.h
@@ -1129,6 +1129,8 @@ static inline void sk_forward_alloc_add(struct sock *sk, int val)
 }
 
 void sk_stream_write_space(struct sock *sk);
+void sk_set_nospace(struct sock *sk);
+void sk_clear_nospace(struct sock *sk);
 
 /* OOB backlog add */
 static inline void __sk_add_backlog(struct sock *sk, struct sk_buff *skb)
diff --git a/net/core/sock.c b/net/core/sock.c
index 2948dffcc3e1b49a9a55e30f1380ec165a88859f..11a22aec7e414152aab115e8d11e30067ab3775f 100644
--- a/net/core/sock.c
+++ b/net/core/sock.c
@@ -2957,6 +2957,39 @@ void sock_kzfree_s(struct sock *sk, void *mem, int size)
 }
 EXPORT_SYMBOL(sock_kzfree_s);
 
+/**
+ *	sk_set_nospace - tell the transport a writer is waiting for space
+ *	@sk: socket
+ *
+ *	Must be called before the final check of the available send space,
+ *	so that the transport can not miss the request and forget to call
+ *	sk->sk_write_space() once space is available again.
+ */
+void sk_set_nospace(struct sock *sk)
+{
+	struct socket *sock = sk->sk_socket;
+
+	if (sock)
+		set_bit(SOCK_NOSPACE, &sock->flags);
+}
+EXPORT_SYMBOL(sk_set_nospace);
+
+/**
+ *	sk_clear_nospace - tell the transport no writer is waiting for space
+ *	@sk: socket
+ *
+ *	Called from ->sk_write_space() handlers, once send space has been
+ *	made available to writers.
+ */
+void sk_clear_nospace(struct sock *sk)
+{
+	struct socket *sock = sk->sk_socket;
+
+	if (sock)
+		clear_bit(SOCK_NOSPACE, &sock->flags);
+}
+EXPORT_SYMBOL(sk_clear_nospace);
+
 /* It is almost wait_for_tcp_memory minus release_sock/lock_sock.
    I think, these locks should be removed for datagram sockets.
  */
@@ -2970,7 +3003,7 @@ static long sock_wait_for_wmem(struct sock *sk, long timeo)
 			break;
 		if (signal_pending(current))
 			break;
-		set_bit(SOCK_NOSPACE, &sk->sk_socket->flags);
+		sk_set_nospace(sk);
 		prepare_to_wait(sk_sleep(sk), &wait, TASK_INTERRUPTIBLE);
 		if (refcount_read(&sk->sk_wmem_alloc) < READ_ONCE(sk->sk_sndbuf))
 			break;
@@ -3011,7 +3044,7 @@ struct sk_buff *sock_alloc_send_pskb(struct sock *sk, unsigned long header_len,
 			break;
 
 		sk_set_bit(SOCKWQ_ASYNC_NOSPACE, sk);
-		set_bit(SOCK_NOSPACE, &sk->sk_socket->flags);
+		sk_set_nospace(sk);
 		err = -EAGAIN;
 		if (!timeo)
 			goto failure;
diff --git a/net/core/stream.c b/net/core/stream.c
index 2d748581862d0eda8523f6f371cd2402e1dd6e01..a853b60afdc35a4735c017eb16c748cf2fab1b99 100644
--- a/net/core/stream.c
+++ b/net/core/stream.c
@@ -37,7 +37,7 @@ void sk_stream_write_space(struct sock *sk)
 	struct socket_wq *wq;
 
 	if (__sk_stream_is_writeable(sk, 1) && sock) {
-		clear_bit(SOCK_NOSPACE, &sock->flags);
+		sk_clear_nospace(sk);
 
 		rcu_read_lock();
 		wq = rcu_dereference(sk->sk_wq);
@@ -143,7 +143,7 @@ int sk_stream_wait_memory(struct sock *sk, long *timeo_p)
 		if (sk_stream_memory_free(sk) && !vm_wait)
 			break;
 
-		set_bit(SOCK_NOSPACE, &sk->sk_socket->flags);
+		sk_set_nospace(sk);
 		sk->sk_write_pending++;
 		ret = sk_wait_event(sk, &current_timeo, READ_ONCE(sk->sk_err) ||
 				    (READ_ONCE(sk->sk_shutdown) & SEND_SHUTDOWN) ||
@@ -177,7 +177,7 @@ int sk_stream_wait_memory(struct sock *sk, long *timeo_p)
 	 * When TCP receives ACK packets that make room, tcp_check_space()
 	 * only calls tcp_new_space() if SOCK_NOSPACE is set.
 	 */
-	set_bit(SOCK_NOSPACE, &sk->sk_socket->flags);
+	sk_set_nospace(sk);
 	err = -EAGAIN;
 	goto out;
 do_interrupted:
diff --git a/net/ipv4/tcp.c b/net/ipv4/tcp.c
index 3ac4856852794736c5d49f042ecd08e4246bdd6d..1cde000cfab4704e6756872f6ddec16851ccc55d 100644
--- a/net/ipv4/tcp.c
+++ b/net/ipv4/tcp.c
@@ -607,7 +607,7 @@ __poll_t tcp_poll(struct file *file, struct socket *sock, poll_table *wait)
 				mask |= EPOLLOUT | EPOLLWRNORM;
 			} else {  /* send SIGIO later */
 				sk_set_bit(SOCKWQ_ASYNC_NOSPACE, sk);
-				set_bit(SOCK_NOSPACE, &sk->sk_socket->flags);
+				sk_set_nospace(sk);
 
 				/* Race breaker. If space is freed after
 				 * wspace test but before the flags are set,
@@ -1398,7 +1398,7 @@ int tcp_sendmsg_locked(struct sock *sk, struct msghdr *msg, size_t size)
 		continue;
 
 wait_for_space:
-		set_bit(SOCK_NOSPACE, &sk->sk_socket->flags);
+		sk_set_nospace(sk);
 		tcp_remove_empty_skb(sk);
 		if (copied)
 			tcp_push(sk, flags & ~MSG_MORE, mss_now,
diff --git a/net/ipv4/tcp_bpf.c b/net/ipv4/tcp_bpf.c
index 2e234d155b5e616d496611d1f367c34ed90f3000..64d74d92af52d1ca2fa3d24910a254a5092466f2 100644
--- a/net/ipv4/tcp_bpf.c
+++ b/net/ipv4/tcp_bpf.c
@@ -602,7 +602,7 @@ static int tcp_bpf_sendmsg(struct sock *sk, struct msghdr *msg, size_t size)
 			goto out_err;
 		continue;
 wait_for_sndbuf:
-		set_bit(SOCK_NOSPACE, &sk->sk_socket->flags);
+		sk_set_nospace(sk);
 wait_for_memory:
 		err = sk_stream_wait_memory(sk, &timeo);
 		if (err) {
diff --git a/net/kcm/kcmsock.c b/net/kcm/kcmsock.c
index 71af69d442f211996e514d8d76a4c9d658f7b352..b29c0ed9caf5220ff741584e49632da45626beb7 100644
--- a/net/kcm/kcmsock.c
+++ b/net/kcm/kcmsock.c
@@ -735,7 +735,7 @@ static void kcm_tx_work(struct work_struct *w)
 	/* Primarily for SOCK_SEQPACKET sockets */
 	if (likely(sk->sk_socket) &&
 	    test_bit(SOCK_NOSPACE, &sk->sk_socket->flags)) {
-		clear_bit(SOCK_NOSPACE, &sk->sk_socket->flags);
+		sk_clear_nospace(sk);
 		sk->sk_write_space(sk);
 	}
 
@@ -779,7 +779,7 @@ static int kcm_sendmsg(struct socket *sock, struct msghdr *msg, size_t len)
 	/* Call the sk_stream functions to manage the sndbuf mem. */
 	if (!sk_stream_memory_free(sk)) {
 		kcm_push(kcm);
-		set_bit(SOCK_NOSPACE, &sk->sk_socket->flags);
+		sk_set_nospace(sk);
 		err = sk_stream_wait_memory(sk, &timeo);
 		if (err)
 			goto out_error;
diff --git a/net/mptcp/protocol.c b/net/mptcp/protocol.c
index e89a69ab927c9139c68c9039327cb0e55c33356b..74b1a512072878a54b229bc80aa76bb3e417c862 100644
--- a/net/mptcp/protocol.c
+++ b/net/mptcp/protocol.c
@@ -2106,7 +2106,7 @@ static int mptcp_sendmsg(struct sock *sk, struct msghdr *msg, size_t len)
 		continue;
 
 wait_for_memory:
-		set_bit(SOCK_NOSPACE, &sk->sk_socket->flags);
+		sk_set_nospace(sk);
 		__mptcp_push_pending(sk, msg->msg_flags);
 		ret = sk_stream_wait_memory(sk, &timeo);
 		if (ret)
@@ -4472,7 +4472,7 @@ static __poll_t mptcp_check_writeable(struct mptcp_sock *msk)
 	if (__mptcp_stream_is_writeable(sk, 1))
 		return EPOLLOUT | EPOLLWRNORM;
 
-	set_bit(SOCK_NOSPACE, &sk->sk_socket->flags);
+	sk_set_nospace(sk);
 	smp_mb__after_atomic(); /* NOSPACE is changed by mptcp_write_space() */
 	if (__mptcp_stream_is_writeable(sk, 1))
 		return EPOLLOUT | EPOLLWRNORM;
diff --git a/net/smc/af_smc.c b/net/smc/af_smc.c
index e9f93b3ab435bbcce0a493f24205bd91bdf15b6a..4f2afbbdef2d7e62431e14562dc26bdfcdaa9f79 100644
--- a/net/smc/af_smc.c
+++ b/net/smc/af_smc.c
@@ -2919,7 +2919,7 @@ __poll_t smc_poll(struct file *file, struct socket *sock,
 				mask |= EPOLLOUT | EPOLLWRNORM;
 			} else {
 				sk_set_bit(SOCKWQ_ASYNC_NOSPACE, sk);
-				set_bit(SOCK_NOSPACE, &sk->sk_socket->flags);
+				sk_set_nospace(sk);
 
 				if (sk->sk_state != SMC_INIT) {
 					/* Race breaker the same way as tcp_poll(). */
diff --git a/net/smc/smc_tx.c b/net/smc/smc_tx.c
index 3144b4b1fe29013cabc63d72671d256761a4c367..52e5395484fd3e0c4b1356226ecdcf4fb190d487 100644
--- a/net/smc/smc_tx.c
+++ b/net/smc/smc_tx.c
@@ -48,7 +48,7 @@ static void smc_tx_write_space(struct sock *sk)
 	if (atomic_read(&smc->conn.sndbuf_space) && sock) {
 		if (test_bit(SOCK_NOSPACE, &sock->flags))
 			SMC_STAT_RMB_TX_FULL(smc, !smc->conn.lnk);
-		clear_bit(SOCK_NOSPACE, &sock->flags);
+		sk_clear_nospace(sk);
 		rcu_read_lock();
 		wq = rcu_dereference(sk->sk_wq);
 		if (skwq_has_sleeper(wq))
@@ -100,7 +100,7 @@ static int smc_tx_wait(struct smc_sock *smc, int flags)
 		}
 		if (!timeo) {
 			/* ensure EPOLLOUT is subsequently generated */
-			set_bit(SOCK_NOSPACE, &sk->sk_socket->flags);
+			sk_set_nospace(sk);
 			rc = -EAGAIN;
 			break;
 		}
@@ -111,7 +111,7 @@ static int smc_tx_wait(struct smc_sock *smc, int flags)
 		sk_clear_bit(SOCKWQ_ASYNC_NOSPACE, sk);
 		if (atomic_read(&conn->sndbuf_space) && !conn->urg_tx_pend)
 			break; /* at least 1 byte of free & no urgent data */
-		set_bit(SOCK_NOSPACE, &sk->sk_socket->flags);
+		sk_set_nospace(sk);
 		sk_wait_event(sk, &timeo,
 			      READ_ONCE(sk->sk_err) ||
 			      (READ_ONCE(sk->sk_shutdown) & SEND_SHUTDOWN) ||
diff --git a/net/tls/tls_sw.c b/net/tls/tls_sw.c
index d1ad31986cf2cee88afbde43a2908791eeabb0fb..312e51270f293dc4df8efbec498473c4c3b48d15 100644
--- a/net/tls/tls_sw.c
+++ b/net/tls/tls_sw.c
@@ -966,7 +966,7 @@ static int tls_sw_sendmsg_locked(struct sock *sk, struct msghdr *msg,
 		continue;
 
 wait_for_sndbuf:
-		set_bit(SOCK_NOSPACE, &sk->sk_socket->flags);
+		sk_set_nospace(sk);
 wait_for_memory:
 		ret = sk_stream_wait_memory(sk, &timeo);
 		if (ret) {
-- 
2.56.0.rc1.315.gc6ed9934b7-goog


^ permalink raw reply related	[flat|nested] 21+ messages in thread

* [PATCH v3 net-next 3/9] sunrpc: use sk_set_nospace() and sk_clear_nospace()
  2026-09-29  7:17 [PATCH v3 net-next 0/9] tcp: avoid struct socket cache line miss in tcp_check_space() Eric Dumazet
  2026-09-29  7:17 ` [PATCH v3 net-next 1/9] dlm: fix send buffer backpressure handling Eric Dumazet
  2026-09-29  7:17 ` [PATCH v3 net-next 2/9] net: add sk_set_nospace() and sk_clear_nospace() Eric Dumazet
@ 2026-09-29  7:17 ` Eric Dumazet
  2026-09-29 15:21   ` Chuck Lever
  2026-09-29  7:17 ` [PATCH v3 net-next 4/9] rds: use sk_set_nospace() Eric Dumazet
                   ` (7 subsequent siblings)
  10 siblings, 1 reply; 21+ messages in thread
From: Eric Dumazet @ 2026-09-29  7:17 UTC (permalink / raw)
  To: David S . Miller, Jakub Kicinski, Paolo Abeni
  Cc: Simon Horman, Neal Cardwell, Kuniyuki Iwashima, edumazet, netdev,
	Alexander Aring, David Teigland, gfs2, John Fastabend,
	Jakub Sitnicki, Sabrina Dubroca, Jiayuan Chen, Matthieu Baerts,
	Mat Martineau, Geliang Tang, mptcp, Wen Gu, Dust Li, D. Wythe,
	Chuck Lever, Jeff Layton, NeilBrown, Olga Kornievskaia, Dai Ngo,
	Tom Talpey, Trond Myklebust, Anna Schumaker, linux-nfs,
	Allison Henderson, rds-devel, Philipp Reisner, Lars Ellenberg,
	Christoph Böhmwalder, Jens Axboe, drbd-dev, Keith Busch,
	Christoph Hellwig, Sagi Grimberg, Chaitanya Kulkarni, linux-nvme,
	Ilya Dryomov, Alex Markuze, Viacheslav Dubeyko, ceph-devel,
	Eric Dumazet

Use the new helpers instead of open coding the SOCK_NOSPACE
manipulation, so that TCP can later maintain a cheaper private
copy of this bit.

No functional change intended.

Reviewed-by: Kuniyuki Iwashima <kuniyu@google.com>
Signed-off-by: Eric Dumazet <edumazet@kernel.org>
---
 net/sunrpc/svcsock.c  | 4 ++--
 net/sunrpc/xprtsock.c | 4 ++--
 2 files changed, 4 insertions(+), 4 deletions(-)

diff --git a/net/sunrpc/svcsock.c b/net/sunrpc/svcsock.c
index 50e5e7f5b762de28c34d0f58cb0c6feb1e3ce79f..2d8cdee0798d540ba8eae589deaac5bd3ff889fa 100644
--- a/net/sunrpc/svcsock.c
+++ b/net/sunrpc/svcsock.c
@@ -792,11 +792,11 @@ static int svc_udp_has_wspace(struct svc_xprt *xprt)
 	 * Set the SOCK_NOSPACE flag before checking the available
 	 * sock space.
 	 */
-	set_bit(SOCK_NOSPACE, &svsk->sk_sock->flags);
+	sk_set_nospace(svsk->sk_sk);
 	required = atomic_read(&svsk->sk_xprt.xpt_reserved) + serv->sv_max_mesg;
 	if (required*2 > sock_wspace(svsk->sk_sk))
 		return 0;
-	clear_bit(SOCK_NOSPACE, &svsk->sk_sock->flags);
+	sk_clear_nospace(svsk->sk_sk);
 	return 1;
 }
 
diff --git a/net/sunrpc/xprtsock.c b/net/sunrpc/xprtsock.c
index 7f60723fa64d84887e260e5e130e44cdbec6f850..b97c70c12f627630509510facfd3c3cf96ab6cb6 100644
--- a/net/sunrpc/xprtsock.c
+++ b/net/sunrpc/xprtsock.c
@@ -858,7 +858,7 @@ static int xs_nospace(struct rpc_rqst *req, struct sock_xprt *transport)
 	if (xprt_connected(xprt)) {
 		/* wait for more buffer space */
 		set_bit(XPRT_SOCK_NOSPACE, &transport->sock_state);
-		set_bit(SOCK_NOSPACE, &sk->sk_socket->flags);
+		sk_set_nospace(sk);
 		sk->sk_write_pending++;
 		xprt_wait_for_buffer_space(xprt);
 	} else
@@ -1615,7 +1615,7 @@ static void xs_write_space(struct sock *sk)
 
 	if (!sk->sk_socket)
 		return;
-	clear_bit(SOCK_NOSPACE, &sk->sk_socket->flags);
+	sk_clear_nospace(sk);
 
 	if (unlikely(!(xprt = xprt_from_sock(sk))))
 		return;
-- 
2.56.0.rc1.315.gc6ed9934b7-goog


^ permalink raw reply related	[flat|nested] 21+ messages in thread

* [PATCH v3 net-next 4/9] rds: use sk_set_nospace()
  2026-09-29  7:17 [PATCH v3 net-next 0/9] tcp: avoid struct socket cache line miss in tcp_check_space() Eric Dumazet
                   ` (2 preceding siblings ...)
  2026-09-29  7:17 ` [PATCH v3 net-next 3/9] sunrpc: use " Eric Dumazet
@ 2026-09-29  7:17 ` Eric Dumazet
  2026-09-30  1:59   ` Allison Henderson
  2026-09-29  7:17 ` [PATCH v3 net-next 5/9] dlm: use sk_set_nospace() and sk_clear_nospace() Eric Dumazet
                   ` (6 subsequent siblings)
  10 siblings, 1 reply; 21+ messages in thread
From: Eric Dumazet @ 2026-09-29  7:17 UTC (permalink / raw)
  To: David S . Miller, Jakub Kicinski, Paolo Abeni
  Cc: Simon Horman, Neal Cardwell, Kuniyuki Iwashima, edumazet, netdev,
	Alexander Aring, David Teigland, gfs2, John Fastabend,
	Jakub Sitnicki, Sabrina Dubroca, Jiayuan Chen, Matthieu Baerts,
	Mat Martineau, Geliang Tang, mptcp, Wen Gu, Dust Li, D. Wythe,
	Chuck Lever, Jeff Layton, NeilBrown, Olga Kornievskaia, Dai Ngo,
	Tom Talpey, Trond Myklebust, Anna Schumaker, linux-nfs,
	Allison Henderson, rds-devel, Philipp Reisner, Lars Ellenberg,
	Christoph Böhmwalder, Jens Axboe, drbd-dev, Keith Busch,
	Christoph Hellwig, Sagi Grimberg, Chaitanya Kulkarni, linux-nvme,
	Ilya Dryomov, Alex Markuze, Viacheslav Dubeyko, ceph-devel,
	Eric Dumazet

Use the new helpers instead of open coding the SOCK_NOSPACE
manipulation, so that TCP can later maintain a cheaper private
copy of this bit.

No functional change intended.

Reviewed-by: Kuniyuki Iwashima <kuniyu@google.com>
Signed-off-by: Eric Dumazet <edumazet@kernel.org>
---
 net/rds/tcp_send.c | 5 ++---
 1 file changed, 2 insertions(+), 3 deletions(-)

diff --git a/net/rds/tcp_send.c b/net/rds/tcp_send.c
index 7c52acc749cf4d4003c749a2646c2b4f87f207fe..824222c475119d91e333123c5afc9291005a8696 100644
--- a/net/rds/tcp_send.c
+++ b/net/rds/tcp_send.c
@@ -100,7 +100,7 @@ int rds_tcp_xmit(struct rds_connection *conn, struct rds_message *rm,
 
 	if (hdr_off < sizeof(struct rds_header)) {
 		/* see rds_tcp_write_space() */
-		set_bit(SOCK_NOSPACE, &tc->t_sock->sk->sk_socket->flags);
+		sk_set_nospace(tc->t_sock->sk);
 
 		ret = rds_tcp_sendmsg(tc->t_sock,
 				      (void *)&rm->m_inc.i_hdr + hdr_off,
@@ -221,6 +221,5 @@ void rds_tcp_write_space(struct sock *sk)
 	 */
 	write_space(sk);
 
-	if (sk->sk_socket)
-		set_bit(SOCK_NOSPACE, &sk->sk_socket->flags);
+	sk_set_nospace(sk);
 }
-- 
2.56.0.rc1.315.gc6ed9934b7-goog


^ permalink raw reply related	[flat|nested] 21+ messages in thread

* [PATCH v3 net-next 5/9] dlm: use sk_set_nospace() and sk_clear_nospace()
  2026-09-29  7:17 [PATCH v3 net-next 0/9] tcp: avoid struct socket cache line miss in tcp_check_space() Eric Dumazet
                   ` (3 preceding siblings ...)
  2026-09-29  7:17 ` [PATCH v3 net-next 4/9] rds: use sk_set_nospace() Eric Dumazet
@ 2026-09-29  7:17 ` Eric Dumazet
  2026-09-29  7:17 ` [PATCH v3 net-next 6/9] drbd: use sk_set_nospace() Eric Dumazet
                   ` (5 subsequent siblings)
  10 siblings, 0 replies; 21+ messages in thread
From: Eric Dumazet @ 2026-09-29  7:17 UTC (permalink / raw)
  To: David S . Miller, Jakub Kicinski, Paolo Abeni
  Cc: Simon Horman, Neal Cardwell, Kuniyuki Iwashima, edumazet, netdev,
	Alexander Aring, David Teigland, gfs2, John Fastabend,
	Jakub Sitnicki, Sabrina Dubroca, Jiayuan Chen, Matthieu Baerts,
	Mat Martineau, Geliang Tang, mptcp, Wen Gu, Dust Li, D. Wythe,
	Chuck Lever, Jeff Layton, NeilBrown, Olga Kornievskaia, Dai Ngo,
	Tom Talpey, Trond Myklebust, Anna Schumaker, linux-nfs,
	Allison Henderson, rds-devel, Philipp Reisner, Lars Ellenberg,
	Christoph Böhmwalder, Jens Axboe, drbd-dev, Keith Busch,
	Christoph Hellwig, Sagi Grimberg, Chaitanya Kulkarni, linux-nvme,
	Ilya Dryomov, Alex Markuze, Viacheslav Dubeyko, ceph-devel,
	Eric Dumazet

Use the new helpers instead of open coding the SOCK_NOSPACE
manipulation, so that TCP can later maintain a cheaper private
copy of this bit.

No functional change intended.

Reviewed-by: Kuniyuki Iwashima <kuniyu@google.com>
Signed-off-by: Eric Dumazet <edumazet@kernel.org>
---
 fs/dlm/lowcomms.c | 4 ++--
 1 file changed, 2 insertions(+), 2 deletions(-)

diff --git a/fs/dlm/lowcomms.c b/fs/dlm/lowcomms.c
index abe9ae4c643fae061d2e168c03219a4f4e1bc007..6aa68a24acab574429664acc3f5ecc25ddd1553d 100644
--- a/fs/dlm/lowcomms.c
+++ b/fs/dlm/lowcomms.c
@@ -519,7 +519,7 @@ static void lowcomms_write_space(struct sock *sk)
 {
 	struct connection *con = sock2con(sk);
 
-	clear_bit(SOCK_NOSPACE, &con->sock->flags);
+	sk_clear_nospace(sk);
 
 	spin_lock_bh(&con->writequeue_lock);
 	if (test_and_clear_bit(CF_APP_LIMITED, &con->flags))
@@ -1394,7 +1394,7 @@ static int send_to_sock(struct connection *con)
 			/* Notify TCP that we're limited by the
 			 * application window size.
 			 */
-			set_bit(SOCK_NOSPACE, &con->sock->sk->sk_socket->flags);
+			sk_set_nospace(con->sock->sk);
 			con->sock->sk->sk_write_pending++;
 
 			clear_bit(CF_SEND_PENDING, &con->flags);
-- 
2.56.0.rc1.315.gc6ed9934b7-goog


^ permalink raw reply related	[flat|nested] 21+ messages in thread

* [PATCH v3 net-next 6/9] drbd: use sk_set_nospace()
  2026-09-29  7:17 [PATCH v3 net-next 0/9] tcp: avoid struct socket cache line miss in tcp_check_space() Eric Dumazet
                   ` (4 preceding siblings ...)
  2026-09-29  7:17 ` [PATCH v3 net-next 5/9] dlm: use sk_set_nospace() and sk_clear_nospace() Eric Dumazet
@ 2026-09-29  7:17 ` Eric Dumazet
  2026-09-29 13:36   ` Christoph Böhmwalder
  2026-09-29  7:17 ` [PATCH v3 net-next 7/9] nvme-tcp: use sk_clear_nospace() Eric Dumazet
                   ` (4 subsequent siblings)
  10 siblings, 1 reply; 21+ messages in thread
From: Eric Dumazet @ 2026-09-29  7:17 UTC (permalink / raw)
  To: David S . Miller, Jakub Kicinski, Paolo Abeni
  Cc: Simon Horman, Neal Cardwell, Kuniyuki Iwashima, edumazet, netdev,
	Alexander Aring, David Teigland, gfs2, John Fastabend,
	Jakub Sitnicki, Sabrina Dubroca, Jiayuan Chen, Matthieu Baerts,
	Mat Martineau, Geliang Tang, mptcp, Wen Gu, Dust Li, D. Wythe,
	Chuck Lever, Jeff Layton, NeilBrown, Olga Kornievskaia, Dai Ngo,
	Tom Talpey, Trond Myklebust, Anna Schumaker, linux-nfs,
	Allison Henderson, rds-devel, Philipp Reisner, Lars Ellenberg,
	Christoph Böhmwalder, Jens Axboe, drbd-dev, Keith Busch,
	Christoph Hellwig, Sagi Grimberg, Chaitanya Kulkarni, linux-nvme,
	Ilya Dryomov, Alex Markuze, Viacheslav Dubeyko, ceph-devel,
	Eric Dumazet

Use the new helpers instead of open coding the SOCK_NOSPACE
manipulation, so that TCP can later maintain a cheaper private
copy of this bit.

No functional change intended.

Reviewed-by: Kuniyuki Iwashima <kuniyu@google.com>
Signed-off-by: Eric Dumazet <edumazet@kernel.org>
---
 drivers/block/drbd/drbd_worker.c | 3 +--
 1 file changed, 1 insertion(+), 2 deletions(-)

diff --git a/drivers/block/drbd/drbd_worker.c b/drivers/block/drbd/drbd_worker.c
index 0697f99fed18d7e70f00b7c559a131517ee8aa42..8e50131b5a5692a8dfb99eafe706602153be94b7 100644
--- a/drivers/block/drbd/drbd_worker.c
+++ b/drivers/block/drbd/drbd_worker.c
@@ -635,8 +635,7 @@ static int make_resync_request(struct drbd_peer_device *const peer_device, int c
 			int sndbuf = sk->sk_sndbuf;
 			if (queued > sndbuf / 2) {
 				requeue = 1;
-				if (sk->sk_socket)
-					set_bit(SOCK_NOSPACE, &sk->sk_socket->flags);
+				sk_set_nospace(sk);
 			}
 		} else
 			requeue = 1;
-- 
2.56.0.rc1.315.gc6ed9934b7-goog


^ permalink raw reply related	[flat|nested] 21+ messages in thread

* [PATCH v3 net-next 7/9] nvme-tcp: use sk_clear_nospace()
  2026-09-29  7:17 [PATCH v3 net-next 0/9] tcp: avoid struct socket cache line miss in tcp_check_space() Eric Dumazet
                   ` (5 preceding siblings ...)
  2026-09-29  7:17 ` [PATCH v3 net-next 6/9] drbd: use sk_set_nospace() Eric Dumazet
@ 2026-09-29  7:17 ` Eric Dumazet
  2026-09-29  7:17 ` [PATCH v3 net-next 8/9] libceph: " Eric Dumazet
                   ` (3 subsequent siblings)
  10 siblings, 0 replies; 21+ messages in thread
From: Eric Dumazet @ 2026-09-29  7:17 UTC (permalink / raw)
  To: David S . Miller, Jakub Kicinski, Paolo Abeni
  Cc: Simon Horman, Neal Cardwell, Kuniyuki Iwashima, edumazet, netdev,
	Alexander Aring, David Teigland, gfs2, John Fastabend,
	Jakub Sitnicki, Sabrina Dubroca, Jiayuan Chen, Matthieu Baerts,
	Mat Martineau, Geliang Tang, mptcp, Wen Gu, Dust Li, D. Wythe,
	Chuck Lever, Jeff Layton, NeilBrown, Olga Kornievskaia, Dai Ngo,
	Tom Talpey, Trond Myklebust, Anna Schumaker, linux-nfs,
	Allison Henderson, rds-devel, Philipp Reisner, Lars Ellenberg,
	Christoph Böhmwalder, Jens Axboe, drbd-dev, Keith Busch,
	Christoph Hellwig, Sagi Grimberg, Chaitanya Kulkarni, linux-nvme,
	Ilya Dryomov, Alex Markuze, Viacheslav Dubeyko, ceph-devel,
	Eric Dumazet

Use the new helpers instead of open coding the SOCK_NOSPACE
manipulation, so that TCP can later maintain a cheaper private
copy of this bit.

No functional change intended.

Reviewed-by: Kuniyuki Iwashima <kuniyu@google.com>
Signed-off-by: Eric Dumazet <edumazet@kernel.org>
---
 drivers/nvme/host/tcp.c   | 2 +-
 drivers/nvme/target/tcp.c | 2 +-
 2 files changed, 2 insertions(+), 2 deletions(-)

diff --git a/drivers/nvme/host/tcp.c b/drivers/nvme/host/tcp.c
index 921934028e0b22c2e4c9fd60d09ed8e50051262d..b853ca2d598fa4cb8ea9cb66513f1e1070cd1fb4 100644
--- a/drivers/nvme/host/tcp.c
+++ b/drivers/nvme/host/tcp.c
@@ -1133,7 +1133,7 @@ static void nvme_tcp_write_space(struct sock *sk)
 	read_lock_bh(&sk->sk_callback_lock);
 	queue = sk->sk_user_data;
 	if (likely(queue && sk_stream_is_writeable(sk))) {
-		clear_bit(SOCK_NOSPACE, &sk->sk_socket->flags);
+		sk_clear_nospace(sk);
 		/* Ensure pending TLS partial records are retried */
 		if (nvme_tcp_queue_tls(queue))
 			queue->write_space(sk);
diff --git a/drivers/nvme/target/tcp.c b/drivers/nvme/target/tcp.c
index e598101752620448423925d2385b606e0b4346e8..c38ea3501ea5def168b6f180b56734b99d21bdea 100644
--- a/drivers/nvme/target/tcp.c
+++ b/drivers/nvme/target/tcp.c
@@ -1686,7 +1686,7 @@ static void nvmet_tcp_write_space(struct sock *sk)
 	}
 
 	if (sk_stream_is_writeable(sk)) {
-		clear_bit(SOCK_NOSPACE, &sk->sk_socket->flags);
+		sk_clear_nospace(sk);
 		queue_work_on(queue_cpu(queue), nvmet_tcp_wq, &queue->io_work);
 	}
 out:
-- 
2.56.0.rc1.315.gc6ed9934b7-goog


^ permalink raw reply related	[flat|nested] 21+ messages in thread

* [PATCH v3 net-next 8/9] libceph: use sk_clear_nospace()
  2026-09-29  7:17 [PATCH v3 net-next 0/9] tcp: avoid struct socket cache line miss in tcp_check_space() Eric Dumazet
                   ` (6 preceding siblings ...)
  2026-09-29  7:17 ` [PATCH v3 net-next 7/9] nvme-tcp: use sk_clear_nospace() Eric Dumazet
@ 2026-09-29  7:17 ` Eric Dumazet
  2026-09-29  7:17 ` [PATCH v3 net-next 9/9] tcp: add tp->tcp_nospace Eric Dumazet
                   ` (2 subsequent siblings)
  10 siblings, 0 replies; 21+ messages in thread
From: Eric Dumazet @ 2026-09-29  7:17 UTC (permalink / raw)
  To: David S . Miller, Jakub Kicinski, Paolo Abeni
  Cc: Simon Horman, Neal Cardwell, Kuniyuki Iwashima, edumazet, netdev,
	Alexander Aring, David Teigland, gfs2, John Fastabend,
	Jakub Sitnicki, Sabrina Dubroca, Jiayuan Chen, Matthieu Baerts,
	Mat Martineau, Geliang Tang, mptcp, Wen Gu, Dust Li, D. Wythe,
	Chuck Lever, Jeff Layton, NeilBrown, Olga Kornievskaia, Dai Ngo,
	Tom Talpey, Trond Myklebust, Anna Schumaker, linux-nfs,
	Allison Henderson, rds-devel, Philipp Reisner, Lars Ellenberg,
	Christoph Böhmwalder, Jens Axboe, drbd-dev, Keith Busch,
	Christoph Hellwig, Sagi Grimberg, Chaitanya Kulkarni, linux-nvme,
	Ilya Dryomov, Alex Markuze, Viacheslav Dubeyko, ceph-devel,
	Eric Dumazet

Use the new helpers instead of open coding the SOCK_NOSPACE
manipulation, so that TCP can later maintain a cheaper private
copy of this bit.

No functional change intended.

Reviewed-by: Kuniyuki Iwashima <kuniyu@google.com>
Signed-off-by: Eric Dumazet <edumazet@kernel.org>
---
 net/ceph/messenger.c | 2 +-
 1 file changed, 1 insertion(+), 1 deletion(-)

diff --git a/net/ceph/messenger.c b/net/ceph/messenger.c
index 212e7797f9e48edde4a4536bba6f1ff293628b5b..cc7db8c61db37a998b698566caa7c936fa83fe4a 100644
--- a/net/ceph/messenger.c
+++ b/net/ceph/messenger.c
@@ -375,7 +375,7 @@ static void ceph_sock_write_space(struct sock *sk)
 	if (ceph_con_flag_test(con, CEPH_CON_F_WRITE_PENDING)) {
 		if (sk_stream_is_writeable(sk)) {
 			dout("%s %p queueing write work\n", __func__, con);
-			clear_bit(SOCK_NOSPACE, &sk->sk_socket->flags);
+			sk_clear_nospace(sk);
 			queue_con(con);
 		}
 	} else {
-- 
2.56.0.rc1.315.gc6ed9934b7-goog


^ permalink raw reply related	[flat|nested] 21+ messages in thread

* [PATCH v3 net-next 9/9] tcp: add tp->tcp_nospace
  2026-09-29  7:17 [PATCH v3 net-next 0/9] tcp: avoid struct socket cache line miss in tcp_check_space() Eric Dumazet
                   ` (7 preceding siblings ...)
  2026-09-29  7:17 ` [PATCH v3 net-next 8/9] libceph: " Eric Dumazet
@ 2026-09-29  7:17 ` Eric Dumazet
  2026-10-02  1:19   ` netdev-bot+sashiko
  2026-09-29  7:24 ` [PATCH v3 net-next 0/9] tcp: avoid struct socket cache line miss in tcp_check_space() netdev-bot+sinfo
  2026-10-05 23:30 ` patchwork-bot+netdevbpf
  10 siblings, 1 reply; 21+ messages in thread
From: Eric Dumazet @ 2026-09-29  7:17 UTC (permalink / raw)
  To: David S . Miller, Jakub Kicinski, Paolo Abeni
  Cc: Simon Horman, Neal Cardwell, Kuniyuki Iwashima, edumazet, netdev,
	Alexander Aring, David Teigland, gfs2, John Fastabend,
	Jakub Sitnicki, Sabrina Dubroca, Jiayuan Chen, Matthieu Baerts,
	Mat Martineau, Geliang Tang, mptcp, Wen Gu, Dust Li, D. Wythe,
	Chuck Lever, Jeff Layton, NeilBrown, Olga Kornievskaia, Dai Ngo,
	Tom Talpey, Trond Myklebust, Anna Schumaker, linux-nfs,
	Allison Henderson, rds-devel, Philipp Reisner, Lars Ellenberg,
	Christoph Böhmwalder, Jens Axboe, drbd-dev, Keith Busch,
	Christoph Hellwig, Sagi Grimberg, Chaitanya Kulkarni, linux-nvme,
	Ilya Dryomov, Alex Markuze, Viacheslav Dubeyko, ceph-devel,
	Eric Dumazet

tcp_check_space() runs for every incoming ACK and for every packet we
send, and reads SOCK_NOSPACE from sk->sk_socket->flags.

struct socket lives in its own cache line, which the TCP fast paths do
not otherwise touch, so testing sk->sk_socket->flags pulls in an extra
cache line that is cold when the working set is large.

Add tp->tcp_nospace, a mirror of SOCK_NOSPACE placed right after
tp->chrono_type in the tcp_sock_write_tx group, in a cache line both
the transmit and the ACK paths already touch (bytes_sent, data_segs_out,
delivered, bytes_acked, chrono_type).  It fits in an existing 3-byte
hole before chrono_start, so no other field moves and sizeof(struct
tcp_sock) is unchanged.

The two flags are now only changed from sk_set_nospace() and
sk_clear_nospace(), which maintain this invariant:

	SOCK_NOSPACE set  =>  tp->tcp_nospace set

sk_set_nospace() sets SOCK_NOSPACE before tp->tcp_nospace (using
smp_mb__after_atomic() and smp_store_mb()), while sk_clear_nospace()
clears tp->tcp_nospace before SOCK_NOSPACE (using
smp_mb__before_atomic()).  If a lockless sk_set_nospace() from
tcp_poll() races with sk_clear_nospace() and SOCK_NOSPACE ends up set,
set_bit(SOCK_NOSPACE) happened after clear_bit(SOCK_NOSPACE), so
tp->tcp_nospace = 1 happens after tp->tcp_nospace = 0 as well.

tcp_check_space() can thus test tp->tcp_nospace alone and leave the
authoritative SOCK_NOSPACE test to __tcp_check_space().  The
invariant is one directional on purpose: a stale tp->tcp_nospace only
costs an extra call to __tcp_check_space(), which is what we do
unconditionally today, while a stale SOCK_NOSPACE would cost a
missed EPOLLOUT.

Note the smp_mb() is kept.  tcp_poll() sets the flag without the
socket lock, and the store-buffer pattern it forms with
tcp_check_space() needs a full barrier on both sides.

MPTCP subflows share the struct socket of their parent, hence its
SOCK_NOSPACE bit, which can not be mirrored in the subflow tcp_sock.
Pin their tp->tcp_nospace in subflow_ulp_init(), so that they always
reach __tcp_check_space() and keep the current behavior.

Microbenchmark on an AMD EPYC 7B13, 64 threads, each thread calling
tcp_check_space() in a loop over a private set of sockets.  Numbers
are cycles per call above a baseline loop that performs the work the
callers already did (the cache lines tcp_write_xmit() and
tcp_clean_rtx_queue() touched) but not tcp_check_space() itself, so
they are the marginal cost of the function.  Median of 11 runs.
The set size controls whether struct socket is still cached:

    sockets/thread	before	 after
    1			  +0.6	  +0.7
    256    (64 KB)	  +8.3	  +3.9
    4096   (1 MB)	 +22.4	  +1.1
    262144 (64 MB)	 +48.4	  -0.8

Signed-off-by: Eric Dumazet <edumazet@kernel.org>
---
 .../networking/net_cachelines/tcp_sock.rst    |  1 +
 include/linux/tcp.h                           |  3 ++
 include/net/tcp.h                             | 30 ++++++++++++++++++-
 net/core/sock.c                               | 17 ++++++++---
 net/ipv4/tcp.c                                |  1 +
 net/ipv4/tcp_input.c                          |  8 ++++-
 net/mptcp/subflow.c                           |  5 ++++
 7 files changed, 59 insertions(+), 6 deletions(-)

diff --git a/Documentation/networking/net_cachelines/tcp_sock.rst b/Documentation/networking/net_cachelines/tcp_sock.rst
index 0f6088c4ab8bb872e7fc86f02479592e84c0247a..103328dc409e4485a3f989c95348c1bc74ab20c6 100644
--- a/Documentation/networking/net_cachelines/tcp_sock.rst
+++ b/Documentation/networking/net_cachelines/tcp_sock.rst
@@ -50,6 +50,7 @@ u8:1                          tcp_usec_ts             read_mostly         read_m
 u32                           chrono_start            read_write                              tcp_chrono_start/stop(tcp_write_xmit,tcp_cwnd_validate,tcp_send_syn_data)
 u32[3]                        chrono_stat             read_write                              tcp_chrono_start/stop(tcp_write_xmit,tcp_cwnd_validate,tcp_send_syn_data)
 u8:2                          chrono_type             read_write                              tcp_chrono_start/stop(tcp_write_xmit,tcp_cwnd_validate,tcp_send_syn_data)
+u8                            tcp_nospace             read_mostly         read_mostly         tcp_check_space(tx);tcp_check_space(rx)
 u8:1                          rate_app_limited                            read_write          tcp_rate_gen
 u8:1                          fastopen_connect
 u8:1                          fastopen_no_cookie
diff --git a/include/linux/tcp.h b/include/linux/tcp.h
index 6a8c77719322f9caee305d954a107892c76d4ef7..e4d1720f49f08f3c12f5ba66fe8fe3379d71b1e0 100644
--- a/include/linux/tcp.h
+++ b/include/linux/tcp.h
@@ -269,6 +269,9 @@ struct tcp_sock {
 				 */
 	u32	snd_sml;	/* Last byte of the most recently transmitted small packet */
 	u8	chrono_type;	/* current chronograph type */
+	u8	tcp_nospace;	/* mirrors SOCK_NOSPACE, must be set whenever
+				 * SOCK_NOSPACE is set.
+				 */
 	u32	chrono_start;	/* Start time in jiffies of a TCP chrono */
 	u32	chrono_stat[3];	/* Time in jiffies for chrono_stat stats */
 	u32	write_seq;	/* Tail(+1) of data held in tcp send buffer */
diff --git a/include/net/tcp.h b/include/net/tcp.h
index 96a1a23e47fac9b2f8717a4e129f468d40c6cfb5..08c04f7b32960e66ae5885dc93710569abda62e5 100644
--- a/include/net/tcp.h
+++ b/include/net/tcp.h
@@ -783,14 +783,42 @@ void tcp_done_with_error(struct sock *sk, int err);
 void tcp_reset(struct sock *sk, struct sk_buff *skb);
 void tcp_fin(struct sock *sk);
 void __tcp_check_space(struct sock *sk);
+
+/* Mirror of SOCK_NOSPACE in tcp_sock, maintained by sk_set_nospace()
+ * and sk_clear_nospace().
+ *
+ * MPTCP subflows share the parent socket, and thus its SOCK_NOSPACE bit.
+ * Keep their mirror always set (see subflow_ulp_init()) so that they
+ * always reach __tcp_check_space() and behave as before.
+ */
+static inline void tcp_set_nospace(struct sock *sk)
+{
+	if (sk_is_tcp(sk)) {
+		/* pairs with smp_mb__before_atomic() in tcp_clear_nospace() */
+		smp_mb__after_atomic();
+		/* pairs with smp_mb() in tcp_check_space() */
+		smp_store_mb(tcp_sk(sk)->tcp_nospace, 1);
+	}
+}
+
+static inline void tcp_clear_nospace(struct sock *sk)
+{
+	if (sk_is_tcp(sk) && !sk_is_mptcp(sk)) {
+		WRITE_ONCE(tcp_sk(sk)->tcp_nospace, 0);
+		/* pairs with smp_mb__after_atomic() in tcp_set_nospace() */
+		smp_mb__before_atomic();
+	}
+}
+
 static inline void tcp_check_space(struct sock *sk)
 {
 	/* pairs with tcp_poll() */
 	smp_mb();
 
-	if (sk->sk_socket && test_bit(SOCK_NOSPACE, &sk->sk_socket->flags))
+	if (unlikely(READ_ONCE(tcp_sk(sk)->tcp_nospace)))
 		__tcp_check_space(sk);
 }
+
 void tcp_sack_compress_send_ack(struct sock *sk);
 
 static inline void tcp_cleanup_skb(struct sk_buff *skb)
diff --git a/net/core/sock.c b/net/core/sock.c
index 11a22aec7e414152aab115e8d11e30067ab3775f..946f614f665b3608d1f310da908d82de622ec5c0 100644
--- a/net/core/sock.c
+++ b/net/core/sock.c
@@ -2969,8 +2969,15 @@ void sk_set_nospace(struct sock *sk)
 {
 	struct socket *sock = sk->sk_socket;
 
-	if (sock)
-		set_bit(SOCK_NOSPACE, &sock->flags);
+	if (!sock)
+		return;
+	/* Set SOCK_NOSPACE before tp->tcp_nospace (paired with
+	 * sk_clear_nospace() clearing tp->tcp_nospace before SOCK_NOSPACE)
+	 * so a concurrent clear cannot leave SOCK_NOSPACE set with
+	 * tp->tcp_nospace cleared.
+	 */
+	set_bit(SOCK_NOSPACE, &sock->flags);
+	tcp_set_nospace(sk);
 }
 EXPORT_SYMBOL(sk_set_nospace);
 
@@ -2985,8 +2992,10 @@ void sk_clear_nospace(struct sock *sk)
 {
 	struct socket *sock = sk->sk_socket;
 
-	if (sock)
-		clear_bit(SOCK_NOSPACE, &sock->flags);
+	if (!sock)
+		return;
+	tcp_clear_nospace(sk);
+	clear_bit(SOCK_NOSPACE, &sock->flags);
 }
 EXPORT_SYMBOL(sk_clear_nospace);
 
diff --git a/net/ipv4/tcp.c b/net/ipv4/tcp.c
index 1cde000cfab4704e6756872f6ddec16851ccc55d..f44cbd3e76178a5ec5b40e420b32836477c29266 100644
--- a/net/ipv4/tcp.c
+++ b/net/ipv4/tcp.c
@@ -5242,6 +5242,7 @@ static void __init tcp_struct_check(void)
 	CACHELINE_ASSERT_GROUP_MEMBER(struct tcp_sock, tcp_sock_write_tx, first_tx_mstamp);
 	CACHELINE_ASSERT_GROUP_MEMBER(struct tcp_sock, tcp_sock_write_tx, delivered_mstamp);
 	CACHELINE_ASSERT_GROUP_MEMBER(struct tcp_sock, tcp_sock_write_tx, snd_sml);
+	CACHELINE_ASSERT_GROUP_MEMBER(struct tcp_sock, tcp_sock_write_tx, tcp_nospace);
 	CACHELINE_ASSERT_GROUP_MEMBER(struct tcp_sock, tcp_sock_write_tx, chrono_start);
 	CACHELINE_ASSERT_GROUP_MEMBER(struct tcp_sock, tcp_sock_write_tx, chrono_stat);
 	CACHELINE_ASSERT_GROUP_MEMBER(struct tcp_sock, tcp_sock_write_tx, write_seq);
diff --git a/net/ipv4/tcp_input.c b/net/ipv4/tcp_input.c
index 892ff256e235272a8483a11949cc98b319dd7cc9..cdea30d07fe2c2a38c3d38da27ad3274221d5958 100644
--- a/net/ipv4/tcp_input.c
+++ b/net/ipv4/tcp_input.c
@@ -6098,8 +6098,14 @@ static void tcp_new_space(struct sock *sk)
  */
 void __tcp_check_space(struct sock *sk)
 {
+	struct socket *sock = sk->sk_socket;
+
+	/* tp->tcp_nospace is only a hint, SOCK_NOSPACE is authoritative. */
+	if (!sock || !test_bit(SOCK_NOSPACE, &sock->flags))
+		return;
+
 	tcp_new_space(sk);
-	if (!test_bit(SOCK_NOSPACE, &sk->sk_socket->flags))
+	if (!test_bit(SOCK_NOSPACE, &sock->flags))
 		tcp_chrono_stop(sk, TCP_CHRONO_SNDBUF_LIMITED);
 }
 
diff --git a/net/mptcp/subflow.c b/net/mptcp/subflow.c
index f0a6725d2c3762def75c879e769e9da37f92215a..e297c88be503b42979886baaf43acfb6b7e78055 100644
--- a/net/mptcp/subflow.c
+++ b/net/mptcp/subflow.c
@@ -2001,6 +2001,11 @@ static int subflow_ulp_init(struct sock *sk)
 	pr_debug("subflow=%p, family=%d\n", ctx, sk->sk_family);
 
 	tp->is_mptcp = 1;
+	/* Subflows share the MPTCP socket, and thus its SOCK_NOSPACE bit,
+	 * which tcp_check_space() can not mirror. Pin the mirror so that
+	 * __tcp_check_space() always tests the shared bit.
+	 */
+	tp->tcp_nospace = 1;
 	ctx->icsk_af_ops = icsk->icsk_af_ops;
 	icsk->icsk_af_ops = subflow_default_af_ops(sk);
 	ctx->tcp_state_change = sk->sk_state_change;
-- 
2.56.0.rc1.315.gc6ed9934b7-goog


^ permalink raw reply related	[flat|nested] 21+ messages in thread

* Re: [PATCH v3 net-next 0/9] tcp: avoid struct socket cache line miss in tcp_check_space()
  2026-09-29  7:17 [PATCH v3 net-next 0/9] tcp: avoid struct socket cache line miss in tcp_check_space() Eric Dumazet
                   ` (8 preceding siblings ...)
  2026-09-29  7:17 ` [PATCH v3 net-next 9/9] tcp: add tp->tcp_nospace Eric Dumazet
@ 2026-09-29  7:24 ` netdev-bot+sinfo
  2026-09-29  7:30   ` Eric Dumazet
  2026-10-05 23:30 ` patchwork-bot+netdevbpf
  10 siblings, 1 reply; 21+ messages in thread
From: netdev-bot+sinfo @ 2026-09-29  7:24 UTC (permalink / raw)
  To: Eric Dumazet
  Cc: David S . Miller, Jakub Kicinski, Paolo Abeni, Simon Horman,
	Neal Cardwell, Kuniyuki Iwashima, edumazet, netdev,
	Alexander Aring, David Teigland, gfs2, John Fastabend,
	Jakub Sitnicki, Sabrina Dubroca, Jiayuan Chen, Matthieu Baerts,
	Mat Martineau, Geliang Tang, mptcp, Wen Gu, Dust Li, D. Wythe,
	Chuck Lever, Jeff Layton, NeilBrown, Olga Kornievskaia, Dai Ngo,
	Tom Talpey, Trond Myklebust, Anna Schumaker, linux-nfs,
	Allison Henderson, rds-devel, Philipp Reisner, Lars Ellenberg,
	Christoph Böhmwalder, Jens Axboe, drbd-dev, Keith Busch,
	Christoph Hellwig, Sagi Grimberg, Chaitanya Kulkarni, linux-nvme,
	Ilya Dryomov, Alex Markuze, Viacheslav Dubeyko, ceph-devel

Hi!

This is an automated message. This series looks like a fix, but its
commit messages seem to be missing some information:

 - Whether the issue was actually triggered, or is only theoretical
   (e.g. found by code inspection). If it was triggered please include
   the symptoms, like the stack trace or error messages.

Please do not repost the series just to address the above. Instead,
reply to this email with the missing information, so that reviewers
can take it into account. If the series needs another revision for
other reasons, please include the information in the commit messages
then.

The evaluation is done by an LLM so it may be wrong, if you think
that is the case please reply and explain.

^ permalink raw reply	[flat|nested] 21+ messages in thread

* Re: [PATCH v3 net-next 0/9] tcp: avoid struct socket cache line miss in tcp_check_space()
  2026-09-29  7:24 ` [PATCH v3 net-next 0/9] tcp: avoid struct socket cache line miss in tcp_check_space() netdev-bot+sinfo
@ 2026-09-29  7:30   ` Eric Dumazet
  0 siblings, 0 replies; 21+ messages in thread
From: Eric Dumazet @ 2026-09-29  7:30 UTC (permalink / raw)
  To: netdev-bot+sinfo
  Cc: Eric Dumazet, David S . Miller, Jakub Kicinski, Paolo Abeni,
	Simon Horman, Neal Cardwell, Kuniyuki Iwashima, netdev,
	Alexander Aring, David Teigland, gfs2, John Fastabend,
	Jakub Sitnicki, Sabrina Dubroca, Jiayuan Chen, Matthieu Baerts,
	Mat Martineau, Geliang Tang, mptcp, Wen Gu, Dust Li, D. Wythe,
	Chuck Lever, Jeff Layton, NeilBrown, Olga Kornievskaia, Dai Ngo,
	Tom Talpey, Trond Myklebust, Anna Schumaker, linux-nfs,
	Allison Henderson, rds-devel, Philipp Reisner, Lars Ellenberg,
	Christoph Böhmwalder, Jens Axboe, drbd-dev, Keith Busch,
	Christoph Hellwig, Sagi Grimberg, Chaitanya Kulkarni, linux-nvme,
	Ilya Dryomov, Alex Markuze, Viacheslav Dubeyko, ceph-devel

On Tue, Sep 29, 2026 at 9:24 AM <netdev-bot+sinfo@kernel.org> wrote:
>
> Hi!
>
> This is an automated message. This series looks like a fix, but its
> commit messages seem to be missing some information:
>
>  - Whether the issue was actually triggered, or is only theoretical
>    (e.g. found by code inspection). If it was triggered please include
>    the symptoms, like the stack trace or error messages.
>

Patch 1/9 fixes a bug found by code inspection.

Other patches are not bug fixes.

> Please do not repost the series just to address the above. Instead,
> reply to this email with the missing information, so that reviewers
> can take it into account. If the series needs another revision for
> other reasons, please include the information in the commit messages
> then.
>
> The evaluation is done by an LLM so it may be wrong, if you think
> that is the case please reply and explain.

^ permalink raw reply	[flat|nested] 21+ messages in thread

* Re: [PATCH v3 net-next 6/9] drbd: use sk_set_nospace()
  2026-09-29  7:17 ` [PATCH v3 net-next 6/9] drbd: use sk_set_nospace() Eric Dumazet
@ 2026-09-29 13:36   ` Christoph Böhmwalder
  0 siblings, 0 replies; 21+ messages in thread
From: Christoph Böhmwalder @ 2026-09-29 13:36 UTC (permalink / raw)
  To: Eric Dumazet, David S . Miller, Jakub Kicinski, Paolo Abeni
  Cc: Simon Horman, Neal Cardwell, Kuniyuki Iwashima, edumazet, netdev,
	Alexander Aring, David Teigland, gfs2, John Fastabend,
	Jakub Sitnicki, Sabrina Dubroca, Jiayuan Chen, Matthieu Baerts,
	Mat Martineau, Geliang Tang, mptcp, Wen Gu, Dust Li, D. Wythe,
	Chuck Lever, Jeff Layton, NeilBrown, Olga Kornievskaia, Dai Ngo,
	Tom Talpey, Trond Myklebust, Anna Schumaker, linux-nfs,
	Allison Henderson, rds-devel, Philipp Reisner, Lars Ellenberg,
	Jens Axboe, drbd-dev, Keith Busch, Christoph Hellwig,
	Sagi Grimberg, Chaitanya Kulkarni, linux-nvme, Ilya Dryomov,
	Alex Markuze, Viacheslav Dubeyko, ceph-devel

On 9/29/26 09:17, Eric Dumazet wrote:
> Use the new helpers instead of open coding the SOCK_NOSPACE
> manipulation, so that TCP can later maintain a cheaper private
> copy of this bit.
> 
> No functional change intended.
> 
> Reviewed-by: Kuniyuki Iwashima <kuniyu@google.com>
> Signed-off-by: Eric Dumazet <edumazet@kernel.org>
> ---
>   drivers/block/drbd/drbd_worker.c | 3 +--
>   1 file changed, 1 insertion(+), 2 deletions(-)
Acked-by: Christoph Böhmwalder <christoph.boehmwalder@linbit.com>

^ permalink raw reply	[flat|nested] 21+ messages in thread

* Re: [PATCH v3 net-next 3/9] sunrpc: use sk_set_nospace() and sk_clear_nospace()
  2026-09-29  7:17 ` [PATCH v3 net-next 3/9] sunrpc: use " Eric Dumazet
@ 2026-09-29 15:21   ` Chuck Lever
  0 siblings, 0 replies; 21+ messages in thread
From: Chuck Lever @ 2026-09-29 15:21 UTC (permalink / raw)
  To: Eric Dumazet, David S . Miller, Jakub Kicinski, Paolo Abeni
  Cc: Simon Horman, Neal Cardwell, Kuniyuki Iwashima, edumazet, netdev,
	Alexander Aring, David Teigland, gfs2, John Fastabend,
	Jakub Sitnicki, Sabrina Dubroca, Jiayuan Chen, Matthieu Baerts,
	Mat Martineau, Geliang Tang, mptcp, Wen Gu, Dust Li, D. Wythe,
	Jeff Layton, NeilBrown, Olga Kornievskaia, Dai Ngo, Tom Talpey,
	Trond Myklebust, Anna Schumaker, linux-nfs, Allison Henderson,
	rds-devel, Philipp Reisner, Lars Ellenberg,
	Christoph Böhmwalder, Jens Axboe, drbd-dev, Keith Busch,
	Christoph Hellwig, Sagi Grimberg, Chaitanya Kulkarni, linux-nvme,
	Ilya Dryomov, Alex Markuze, Viacheslav Dubeyko, ceph-devel



On Tue, Sep 29, 2026, at 12:17 AM, Eric Dumazet wrote:
> Use the new helpers instead of open coding the SOCK_NOSPACE
> manipulation, so that TCP can later maintain a cheaper private
> copy of this bit.
>
> No functional change intended.
>
> Reviewed-by: Kuniyuki Iwashima <kuniyu@google.com>
> Signed-off-by: Eric Dumazet <edumazet@kernel.org>
> ---
>  net/sunrpc/svcsock.c  | 4 ++--
>  net/sunrpc/xprtsock.c | 4 ++--
>  2 files changed, 4 insertions(+), 4 deletions(-)
>
> diff --git a/net/sunrpc/svcsock.c b/net/sunrpc/svcsock.c
> index 
> 50e5e7f5b762de28c34d0f58cb0c6feb1e3ce79f..2d8cdee0798d540ba8eae589deaac5bd3ff889fa 
> 100644
> --- a/net/sunrpc/svcsock.c
> +++ b/net/sunrpc/svcsock.c
> @@ -792,11 +792,11 @@ static int svc_udp_has_wspace(struct svc_xprt 
> *xprt)
>  	 * Set the SOCK_NOSPACE flag before checking the available
>  	 * sock space.
>  	 */
> -	set_bit(SOCK_NOSPACE, &svsk->sk_sock->flags);
> +	sk_set_nospace(svsk->sk_sk);
>  	required = atomic_read(&svsk->sk_xprt.xpt_reserved) + 
> serv->sv_max_mesg;
>  	if (required*2 > sock_wspace(svsk->sk_sk))
>  		return 0;
> -	clear_bit(SOCK_NOSPACE, &svsk->sk_sock->flags);
> +	sk_clear_nospace(svsk->sk_sk);
>  	return 1;
>  }
> 
> diff --git a/net/sunrpc/xprtsock.c b/net/sunrpc/xprtsock.c
> index 
> 7f60723fa64d84887e260e5e130e44cdbec6f850..b97c70c12f627630509510facfd3c3cf96ab6cb6 
> 100644
> --- a/net/sunrpc/xprtsock.c
> +++ b/net/sunrpc/xprtsock.c
> @@ -858,7 +858,7 @@ static int xs_nospace(struct rpc_rqst *req, struct 
> sock_xprt *transport)
>  	if (xprt_connected(xprt)) {
>  		/* wait for more buffer space */
>  		set_bit(XPRT_SOCK_NOSPACE, &transport->sock_state);
> -		set_bit(SOCK_NOSPACE, &sk->sk_socket->flags);
> +		sk_set_nospace(sk);
>  		sk->sk_write_pending++;
>  		xprt_wait_for_buffer_space(xprt);
>  	} else
> @@ -1615,7 +1615,7 @@ static void xs_write_space(struct sock *sk)
> 
>  	if (!sk->sk_socket)
>  		return;
> -	clear_bit(SOCK_NOSPACE, &sk->sk_socket->flags);
> +	sk_clear_nospace(sk);
> 
>  	if (unlikely(!(xprt = xprt_from_sock(sk))))
>  		return;
> -- 
> 2.56.0.rc1.315.gc6ed9934b7-goog

For the svcsock.c hunks of this patch:

Acked-by: Chuck Lever <cel@kernel.org>


-- 
Chuck Lever (Come to NFS bake-a-thon! https://nfsv4bat.org)

^ permalink raw reply	[flat|nested] 21+ messages in thread

* Re: [PATCH v3 net-next 4/9] rds: use sk_set_nospace()
  2026-09-29  7:17 ` [PATCH v3 net-next 4/9] rds: use sk_set_nospace() Eric Dumazet
@ 2026-09-30  1:59   ` Allison Henderson
  0 siblings, 0 replies; 21+ messages in thread
From: Allison Henderson @ 2026-09-30  1:59 UTC (permalink / raw)
  To: Eric Dumazet, David S . Miller, Jakub Kicinski, Paolo Abeni
  Cc: Simon Horman, Neal Cardwell, Kuniyuki Iwashima, edumazet, netdev,
	Alexander Aring, David Teigland, gfs2, John Fastabend,
	Jakub Sitnicki, Sabrina Dubroca, Jiayuan Chen, Matthieu Baerts,
	Mat Martineau, Geliang Tang, mptcp, Wen Gu, Dust Li, D. Wythe,
	Chuck Lever, Jeff Layton, NeilBrown, Olga Kornievskaia, Dai Ngo,
	Tom Talpey, Trond Myklebust, Anna Schumaker, linux-nfs, rds-devel,
	Philipp Reisner, Lars Ellenberg, Christoph Böhmwalder,
	Jens Axboe, drbd-dev, Keith Busch, Christoph Hellwig,
	Sagi Grimberg, Chaitanya Kulkarni, linux-nvme, Ilya Dryomov,
	Alex Markuze, Viacheslav Dubeyko, ceph-devel

On Tue, 2026-09-29 at 07:17 +0000, Eric Dumazet wrote:
> Use the new helpers instead of open coding the SOCK_NOSPACE
> manipulation, so that TCP can later maintain a cheaper private
> copy of this bit.
> 
> No functional change intended.
> 
> Reviewed-by: Kuniyuki Iwashima <kuniyu@google.com>
> Signed-off-by: Eric Dumazet <edumazet@kernel.org>

Thank you!
Acked-by: Allison Henderson <achender@kernel.org>

> ---
>  net/rds/tcp_send.c | 5 ++---
>  1 file changed, 2 insertions(+), 3 deletions(-)
> 
> diff --git a/net/rds/tcp_send.c b/net/rds/tcp_send.c
> index 7c52acc749cf4d4003c749a2646c2b4f87f207fe..824222c475119d91e333123c5afc9291005a8696 100644
> --- a/net/rds/tcp_send.c
> +++ b/net/rds/tcp_send.c
> @@ -100,7 +100,7 @@ int rds_tcp_xmit(struct rds_connection *conn, struct rds_message *rm,
>  
>  	if (hdr_off < sizeof(struct rds_header)) {
>  		/* see rds_tcp_write_space() */
> -		set_bit(SOCK_NOSPACE, &tc->t_sock->sk->sk_socket->flags);
> +		sk_set_nospace(tc->t_sock->sk);
>  
>  		ret = rds_tcp_sendmsg(tc->t_sock,
>  				      (void *)&rm->m_inc.i_hdr + hdr_off,
> @@ -221,6 +221,5 @@ void rds_tcp_write_space(struct sock *sk)
>  	 */
>  	write_space(sk);
>  
> -	if (sk->sk_socket)
> -		set_bit(SOCK_NOSPACE, &sk->sk_socket->flags);
> +	sk_set_nospace(sk);
>  }


^ permalink raw reply	[flat|nested] 21+ messages in thread

* Re: [PATCH v3 net-next 1/9] dlm: fix send buffer backpressure handling
  2026-09-29  7:17 ` [PATCH v3 net-next 1/9] dlm: fix send buffer backpressure handling Eric Dumazet
@ 2026-10-02  1:19   ` netdev-bot+sashiko
  2026-10-02  8:33     ` Eric Dumazet
  0 siblings, 1 reply; 21+ messages in thread
From: netdev-bot+sashiko @ 2026-10-02  1:19 UTC (permalink / raw)
  To: edumazet
  Cc: davem, kuba, pabeni, horms, ncardwell, kuniyu, edumazet, netdev,
	aahringo, teigland, gfs2, john.fastabend, jakub, sd, jiayuan.chen,
	matttbe, martineau, geliang, mptcp, guwen, dust.li, alibuda, cel,
	jlayton, neil, okorniev, Dai.Ngo, tom, trondmy, anna, linux-nfs,
	achender, rds-devel, philipp.reisner, lars.ellenberg,
	christoph.boehmwalder, axboe, drbd-dev, kbusch, hch, sagi, kch,
	linux-nvme, idryomov, amarkuze, slava, ceph-devel

Thank you for your contribution! Sashiko AI review found 2 potential
issue(s) to consider.

Critical: 0 · High: 1 · Medium: 0 · Low: 1

If you believe none of the issues are legitimate, please reply to each
issue of Medium severity or higher and explain why the AI is wrong.
If any of them are legitimate and you plan to respin, please let patchwork
know by including "pw-bot: cr" as a separate line at the end of your reply
(one such reply per series is enough).

- [High] A TCP reset or TCP timeout while the connection is waiting for
  send space can stall the dlm connection permanently.
- [Low] In send_to_sock() (fs/dlm/lowcomms.c:1392-1397), the
  set_bit(SOCK_NOSPACE, &con->sock->sk->sk_socket->flags) under the…

--

dlm: fix send buffer backpressure handling

This commit changes send_to_sock() to test SOCK_NOSPACE instead of
SOCKWQ_ASYNC_NOSPACE in con->sock->flags. With that change the
CF_APP_LIMITED wait path is taken again when TCP returns -EAGAIN because
the send buffer is full. It also removes the clear_bit() of
SOCKWQ_ASYNC_NOSPACE from lowcomms_write_space(), since that bit no
longer lives in con->sock->flags.

> diff --git a/fs/dlm/lowcomms.c b/fs/dlm/lowcomms.c
> index 2aff1c7c17de4..abe9ae4c643fa 100644
> --- a/fs/dlm/lowcomms.c
> +++ b/fs/dlm/lowcomms.c

[ ... ]

> @@ -1391,7 +1389,7 @@ static int send_to_sock(struct connection *con)
>  	if (ret == -EAGAIN || ret == 0) {
>  		lock_sock(con->sock->sk);
>  		spin_lock_bh(&con->writequeue_lock);
> -		if (test_bit(SOCKWQ_ASYNC_NOSPACE, &con->sock->flags) &&
> +		if (test_bit(SOCK_NOSPACE, &con->sock->flags) &&
>  		    !test_and_set_bit(CF_APP_LIMITED, &con->flags)) {

[Severity: High]
Can a TCP reset or timeout that arrives while the connection is in this
state leave it stuck for good?

With this change, send_to_sock() sets CF_APP_LIMITED, clears
CF_SEND_PENDING and returns DLM_IO_END. After that,
lowcomms_queue_swork() will not queue swork:

fs/dlm/lowcomms.c:lowcomms_queue_swork() {
	if (!test_bit(CF_IO_STOP, &con->flags) &&
	    !test_bit(CF_APP_LIMITED, &con->flags) &&
	    !test_and_set_bit(CF_SEND_PENDING, &con->flags))
		queue_work(io_workqueue, &con->swork);
}

>From then on, only lowcomms_write_space() or close_connection() clears
CF_APP_LIMITED.

tcp_reset() and tcp_write_err() (RTO or keepalive timeout) both go
through tcp_done_with_error():

net/ipv4/tcp_input.c:tcp_done_with_error() {
    ...
	tcp_write_queue_purge(sk);
	tcp_done(sk);

	if (!sock_flag(sk, SOCK_DEAD))
		sk_error_report(sk);
}

Neither path calls sk_write_space(). tcp_done() calls sk_state_change,
but add_sock() only installs lowcomms_state_change for SCTP.

Every dlm recovery path appears to end in lowcomms_queue_swork(), which
returns early here:

lowcomms_error_report()
  dlm_midcomms_unack_msg_resend()
    dlm_lowcomms_resend_msg()
      _dlm_lowcomms_commit_msg()
        lowcomms_queue_swork()

On a receive error on the main connection, process_recv_sockets() also
only calls lowcomms_queue_swork():

		if (ret < 0) {
			if (test_bit(CF_IS_OTHERCON, &con->flags)) {
				close_connection(con, false);
			} else {
				spin_lock_bh(&con->writequeue_lock);
				lowcomms_queue_swork(con);
				spin_unlock_bh(&con->writequeue_lock);
			}

A RST does not call sk_data_ready, so the receive worker may not run at
all. New outgoing messages hit the same CF_APP_LIMITED check. If the peer
reconnects, accept_from_sock() puts the new socket on othercon because
the main connection's sock is still set.

Before this patch, the send worker kept requeueing itself with
DLM_IO_RESCHED. The next sock_sendmsg() then returned the socket error,
and process_send_sockets() called close_connection() and reconnected.

A peer RST seems realistic here. It happens when the peer's dlm closes
its socket with unread data in the receive queue, which is likely when
the peer is slow to drain. It also happens with ss -K or a middlebox RST.

If this happens, won't traffic to that node stay queued on a dead socket
until dlm_lowcomms_close() or lowcomms shutdown?

The commit message says that for SCTP, returning DLM_IO_END "would wait
for a callback that never comes". Doesn't the TCP error path have the
same problem?

Later patches in the series convert these bit operations to
sk_set_nospace()/sk_clear_nospace() and add tp->tcp_nospace. None of
them seem to change this behavior.

>  			/* Notify TCP that we're limited by the
>  			 * application window size.

[Severity: Low]
This isn't a bug, but the statement under this comment now does
nothing:

			set_bit(SOCK_NOSPACE, &con->sock->sk->sk_socket->flags);

The branch is only entered when test_bit(SOCK_NOSPACE, &con->sock->flags)
is already true. con->sock->sk->sk_socket is con->sock, so this sets a
bit that is already set.

TCP already sets it on its -EAGAIN paths (tcp_sendmsg_locked() and
sk_stream_wait_memory()). While lock_sock() is held, TCP input that could
run lowcomms_write_space() is deferred to the backlog, so nothing can
clear the bit between the test and the set.

Should the set_bit() and the "Notify TCP" comment be removed or
reworded?

The same applies later in the series, where this line becomes
sk_set_nospace(con->sock->sk).

-- 
Sashiko AI review · https://netdev-ai.bots.linux.dev/sashiko/#/patchset/20260929071743.23624-1-edumazet%40kernel.org

^ permalink raw reply	[flat|nested] 21+ messages in thread

* Re: [PATCH v3 net-next 2/9] net: add sk_set_nospace() and sk_clear_nospace()
  2026-09-29  7:17 ` [PATCH v3 net-next 2/9] net: add sk_set_nospace() and sk_clear_nospace() Eric Dumazet
@ 2026-10-02  1:19   ` netdev-bot+sashiko
  0 siblings, 0 replies; 21+ messages in thread
From: netdev-bot+sashiko @ 2026-10-02  1:19 UTC (permalink / raw)
  To: edumazet
  Cc: davem, kuba, pabeni, horms, ncardwell, kuniyu, edumazet, netdev,
	aahringo, teigland, gfs2, john.fastabend, jakub, sd, jiayuan.chen,
	matttbe, martineau, geliang, mptcp, guwen, dust.li, alibuda, cel,
	jlayton, neil, okorniev, Dai.Ngo, tom, trondmy, anna, linux-nfs,
	achender, rds-devel, philipp.reisner, lars.ellenberg,
	christoph.boehmwalder, axboe, drbd-dev, kbusch, hch, sagi, kch,
	linux-nvme, idryomov, amarkuze, slava, ceph-devel

Thank you for your contribution! Sashiko AI review found more severe
issue(s) in other patches of this series. This mail is only an FYI, in case
you want to address the nit picks below when you respin.

Low priority / nit picks for your consideration:

- [Low] The kernel-doc added in net/core/sock.c for the new helpers does
  not match how the helpers are used, both in this patch and later in…
- [Low] sk_set_nospace() and sk_clear_nospace() (net/core/sock.c) read
  sk->sk_socket with a plain load (`struct socket *sock = sk->sk_socket;…

--

net: add sk_set_nospace() and sk_clear_nospace()

This commit adds two helpers, sk_set_nospace() and sk_clear_nospace().
They set or clear SOCK_NOSPACE in sk->sk_socket->flags after a NULL
check, and the commit converts the core networking callers to use them.
It prepares for a later patch that gives TCP a private copy of the bit.

> diff --git a/net/core/sock.c b/net/core/sock.c
> index 2948dffcc3e1b..11a22aec7e414 100644
> --- a/net/core/sock.c
> +++ b/net/core/sock.c
> @@ -2957,6 +2957,39 @@ void sock_kzfree_s(struct sock *sk, void *mem, int size)
>  }
>  EXPORT_SYMBOL(sock_kzfree_s);
>  
> +/**
> + *	sk_set_nospace - tell the transport a writer is waiting for space
> + *	@sk: socket
> + *
> + *	Must be called before the final check of the available send space,
> + *	so that the transport can not miss the request and forget to call
> + *	sk->sk_write_space() once space is available again.
> + */
> +void sk_set_nospace(struct sock *sk)
> +{
> +	struct socket *sock = sk->sk_socket;
> +
> +	if (sock)
> +		set_bit(SOCK_NOSPACE, &sock->flags);
> +}

[Severity: Low]

Should this be READ_ONCE(sk->sk_socket)?

The writer side uses a marked store:

include/net/sock.h:sk_set_socket() {
	WRITE_ONCE(sk->sk_socket, sock);
}

sock_orphan() makes that store holding only sk_callback_lock:

	write_lock_bh(&sk->sk_callback_lock);
	sock_set_flag(sk, SOCK_DEAD);
	sk_set_socket(sk, NULL);

At least one path can reach these helpers at the same time as that store.
It does not take sk_callback_lock and does not check socket ownership:

CPU1 (SMC CDC receive tasklet)
smc_cdc_msg_recv()
  bh_lock_sock(&smc->sk)
  smc_cdc_msg_recv_action()
    ...
      smc_tx_write_space()
        sk_clear_nospace()
          sock = sk->sk_socket;

CPU2
close()
  smc_release()
    lock_sock(sk)
    sock_orphan(sk)
      WRITE_ONCE(sk->sk_socket, NULL)

Won't KCSAN report this plain read racing with the marked write?

The compiler is also allowed to reload sk->sk_socket after the NULL
check when it computes &sock->flags. That would undo the NULL check that
the commit message relies on to make the helpers "more robust".

This patch also adds a second plain read of sk->sk_socket to
sk_stream_write_space() and smc_tx_write_space(). Both functions already
had the pointer in a local sock variable.

sk_clear_nospace() below has the same plain load. It is still there at
the end of the series, after "tcp: add tp->tcp_nospace".

> +EXPORT_SYMBOL(sk_set_nospace);
> +
> +/**
> + *	sk_clear_nospace - tell the transport no writer is waiting for space
> + *	@sk: socket
> + *
> + *	Called from ->sk_write_space() handlers, once send space has been
> + *	made available to writers.
> + */

[Severity: Low]

Does this kernel-doc, and the one for sk_set_nospace() above, match how
the helpers are actually used?

This comment says the helper is "Called from ->sk_write_space() handlers".
kcm_tx_work() is a workqueue handler, though, and it clears the bit
before it calls ->sk_write_space():

net/kcm/kcmsock.c:kcm_tx_work() {
    ...
	if (likely(sk->sk_socket) &&
	    test_bit(SOCK_NOSPACE, &sk->sk_socket->flags)) {
		sk_clear_nospace(sk);
		sk->sk_write_space(sk);
	}
    ...
}

Later in the series, "sunrpc: use sk_set_nospace() and
sk_clear_nospace()" adds a call to it from svc_tcp_has_wspace() in
net/sunrpc/svcsock.c. That is an ->xpo_has_wspace check, not a
write_space callback.

The sk_set_nospace() comment says it "Must be called before the final
check of the available send space". Several converted callers, however,
set the bit on an -EAGAIN exit and never check again. There the point is
to arm a later ->sk_write_space() or EPOLLOUT:

net/core/stream.c:sk_stream_wait_memory() {
    ...
do_eagain:
    ...
	sk_set_nospace(sk);
	err = -EAGAIN;
	goto out;
}

The !timeo branches in smc_tx_wait() and sock_alloc_send_pskb() work
the same way.

The comment also does not say that sk_set_nospace() is only a set_bit()
with no barrier. Callers have to provide a full barrier between setting
the bit and re-checking. Today that comes from smp_mb__after_atomic() in
tcp_poll(), smc_poll() and mptcp_check_writeable(), and from
prepare_to_wait() in sock_wait_for_wmem().

Could the kernel-doc also cover the -EAGAIN use, the callers outside
write_space handlers, and the barrier the caller must provide?

> +void sk_clear_nospace(struct sock *sk)
> +{
> +	struct socket *sock = sk->sk_socket;
> +
> +	if (sock)
> +		clear_bit(SOCK_NOSPACE, &sock->flags);
> +}
> +EXPORT_SYMBOL(sk_clear_nospace);

[ ... ]

-- 
Sashiko AI review · https://netdev-ai.bots.linux.dev/sashiko/#/patchset/20260929071743.23624-1-edumazet%40kernel.org

^ permalink raw reply	[flat|nested] 21+ messages in thread

* Re: [PATCH v3 net-next 9/9] tcp: add tp->tcp_nospace
  2026-09-29  7:17 ` [PATCH v3 net-next 9/9] tcp: add tp->tcp_nospace Eric Dumazet
@ 2026-10-02  1:19   ` netdev-bot+sashiko
  0 siblings, 0 replies; 21+ messages in thread
From: netdev-bot+sashiko @ 2026-10-02  1:19 UTC (permalink / raw)
  To: edumazet
  Cc: davem, kuba, pabeni, horms, ncardwell, kuniyu, edumazet, netdev,
	aahringo, teigland, gfs2, john.fastabend, jakub, sd, jiayuan.chen,
	matttbe, martineau, geliang, mptcp, guwen, dust.li, alibuda, cel,
	jlayton, neil, okorniev, Dai.Ngo, tom, trondmy, anna, linux-nfs,
	achender, rds-devel, philipp.reisner, lars.ellenberg,
	christoph.boehmwalder, axboe, drbd-dev, kbusch, hch, sagi, kch,
	linux-nvme, idryomov, amarkuze, slava, ceph-devel

Thank you for your contribution! Sashiko AI review found more severe
issue(s) in other patches of this series. This mail is only an FYI, in case
you want to address the nit picks below when you respin.

Low priority / nit picks for your consideration:

- [Low] The commit message says struct socket lives in a cache line
  "which the TCP fast paths do not otherwise touch".
- [Low] A child socket inherits the listener's tp->tcp_nospace and
  nothing resets it.

--

tcp: add tp->tcp_nospace

This commit adds tp->tcp_nospace, a copy of SOCK_NOSPACE that lives in the
tcp_sock_write_tx cache line group. With it, tcp_check_space() no longer
has to read sk->sk_socket->flags on every ACK and every transmit.
sk_set_nospace() and sk_clear_nospace() keep the rule that SOCK_NOSPACE
set implies tp->tcp_nospace set. MPTCP subflows pin tp->tcp_nospace to 1
in subflow_ulp_init().

> struct socket lives in its own cache line, which the TCP fast paths do
> not otherwise touch, so testing sk->sk_socket->flags pulls in an extra
> cache line that is cold when the working set is large.

[Severity: Low]
Is this still accurate after this patch? Only tcp_check_space() was
converted to use the new field.

tcp_write_xmit() calls tcp_cwnd_validate() whenever it has sent packets.
If the flow is not cwnd limited and the write queue is empty,
tcp_cwnd_validate() still reads sk->sk_socket->flags:

net/ipv4/tcp_output.c:tcp_cwnd_validate() {
    ...
		if (tcp_write_queue_empty(sk) && sk->sk_socket &&
		    test_bit(SOCK_NOSPACE, &sk->sk_socket->flags) &&
		    (1 << sk->sk_state) & (TCPF_ESTABLISHED | TCPF_CLOSE_WAIT))
			tcp_chrono_start(sk, TCP_CHRONO_SNDBUF_LIMITED);
    ...
}

tcp_cwnd_application_limited() also reads it on the
slow_start_after_idle path:

net/ipv4/tcp_output.c:tcp_cwnd_application_limited() {
    ...
	if (inet_csk(sk)->icsk_ca_state == TCP_CA_Open &&
	    sk->sk_socket && !test_bit(SOCK_NOSPACE, &sk->sk_socket->flags)) {
    ...
}

Request/response senders limited by the application empty the write
queue on every transmit. Their transmits would still pull in the struct
socket cache line. The microbenchmark only times tcp_check_space(), so it
would not show this.

The new rule (SOCK_NOSPACE set => tp->tcp_nospace set) means both sites
could check READ_ONCE(tp->tcp_nospace) first without changing behavior.
Would that be worth doing? If not, could the commit message avoid saying
the fast paths no longer touch struct socket?

[ ... ]

> diff --git a/net/mptcp/subflow.c b/net/mptcp/subflow.c
> index f0a6725d2c376..e297c88be503b 100644
> --- a/net/mptcp/subflow.c
> +++ b/net/mptcp/subflow.c
> @@ -2001,6 +2001,11 @@ static int subflow_ulp_init(struct sock *sk)
>  	pr_debug("subflow=%p, family=%d\n", ctx, sk->sk_family);
>  
>  	tp->is_mptcp = 1;
> +	/* Subflows share the MPTCP socket, and thus its SOCK_NOSPACE bit,
> +	 * which tcp_check_space() can not mirror. Pin the mirror so that
> +	 * __tcp_check_space() always tests the shared bit.
> +	 */
> +	tp->tcp_nospace = 1;

[Severity: Low]
Do passive children of an MPTCP listener inherit this pinned value, even
when they fall back to plain TCP?

The listener subflow gets tcp_nospace = 1 here through
mptcp_subflow_create_socket()->tcp_set_ulp()->subflow_ulp_init(). Its
is_mptcp stays 1, so tcp_clear_nospace() never clears it.

tcp_create_openreq_child()->inet_csk_clone_lock()->sk_clone_lock() then
copies the tcp_sock, and nothing resets tcp_nospace.

When the SYN comes from a peer without MPTCP,
subflow_ulp_clone()->subflow_ulp_fallback() makes the child plain TCP:

	tcp_sk(sk)->is_mptcp = 0;

mptcp_stream_accept() then grafts it as a plain TCP socket. Its new
struct socket has SOCK_NOSPACE clear. This new early return in
__tcp_check_space() then applies:

	if (!sock || !test_bit(SOCK_NOSPACE, &sock->flags))
		return;

So tcp_new_space()->sk_stream_write_space()->sk_clear_nospace() is never
reached, and that is the only place the hint gets cleared. It stays set
until the application first fills its send buffer.

Until then, every ACK and every transmit calls __tcp_check_space() out of
line and reads sk->sk_socket->flags. These connections would not get the
optimization.

The rule still holds, so no EPOLLOUT is lost. Could tcp_nospace be reset
for the child, for example in tcp_create_openreq_child() or in
subflow_ulp_fallback()?

The same can happen with a plain TCP listener. If a socket was connected,
ended up with tcp_nospace = 1, was disconnected and then put into
listen(), its children inherit the stale value too. tcp_disconnect() does
not reset it either.

[ ... ]

-- 
Sashiko AI review · https://netdev-ai.bots.linux.dev/sashiko/#/patchset/20260929071743.23624-1-edumazet%40kernel.org

^ permalink raw reply	[flat|nested] 21+ messages in thread

* Re: [PATCH v3 net-next 1/9] dlm: fix send buffer backpressure handling
  2026-10-02  1:19   ` netdev-bot+sashiko
@ 2026-10-02  8:33     ` Eric Dumazet
  2026-10-02 14:07       ` Alexander Aring
  0 siblings, 1 reply; 21+ messages in thread
From: Eric Dumazet @ 2026-10-02  8:33 UTC (permalink / raw)
  To: netdev-bot+sashiko
  Cc: davem, kuba, pabeni, horms, ncardwell, kuniyu, edumazet, netdev,
	aahringo, teigland, gfs2, john.fastabend, jakub, sd, jiayuan.chen,
	matttbe, martineau, geliang, mptcp, guwen, dust.li, alibuda, cel,
	jlayton, neil, okorniev, Dai.Ngo, tom, trondmy, anna, linux-nfs,
	achender, rds-devel, philipp.reisner, lars.ellenberg,
	christoph.boehmwalder, axboe, drbd-dev, kbusch, hch, sagi, kch,
	linux-nvme, idryomov, amarkuze, slava, ceph-devel

On Fri, Oct 2, 2026 at 3:19 AM <netdev-bot+sashiko@kernel.org> wrote:
>
> Thank you for your contribution! Sashiko AI review found 2 potential
> issue(s) to consider.
>
> Critical: 0 · High: 1 · Medium: 0 · Low: 1
>
> If you believe none of the issues are legitimate, please reply to each
> issue of Medium severity or higher and explain why the AI is wrong.
> If any of them are legitimate and you plan to respin, please let patchwork
> know by including "pw-bot: cr" as a separate line at the end of your reply
> (one such reply per series is enough).
>
> - [High] A TCP reset or TCP timeout while the connection is waiting for
>   send space can stall the dlm connection permanently.

I dunno if I want to spend time fixing all dlm bugs.

It seems TCP changes are now stuck because ood pre-existing bugs in
in-kernel-users.

Sad.

^ permalink raw reply	[flat|nested] 21+ messages in thread

* Re: [PATCH v3 net-next 1/9] dlm: fix send buffer backpressure handling
  2026-10-02  8:33     ` Eric Dumazet
@ 2026-10-02 14:07       ` Alexander Aring
  0 siblings, 0 replies; 21+ messages in thread
From: Alexander Aring @ 2026-10-02 14:07 UTC (permalink / raw)
  To: Eric Dumazet
  Cc: netdev-bot+sashiko, davem, kuba, pabeni, horms, ncardwell, kuniyu,
	edumazet, netdev, teigland, gfs2, john.fastabend, jakub, sd,
	jiayuan.chen, matttbe, martineau, geliang, mptcp, guwen, dust.li,
	alibuda, cel, jlayton, neil, okorniev, Dai.Ngo, tom, trondmy,
	anna, linux-nfs, achender, rds-devel, philipp.reisner,
	lars.ellenberg, christoph.boehmwalder, axboe, drbd-dev, kbusch,
	hch, sagi, kch, linux-nvme, idryomov, amarkuze, slava, ceph-devel

Hi,

On Fri, Oct 2, 2026 at 4:33 AM Eric Dumazet <edumazet@kernel.org> wrote:
>
> On Fri, Oct 2, 2026 at 3:19 AM <netdev-bot+sashiko@kernel.org> wrote:
> >
> > Thank you for your contribution! Sashiko AI review found 2 potential
> > issue(s) to consider.
> >
> > Critical: 0 · High: 1 · Medium: 0 · Low: 1
> >
> > If you believe none of the issues are legitimate, please reply to each
> > issue of Medium severity or higher and explain why the AI is wrong.
> > If any of them are legitimate and you plan to respin, please let patchwork
> > know by including "pw-bot: cr" as a separate line at the end of your reply
> > (one such reply per series is enough).
> >
> > - [High] A TCP reset or TCP timeout while the connection is waiting for
> >   send space can stall the dlm connection permanently.
>
> I dunno if I want to spend time fixing all dlm bugs.
>
> It seems TCP changes are now stuck because ood pre-existing bugs in
> in-kernel-users.
>
> Sad.

I will take a look into it and cc netdev people.

Thanks.

- Alex


^ permalink raw reply	[flat|nested] 21+ messages in thread

* Re: [PATCH v3 net-next 0/9] tcp: avoid struct socket cache line miss in tcp_check_space()
  2026-09-29  7:17 [PATCH v3 net-next 0/9] tcp: avoid struct socket cache line miss in tcp_check_space() Eric Dumazet
                   ` (9 preceding siblings ...)
  2026-09-29  7:24 ` [PATCH v3 net-next 0/9] tcp: avoid struct socket cache line miss in tcp_check_space() netdev-bot+sinfo
@ 2026-10-05 23:30 ` patchwork-bot+netdevbpf
  10 siblings, 0 replies; 21+ messages in thread
From: patchwork-bot+netdevbpf @ 2026-10-05 23:30 UTC (permalink / raw)
  To: Eric Dumazet
  Cc: davem, kuba, pabeni, horms, ncardwell, kuniyu, edumazet, netdev,
	aahringo, teigland, gfs2, john.fastabend, jakub, sd, jiayuan.chen,
	matttbe, martineau, geliang, mptcp, guwen, dust.li, alibuda, cel,
	jlayton, neil, okorniev, Dai.Ngo, tom, trondmy, anna, linux-nfs,
	achender, rds-devel, philipp.reisner, lars.ellenberg,
	christoph.boehmwalder, axboe, drbd-dev, kbusch, hch, sagi, kch,
	linux-nvme, idryomov, amarkuze, slava, ceph-devel

Hello:

This series was applied to netdev/net-next.git (main)
by Jakub Kicinski <kuba@kernel.org>:

On Tue, 29 Sep 2026 07:17:34 +0000 you wrote:
> tcp_check_space() is called on every incoming ACK and every transmitted
> packet, and tests SOCK_NOSPACE in sk->sk_socket->flags.  Because struct
> socket lives in its own cache line and the TCP fast paths do not touch
> it for anything else, that test pulls in an extra cache line that is
> cold when the working set of active sockets is large.
> 
> This series mirrors SOCK_NOSPACE into a new u8 field, tp->tcp_nospace,
> placed right after tp->chrono_type in the tcp_sock_write_tx cache line
> group (fitting in an existing 3-byte hole before chrono_start, so no
> other field moves and sizeof(struct tcp_sock) is unchanged).  Both the
> transmit and ACK fast paths already touch that cache line, making the
> fast-path test in tcp_check_space() free of extra cache misses while
> keeping __tcp_check_space() authoritative on SOCK_NOSPACE.
> 
> [...]

Here is the summary with links:
  - [v3,net-next,1/9] dlm: fix send buffer backpressure handling
    https://git.kernel.org/netdev/net-next/c/e234d7625598
  - [v3,net-next,2/9] net: add sk_set_nospace() and sk_clear_nospace()
    https://git.kernel.org/netdev/net-next/c/6ba48a271bdb
  - [v3,net-next,3/9] sunrpc: use sk_set_nospace() and sk_clear_nospace()
    https://git.kernel.org/netdev/net-next/c/28613c4d4757
  - [v3,net-next,4/9] rds: use sk_set_nospace()
    https://git.kernel.org/netdev/net-next/c/b747807ccaf9
  - [v3,net-next,5/9] dlm: use sk_set_nospace() and sk_clear_nospace()
    https://git.kernel.org/netdev/net-next/c/0e0be63cb35d
  - [v3,net-next,6/9] drbd: use sk_set_nospace()
    https://git.kernel.org/netdev/net-next/c/24280d1d2832
  - [v3,net-next,7/9] nvme-tcp: use sk_clear_nospace()
    https://git.kernel.org/netdev/net-next/c/f80fa2ddfe95
  - [v3,net-next,8/9] libceph: use sk_clear_nospace()
    https://git.kernel.org/netdev/net-next/c/5181ac617537
  - [v3,net-next,9/9] tcp: add tp->tcp_nospace
    https://git.kernel.org/netdev/net-next/c/4eee6d58ad14

You are awesome, thank you!
-- 
Deet-doot-dot, I am a bot.
https://korg.docs.kernel.org/patchwork/pwbot.html



^ permalink raw reply	[flat|nested] 21+ messages in thread

end of thread, other threads:[~2026-10-05 23:30 UTC | newest]

Thread overview: 21+ messages (download: mbox.gz follow: Atom feed
-- links below jump to the message on this page --
2026-09-29  7:17 [PATCH v3 net-next 0/9] tcp: avoid struct socket cache line miss in tcp_check_space() Eric Dumazet
2026-09-29  7:17 ` [PATCH v3 net-next 1/9] dlm: fix send buffer backpressure handling Eric Dumazet
2026-10-02  1:19   ` netdev-bot+sashiko
2026-10-02  8:33     ` Eric Dumazet
2026-10-02 14:07       ` Alexander Aring
2026-09-29  7:17 ` [PATCH v3 net-next 2/9] net: add sk_set_nospace() and sk_clear_nospace() Eric Dumazet
2026-10-02  1:19   ` netdev-bot+sashiko
2026-09-29  7:17 ` [PATCH v3 net-next 3/9] sunrpc: use " Eric Dumazet
2026-09-29 15:21   ` Chuck Lever
2026-09-29  7:17 ` [PATCH v3 net-next 4/9] rds: use sk_set_nospace() Eric Dumazet
2026-09-30  1:59   ` Allison Henderson
2026-09-29  7:17 ` [PATCH v3 net-next 5/9] dlm: use sk_set_nospace() and sk_clear_nospace() Eric Dumazet
2026-09-29  7:17 ` [PATCH v3 net-next 6/9] drbd: use sk_set_nospace() Eric Dumazet
2026-09-29 13:36   ` Christoph Böhmwalder
2026-09-29  7:17 ` [PATCH v3 net-next 7/9] nvme-tcp: use sk_clear_nospace() Eric Dumazet
2026-09-29  7:17 ` [PATCH v3 net-next 8/9] libceph: " Eric Dumazet
2026-09-29  7:17 ` [PATCH v3 net-next 9/9] tcp: add tp->tcp_nospace Eric Dumazet
2026-10-02  1:19   ` netdev-bot+sashiko
2026-09-29  7:24 ` [PATCH v3 net-next 0/9] tcp: avoid struct socket cache line miss in tcp_check_space() netdev-bot+sinfo
2026-09-29  7:30   ` Eric Dumazet
2026-10-05 23:30 ` patchwork-bot+netdevbpf

This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox