Netdev List
 help / color / mirror / Atom feed
* [PATCH v2 net-next 0/9] tcp: avoid struct socket cache line miss in tcp_check_space()
@ 2026-09-24 13:47 Eric Dumazet
  2026-09-24 13:47 ` [PATCH v2 net-next 1/9] dlm: fix send buffer backpressure handling Eric Dumazet
                   ` (9 more replies)
  0 siblings, 10 replies; 13+ messages in thread
From: Eric Dumazet @ 2026-09-24 13:47 UTC (permalink / raw)
  To: David S . Miller, Jakub Kicinski, Paolo Abeni
  Cc: Simon Horman, Neal Cardwell, Kuniyuki Iwashima, Willem de Bruijn,
	netdev, eric.dumazet, Eric Dumazet

tcp_check_space() is called on every incoming ACK and every transmitted
packet, and tests SOCK_NOSPACE in sk->sk_socket->flags.  Because struct
socket lives in its own cache line and the TCP fast paths do not touch
it for anything else, that test pulls in an extra cache line that is
cold when the working set of active sockets is large.

This series mirrors SOCK_NOSPACE into a new u8 field, tp->tcp_nospace,
placed right after tp->chrono_type in the tcp_sock_write_tx cache line
group (fitting in an existing 3-byte hole before chrono_start, so no
other field moves and sizeof(struct tcp_sock) is unchanged).  Both the
transmit and ACK fast paths already touch that cache line, making the
fast-path test in tcp_check_space() free of extra cache misses while
keeping __tcp_check_space() authoritative on SOCK_NOSPACE.

To maintain the invariant (SOCK_NOSPACE set => tp->tcp_nospace set) from
a single choke point:

 - Patch 1 fixes a long-standing bug in dlm where SOCKWQ_ASYNC_NOSPACE
   was tested and cleared on con->sock->flags instead of SOCK_NOSPACE.
 - Patch 2 introduces sk_set_nospace() and sk_clear_nospace() and
   converts the core networking callers.
 - Patches 3-8 convert the remaining in-kernel callers (sunrpc, rds,
   dlm, drbd, nvme-tcp, libceph) so that no open-coded set_bit() or
   clear_bit() of SOCK_NOSPACE remains in the tree.
 - Patch 9 adds tp->tcp_nospace, wires it into sk_set_nospace() and
   sk_clear_nospace(), and switches tcp_check_space() to test it.

v2:
 - Order set_bit(SOCK_NOSPACE) before tp->tcp_nospace = 1 in
   sk_set_nospace() and tp->tcp_nospace = 0 before
   clear_bit(SOCK_NOSPACE) in sk_clear_nospace() so a concurrent
   lockless tcp_poll() cannot leave SOCK_NOSPACE set with
   tp->tcp_nospace cleared (Sashiko)
 - Move tp->tcp_nospace to the 3-byte hole after tp->chrono_type so no
   field in struct tcp_sock shifts after the AccECN bitfield additions
   (Sashiko)
 - Clarify comment and Patch 2 changelog wording (Sashiko)
 - Link to v1: https://lore.kernel.org/netdev/20260922122721.3568295-1-edumazet@google.com/

Eric Dumazet (9):
  dlm: fix send buffer backpressure handling
  net: add sk_set_nospace() and sk_clear_nospace()
  sunrpc: use sk_set_nospace() and sk_clear_nospace()
  rds: use sk_set_nospace()
  dlm: use sk_set_nospace() and sk_clear_nospace()
  drbd: use sk_set_nospace()
  nvme-tcp: use sk_clear_nospace()
  libceph: use sk_clear_nospace()
  tcp: add tp->tcp_nospace

 .../networking/net_cachelines/tcp_sock.rst    |  1 +
 drivers/block/drbd/drbd_worker.c              |  3 +-
 drivers/nvme/host/tcp.c                       |  2 +-
 drivers/nvme/target/tcp.c                     |  2 +-
 fs/dlm/lowcomms.c                             | 10 ++--
 include/linux/tcp.h                           |  3 ++
 include/net/sock.h                            |  2 +
 include/net/tcp.h                             | 30 +++++++++++-
 net/ceph/messenger.c                          |  2 +-
 net/core/sock.c                               | 46 ++++++++++++++++++-
 net/core/stream.c                             |  6 +--
 net/ipv4/tcp.c                                |  5 +-
 net/ipv4/tcp_bpf.c                            |  2 +-
 net/ipv4/tcp_input.c                          |  8 +++-
 net/kcm/kcmsock.c                             |  4 +-
 net/mptcp/protocol.c                          |  4 +-
 net/mptcp/subflow.c                           |  5 ++
 net/rds/tcp_send.c                            |  5 +-
 net/smc/af_smc.c                              |  2 +-
 net/smc/smc_tx.c                              |  6 +--
 net/sunrpc/svcsock.c                          |  4 +-
 net/sunrpc/xprtsock.c                         |  4 +-
 net/tls/tls_sw.c                              |  2 +-
 23 files changed, 121 insertions(+), 37 deletions(-)

-- 
2.56.0.rc1.310.g51773c2048-goog


^ permalink raw reply	[flat|nested] 13+ messages in thread

* [PATCH v2 net-next 1/9] dlm: fix send buffer backpressure handling
  2026-09-24 13:47 [PATCH v2 net-next 0/9] tcp: avoid struct socket cache line miss in tcp_check_space() Eric Dumazet
@ 2026-09-24 13:47 ` Eric Dumazet
  2026-09-25 13:49   ` netdev-bot+sashiko
  2026-09-24 13:47 ` [PATCH v2 net-next 2/9] net: add sk_set_nospace() and sk_clear_nospace() Eric Dumazet
                   ` (8 subsequent siblings)
  9 siblings, 1 reply; 13+ messages in thread
From: Eric Dumazet @ 2026-09-24 13:47 UTC (permalink / raw)
  To: David S . Miller, Jakub Kicinski, Paolo Abeni
  Cc: Simon Horman, Neal Cardwell, Kuniyuki Iwashima, Willem de Bruijn,
	netdev, eric.dumazet, Eric Dumazet, stable, Alexander Aring,
	David Teigland

lowcomms.c tests and clears SOCKWQ_ASYNC_NOSPACE in con->sock->flags,
but this bit has not been stored there for ten years.

Commit 9cd3e072b0be ("net: rename SOCK_ASYNC_NOSPACE and
SOCK_ASYNC_WAITDATA") mechanically renamed the two dlm users, then
commit ceb5d58b2170 ("net: fix sock_wake_async() rcu protection") moved
the bit from socket->flags to the RCU protected socket_wq->flags, where
it is reachable only through sk_set_bit() and sk_clear_bit(), and is
only maintained for sockets having SOCK_FASYNC set.  dlm uses kernel
sockets, which never have SOCK_FASYNC set, and never sets the bit
itself, so the test in send_to_sock() has been false ever since.

The consequence is that when sock_sendmsg() returns -EAGAIN because the
socket send buffer is full, dlm no longer sets CF_APP_LIMITED, does not
increment sk_write_pending, and does not return DLM_IO_END to wait for
lowcomms_write_space().  It returns DLM_IO_RESCHED instead, and
process_send_sockets() immediately requeues the send work.  A connection
to a peer that is slow to drain thus keeps cycling through
sock_sendmsg() and -EAGAIN, burning CPU, instead of sleeping until TCP
reports that space is available again.

Test SOCK_NOSPACE instead.  This is the bit that lives in socket->flags,
that TCP sets whenever sendmsg() returns -EAGAIN for lack of send buffer
space (tcp_sendmsg_locked() and sk_stream_wait_memory()), and that
lowcomms_write_space() already clears.  This restores the semantics dlm
had before the bit moved.

Also remove the clear_bit() of SOCKWQ_ASYNC_NOSPACE from
lowcomms_write_space(), for the same reason.

SCTP connections are deliberately left as they are.  SCTP does not set
SOCK_NOSPACE, and it never calls sk->sk_write_space(): sctp_wfree() ends
up in sctp_wake_up_waiters(), which calls sctp_write_space() directly.
lowcomms_write_space() is thus never invoked for an SCTP connection, and
the test added here stays false, so send_to_sock() keeps returning
DLM_IO_RESCHED as it does today.  This is the only safe behavior, as
returning DLM_IO_END would wait for a callback that never comes.

Fixes: ceb5d58b2170 ("net: fix sock_wake_async() rcu protection")
Cc: stable@vger.kernel.org
Cc: Alexander Aring <aahringo@redhat.com>
Cc: David Teigland <teigland@redhat.com>
Acked-by: Alexander Aring <aahringo@redhat.com>
Reviewed-by: Kuniyuki Iwashima <kuniyu@google.com>
Signed-off-by: Eric Dumazet <edumazet@google.com>
---
 fs/dlm/lowcomms.c | 6 ++----
 1 file changed, 2 insertions(+), 4 deletions(-)

diff --git a/fs/dlm/lowcomms.c b/fs/dlm/lowcomms.c
index 2aff1c7c17de49c41f9fd02a1fae00bcc9e6af5b..abe9ae4c643fae061d2e168c03219a4f4e1bc007 100644
--- a/fs/dlm/lowcomms.c
+++ b/fs/dlm/lowcomms.c
@@ -522,10 +522,8 @@ static void lowcomms_write_space(struct sock *sk)
 	clear_bit(SOCK_NOSPACE, &con->sock->flags);
 
 	spin_lock_bh(&con->writequeue_lock);
-	if (test_and_clear_bit(CF_APP_LIMITED, &con->flags)) {
+	if (test_and_clear_bit(CF_APP_LIMITED, &con->flags))
 		con->sock->sk->sk_write_pending--;
-		clear_bit(SOCKWQ_ASYNC_NOSPACE, &con->sock->flags);
-	}
 
 	lowcomms_queue_swork(con);
 	spin_unlock_bh(&con->writequeue_lock);
@@ -1391,7 +1389,7 @@ static int send_to_sock(struct connection *con)
 	if (ret == -EAGAIN || ret == 0) {
 		lock_sock(con->sock->sk);
 		spin_lock_bh(&con->writequeue_lock);
-		if (test_bit(SOCKWQ_ASYNC_NOSPACE, &con->sock->flags) &&
+		if (test_bit(SOCK_NOSPACE, &con->sock->flags) &&
 		    !test_and_set_bit(CF_APP_LIMITED, &con->flags)) {
 			/* Notify TCP that we're limited by the
 			 * application window size.
-- 
2.56.0.rc1.310.g51773c2048-goog


^ permalink raw reply related	[flat|nested] 13+ messages in thread

* [PATCH v2 net-next 2/9] net: add sk_set_nospace() and sk_clear_nospace()
  2026-09-24 13:47 [PATCH v2 net-next 0/9] tcp: avoid struct socket cache line miss in tcp_check_space() Eric Dumazet
  2026-09-24 13:47 ` [PATCH v2 net-next 1/9] dlm: fix send buffer backpressure handling Eric Dumazet
@ 2026-09-24 13:47 ` Eric Dumazet
  2026-09-25 13:49   ` netdev-bot+sashiko
  2026-09-24 13:47 ` [PATCH v2 net-next 3/9] sunrpc: use " Eric Dumazet
                   ` (7 subsequent siblings)
  9 siblings, 1 reply; 13+ messages in thread
From: Eric Dumazet @ 2026-09-24 13:47 UTC (permalink / raw)
  To: David S . Miller, Jakub Kicinski, Paolo Abeni
  Cc: Simon Horman, Neal Cardwell, Kuniyuki Iwashima, Willem de Bruijn,
	netdev, eric.dumazet, Eric Dumazet

SOCK_NOSPACE lives in sk->sk_socket->flags and is manipulated from
about thirty places in the tree, all of them open coding the
sk->sk_socket dereference, some with a NULL check, some without,
and some reaching the struct socket by yet another path.

Add sk_set_nospace() and sk_clear_nospace() helpers and convert the
core networking setters and clearers to them; the following patches
convert the remaining subsystem callers so that
"git grep _bit(SOCK_NOSPACE" only reports the two helpers and the
remaining test_bit() sites.

Callers that had no NULL check are all called from user context with a
socket attached, so folding the check into the helpers only makes them
more robust.

No functional change intended.

This is a preparation patch: a following one gives TCP a cheaper
private copy of this bit, and needs a single choke point to keep it
in sync.

Reviewed-by: Kuniyuki Iwashima <kuniyu@google.com>
Signed-off-by: Eric Dumazet <edumazet@google.com>
---
 include/net/sock.h   |  2 ++
 net/core/sock.c      | 37 +++++++++++++++++++++++++++++++++++--
 net/core/stream.c    |  6 +++---
 net/ipv4/tcp.c       |  4 ++--
 net/ipv4/tcp_bpf.c   |  2 +-
 net/kcm/kcmsock.c    |  4 ++--
 net/mptcp/protocol.c |  4 ++--
 net/smc/af_smc.c     |  2 +-
 net/smc/smc_tx.c     |  6 +++---
 net/tls/tls_sw.c     |  2 +-
 10 files changed, 52 insertions(+), 17 deletions(-)

diff --git a/include/net/sock.h b/include/net/sock.h
index 60ea55dc18854a9759f5df618cc8c904d2323e95..f4dd2e105171386f1b68f808175457c51aff183a 100644
--- a/include/net/sock.h
+++ b/include/net/sock.h
@@ -1129,6 +1129,8 @@ static inline void sk_forward_alloc_add(struct sock *sk, int val)
 }
 
 void sk_stream_write_space(struct sock *sk);
+void sk_set_nospace(struct sock *sk);
+void sk_clear_nospace(struct sock *sk);
 
 /* OOB backlog add */
 static inline void __sk_add_backlog(struct sock *sk, struct sk_buff *skb)
diff --git a/net/core/sock.c b/net/core/sock.c
index 2948dffcc3e1b49a9a55e30f1380ec165a88859f..11a22aec7e414152aab115e8d11e30067ab3775f 100644
--- a/net/core/sock.c
+++ b/net/core/sock.c
@@ -2957,6 +2957,39 @@ void sock_kzfree_s(struct sock *sk, void *mem, int size)
 }
 EXPORT_SYMBOL(sock_kzfree_s);
 
+/**
+ *	sk_set_nospace - tell the transport a writer is waiting for space
+ *	@sk: socket
+ *
+ *	Must be called before the final check of the available send space,
+ *	so that the transport can not miss the request and forget to call
+ *	sk->sk_write_space() once space is available again.
+ */
+void sk_set_nospace(struct sock *sk)
+{
+	struct socket *sock = sk->sk_socket;
+
+	if (sock)
+		set_bit(SOCK_NOSPACE, &sock->flags);
+}
+EXPORT_SYMBOL(sk_set_nospace);
+
+/**
+ *	sk_clear_nospace - tell the transport no writer is waiting for space
+ *	@sk: socket
+ *
+ *	Called from ->sk_write_space() handlers, once send space has been
+ *	made available to writers.
+ */
+void sk_clear_nospace(struct sock *sk)
+{
+	struct socket *sock = sk->sk_socket;
+
+	if (sock)
+		clear_bit(SOCK_NOSPACE, &sock->flags);
+}
+EXPORT_SYMBOL(sk_clear_nospace);
+
 /* It is almost wait_for_tcp_memory minus release_sock/lock_sock.
    I think, these locks should be removed for datagram sockets.
  */
@@ -2970,7 +3003,7 @@ static long sock_wait_for_wmem(struct sock *sk, long timeo)
 			break;
 		if (signal_pending(current))
 			break;
-		set_bit(SOCK_NOSPACE, &sk->sk_socket->flags);
+		sk_set_nospace(sk);
 		prepare_to_wait(sk_sleep(sk), &wait, TASK_INTERRUPTIBLE);
 		if (refcount_read(&sk->sk_wmem_alloc) < READ_ONCE(sk->sk_sndbuf))
 			break;
@@ -3011,7 +3044,7 @@ struct sk_buff *sock_alloc_send_pskb(struct sock *sk, unsigned long header_len,
 			break;
 
 		sk_set_bit(SOCKWQ_ASYNC_NOSPACE, sk);
-		set_bit(SOCK_NOSPACE, &sk->sk_socket->flags);
+		sk_set_nospace(sk);
 		err = -EAGAIN;
 		if (!timeo)
 			goto failure;
diff --git a/net/core/stream.c b/net/core/stream.c
index 2d748581862d0eda8523f6f371cd2402e1dd6e01..a853b60afdc35a4735c017eb16c748cf2fab1b99 100644
--- a/net/core/stream.c
+++ b/net/core/stream.c
@@ -37,7 +37,7 @@ void sk_stream_write_space(struct sock *sk)
 	struct socket_wq *wq;
 
 	if (__sk_stream_is_writeable(sk, 1) && sock) {
-		clear_bit(SOCK_NOSPACE, &sock->flags);
+		sk_clear_nospace(sk);
 
 		rcu_read_lock();
 		wq = rcu_dereference(sk->sk_wq);
@@ -143,7 +143,7 @@ int sk_stream_wait_memory(struct sock *sk, long *timeo_p)
 		if (sk_stream_memory_free(sk) && !vm_wait)
 			break;
 
-		set_bit(SOCK_NOSPACE, &sk->sk_socket->flags);
+		sk_set_nospace(sk);
 		sk->sk_write_pending++;
 		ret = sk_wait_event(sk, &current_timeo, READ_ONCE(sk->sk_err) ||
 				    (READ_ONCE(sk->sk_shutdown) & SEND_SHUTDOWN) ||
@@ -177,7 +177,7 @@ int sk_stream_wait_memory(struct sock *sk, long *timeo_p)
 	 * When TCP receives ACK packets that make room, tcp_check_space()
 	 * only calls tcp_new_space() if SOCK_NOSPACE is set.
 	 */
-	set_bit(SOCK_NOSPACE, &sk->sk_socket->flags);
+	sk_set_nospace(sk);
 	err = -EAGAIN;
 	goto out;
 do_interrupted:
diff --git a/net/ipv4/tcp.c b/net/ipv4/tcp.c
index 3ac4856852794736c5d49f042ecd08e4246bdd6d..1cde000cfab4704e6756872f6ddec16851ccc55d 100644
--- a/net/ipv4/tcp.c
+++ b/net/ipv4/tcp.c
@@ -607,7 +607,7 @@ __poll_t tcp_poll(struct file *file, struct socket *sock, poll_table *wait)
 				mask |= EPOLLOUT | EPOLLWRNORM;
 			} else {  /* send SIGIO later */
 				sk_set_bit(SOCKWQ_ASYNC_NOSPACE, sk);
-				set_bit(SOCK_NOSPACE, &sk->sk_socket->flags);
+				sk_set_nospace(sk);
 
 				/* Race breaker. If space is freed after
 				 * wspace test but before the flags are set,
@@ -1398,7 +1398,7 @@ int tcp_sendmsg_locked(struct sock *sk, struct msghdr *msg, size_t size)
 		continue;
 
 wait_for_space:
-		set_bit(SOCK_NOSPACE, &sk->sk_socket->flags);
+		sk_set_nospace(sk);
 		tcp_remove_empty_skb(sk);
 		if (copied)
 			tcp_push(sk, flags & ~MSG_MORE, mss_now,
diff --git a/net/ipv4/tcp_bpf.c b/net/ipv4/tcp_bpf.c
index 2e234d155b5e616d496611d1f367c34ed90f3000..64d74d92af52d1ca2fa3d24910a254a5092466f2 100644
--- a/net/ipv4/tcp_bpf.c
+++ b/net/ipv4/tcp_bpf.c
@@ -602,7 +602,7 @@ static int tcp_bpf_sendmsg(struct sock *sk, struct msghdr *msg, size_t size)
 			goto out_err;
 		continue;
 wait_for_sndbuf:
-		set_bit(SOCK_NOSPACE, &sk->sk_socket->flags);
+		sk_set_nospace(sk);
 wait_for_memory:
 		err = sk_stream_wait_memory(sk, &timeo);
 		if (err) {
diff --git a/net/kcm/kcmsock.c b/net/kcm/kcmsock.c
index 71af69d442f211996e514d8d76a4c9d658f7b352..b29c0ed9caf5220ff741584e49632da45626beb7 100644
--- a/net/kcm/kcmsock.c
+++ b/net/kcm/kcmsock.c
@@ -735,7 +735,7 @@ static void kcm_tx_work(struct work_struct *w)
 	/* Primarily for SOCK_SEQPACKET sockets */
 	if (likely(sk->sk_socket) &&
 	    test_bit(SOCK_NOSPACE, &sk->sk_socket->flags)) {
-		clear_bit(SOCK_NOSPACE, &sk->sk_socket->flags);
+		sk_clear_nospace(sk);
 		sk->sk_write_space(sk);
 	}
 
@@ -779,7 +779,7 @@ static int kcm_sendmsg(struct socket *sock, struct msghdr *msg, size_t len)
 	/* Call the sk_stream functions to manage the sndbuf mem. */
 	if (!sk_stream_memory_free(sk)) {
 		kcm_push(kcm);
-		set_bit(SOCK_NOSPACE, &sk->sk_socket->flags);
+		sk_set_nospace(sk);
 		err = sk_stream_wait_memory(sk, &timeo);
 		if (err)
 			goto out_error;
diff --git a/net/mptcp/protocol.c b/net/mptcp/protocol.c
index e89a69ab927c9139c68c9039327cb0e55c33356b..74b1a512072878a54b229bc80aa76bb3e417c862 100644
--- a/net/mptcp/protocol.c
+++ b/net/mptcp/protocol.c
@@ -2106,7 +2106,7 @@ static int mptcp_sendmsg(struct sock *sk, struct msghdr *msg, size_t len)
 		continue;
 
 wait_for_memory:
-		set_bit(SOCK_NOSPACE, &sk->sk_socket->flags);
+		sk_set_nospace(sk);
 		__mptcp_push_pending(sk, msg->msg_flags);
 		ret = sk_stream_wait_memory(sk, &timeo);
 		if (ret)
@@ -4472,7 +4472,7 @@ static __poll_t mptcp_check_writeable(struct mptcp_sock *msk)
 	if (__mptcp_stream_is_writeable(sk, 1))
 		return EPOLLOUT | EPOLLWRNORM;
 
-	set_bit(SOCK_NOSPACE, &sk->sk_socket->flags);
+	sk_set_nospace(sk);
 	smp_mb__after_atomic(); /* NOSPACE is changed by mptcp_write_space() */
 	if (__mptcp_stream_is_writeable(sk, 1))
 		return EPOLLOUT | EPOLLWRNORM;
diff --git a/net/smc/af_smc.c b/net/smc/af_smc.c
index e9f93b3ab435bbcce0a493f24205bd91bdf15b6a..4f2afbbdef2d7e62431e14562dc26bdfcdaa9f79 100644
--- a/net/smc/af_smc.c
+++ b/net/smc/af_smc.c
@@ -2919,7 +2919,7 @@ __poll_t smc_poll(struct file *file, struct socket *sock,
 				mask |= EPOLLOUT | EPOLLWRNORM;
 			} else {
 				sk_set_bit(SOCKWQ_ASYNC_NOSPACE, sk);
-				set_bit(SOCK_NOSPACE, &sk->sk_socket->flags);
+				sk_set_nospace(sk);
 
 				if (sk->sk_state != SMC_INIT) {
 					/* Race breaker the same way as tcp_poll(). */
diff --git a/net/smc/smc_tx.c b/net/smc/smc_tx.c
index 3144b4b1fe29013cabc63d72671d256761a4c367..52e5395484fd3e0c4b1356226ecdcf4fb190d487 100644
--- a/net/smc/smc_tx.c
+++ b/net/smc/smc_tx.c
@@ -48,7 +48,7 @@ static void smc_tx_write_space(struct sock *sk)
 	if (atomic_read(&smc->conn.sndbuf_space) && sock) {
 		if (test_bit(SOCK_NOSPACE, &sock->flags))
 			SMC_STAT_RMB_TX_FULL(smc, !smc->conn.lnk);
-		clear_bit(SOCK_NOSPACE, &sock->flags);
+		sk_clear_nospace(sk);
 		rcu_read_lock();
 		wq = rcu_dereference(sk->sk_wq);
 		if (skwq_has_sleeper(wq))
@@ -100,7 +100,7 @@ static int smc_tx_wait(struct smc_sock *smc, int flags)
 		}
 		if (!timeo) {
 			/* ensure EPOLLOUT is subsequently generated */
-			set_bit(SOCK_NOSPACE, &sk->sk_socket->flags);
+			sk_set_nospace(sk);
 			rc = -EAGAIN;
 			break;
 		}
@@ -111,7 +111,7 @@ static int smc_tx_wait(struct smc_sock *smc, int flags)
 		sk_clear_bit(SOCKWQ_ASYNC_NOSPACE, sk);
 		if (atomic_read(&conn->sndbuf_space) && !conn->urg_tx_pend)
 			break; /* at least 1 byte of free & no urgent data */
-		set_bit(SOCK_NOSPACE, &sk->sk_socket->flags);
+		sk_set_nospace(sk);
 		sk_wait_event(sk, &timeo,
 			      READ_ONCE(sk->sk_err) ||
 			      (READ_ONCE(sk->sk_shutdown) & SEND_SHUTDOWN) ||
diff --git a/net/tls/tls_sw.c b/net/tls/tls_sw.c
index d1ad31986cf2cee88afbde43a2908791eeabb0fb..312e51270f293dc4df8efbec498473c4c3b48d15 100644
--- a/net/tls/tls_sw.c
+++ b/net/tls/tls_sw.c
@@ -966,7 +966,7 @@ static int tls_sw_sendmsg_locked(struct sock *sk, struct msghdr *msg,
 		continue;
 
 wait_for_sndbuf:
-		set_bit(SOCK_NOSPACE, &sk->sk_socket->flags);
+		sk_set_nospace(sk);
 wait_for_memory:
 		ret = sk_stream_wait_memory(sk, &timeo);
 		if (ret) {
-- 
2.56.0.rc1.310.g51773c2048-goog


^ permalink raw reply related	[flat|nested] 13+ messages in thread

* [PATCH v2 net-next 3/9] sunrpc: use sk_set_nospace() and sk_clear_nospace()
  2026-09-24 13:47 [PATCH v2 net-next 0/9] tcp: avoid struct socket cache line miss in tcp_check_space() Eric Dumazet
  2026-09-24 13:47 ` [PATCH v2 net-next 1/9] dlm: fix send buffer backpressure handling Eric Dumazet
  2026-09-24 13:47 ` [PATCH v2 net-next 2/9] net: add sk_set_nospace() and sk_clear_nospace() Eric Dumazet
@ 2026-09-24 13:47 ` Eric Dumazet
  2026-09-24 13:47 ` [PATCH v2 net-next 4/9] rds: use sk_set_nospace() Eric Dumazet
                   ` (6 subsequent siblings)
  9 siblings, 0 replies; 13+ messages in thread
From: Eric Dumazet @ 2026-09-24 13:47 UTC (permalink / raw)
  To: David S . Miller, Jakub Kicinski, Paolo Abeni
  Cc: Simon Horman, Neal Cardwell, Kuniyuki Iwashima, Willem de Bruijn,
	netdev, eric.dumazet, Eric Dumazet

Use the new helpers instead of open coding the SOCK_NOSPACE
manipulation, so that TCP can later maintain a cheaper private
copy of this bit.

No functional change intended.

Reviewed-by: Kuniyuki Iwashima <kuniyu@google.com>
Signed-off-by: Eric Dumazet <edumazet@google.com>
---
 net/sunrpc/svcsock.c  | 4 ++--
 net/sunrpc/xprtsock.c | 4 ++--
 2 files changed, 4 insertions(+), 4 deletions(-)

diff --git a/net/sunrpc/svcsock.c b/net/sunrpc/svcsock.c
index 50e5e7f5b762de28c34d0f58cb0c6feb1e3ce79f..2d8cdee0798d540ba8eae589deaac5bd3ff889fa 100644
--- a/net/sunrpc/svcsock.c
+++ b/net/sunrpc/svcsock.c
@@ -792,11 +792,11 @@ static int svc_udp_has_wspace(struct svc_xprt *xprt)
 	 * Set the SOCK_NOSPACE flag before checking the available
 	 * sock space.
 	 */
-	set_bit(SOCK_NOSPACE, &svsk->sk_sock->flags);
+	sk_set_nospace(svsk->sk_sk);
 	required = atomic_read(&svsk->sk_xprt.xpt_reserved) + serv->sv_max_mesg;
 	if (required*2 > sock_wspace(svsk->sk_sk))
 		return 0;
-	clear_bit(SOCK_NOSPACE, &svsk->sk_sock->flags);
+	sk_clear_nospace(svsk->sk_sk);
 	return 1;
 }
 
diff --git a/net/sunrpc/xprtsock.c b/net/sunrpc/xprtsock.c
index 7f60723fa64d84887e260e5e130e44cdbec6f850..b97c70c12f627630509510facfd3c3cf96ab6cb6 100644
--- a/net/sunrpc/xprtsock.c
+++ b/net/sunrpc/xprtsock.c
@@ -858,7 +858,7 @@ static int xs_nospace(struct rpc_rqst *req, struct sock_xprt *transport)
 	if (xprt_connected(xprt)) {
 		/* wait for more buffer space */
 		set_bit(XPRT_SOCK_NOSPACE, &transport->sock_state);
-		set_bit(SOCK_NOSPACE, &sk->sk_socket->flags);
+		sk_set_nospace(sk);
 		sk->sk_write_pending++;
 		xprt_wait_for_buffer_space(xprt);
 	} else
@@ -1615,7 +1615,7 @@ static void xs_write_space(struct sock *sk)
 
 	if (!sk->sk_socket)
 		return;
-	clear_bit(SOCK_NOSPACE, &sk->sk_socket->flags);
+	sk_clear_nospace(sk);
 
 	if (unlikely(!(xprt = xprt_from_sock(sk))))
 		return;
-- 
2.56.0.rc1.310.g51773c2048-goog


^ permalink raw reply related	[flat|nested] 13+ messages in thread

* [PATCH v2 net-next 4/9] rds: use sk_set_nospace()
  2026-09-24 13:47 [PATCH v2 net-next 0/9] tcp: avoid struct socket cache line miss in tcp_check_space() Eric Dumazet
                   ` (2 preceding siblings ...)
  2026-09-24 13:47 ` [PATCH v2 net-next 3/9] sunrpc: use " Eric Dumazet
@ 2026-09-24 13:47 ` Eric Dumazet
  2026-09-24 13:47 ` [PATCH v2 net-next 5/9] dlm: use sk_set_nospace() and sk_clear_nospace() Eric Dumazet
                   ` (5 subsequent siblings)
  9 siblings, 0 replies; 13+ messages in thread
From: Eric Dumazet @ 2026-09-24 13:47 UTC (permalink / raw)
  To: David S . Miller, Jakub Kicinski, Paolo Abeni
  Cc: Simon Horman, Neal Cardwell, Kuniyuki Iwashima, Willem de Bruijn,
	netdev, eric.dumazet, Eric Dumazet

Use the new helpers instead of open coding the SOCK_NOSPACE
manipulation, so that TCP can later maintain a cheaper private
copy of this bit.

No functional change intended.

Reviewed-by: Kuniyuki Iwashima <kuniyu@google.com>
Signed-off-by: Eric Dumazet <edumazet@google.com>
---
 net/rds/tcp_send.c | 5 ++---
 1 file changed, 2 insertions(+), 3 deletions(-)

diff --git a/net/rds/tcp_send.c b/net/rds/tcp_send.c
index 7c52acc749cf4d4003c749a2646c2b4f87f207fe..824222c475119d91e333123c5afc9291005a8696 100644
--- a/net/rds/tcp_send.c
+++ b/net/rds/tcp_send.c
@@ -100,7 +100,7 @@ int rds_tcp_xmit(struct rds_connection *conn, struct rds_message *rm,
 
 	if (hdr_off < sizeof(struct rds_header)) {
 		/* see rds_tcp_write_space() */
-		set_bit(SOCK_NOSPACE, &tc->t_sock->sk->sk_socket->flags);
+		sk_set_nospace(tc->t_sock->sk);
 
 		ret = rds_tcp_sendmsg(tc->t_sock,
 				      (void *)&rm->m_inc.i_hdr + hdr_off,
@@ -221,6 +221,5 @@ void rds_tcp_write_space(struct sock *sk)
 	 */
 	write_space(sk);
 
-	if (sk->sk_socket)
-		set_bit(SOCK_NOSPACE, &sk->sk_socket->flags);
+	sk_set_nospace(sk);
 }
-- 
2.56.0.rc1.310.g51773c2048-goog


^ permalink raw reply related	[flat|nested] 13+ messages in thread

* [PATCH v2 net-next 5/9] dlm: use sk_set_nospace() and sk_clear_nospace()
  2026-09-24 13:47 [PATCH v2 net-next 0/9] tcp: avoid struct socket cache line miss in tcp_check_space() Eric Dumazet
                   ` (3 preceding siblings ...)
  2026-09-24 13:47 ` [PATCH v2 net-next 4/9] rds: use sk_set_nospace() Eric Dumazet
@ 2026-09-24 13:47 ` Eric Dumazet
  2026-09-24 13:47 ` [PATCH v2 net-next 6/9] drbd: use sk_set_nospace() Eric Dumazet
                   ` (4 subsequent siblings)
  9 siblings, 0 replies; 13+ messages in thread
From: Eric Dumazet @ 2026-09-24 13:47 UTC (permalink / raw)
  To: David S . Miller, Jakub Kicinski, Paolo Abeni
  Cc: Simon Horman, Neal Cardwell, Kuniyuki Iwashima, Willem de Bruijn,
	netdev, eric.dumazet, Eric Dumazet

Use the new helpers instead of open coding the SOCK_NOSPACE
manipulation, so that TCP can later maintain a cheaper private
copy of this bit.

No functional change intended.

Reviewed-by: Kuniyuki Iwashima <kuniyu@google.com>
Signed-off-by: Eric Dumazet <edumazet@google.com>
---
 fs/dlm/lowcomms.c | 4 ++--
 1 file changed, 2 insertions(+), 2 deletions(-)

diff --git a/fs/dlm/lowcomms.c b/fs/dlm/lowcomms.c
index abe9ae4c643fae061d2e168c03219a4f4e1bc007..6aa68a24acab574429664acc3f5ecc25ddd1553d 100644
--- a/fs/dlm/lowcomms.c
+++ b/fs/dlm/lowcomms.c
@@ -519,7 +519,7 @@ static void lowcomms_write_space(struct sock *sk)
 {
 	struct connection *con = sock2con(sk);
 
-	clear_bit(SOCK_NOSPACE, &con->sock->flags);
+	sk_clear_nospace(sk);
 
 	spin_lock_bh(&con->writequeue_lock);
 	if (test_and_clear_bit(CF_APP_LIMITED, &con->flags))
@@ -1394,7 +1394,7 @@ static int send_to_sock(struct connection *con)
 			/* Notify TCP that we're limited by the
 			 * application window size.
 			 */
-			set_bit(SOCK_NOSPACE, &con->sock->sk->sk_socket->flags);
+			sk_set_nospace(con->sock->sk);
 			con->sock->sk->sk_write_pending++;
 
 			clear_bit(CF_SEND_PENDING, &con->flags);
-- 
2.56.0.rc1.310.g51773c2048-goog


^ permalink raw reply related	[flat|nested] 13+ messages in thread

* [PATCH v2 net-next 6/9] drbd: use sk_set_nospace()
  2026-09-24 13:47 [PATCH v2 net-next 0/9] tcp: avoid struct socket cache line miss in tcp_check_space() Eric Dumazet
                   ` (4 preceding siblings ...)
  2026-09-24 13:47 ` [PATCH v2 net-next 5/9] dlm: use sk_set_nospace() and sk_clear_nospace() Eric Dumazet
@ 2026-09-24 13:47 ` Eric Dumazet
  2026-09-24 13:47 ` [PATCH v2 net-next 7/9] nvme-tcp: use sk_clear_nospace() Eric Dumazet
                   ` (3 subsequent siblings)
  9 siblings, 0 replies; 13+ messages in thread
From: Eric Dumazet @ 2026-09-24 13:47 UTC (permalink / raw)
  To: David S . Miller, Jakub Kicinski, Paolo Abeni
  Cc: Simon Horman, Neal Cardwell, Kuniyuki Iwashima, Willem de Bruijn,
	netdev, eric.dumazet, Eric Dumazet

Use the new helpers instead of open coding the SOCK_NOSPACE
manipulation, so that TCP can later maintain a cheaper private
copy of this bit.

No functional change intended.

Reviewed-by: Kuniyuki Iwashima <kuniyu@google.com>
Signed-off-by: Eric Dumazet <edumazet@google.com>
---
 drivers/block/drbd/drbd_worker.c | 3 +--
 1 file changed, 1 insertion(+), 2 deletions(-)

diff --git a/drivers/block/drbd/drbd_worker.c b/drivers/block/drbd/drbd_worker.c
index 0697f99fed18d7e70f00b7c559a131517ee8aa42..8e50131b5a5692a8dfb99eafe706602153be94b7 100644
--- a/drivers/block/drbd/drbd_worker.c
+++ b/drivers/block/drbd/drbd_worker.c
@@ -635,8 +635,7 @@ static int make_resync_request(struct drbd_peer_device *const peer_device, int c
 			int sndbuf = sk->sk_sndbuf;
 			if (queued > sndbuf / 2) {
 				requeue = 1;
-				if (sk->sk_socket)
-					set_bit(SOCK_NOSPACE, &sk->sk_socket->flags);
+				sk_set_nospace(sk);
 			}
 		} else
 			requeue = 1;
-- 
2.56.0.rc1.310.g51773c2048-goog


^ permalink raw reply related	[flat|nested] 13+ messages in thread

* [PATCH v2 net-next 7/9] nvme-tcp: use sk_clear_nospace()
  2026-09-24 13:47 [PATCH v2 net-next 0/9] tcp: avoid struct socket cache line miss in tcp_check_space() Eric Dumazet
                   ` (5 preceding siblings ...)
  2026-09-24 13:47 ` [PATCH v2 net-next 6/9] drbd: use sk_set_nospace() Eric Dumazet
@ 2026-09-24 13:47 ` Eric Dumazet
  2026-09-24 13:47 ` [PATCH v2 net-next 8/9] libceph: " Eric Dumazet
                   ` (2 subsequent siblings)
  9 siblings, 0 replies; 13+ messages in thread
From: Eric Dumazet @ 2026-09-24 13:47 UTC (permalink / raw)
  To: David S . Miller, Jakub Kicinski, Paolo Abeni
  Cc: Simon Horman, Neal Cardwell, Kuniyuki Iwashima, Willem de Bruijn,
	netdev, eric.dumazet, Eric Dumazet

Use the new helpers instead of open coding the SOCK_NOSPACE
manipulation, so that TCP can later maintain a cheaper private
copy of this bit.

No functional change intended.

Reviewed-by: Kuniyuki Iwashima <kuniyu@google.com>
Signed-off-by: Eric Dumazet <edumazet@google.com>
---
 drivers/nvme/host/tcp.c   | 2 +-
 drivers/nvme/target/tcp.c | 2 +-
 2 files changed, 2 insertions(+), 2 deletions(-)

diff --git a/drivers/nvme/host/tcp.c b/drivers/nvme/host/tcp.c
index 921934028e0b22c2e4c9fd60d09ed8e50051262d..b853ca2d598fa4cb8ea9cb66513f1e1070cd1fb4 100644
--- a/drivers/nvme/host/tcp.c
+++ b/drivers/nvme/host/tcp.c
@@ -1133,7 +1133,7 @@ static void nvme_tcp_write_space(struct sock *sk)
 	read_lock_bh(&sk->sk_callback_lock);
 	queue = sk->sk_user_data;
 	if (likely(queue && sk_stream_is_writeable(sk))) {
-		clear_bit(SOCK_NOSPACE, &sk->sk_socket->flags);
+		sk_clear_nospace(sk);
 		/* Ensure pending TLS partial records are retried */
 		if (nvme_tcp_queue_tls(queue))
 			queue->write_space(sk);
diff --git a/drivers/nvme/target/tcp.c b/drivers/nvme/target/tcp.c
index e598101752620448423925d2385b606e0b4346e8..c38ea3501ea5def168b6f180b56734b99d21bdea 100644
--- a/drivers/nvme/target/tcp.c
+++ b/drivers/nvme/target/tcp.c
@@ -1686,7 +1686,7 @@ static void nvmet_tcp_write_space(struct sock *sk)
 	}
 
 	if (sk_stream_is_writeable(sk)) {
-		clear_bit(SOCK_NOSPACE, &sk->sk_socket->flags);
+		sk_clear_nospace(sk);
 		queue_work_on(queue_cpu(queue), nvmet_tcp_wq, &queue->io_work);
 	}
 out:
-- 
2.56.0.rc1.310.g51773c2048-goog


^ permalink raw reply related	[flat|nested] 13+ messages in thread

* [PATCH v2 net-next 8/9] libceph: use sk_clear_nospace()
  2026-09-24 13:47 [PATCH v2 net-next 0/9] tcp: avoid struct socket cache line miss in tcp_check_space() Eric Dumazet
                   ` (6 preceding siblings ...)
  2026-09-24 13:47 ` [PATCH v2 net-next 7/9] nvme-tcp: use sk_clear_nospace() Eric Dumazet
@ 2026-09-24 13:47 ` Eric Dumazet
  2026-09-24 13:47 ` [PATCH v2 net-next 9/9] tcp: add tp->tcp_nospace Eric Dumazet
  2026-09-29  1:18 ` [PATCH v2 net-next 0/9] tcp: avoid struct socket cache line miss in tcp_check_space() Jakub Kicinski
  9 siblings, 0 replies; 13+ messages in thread
From: Eric Dumazet @ 2026-09-24 13:47 UTC (permalink / raw)
  To: David S . Miller, Jakub Kicinski, Paolo Abeni
  Cc: Simon Horman, Neal Cardwell, Kuniyuki Iwashima, Willem de Bruijn,
	netdev, eric.dumazet, Eric Dumazet

Use the new helpers instead of open coding the SOCK_NOSPACE
manipulation, so that TCP can later maintain a cheaper private
copy of this bit.

No functional change intended.

Reviewed-by: Kuniyuki Iwashima <kuniyu@google.com>
Signed-off-by: Eric Dumazet <edumazet@google.com>
---
 net/ceph/messenger.c | 2 +-
 1 file changed, 1 insertion(+), 1 deletion(-)

diff --git a/net/ceph/messenger.c b/net/ceph/messenger.c
index 212e7797f9e48edde4a4536bba6f1ff293628b5b..cc7db8c61db37a998b698566caa7c936fa83fe4a 100644
--- a/net/ceph/messenger.c
+++ b/net/ceph/messenger.c
@@ -375,7 +375,7 @@ static void ceph_sock_write_space(struct sock *sk)
 	if (ceph_con_flag_test(con, CEPH_CON_F_WRITE_PENDING)) {
 		if (sk_stream_is_writeable(sk)) {
 			dout("%s %p queueing write work\n", __func__, con);
-			clear_bit(SOCK_NOSPACE, &sk->sk_socket->flags);
+			sk_clear_nospace(sk);
 			queue_con(con);
 		}
 	} else {
-- 
2.56.0.rc1.310.g51773c2048-goog


^ permalink raw reply related	[flat|nested] 13+ messages in thread

* [PATCH v2 net-next 9/9] tcp: add tp->tcp_nospace
  2026-09-24 13:47 [PATCH v2 net-next 0/9] tcp: avoid struct socket cache line miss in tcp_check_space() Eric Dumazet
                   ` (7 preceding siblings ...)
  2026-09-24 13:47 ` [PATCH v2 net-next 8/9] libceph: " Eric Dumazet
@ 2026-09-24 13:47 ` Eric Dumazet
  2026-09-29  1:18 ` [PATCH v2 net-next 0/9] tcp: avoid struct socket cache line miss in tcp_check_space() Jakub Kicinski
  9 siblings, 0 replies; 13+ messages in thread
From: Eric Dumazet @ 2026-09-24 13:47 UTC (permalink / raw)
  To: David S . Miller, Jakub Kicinski, Paolo Abeni
  Cc: Simon Horman, Neal Cardwell, Kuniyuki Iwashima, Willem de Bruijn,
	netdev, eric.dumazet, Eric Dumazet

tcp_check_space() runs for every incoming ACK and for every packet we
send, and reads SOCK_NOSPACE from sk->sk_socket->flags.

struct socket lives in its own cache line, which the TCP fast paths do
not otherwise touch, so testing sk->sk_socket->flags pulls in an extra
cache line that is cold when the working set is large.

Add tp->tcp_nospace, a mirror of SOCK_NOSPACE placed right after
tp->chrono_type in the tcp_sock_write_tx group, in a cache line both
the transmit and the ACK paths already touch (bytes_sent, data_segs_out,
delivered, bytes_acked, chrono_type).  It fits in an existing 3-byte
hole before chrono_start, so no other field moves and sizeof(struct
tcp_sock) is unchanged.

The two flags are now only changed from sk_set_nospace() and
sk_clear_nospace(), which maintain this invariant:

	SOCK_NOSPACE set  =>  tp->tcp_nospace set

sk_set_nospace() sets SOCK_NOSPACE before tp->tcp_nospace (using
smp_mb__after_atomic() and smp_store_mb()), while sk_clear_nospace()
clears tp->tcp_nospace before SOCK_NOSPACE (using
smp_mb__before_atomic()).  If a lockless sk_set_nospace() from
tcp_poll() races with sk_clear_nospace() and SOCK_NOSPACE ends up set,
set_bit(SOCK_NOSPACE) happened after clear_bit(SOCK_NOSPACE), so
tp->tcp_nospace = 1 happens after tp->tcp_nospace = 0 as well.

tcp_check_space() can thus test tp->tcp_nospace alone and leave the
authoritative SOCK_NOSPACE test to __tcp_check_space().  The
invariant is one directional on purpose: a stale tp->tcp_nospace only
costs an extra call to __tcp_check_space(), which is what we do
unconditionally today, while a stale SOCK_NOSPACE would cost a
missed EPOLLOUT.

Note the smp_mb() is kept.  tcp_poll() sets the flag without the
socket lock, and the store-buffer pattern it forms with
tcp_check_space() needs a full barrier on both sides.

MPTCP subflows share the struct socket of their parent, hence its
SOCK_NOSPACE bit, which can not be mirrored in the subflow tcp_sock.
Pin their tp->tcp_nospace in subflow_ulp_init(), so that they always
reach __tcp_check_space() and keep the current behavior.

Microbenchmark on an AMD EPYC 7B13, 64 threads, each thread calling
tcp_check_space() in a loop over a private set of sockets.  Numbers
are cycles per call above a baseline loop that performs the work the
callers already did (the cache lines tcp_write_xmit() and
tcp_clean_rtx_queue() touched) but not tcp_check_space() itself, so
they are the marginal cost of the function.  Median of 11 runs.
The set size controls whether struct socket is still cached:

    sockets/thread	before	 after
    1			  +0.6	  +0.7
    256    (64 KB)	  +8.3	  +3.9
    4096   (1 MB)	 +22.4	  +1.1
    262144 (64 MB)	 +48.4	  -0.8

Signed-off-by: Eric Dumazet <edumazet@google.com>
---
 .../networking/net_cachelines/tcp_sock.rst    |  1 +
 include/linux/tcp.h                           |  3 ++
 include/net/tcp.h                             | 30 ++++++++++++++++++-
 net/core/sock.c                               | 17 ++++++++---
 net/ipv4/tcp.c                                |  1 +
 net/ipv4/tcp_input.c                          |  8 ++++-
 net/mptcp/subflow.c                           |  5 ++++
 7 files changed, 59 insertions(+), 6 deletions(-)

diff --git a/Documentation/networking/net_cachelines/tcp_sock.rst b/Documentation/networking/net_cachelines/tcp_sock.rst
index 0f6088c4ab8bb872e7fc86f02479592e84c0247a..103328dc409e4485a3f989c95348c1bc74ab20c6 100644
--- a/Documentation/networking/net_cachelines/tcp_sock.rst
+++ b/Documentation/networking/net_cachelines/tcp_sock.rst
@@ -50,6 +50,7 @@ u8:1                          tcp_usec_ts             read_mostly         read_m
 u32                           chrono_start            read_write                              tcp_chrono_start/stop(tcp_write_xmit,tcp_cwnd_validate,tcp_send_syn_data)
 u32[3]                        chrono_stat             read_write                              tcp_chrono_start/stop(tcp_write_xmit,tcp_cwnd_validate,tcp_send_syn_data)
 u8:2                          chrono_type             read_write                              tcp_chrono_start/stop(tcp_write_xmit,tcp_cwnd_validate,tcp_send_syn_data)
+u8                            tcp_nospace             read_mostly         read_mostly         tcp_check_space(tx);tcp_check_space(rx)
 u8:1                          rate_app_limited                            read_write          tcp_rate_gen
 u8:1                          fastopen_connect
 u8:1                          fastopen_no_cookie
diff --git a/include/linux/tcp.h b/include/linux/tcp.h
index 6a8c77719322f9caee305d954a107892c76d4ef7..e4d1720f49f08f3c12f5ba66fe8fe3379d71b1e0 100644
--- a/include/linux/tcp.h
+++ b/include/linux/tcp.h
@@ -269,6 +269,9 @@ struct tcp_sock {
 				 */
 	u32	snd_sml;	/* Last byte of the most recently transmitted small packet */
 	u8	chrono_type;	/* current chronograph type */
+	u8	tcp_nospace;	/* mirrors SOCK_NOSPACE, must be set whenever
+				 * SOCK_NOSPACE is set.
+				 */
 	u32	chrono_start;	/* Start time in jiffies of a TCP chrono */
 	u32	chrono_stat[3];	/* Time in jiffies for chrono_stat stats */
 	u32	write_seq;	/* Tail(+1) of data held in tcp send buffer */
diff --git a/include/net/tcp.h b/include/net/tcp.h
index 5e5f5f9b89a386568fc5efebfa3d3c7e1ff62683..3389c51790e6b03714eda8fd89c05cf5c2999809 100644
--- a/include/net/tcp.h
+++ b/include/net/tcp.h
@@ -783,14 +783,42 @@ void tcp_done_with_error(struct sock *sk, int err);
 void tcp_reset(struct sock *sk, struct sk_buff *skb);
 void tcp_fin(struct sock *sk);
 void __tcp_check_space(struct sock *sk);
+
+/* Mirror of SOCK_NOSPACE in tcp_sock, maintained by sk_set_nospace()
+ * and sk_clear_nospace().
+ *
+ * MPTCP subflows share the parent socket, and thus its SOCK_NOSPACE bit.
+ * Keep their mirror always set (see subflow_ulp_init()) so that they
+ * always reach __tcp_check_space() and behave as before.
+ */
+static inline void tcp_set_nospace(struct sock *sk)
+{
+	if (sk_is_tcp(sk)) {
+		/* pairs with smp_mb__before_atomic() in tcp_clear_nospace() */
+		smp_mb__after_atomic();
+		/* pairs with smp_mb() in tcp_check_space() */
+		smp_store_mb(tcp_sk(sk)->tcp_nospace, 1);
+	}
+}
+
+static inline void tcp_clear_nospace(struct sock *sk)
+{
+	if (sk_is_tcp(sk) && !sk_is_mptcp(sk)) {
+		WRITE_ONCE(tcp_sk(sk)->tcp_nospace, 0);
+		/* pairs with smp_mb__after_atomic() in tcp_set_nospace() */
+		smp_mb__before_atomic();
+	}
+}
+
 static inline void tcp_check_space(struct sock *sk)
 {
 	/* pairs with tcp_poll() */
 	smp_mb();
 
-	if (sk->sk_socket && test_bit(SOCK_NOSPACE, &sk->sk_socket->flags))
+	if (unlikely(READ_ONCE(tcp_sk(sk)->tcp_nospace)))
 		__tcp_check_space(sk);
 }
+
 void tcp_sack_compress_send_ack(struct sock *sk);
 
 static inline void tcp_cleanup_skb(struct sk_buff *skb)
diff --git a/net/core/sock.c b/net/core/sock.c
index 11a22aec7e414152aab115e8d11e30067ab3775f..946f614f665b3608d1f310da908d82de622ec5c0 100644
--- a/net/core/sock.c
+++ b/net/core/sock.c
@@ -2969,8 +2969,15 @@ void sk_set_nospace(struct sock *sk)
 {
 	struct socket *sock = sk->sk_socket;
 
-	if (sock)
-		set_bit(SOCK_NOSPACE, &sock->flags);
+	if (!sock)
+		return;
+	/* Set SOCK_NOSPACE before tp->tcp_nospace (paired with
+	 * sk_clear_nospace() clearing tp->tcp_nospace before SOCK_NOSPACE)
+	 * so a concurrent clear cannot leave SOCK_NOSPACE set with
+	 * tp->tcp_nospace cleared.
+	 */
+	set_bit(SOCK_NOSPACE, &sock->flags);
+	tcp_set_nospace(sk);
 }
 EXPORT_SYMBOL(sk_set_nospace);
 
@@ -2985,8 +2992,10 @@ void sk_clear_nospace(struct sock *sk)
 {
 	struct socket *sock = sk->sk_socket;
 
-	if (sock)
-		clear_bit(SOCK_NOSPACE, &sock->flags);
+	if (!sock)
+		return;
+	tcp_clear_nospace(sk);
+	clear_bit(SOCK_NOSPACE, &sock->flags);
 }
 EXPORT_SYMBOL(sk_clear_nospace);
 
diff --git a/net/ipv4/tcp.c b/net/ipv4/tcp.c
index 1cde000cfab4704e6756872f6ddec16851ccc55d..f44cbd3e76178a5ec5b40e420b32836477c29266 100644
--- a/net/ipv4/tcp.c
+++ b/net/ipv4/tcp.c
@@ -5242,6 +5242,7 @@ static void __init tcp_struct_check(void)
 	CACHELINE_ASSERT_GROUP_MEMBER(struct tcp_sock, tcp_sock_write_tx, first_tx_mstamp);
 	CACHELINE_ASSERT_GROUP_MEMBER(struct tcp_sock, tcp_sock_write_tx, delivered_mstamp);
 	CACHELINE_ASSERT_GROUP_MEMBER(struct tcp_sock, tcp_sock_write_tx, snd_sml);
+	CACHELINE_ASSERT_GROUP_MEMBER(struct tcp_sock, tcp_sock_write_tx, tcp_nospace);
 	CACHELINE_ASSERT_GROUP_MEMBER(struct tcp_sock, tcp_sock_write_tx, chrono_start);
 	CACHELINE_ASSERT_GROUP_MEMBER(struct tcp_sock, tcp_sock_write_tx, chrono_stat);
 	CACHELINE_ASSERT_GROUP_MEMBER(struct tcp_sock, tcp_sock_write_tx, write_seq);
diff --git a/net/ipv4/tcp_input.c b/net/ipv4/tcp_input.c
index 892ff256e235272a8483a11949cc98b319dd7cc9..cdea30d07fe2c2a38c3d38da27ad3274221d5958 100644
--- a/net/ipv4/tcp_input.c
+++ b/net/ipv4/tcp_input.c
@@ -6098,8 +6098,14 @@ static void tcp_new_space(struct sock *sk)
  */
 void __tcp_check_space(struct sock *sk)
 {
+	struct socket *sock = sk->sk_socket;
+
+	/* tp->tcp_nospace is only a hint, SOCK_NOSPACE is authoritative. */
+	if (!sock || !test_bit(SOCK_NOSPACE, &sock->flags))
+		return;
+
 	tcp_new_space(sk);
-	if (!test_bit(SOCK_NOSPACE, &sk->sk_socket->flags))
+	if (!test_bit(SOCK_NOSPACE, &sock->flags))
 		tcp_chrono_stop(sk, TCP_CHRONO_SNDBUF_LIMITED);
 }
 
diff --git a/net/mptcp/subflow.c b/net/mptcp/subflow.c
index f0a6725d2c3762def75c879e769e9da37f92215a..e297c88be503b42979886baaf43acfb6b7e78055 100644
--- a/net/mptcp/subflow.c
+++ b/net/mptcp/subflow.c
@@ -2001,6 +2001,11 @@ static int subflow_ulp_init(struct sock *sk)
 	pr_debug("subflow=%p, family=%d\n", ctx, sk->sk_family);
 
 	tp->is_mptcp = 1;
+	/* Subflows share the MPTCP socket, and thus its SOCK_NOSPACE bit,
+	 * which tcp_check_space() can not mirror. Pin the mirror so that
+	 * __tcp_check_space() always tests the shared bit.
+	 */
+	tp->tcp_nospace = 1;
 	ctx->icsk_af_ops = icsk->icsk_af_ops;
 	icsk->icsk_af_ops = subflow_default_af_ops(sk);
 	ctx->tcp_state_change = sk->sk_state_change;
-- 
2.56.0.rc1.310.g51773c2048-goog


^ permalink raw reply related	[flat|nested] 13+ messages in thread

* Re: [PATCH v2 net-next 1/9] dlm: fix send buffer backpressure handling
  2026-09-24 13:47 ` [PATCH v2 net-next 1/9] dlm: fix send buffer backpressure handling Eric Dumazet
@ 2026-09-25 13:49   ` netdev-bot+sashiko
  0 siblings, 0 replies; 13+ messages in thread
From: netdev-bot+sashiko @ 2026-09-25 13:49 UTC (permalink / raw)
  To: edumazet
  Cc: davem, kuba, pabeni, horms, ncardwell, kuniyu, willemb, netdev,
	eric.dumazet

Thank you for your contribution! Sashiko AI review found 1 potential
issue(s) to consider.

Critical: 0 · High: 1 · Medium: 0 · Low: 0

If you believe none of the issues are legitimate, please reply to each
issue of Medium severity or higher and explain why the AI is wrong.
If any of them are legitimate and you plan to respin, please let patchwork
know by including "pw-bot: cr" as a separate line at the end of your reply
(one such reply per series is enough).

- [High] This patch turns on the CF_APP_LIMITED / DLM_IO_END wait in
  send_to_sock() (fs/dlm/lowcomms.c).

--

dlm: fix send buffer backpressure handling

This patch changes send_to_sock() to test SOCK_NOSPACE instead of
SOCKWQ_ASYNC_NOSPACE. That bit has not lived in socket->flags since it
moved to socket_wq->flags.

With the patch, a TCP -EAGAIN sets CF_APP_LIMITED and returns
DLM_IO_END to wait for lowcomms_write_space(). Before, the send work was
requeued right away.

The patch also removes the stale clear_bit() of SOCKWQ_ASYNC_NOSPACE from
lowcomms_write_space().

> diff --git a/fs/dlm/lowcomms.c b/fs/dlm/lowcomms.c
> index 2aff1c7c17de4..abe9ae4c643fa 100644
> --- a/fs/dlm/lowcomms.c
> +++ b/fs/dlm/lowcomms.c

[ ... ]

> @@ -1391,7 +1389,7 @@ static int send_to_sock(struct connection *con)
>  	if (ret == -EAGAIN || ret == 0) {
>  		lock_sock(con->sock->sk);
>  		spin_lock_bh(&con->writequeue_lock);
> -		if (test_bit(SOCKWQ_ASYNC_NOSPACE, &con->sock->flags) &&
> +		if (test_bit(SOCK_NOSPACE, &con->sock->flags) &&
>  		    !test_and_set_bit(CF_APP_LIMITED, &con->flags)) {
>  			/* Notify TCP that we're limited by the
>  			 * application window size.

[Severity: High]
If the TCP connection is aborted while CF_APP_LIMITED is set, can the
send worker for that node stay stuck for good?

With this change the CF_APP_LIMITED / DLM_IO_END path is now live for
TCP. Once the flag is set, only lowcomms_write_space() or
close_connection() clear it.

A connection can be aborted by an RST from the peer (tcp_reset()), or by
tcp_write_err() after a retransmit, zero window probe or keepalive
timeout. All of these end up in tcp_done_with_error():

net/ipv4/tcp_input.c:tcp_done_with_error() {
    ...
	WRITE_ONCE(sk->sk_err, err);
	smp_wmb();
	tcp_write_queue_purge(sk);
	tcp_done(sk);
	if (!sock_flag(sk, SOCK_DEAD))
		sk_error_report(sk);
    ...
}

That path never calls sk->sk_write_space(), so lowcomms_write_space()
does not run.

lowcomms_error_report() does run, and it resends like this:

lowcomms_error_report()
  dlm_midcomms_unack_msg_resend()
    dlm_lowcomms_resend_msg()
      _dlm_lowcomms_commit_msg()
        lowcomms_queue_swork()

However, lowcomms_queue_swork() does nothing while the flag is set:

	if (!test_bit(CF_IO_STOP, &con->flags) &&
	    !test_bit(CF_APP_LIMITED, &con->flags) &&
	    !test_and_set_bit(CF_SEND_PENDING, &con->flags))
		queue_work(io_workqueue, &con->swork);

An RST does not call sk_data_ready, so process_recv_sockets() usually
does not run. When it does, its ret < 0 path for the main connection also
goes through lowcomms_queue_swork(), and is blocked in the same way.

So send_to_sock() never calls sock_sendmsg() again and never sees the
socket error. As a result, process_send_sockets() never reaches:

	default:
		if (ret < 0) {
			close_connection(con, false);

After that, every dlm_lowcomms_commit_msg() to that node is queued but
never sent.

If the peer reconnects, accept_from_sock() still finds con->sock set.
It attaches the new socket to othercon, which only receives. The send
side stays stuck until membership recovery calls dlm_lowcomms_close().

There is also a race here. Suppose the RST arrives after sock_sendmsg()
returns -EAGAIN but before lock_sock(). The error_report kick is lost,
because CF_SEND_PENDING is still set.

This code then sees SOCK_NOSPACE still set, and arms CF_APP_LIMITED on a
socket that is already dead. Neither sk->sk_err nor sk_state is checked
before it arms the flag.

Before this patch, send_to_sock() always returned DLM_IO_RESCHED. The
next sendmsg() on the dead socket failed with -ECONNRESET or -EPIPE, and
process_send_sockets() then closed the connection and reconnected.

The commit message says:

    This restores the semantics dlm had before the bit moved.

The CF_APP_LIMITED check in lowcomms_queue_swork() was added in 2022 by
commit dbb751ffab0b ("fs: dlm: parallelize lowcomms socket handling").
That was long after the bit stopped working.

Has that check ever been exercised with a working CF_APP_LIMITED path?

The later patches in this series do not seem to change this. They only
replace the bit operations with sk_set_nospace() / sk_clear_nospace().
lowcomms_error_report(), lowcomms_queue_swork() and this arming check
stay the same.

-- 
Sashiko AI review · https://netdev-ai.bots.linux.dev/sashiko/#/patchset/20260924134729.2047213-1-edumazet%40google.com

^ permalink raw reply	[flat|nested] 13+ messages in thread

* Re: [PATCH v2 net-next 2/9] net: add sk_set_nospace() and sk_clear_nospace()
  2026-09-24 13:47 ` [PATCH v2 net-next 2/9] net: add sk_set_nospace() and sk_clear_nospace() Eric Dumazet
@ 2026-09-25 13:49   ` netdev-bot+sashiko
  0 siblings, 0 replies; 13+ messages in thread
From: netdev-bot+sashiko @ 2026-09-25 13:49 UTC (permalink / raw)
  To: edumazet
  Cc: davem, kuba, pabeni, horms, ncardwell, kuniyu, willemb, netdev,
	eric.dumazet

Thank you for your contribution! Sashiko AI review found more severe
issue(s) in other patches of this series. This mail is only an FYI, in case
you want to address the nit picks below when you respin.

Low priority / nit picks for your consideration:

- [Low] The kernel-doc added for sk_clear_nospace() in net/core/sock.c
  says the helper is "Called from ->sk_write_space() handlers, once send…

--

net: add sk_set_nospace() and sk_clear_nospace()

This patch adds sk_set_nospace() and sk_clear_nospace() helpers that wrap
the SOCK_NOSPACE bit operations on sk->sk_socket->flags with a NULL check.
It also converts the core networking, tcp, kcm, mptcp, smc and tls setters
and clearers to use them. It prepares for a later patch that gives TCP a
private copy of the bit and needs a single choke point.

> diff --git a/net/core/sock.c b/net/core/sock.c
> index 2948dffcc3e1b..11a22aec7e414 100644
> --- a/net/core/sock.c
> +++ b/net/core/sock.c

[ ... ]

> +/**
> + *	sk_clear_nospace - tell the transport no writer is waiting for space
> + *	@sk: socket
> + *
> + *	Called from ->sk_write_space() handlers, once send space has been
> + *	made available to writers.
> + */

[Severity: Low]
This isn't a bug, but does this kernel-doc match how the helper is used?
One caller converted in this same patch doesn't follow the documented
calling context.

kcm_tx_work() in net/kcm/kcmsock.c is a work item handler
(INIT_WORK(&kcm->tx_work, kcm_tx_work)), not a ->sk_write_space() handler.
It only tests SOCK_NOSPACE and clears it without checking whether send
space is available. After that it calls ->sk_write_space():

net/kcm/kcmsock.c:kcm_tx_work() {
    ...
	if (likely(sk->sk_socket) &&
	    test_bit(SOCK_NOSPACE, &sk->sk_socket->flags)) {
		sk_clear_nospace(sk);
		sk->sk_write_space(sk);
	}
    ...
}

Other callers converted later in the series are also outside the
documented context. For example, the sunrpc svcsock code clears the bit
from a helper that checks wspace. The doc text is still the same at the
end of the series.

The commit message calls these helpers the single choke point that the
follow-up tcp_nospace patch depends on. Could the kernel-doc describe the
actual calling contexts, so nobody relies on a stricter contract when
adding ordering or locking to the helper later?

> +void sk_clear_nospace(struct sock *sk)
> +{
> +	struct socket *sock = sk->sk_socket;
> +
> +	if (sock)
> +		clear_bit(SOCK_NOSPACE, &sock->flags);
> +}
> +EXPORT_SYMBOL(sk_clear_nospace);

[ ... ]

-- 
Sashiko AI review · https://netdev-ai.bots.linux.dev/sashiko/#/patchset/20260924134729.2047213-1-edumazet%40google.com

^ permalink raw reply	[flat|nested] 13+ messages in thread

* Re: [PATCH v2 net-next 0/9] tcp: avoid struct socket cache line miss in tcp_check_space()
  2026-09-24 13:47 [PATCH v2 net-next 0/9] tcp: avoid struct socket cache line miss in tcp_check_space() Eric Dumazet
                   ` (8 preceding siblings ...)
  2026-09-24 13:47 ` [PATCH v2 net-next 9/9] tcp: add tp->tcp_nospace Eric Dumazet
@ 2026-09-29  1:18 ` Jakub Kicinski
  9 siblings, 0 replies; 13+ messages in thread
From: Jakub Kicinski @ 2026-09-29  1:18 UTC (permalink / raw)
  To: Eric Dumazet
  Cc: David S . Miller, Paolo Abeni, Simon Horman, Neal Cardwell,
	Kuniyuki Iwashima, Willem de Bruijn, netdev, eric.dumazet

On Thu, 24 Sep 2026 13:47:20 +0000 Eric Dumazet wrote:
> tcp_check_space() is called on every incoming ACK and every transmitted
> packet, and tests SOCK_NOSPACE in sk->sk_socket->flags.  Because struct
> socket lives in its own cache line and the TCP fast paths do not touch
> it for anything else, that test pulls in an extra cache line that is
> cold when the working set of active sockets is large.
> 
> This series mirrors SOCK_NOSPACE into a new u8 field, tp->tcp_nospace,
> placed right after tp->chrono_type in the tcp_sock_write_tx cache line
> group (fitting in an existing 3-byte hole before chrono_start, so no
> other field moves and sizeof(struct tcp_sock) is unchanged).  Both the
> transmit and ACK fast paths already touch that cache line, making the
> fast-path test in tcp_check_space() free of extra cache misses while
> keeping __tcp_check_space() authoritative on SOCK_NOSPACE.

Please resend this and CC the maintainers of the code we're about 
to change.
-- 
pw-bot: cr

^ permalink raw reply	[flat|nested] 13+ messages in thread

end of thread, other threads:[~2026-09-29  1:18 UTC | newest]

Thread overview: 13+ messages (download: mbox.gz follow: Atom feed
-- links below jump to the message on this page --
2026-09-24 13:47 [PATCH v2 net-next 0/9] tcp: avoid struct socket cache line miss in tcp_check_space() Eric Dumazet
2026-09-24 13:47 ` [PATCH v2 net-next 1/9] dlm: fix send buffer backpressure handling Eric Dumazet
2026-09-25 13:49   ` netdev-bot+sashiko
2026-09-24 13:47 ` [PATCH v2 net-next 2/9] net: add sk_set_nospace() and sk_clear_nospace() Eric Dumazet
2026-09-25 13:49   ` netdev-bot+sashiko
2026-09-24 13:47 ` [PATCH v2 net-next 3/9] sunrpc: use " Eric Dumazet
2026-09-24 13:47 ` [PATCH v2 net-next 4/9] rds: use sk_set_nospace() Eric Dumazet
2026-09-24 13:47 ` [PATCH v2 net-next 5/9] dlm: use sk_set_nospace() and sk_clear_nospace() Eric Dumazet
2026-09-24 13:47 ` [PATCH v2 net-next 6/9] drbd: use sk_set_nospace() Eric Dumazet
2026-09-24 13:47 ` [PATCH v2 net-next 7/9] nvme-tcp: use sk_clear_nospace() Eric Dumazet
2026-09-24 13:47 ` [PATCH v2 net-next 8/9] libceph: " Eric Dumazet
2026-09-24 13:47 ` [PATCH v2 net-next 9/9] tcp: add tp->tcp_nospace Eric Dumazet
2026-09-29  1:18 ` [PATCH v2 net-next 0/9] tcp: avoid struct socket cache line miss in tcp_check_space() Jakub Kicinski

This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox