Netdev List
 help / color / mirror / Atom feed
From: Eric Dumazet <edumazet@google.com>
To: "David S . Miller" <davem@davemloft.net>,
	Jakub Kicinski <kuba@kernel.org>,
	 Paolo Abeni <pabeni@redhat.com>
Cc: Simon Horman <horms@kernel.org>,
	Neal Cardwell <ncardwell@google.com>,
	 Kuniyuki Iwashima <kuniyu@google.com>,
	Willem de Bruijn <willemb@google.com>,
	netdev@vger.kernel.org,  eric.dumazet@gmail.com,
	Eric Dumazet <edumazet@google.com>
Subject: [PATCH net-next 9/9] tcp: add tp->tcp_nospace
Date: Tue, 22 Sep 2026 12:27:20 +0000	[thread overview]
Message-ID: <20260922122721.3568295-10-edumazet@google.com> (raw)
In-Reply-To: <20260922122721.3568295-1-edumazet@google.com>

tcp_check_space() runs for every incoming ACK and for every packet we
send, and reads SOCK_NOSPACE from sk->sk_socket->flags.

struct socket lives in its own cache line, which the TCP fast paths do
not otherwise touch, so testing sk->sk_socket->flags pulls in an extra
cache line that is cold when the working set is large.

Add tp->tcp_nospace, a mirror of SOCK_NOSPACE placed in the
tcp_sock_write_txrx group, that is in a cache line both the transmit
and the receive paths already have to touch.  It fits in an existing
hole, sizeof(struct tcp_sock) is unchanged.

The two flags are now only changed from sk_set_nospace() and
sk_clear_nospace(), which maintain this invariant:

	SOCK_NOSPACE set  =>  tp->tcp_nospace set

tcp_check_space() can thus test tp->tcp_nospace alone and leave the
authoritative SOCK_NOSPACE test to __tcp_check_space().  The
invariant is one directional on purpose: a stale tp->tcp_nospace only
costs an extra call to __tcp_check_space(), which is what we do
unconditionally today, while a stale SOCK_NOSPACE would cost a
missed EPOLLOUT.

Note the smp_mb() is kept.  tcp_poll() sets the flag without the
socket lock, and the store-buffer pattern it forms with
tcp_check_space() needs a full barrier on both sides.

MPTCP subflows share the struct socket of their parent, hence its
SOCK_NOSPACE bit, which can not be mirrored in the subflow tcp_sock.
Pin their tp->tcp_nospace in subflow_ulp_init(), so that they always
reach __tcp_check_space() and keep the current behavior.

Microbenchmark on an AMD EPYC 7B13, 64 threads, each thread calling
tcp_check_space() in a loop over a private set of sockets.  Numbers
are cycles per call above a baseline loop that performs the work the
callers already did (the cache lines tcp_write_xmit() and
tcp_clean_rtx_queue() touched) but not tcp_check_space() itself, so
they are the marginal cost of the function.  Median of 11 runs.
The set size controls whether struct socket is still cached:

    sockets/thread	before	 after
    1			  +0.6	  +0.7
    256    (64 KB)	  +8.3	  +3.9
    4096   (1 MB)	 +22.4	  +1.1
    262144 (64 MB)	 +48.4	  -0.8

Signed-off-by: Eric Dumazet <edumazet@google.com>
---
 .../networking/net_cachelines/tcp_sock.rst    |  1 +
 include/linux/tcp.h                           |  4 ++++
 include/net/tcp.h                             | 23 ++++++++++++++++++-
 net/core/sock.c                               | 18 +++++++++++----
 net/ipv4/tcp.c                                |  1 +
 net/ipv4/tcp_input.c                          |  8 ++++++-
 net/mptcp/subflow.c                           |  5 ++++
 7 files changed, 54 insertions(+), 6 deletions(-)

diff --git a/Documentation/networking/net_cachelines/tcp_sock.rst b/Documentation/networking/net_cachelines/tcp_sock.rst
index 0f6088c4ab8bb872e7fc86f02479592e84c0247a..3f225cdf12983d00c5347cd3ab16049b14b59047 100644
--- a/Documentation/networking/net_cachelines/tcp_sock.rst
+++ b/Documentation/networking/net_cachelines/tcp_sock.rst
@@ -12,6 +12,7 @@ struct inet_connection_sock   inet_conn
 u16                           tcp_header_len          read_mostly         read_mostly         tcp_bound_to_half_wnd,tcp_current_mss(tx);tcp_rcv_established(rx)
 u16                           gso_segs                read_mostly                             tcp_xmit_size_goal
 __be32                        pred_flags              read_write          read_mostly         tcp_select_window(tx);tcp_rcv_established(rx)
+u8                            tcp_nospace             read_mostly         read_mostly         tcp_check_space(tx);tcp_check_space(rx)
 u64                           bytes_received                              read_write          tcp_rcv_nxt_update(rx)
 u32                           segs_in                 read_write          read_write          tcp_segs_in(),tcp_v6_rcv(rx),tcp_v4_rcv()
 u32                           data_segs_in                                read_write          tcp_v6_rcv(rx)
diff --git a/include/linux/tcp.h b/include/linux/tcp.h
index 6a8c77719322f9caee305d954a107892c76d4ef7..d51aae60aa45bba31ae2f06b004383733caaf599 100644
--- a/include/linux/tcp.h
+++ b/include/linux/tcp.h
@@ -307,6 +307,10 @@ struct tcp_sock {
 		accecn_opt_demand:2,/* Demand AccECN option for n next ACKs */
 		prev_ecnfield:2; /* ECN bits from the previous segment */
 	__be32	pred_flags;
+	u8	tcp_nospace;	/* mirrors SOCK_NOSPACE, but in a cache line
+				 * that tcp_check_space() already needs.
+				 * Can only be set if SOCK_NOSPACE is set.
+				 */
 	u64	tcp_clock_cache; /* cache last tcp_clock_ns() (see tcp_mstamp_refresh()) */
 	u64	tcp_mstamp;	/* most recent packet received/sent */
 	u32	rcv_nxt;	/* What we want to receive next		*/
diff --git a/include/net/tcp.h b/include/net/tcp.h
index 5e5f5f9b89a386568fc5efebfa3d3c7e1ff62683..1e1950dd184ec3daecd5f691a0c099382e873937 100644
--- a/include/net/tcp.h
+++ b/include/net/tcp.h
@@ -783,14 +783,35 @@ void tcp_done_with_error(struct sock *sk, int err);
 void tcp_reset(struct sock *sk, struct sk_buff *skb);
 void tcp_fin(struct sock *sk);
 void __tcp_check_space(struct sock *sk);
+
+/* Mirror of SOCK_NOSPACE in tcp_sock, maintained by sk_set_nospace()
+ * and sk_clear_nospace().
+ *
+ * MPTCP subflows share the parent socket, and thus its SOCK_NOSPACE bit.
+ * Keep their mirror always set (see subflow_ulp_init()) so that they
+ * always reach __tcp_check_space() and behave as before.
+ */
+static inline void tcp_set_nospace(struct sock *sk)
+{
+	if (sk_is_tcp(sk))
+		WRITE_ONCE(tcp_sk(sk)->tcp_nospace, 1);
+}
+
+static inline void tcp_clear_nospace(struct sock *sk)
+{
+	if (sk_is_tcp(sk) && !sk_is_mptcp(sk))
+		WRITE_ONCE(tcp_sk(sk)->tcp_nospace, 0);
+}
+
 static inline void tcp_check_space(struct sock *sk)
 {
 	/* pairs with tcp_poll() */
 	smp_mb();
 
-	if (sk->sk_socket && test_bit(SOCK_NOSPACE, &sk->sk_socket->flags))
+	if (unlikely(READ_ONCE(tcp_sk(sk)->tcp_nospace)))
 		__tcp_check_space(sk);
 }
+
 void tcp_sack_compress_send_ack(struct sock *sk);
 
 static inline void tcp_cleanup_skb(struct sk_buff *skb)
diff --git a/net/core/sock.c b/net/core/sock.c
index 11a22aec7e414152aab115e8d11e30067ab3775f..dce4e8e4e2bbef710eae641a270395a315bef56e 100644
--- a/net/core/sock.c
+++ b/net/core/sock.c
@@ -2969,8 +2969,13 @@ void sk_set_nospace(struct sock *sk)
 {
 	struct socket *sock = sk->sk_socket;
 
-	if (sock)
-		set_bit(SOCK_NOSPACE, &sock->flags);
+	if (!sock)
+		return;
+	/* Mirror first: callers relying on the barrier implied by
+	 * set_bit() + smp_mb__after_atomic() are then also covered.
+	 */
+	tcp_set_nospace(sk);
+	set_bit(SOCK_NOSPACE, &sock->flags);
 }
 EXPORT_SYMBOL(sk_set_nospace);
 
@@ -2985,8 +2990,13 @@ void sk_clear_nospace(struct sock *sk)
 {
 	struct socket *sock = sk->sk_socket;
 
-	if (sock)
-		clear_bit(SOCK_NOSPACE, &sock->flags);
+	if (!sock)
+		return;
+	clear_bit(SOCK_NOSPACE, &sock->flags);
+	/* Mirror last: a stale mirror only costs a slow path, while a
+	 * stale SOCK_NOSPACE would cost a missed EPOLLOUT.
+	 */
+	tcp_clear_nospace(sk);
 }
 EXPORT_SYMBOL(sk_clear_nospace);
 
diff --git a/net/ipv4/tcp.c b/net/ipv4/tcp.c
index 1cde000cfab4704e6756872f6ddec16851ccc55d..92728e4c1e3df928cc7aa1bcbcf93fb15afaa28e 100644
--- a/net/ipv4/tcp.c
+++ b/net/ipv4/tcp.c
@@ -5261,6 +5261,7 @@ static void __init tcp_struct_check(void)
 
 	/* TXRX read-write hotpath cache lines */
 	CACHELINE_ASSERT_GROUP_MEMBER(struct tcp_sock, tcp_sock_write_txrx, pred_flags);
+	CACHELINE_ASSERT_GROUP_MEMBER(struct tcp_sock, tcp_sock_write_txrx, tcp_nospace);
 	CACHELINE_ASSERT_GROUP_MEMBER(struct tcp_sock, tcp_sock_write_txrx, tcp_clock_cache);
 	CACHELINE_ASSERT_GROUP_MEMBER(struct tcp_sock, tcp_sock_write_txrx, tcp_mstamp);
 	CACHELINE_ASSERT_GROUP_MEMBER(struct tcp_sock, tcp_sock_write_txrx, rcv_nxt);
diff --git a/net/ipv4/tcp_input.c b/net/ipv4/tcp_input.c
index 92bc60716f33d81e9ce90d9de2e5d989ba71c8a3..914d708326bf841641a4519341bfd77deba9e7e0 100644
--- a/net/ipv4/tcp_input.c
+++ b/net/ipv4/tcp_input.c
@@ -6098,8 +6098,14 @@ static void tcp_new_space(struct sock *sk)
  */
 void __tcp_check_space(struct sock *sk)
 {
+	struct socket *sock = sk->sk_socket;
+
+	/* tp->tcp_nospace is only a hint, SOCK_NOSPACE is authoritative. */
+	if (!sock || !test_bit(SOCK_NOSPACE, &sock->flags))
+		return;
+
 	tcp_new_space(sk);
-	if (!test_bit(SOCK_NOSPACE, &sk->sk_socket->flags))
+	if (!test_bit(SOCK_NOSPACE, &sock->flags))
 		tcp_chrono_stop(sk, TCP_CHRONO_SNDBUF_LIMITED);
 }
 
diff --git a/net/mptcp/subflow.c b/net/mptcp/subflow.c
index f0a6725d2c3762def75c879e769e9da37f92215a..e297c88be503b42979886baaf43acfb6b7e78055 100644
--- a/net/mptcp/subflow.c
+++ b/net/mptcp/subflow.c
@@ -2001,6 +2001,11 @@ static int subflow_ulp_init(struct sock *sk)
 	pr_debug("subflow=%p, family=%d\n", ctx, sk->sk_family);
 
 	tp->is_mptcp = 1;
+	/* Subflows share the MPTCP socket, and thus its SOCK_NOSPACE bit,
+	 * which tcp_check_space() can not mirror. Pin the mirror so that
+	 * __tcp_check_space() always tests the shared bit.
+	 */
+	tp->tcp_nospace = 1;
 	ctx->icsk_af_ops = icsk->icsk_af_ops;
 	icsk->icsk_af_ops = subflow_default_af_ops(sk);
 	ctx->tcp_state_change = sk->sk_state_change;
-- 
2.55.0.1082.g2b9226bbc0-goog


  parent reply	other threads:[~2026-09-22 12:27 UTC|newest]

Thread overview: 25+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-09-22 12:27 [PATCH net-next 0/9] tcp: avoid struct socket cache line miss in tcp_check_space() Eric Dumazet
2026-09-22 12:27 ` [PATCH net-next 1/9] dlm: fix send buffer backpressure handling Eric Dumazet
2026-09-22 13:13   ` Alexander Aring
2026-09-23 21:46   ` Kuniyuki Iwashima
2026-09-24  0:27   ` netdev-bot+sashiko
2026-09-24  0:37     ` Eric Dumazet
2026-09-22 12:27 ` [PATCH net-next 2/9] net: add sk_set_nospace() and sk_clear_nospace() Eric Dumazet
2026-09-23 21:54   ` Kuniyuki Iwashima
2026-09-24  0:27   ` netdev-bot+sashiko
2026-09-22 12:27 ` [PATCH net-next 3/9] sunrpc: use " Eric Dumazet
2026-09-23 21:54   ` Kuniyuki Iwashima
2026-09-22 12:27 ` [PATCH net-next 4/9] rds: use sk_set_nospace() Eric Dumazet
2026-09-23 21:55   ` Kuniyuki Iwashima
2026-09-22 12:27 ` [PATCH net-next 5/9] dlm: use sk_set_nospace() and sk_clear_nospace() Eric Dumazet
2026-09-23 21:55   ` Kuniyuki Iwashima
2026-09-22 12:27 ` [PATCH net-next 6/9] drbd: use sk_set_nospace() Eric Dumazet
2026-09-23 21:55   ` Kuniyuki Iwashima
2026-09-22 12:27 ` [PATCH net-next 7/9] nvme-tcp: use sk_clear_nospace() Eric Dumazet
2026-09-23 21:56   ` Kuniyuki Iwashima
2026-09-22 12:27 ` [PATCH net-next 8/9] libceph: " Eric Dumazet
2026-09-23 21:56   ` Kuniyuki Iwashima
2026-09-22 12:27 ` Eric Dumazet [this message]
2026-09-23 22:01   ` [PATCH net-next 9/9] tcp: add tp->tcp_nospace Kuniyuki Iwashima
2026-09-24  0:27   ` netdev-bot+sashiko
2026-09-24 13:01     ` Eric Dumazet

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=20260922122721.3568295-10-edumazet@google.com \
    --to=edumazet@google.com \
    --cc=davem@davemloft.net \
    --cc=eric.dumazet@gmail.com \
    --cc=horms@kernel.org \
    --cc=kuba@kernel.org \
    --cc=kuniyu@google.com \
    --cc=ncardwell@google.com \
    --cc=netdev@vger.kernel.org \
    --cc=pabeni@redhat.com \
    --cc=willemb@google.com \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox