Netdev List
 help / color / mirror / Atom feed
* [PATCH net-next v2 0/2] net: annotate remaining lockless sk->sk_err accesses
@ 2026-10-04  4:44 Quanye Yang via B4 Relay
  2026-10-04  4:44 ` [PATCH net-next v2 1/2] tls: annotate lockless access to sk->sk_err Quanye Yang via B4 Relay
  2026-10-04  4:44 ` [PATCH net-next v2 2/2] net: annotate lockless writes " Quanye Yang via B4 Relay
  0 siblings, 2 replies; 6+ messages in thread
From: Quanye Yang via B4 Relay @ 2026-10-04  4:44 UTC (permalink / raw)
  To: John Fastabend, Jakub Kicinski, Sabrina Dubroca, David S. Miller,
	Eric Dumazet, Paolo Abeni, Simon Horman, Jakub Sitnicki,
	Jiayuan Chen
  Cc: netdev, linux-kernel, bpf

The TCP/MPTCP series that annotated lockless sk_err peeks and
consumes is in net-next. Paolo asked for the same treatment on
kTLS and on the unmarked writers that still race those readers.

do_recvmmsg() and getsockopt(SO_ERROR) still call sock_error()
without the socket lock. kTLS is a ULP on that same struct sock,
so tls_rx_rec_wait() has the same peek-versus-consume split, and
the send path still does a double unmarked load.

sock_dequeue_err_skb() can store sk_err from MSG_ERRQUEUE before
lock_sock(). strp_abort_strp() and sk_psock_report_error() write
the same field on the TCP/TLS and sockmap paths.

Patch 1 annotates the TLS readers and consumes sk_err once on the
no-data path. Patch 2 pairs the remaining writers with WRITE_ONCE().
No extra ordering is added; ICMP error-queue overwrite semantics
are unchanged.

Link: https://lore.kernel.org/netdev/3d9d442f-f168-43da-87b0-010ad5a78365@redhat.com/

Signed-off-by: Quanye Yang <quanyeyang@proton.me>
---
Changes in v2:
- tls_encrypt_done(): fold the three unmarked sk_err loads into one
  READ_ONCE()
- Link to v1: https://patch.msgid.link/20261002-tls-fix-sk-kcsan-err-v1-0-baa0ba056323@proton.me

---
Quanye Yang (2):
      tls: annotate lockless access to sk->sk_err
      net: annotate lockless writes to sk->sk_err

 include/linux/skmsg.h     |  2 +-
 net/core/skbuff.c         |  5 +++--
 net/strparser/strparser.c |  2 +-
 net/tls/tls_device.c      |  5 +++--
 net/tls/tls_sw.c          | 54 +++++++++++++++++++++++++++++++----------------
 5 files changed, 44 insertions(+), 24 deletions(-)
---
base-commit: 071876fd50482a68603a9460d80dd6dd58827ee1
change-id: 20261002-tls-fix-sk-kcsan-err-6fef32744f15

Best regards,
--  
Quanye Yang <quanyeyang@proton.me>



^ permalink raw reply	[flat|nested] 6+ messages in thread

* [PATCH net-next v2 1/2] tls: annotate lockless access to sk->sk_err
  2026-10-04  4:44 [PATCH net-next v2 0/2] net: annotate remaining lockless sk->sk_err accesses Quanye Yang via B4 Relay
@ 2026-10-04  4:44 ` Quanye Yang via B4 Relay
  2026-10-04  4:44 ` [PATCH net-next v2 2/2] net: annotate lockless writes " Quanye Yang via B4 Relay
  1 sibling, 0 replies; 6+ messages in thread
From: Quanye Yang via B4 Relay @ 2026-10-04  4:44 UTC (permalink / raw)
  To: John Fastabend, Jakub Kicinski, Sabrina Dubroca, David S. Miller,
	Eric Dumazet, Paolo Abeni, Simon Horman, Jakub Sitnicki,
	Jiayuan Chen
  Cc: netdev, linux-kernel, bpf

From: Quanye Yang <quanyeyang@proton.me>

kTLS sits on the same struct sock as TCP. do_recvmmsg() and
getsockopt(SO_ERROR) still call sock_error() without the socket lock
and clear sk_err with xchg().

tls_rx_rec_wait() already peeks when data has been copied and consumes
otherwise, but the outer if (sk_err) is an unmarked load. On the
no-data path that check-then-sock_error() window can return 0 after
another thread consumes the error. Call sock_error() once and only
return when it is non-zero; keep READ_ONCE() on the peek path.

tls_sw_sendmsg_locked(), tls_push_data(), bpf_exec_tx_verdict() and
tls_encrypt_done() read sk_err more than once. Fold those unmarked loads
into one READ_ONCE() and use that value as the returned errno. The field is
still not consumed there.

Link: https://lore.kernel.org/netdev/3d9d442f-f168-43da-87b0-010ad5a78365@redhat.com/
Signed-off-by: Quanye Yang <quanyeyang@proton.me>
---
 net/tls/tls_device.c |  5 +++--
 net/tls/tls_sw.c     | 54 ++++++++++++++++++++++++++++++++++------------------
 2 files changed, 39 insertions(+), 20 deletions(-)

diff --git a/net/tls/tls_device.c b/net/tls/tls_device.c
index f11d0528fc43..03ce83a9d4e9 100644
--- a/net/tls/tls_device.c
+++ b/net/tls/tls_device.c
@@ -444,8 +444,9 @@ static int tls_push_data(struct sock *sk,
 	if ((flags & (MSG_MORE | MSG_EOR)) == (MSG_MORE | MSG_EOR))
 		return -EINVAL;
 
-	if (unlikely(sk->sk_err))
-		return -sk->sk_err;
+	rc = -READ_ONCE(sk->sk_err);
+	if (unlikely(rc))
+		return rc;
 
 	flags |= MSG_SENDPAGE_DECRYPTED;
 	tls_push_record_flags = flags | MSG_MORE;
diff --git a/net/tls/tls_sw.c b/net/tls/tls_sw.c
index d1ad31986cf2..d85d0c1ff546 100644
--- a/net/tls/tls_sw.c
+++ b/net/tls/tls_sw.c
@@ -473,6 +473,7 @@ static void tls_encrypt_done(void *data, int err)
 	struct scatterlist *sge;
 	struct sk_msg *msg_en;
 	struct sock *sk;
+	int skerr;
 
 	if (err == -EINPROGRESS) /* see the comment in tls_decrypt_done() */
 		return;
@@ -488,13 +489,15 @@ static void tls_encrypt_done(void *data, int err)
 	sge->offset -= prot->prepend_size;
 	sge->length += prot->prepend_size;
 
+	skerr = READ_ONCE(sk->sk_err);
+
 	/* Check if error is previously set on socket */
-	if (err || sk->sk_err) {
+	if (err || skerr) {
 		rec = NULL;
 
 		/* If err is already set on socket, return the same code */
-		if (sk->sk_err) {
-			ctx->async_wait.err = -sk->sk_err;
+		if (skerr) {
+			ctx->async_wait.err = -skerr;
 		} else {
 			ctx->async_wait.err = err;
 			tls_err_abort(sk, err);
@@ -704,10 +707,14 @@ static int bpf_exec_tx_verdict(struct sk_msg *msg, struct sock *sk,
 	int err;
 
 	err = tls_push_record(sk, flags, record_type);
-	if (err && err != -EINPROGRESS && sk->sk_err == EBADMSG) {
-		*copied -= sk_msg_free(sk, msg);
-		tls_free_open_rec(sk);
-		err = -sk->sk_err;
+	if (err && err != -EINPROGRESS) {
+		int skerr = READ_ONCE(sk->sk_err);
+
+		if (skerr == EBADMSG) {
+			*copied -= sk_msg_free(sk, msg);
+			tls_free_open_rec(sk);
+			err = -skerr;
+		}
 	}
 	return err;
 }
@@ -800,10 +807,9 @@ static int tls_sw_sendmsg_locked(struct sock *sk, struct msghdr *msg,
 	}
 
 	while (msg_data_left(msg)) {
-		if (sk->sk_err) {
-			ret = -sk->sk_err;
+		ret = -READ_ONCE(sk->sk_err);
+		if (ret)
 			goto send_end;
-		}
 
 		if (ctx->open_rec)
 			rec = ctx->open_rec;
@@ -1107,10 +1113,16 @@ tls_rx_rec_wait(struct sock *sk, bool nonblock, bool released, bool has_copied)
 	timeo = sock_rcvtimeo(sk, nonblock);
 
 	while (!tls_strp_msg_ready(ctx)) {
-		if (sk->sk_err) {
-			if (has_copied)
-				return -READ_ONCE(sk->sk_err);
-			return sock_error(sk);
+		if (has_copied) {
+			int err = READ_ONCE(sk->sk_err);
+
+			if (err)
+				return -err;
+		} else {
+			int err = sock_error(sk);
+
+			if (err)
+				return err;
 		}
 
 		if (ret < 0)
@@ -1132,10 +1144,16 @@ tls_rx_rec_wait(struct sock *sk, bool nonblock, bool released, bool has_copied)
 		 * sk_err here so a connection abort surfaces as the
 		 * actual error rather than a clean EOF.
 		 */
-		if (sk->sk_err) {
-			if (has_copied)
-				return -READ_ONCE(sk->sk_err);
-			return sock_error(sk);
+		if (has_copied) {
+			int err = READ_ONCE(sk->sk_err);
+
+			if (err)
+				return -err;
+		} else {
+			int err = sock_error(sk);
+
+			if (err)
+				return err;
 		}
 		if (sk->sk_shutdown & RCV_SHUTDOWN)
 			return 0;

-- 
2.55.0



^ permalink raw reply related	[flat|nested] 6+ messages in thread

* [PATCH net-next v2 2/2] net: annotate lockless writes to sk->sk_err
  2026-10-04  4:44 [PATCH net-next v2 0/2] net: annotate remaining lockless sk->sk_err accesses Quanye Yang via B4 Relay
  2026-10-04  4:44 ` [PATCH net-next v2 1/2] tls: annotate lockless access to sk->sk_err Quanye Yang via B4 Relay
@ 2026-10-04  4:44 ` Quanye Yang via B4 Relay
  2026-10-04  7:14   ` Eric Dumazet
  2026-10-08 23:44   ` Jakub Kicinski
  1 sibling, 2 replies; 6+ messages in thread
From: Quanye Yang via B4 Relay @ 2026-10-04  4:44 UTC (permalink / raw)
  To: John Fastabend, Jakub Kicinski, Sabrina Dubroca, David S. Miller,
	Eric Dumazet, Paolo Abeni, Simon Horman, Jakub Sitnicki,
	Jiayuan Chen
  Cc: netdev, linux-kernel, bpf

From: Quanye Yang <quanyeyang@proton.me>

do_recvmmsg() and getsockopt(SO_ERROR) clear sk_err with xchg()
without the socket lock. TCP, MPTCP and kTLS already peek the same
field with READ_ONCE() or consume it via sock_error().

sock_dequeue_err_skb() still uses plain stores. tcp_recvmsg() can
call it via MSG_ERRQUEUE before lock_sock(), so those writes race
with the annotated readers and with sock_error(). The same unmarked
stores exist in strp_abort_strp() and sk_psock_report_error(), which
run on the TCP/TLS socket.

Annotate those writers with WRITE_ONCE(). No extra ordering is
needed; this does not change who wins when ICMP error-queue entries
overwrite sk_err.

Link: https://lore.kernel.org/netdev/3d9d442f-f168-43da-87b0-010ad5a78365@redhat.com/
Signed-off-by: Quanye Yang <quanyeyang@proton.me>
---
 include/linux/skmsg.h     | 2 +-
 net/core/skbuff.c         | 5 +++--
 net/strparser/strparser.c | 2 +-
 3 files changed, 5 insertions(+), 4 deletions(-)

diff --git a/include/linux/skmsg.h b/include/linux/skmsg.h
index d5e35f24738d..52ce45f25f5a 100644
--- a/include/linux/skmsg.h
+++ b/include/linux/skmsg.h
@@ -429,7 +429,7 @@ static inline void sk_psock_report_error(struct sk_psock *psock, int err)
 {
 	struct sock *sk = psock->sk;
 
-	sk->sk_err = err;
+	WRITE_ONCE(sk->sk_err, err);
 	sk_error_report(sk);
 }
 
diff --git a/net/core/skbuff.c b/net/core/skbuff.c
index 5c4024a03e10..51e3cf1ea985 100644
--- a/net/core/skbuff.c
+++ b/net/core/skbuff.c
@@ -5535,12 +5535,13 @@ struct sk_buff *sock_dequeue_err_skb(struct sock *sk)
 	if (skb && (skb_next = skb_peek(q))) {
 		icmp_next = is_icmp_err_skb(skb_next);
 		if (icmp_next)
-			sk->sk_err = SKB_EXT_ERR(skb_next)->ee.ee_errno;
+			WRITE_ONCE(sk->sk_err,
+				   SKB_EXT_ERR(skb_next)->ee.ee_errno);
 	}
 	spin_unlock_irqrestore(&q->lock, flags);
 
 	if (is_icmp_err_skb(skb) && !icmp_next)
-		sk->sk_err = 0;
+		WRITE_ONCE(sk->sk_err, 0);
 
 	if (skb_next)
 		sk_error_report(sk);
diff --git a/net/strparser/strparser.c b/net/strparser/strparser.c
index a23f4b4dfc67..e5d5d755e532 100644
--- a/net/strparser/strparser.c
+++ b/net/strparser/strparser.c
@@ -57,7 +57,7 @@ static void strp_abort_strp(struct strparser *strp, int err)
 		struct sock *sk = strp->sk;
 
 		/* Report an error on the lower socket */
-		sk->sk_err = -err;
+		WRITE_ONCE(sk->sk_err, -err);
 		sk_error_report(sk);
 	}
 }

-- 
2.55.0



^ permalink raw reply related	[flat|nested] 6+ messages in thread

* Re: [PATCH net-next v2 2/2] net: annotate lockless writes to sk->sk_err
  2026-10-04  4:44 ` [PATCH net-next v2 2/2] net: annotate lockless writes " Quanye Yang via B4 Relay
@ 2026-10-04  7:14   ` Eric Dumazet
  2026-10-04  7:42     ` quanyeyang
  2026-10-08 23:44   ` Jakub Kicinski
  1 sibling, 1 reply; 6+ messages in thread
From: Eric Dumazet @ 2026-10-04  7:14 UTC (permalink / raw)
  To: quanyeyang
  Cc: John Fastabend, Jakub Kicinski, Sabrina Dubroca, David S. Miller,
	Paolo Abeni, Simon Horman, Jakub Sitnicki, Jiayuan Chen, netdev,
	linux-kernel, bpf

Le dim. 4 oct. 2026 à 06:44, Quanye Yang via B4 Relay
<devnull+quanyeyang.proton.me@kernel.org> a écrit :
>
> From: Quanye Yang <quanyeyang@proton.me>
>
> do_recvmmsg() and getsockopt(SO_ERROR) clear sk_err with xchg()
> without the socket lock. TCP, MPTCP and kTLS already peek the same
> field with READ_ONCE() or consume it via sock_error().
>
> sock_dequeue_err_skb() still uses plain stores. tcp_recvmsg() can
> call it via MSG_ERRQUEUE before lock_sock(), so those writes race
> with the annotated readers and with sock_error(). The same unmarked
> stores exist in strp_abort_strp() and sk_psock_report_error(), which
> run on the TCP/TLS socket.
>
> Annotate those writers with WRITE_ONCE(). No extra ordering is
> needed; this does not change who wins when ICMP error-queue entries
> overwrite sk_err.
>
> Link: https://lore.kernel.org/netdev/3d9d442f-f168-43da-87b0-010ad5a78365@redhat.com/
> Signed-off-by: Quanye Yang <quanyeyang@proton.me>
> ---
>  include/linux/skmsg.h     | 2 +-
>  net/core/skbuff.c         | 5 +++--
>  net/strparser/strparser.c | 2 +-
>  3 files changed, 5 insertions(+), 4 deletions(-)

Has this patch changed between V1 and V2 ?

You are supposed to carry the Acked-by and Reviewed-by tags collected
during prior iterations.

Please help reviewers, they need to recover their precious time.

Reviewed-by: Eric Dumazet <edumazet@kernel.org>

Thank you.

^ permalink raw reply	[flat|nested] 6+ messages in thread

* Re: [PATCH net-next v2 2/2] net: annotate lockless writes to sk->sk_err
  2026-10-04  7:14   ` Eric Dumazet
@ 2026-10-04  7:42     ` quanyeyang
  0 siblings, 0 replies; 6+ messages in thread
From: quanyeyang @ 2026-10-04  7:42 UTC (permalink / raw)
  To: Eric Dumazet
  Cc: John Fastabend, Jakub Kicinski, Sabrina Dubroca, David S. Miller,
	Paolo Abeni, Simon Horman, Jakub Sitnicki, Jiayuan Chen, netdev,
	linux-kernel, bpf

On Sunday, October 4th, 2026 at AM 12:14, Eric Dumazet <edumazet@kernel.org> wrote:

> 
> Has this patch changed between V1 and V2 ?
> 
> You are supposed to carry the Acked-by and Reviewed-by tags collected
> during prior iterations.
On Sun, Oct 4, 2026 at ... Eric Dumazet wrote:
> Has this patch changed between V1 and V2 ?
>
> You are supposed to carry the Acked-by and Reviewed-by tags collected
> during prior iterations.

No, 2/2 is identical to v1. Sorry for the extra pass — I should have
said "Patch 2: unchanged" in the v2 changelog and kept your tag.

Thanks for the Reviewed-by.

> 
> Please help reviewers, they need to recover their precious time.

Thanks, I'll pay attention to this next time.

> 
> Reviewed-by: Eric Dumazet <edumazet@kernel.org>
> 
> Thank you.
>

^ permalink raw reply	[flat|nested] 6+ messages in thread

* Re: [PATCH net-next v2 2/2] net: annotate lockless writes to sk->sk_err
  2026-10-04  4:44 ` [PATCH net-next v2 2/2] net: annotate lockless writes " Quanye Yang via B4 Relay
  2026-10-04  7:14   ` Eric Dumazet
@ 2026-10-08 23:44   ` Jakub Kicinski
  1 sibling, 0 replies; 6+ messages in thread
From: Jakub Kicinski @ 2026-10-08 23:44 UTC (permalink / raw)
  To: Quanye Yang via B4 Relay
  Cc: quanyeyang, John Fastabend, Sabrina Dubroca, David S. Miller,
	Eric Dumazet, Paolo Abeni, Simon Horman, Jakub Sitnicki,
	Jiayuan Chen, netdev, linux-kernel, bpf

On Sat, 03 Oct 2026 21:44:29 -0700 Quanye Yang via B4 Relay wrote:
> diff --git a/net/strparser/strparser.c b/net/strparser/strparser.c
> index a23f4b4dfc67..e5d5d755e532 100644
> --- a/net/strparser/strparser.c
> +++ b/net/strparser/strparser.c
> @@ -57,7 +57,7 @@ static void strp_abort_strp(struct strparser *strp, int err)
>  		struct sock *sk = strp->sk;
>  
>  		/* Report an error on the lower socket */
> -		sk->sk_err = -err;
> +		WRITE_ONCE(sk->sk_err, -err);
>  		sk_error_report(sk);
>  	}

This chunk doesn't apply any more, please wait a day (for net->net-next
merge), then rebase & repost.
-- 
pw-bot: cr

^ permalink raw reply	[flat|nested] 6+ messages in thread

end of thread, other threads:[~2026-10-08 23:44 UTC | newest]

Thread overview: 6+ messages (download: mbox.gz follow: Atom feed
-- links below jump to the message on this page --
2026-10-04  4:44 [PATCH net-next v2 0/2] net: annotate remaining lockless sk->sk_err accesses Quanye Yang via B4 Relay
2026-10-04  4:44 ` [PATCH net-next v2 1/2] tls: annotate lockless access to sk->sk_err Quanye Yang via B4 Relay
2026-10-04  4:44 ` [PATCH net-next v2 2/2] net: annotate lockless writes " Quanye Yang via B4 Relay
2026-10-04  7:14   ` Eric Dumazet
2026-10-04  7:42     ` quanyeyang
2026-10-08 23:44   ` Jakub Kicinski

This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox