* [PATCH mptcp-net v3 0/3] mptcp: fix reported data-races
@ 2026-09-07 10:15 quanyeyang
2026-09-07 10:15 ` [PATCH mptcp-net v3 1/3] mptcp: fix data-race in __mptcp_retrans / mptcp_incoming_options quanyeyang
` (3 more replies)
0 siblings, 4 replies; 13+ messages in thread
From: quanyeyang @ 2026-09-07 10:15 UTC (permalink / raw)
To: MPTCP Linux; +Cc: geliang, matttbe, pabeni
Fix three KCSAN reports involving the MPTCP receive and retransmit
paths.
Patch 1 picks up Matthieu's v2 fixing the lockless snd_una read.
Patch 2 fixes the lockless access to the TCP subflow's icsk_pending
field when calculating the MPTCP retransmission timeout.
Patch 3 fixes the lockless sk_err access in tcp_recv_should_stop(),
as sock_error() can clear the field without holding the socket lock.
Closes: https://github.com/multipath-tcp/mptcp_net-next/issues/627
Link: https://lore.kernel.org/mptcp/20260317-mptcp-data-race-snd_una-v2-1-2caac60de92a@kernel.org/
---
v3:
- rebase on the MPTCP export branch
- wrap the snd_una field comment to satisfy checkpatch
- add fixes for the reported icsk_pending and sk_err races
- exclude the unrelated workqueue and timekeeping reports
v2:
- mention under which locks the field is updated and read
- add a comment where snd_una is defined
To: mptcp@lists.linux.dev
---
Matthieu Baerts (NGI0) (1):
mptcp: fix data-race in __mptcp_retrans / mptcp_incoming_options
Quanye Yang (2):
mptcp: fix data-race in mptcp_subflow_get_send / tcp_ack
tcp: fix data-race in do_recvmmsg / mptcp_recvmsg
include/net/tcp.h | 3 ++-
net/mptcp/protocol.c | 4 +++-
net/mptcp/protocol.h | 6 ++++--
3 files changed, 9 insertions(+), 4 deletions(-)
---
base-commit: 45f7c939f8b155208c7447517e588c9fb133388c
change-id: 20260906-mptcp-snd-una-race-be0262e82b97
Best regards,
--
Quanye Yang <quanyeyang@proton.me>
^ permalink raw reply [flat|nested] 13+ messages in thread
* [PATCH mptcp-net v3 1/3] mptcp: fix data-race in __mptcp_retrans / mptcp_incoming_options
2026-09-07 10:15 [PATCH mptcp-net v3 0/3] mptcp: fix reported data-races quanyeyang
@ 2026-09-07 10:15 ` quanyeyang
2026-09-07 13:38 ` Paolo Abeni
2026-09-07 10:15 ` Quanye Yang
` (2 subsequent siblings)
3 siblings, 1 reply; 13+ messages in thread
From: quanyeyang @ 2026-09-07 10:15 UTC (permalink / raw)
To: MPTCP Linux; +Cc: geliang, matttbe, pabeni
From: "Matthieu Baerts (NGI0)" <matttbe@kernel.org>
SyzKaller found this data-race:
BUG: KCSAN: data-race in __mptcp_retrans / mptcp_incoming_options
write (marked) to 0xffff888015e8e5f0 of 8 bytes by interrupt on cpu 0:
__mptcp_snd_una_update net/mptcp/options.c:1055 [inline]
mptcp_incoming_options+0x6a3/0x1ac0 net/mptcp/options.c:1183
tcp_data_queue+0x101b/0x2440 net/ipv4/tcp_input.c:5583
tcp_rcv_established+0x684/0x1fc0 net/ipv4/tcp_input.c:6654
tcp_v4_do_rcv+0x35c/0x690 net/ipv4/tcp_ipv4.c:1866
tcp_v4_rcv+0x1d91/0x25a0 net/ipv4/tcp_ipv4.c:2263
ip_protocol_deliver_rcu+0x46/0x280 net/ipv4/ip_input.c:207
ip_local_deliver_finish+0x190/0x270 net/ipv4/ip_input.c:241
NF_HOOK include/linux/netfilter.h:318 [inline]
NF_HOOK include/linux/netfilter.h:312 [inline]
ip_local_deliver+0xe3/0x210 net/ipv4/ip_input.c:262
dst_input include/net/dst.h:480 [inline]
ip_rcv_finish net/ipv4/ip_input.c:492 [inline]
NF_HOOK include/linux/netfilter.h:318 [inline]
NF_HOOK include/linux/netfilter.h:312 [inline]
ip_rcv+0x200/0x220 net/ipv4/ip_input.c:612
__netif_receive_skb_one_core+0xeb/0x110 net/core/dev.c:6178
__netif_receive_skb+0x1f/0xc0 net/core/dev.c:6291
process_backlog+0x168/0x360 net/core/dev.c:6642
__napi_poll+0x71/0x460 net/core/dev.c:7706
napi_poll net/core/dev.c:7769 [inline]
net_rx_action+0x6f8/0x810 net/core/dev.c:7926
handle_softirqs+0xc9/0x2e0 kernel/softirq.c:622
run_ksoftirqd kernel/softirq.c:1063 [inline]
run_ksoftirqd+0x20/0x30 kernel/softirq.c:1055
smpboot_thread_fn+0x287/0x520 kernel/smpboot.c:160
kthread+0x1f2/0x240 kernel/kthread.c:436
ret_from_fork+0x321/0x440 arch/x86/kernel/process.c:158
ret_from_fork_asm+0x1a/0x30 arch/x86/entry/entry_64.S:245
read to 0xffff888015e8e5f0 of 8 bytes by task 24 on cpu 1:
mptcp_rtx_head net/mptcp/protocol.h:487 [inline]
__mptcp_retrans+0x169/0x8f0 net/mptcp/protocol.c:2759
mptcp_worker+0x6a6/0xb30 net/mptcp/protocol.c:2980
process_one_work+0x3ee/0x970 kernel/workqueue.c:3275
process_scheduled_works kernel/workqueue.c:3358 [inline]
worker_thread+0x3c3/0x730 kernel/workqueue.c:3439
kthread+0x1f2/0x240 kernel/kthread.c:436
ret_from_fork+0x321/0x440 arch/x86/kernel/process.c:158
ret_from_fork_asm+0x1a/0x30 arch/x86/entry/entry_64.S:245
value changed: 0x0b17a4285ae6137d -> 0x0b17a4285b078905
It looks like msk->snd_una was being modified in __mptcp_snd_una_update
under the msk data lock (spin lock), while being accessed in
mptcp_rtx_head() under a different lock: the msk socket lock.
Annotate access to msk->snd_una in mptcp_rtx_head() to prevent such
issue.
Fixes: 64b9cea7a0af ("mptcp: fix spurious retransmissions")
Signed-off-by: Matthieu Baerts (NGI0) <matttbe@kernel.org>
Reviewed-by: Geliang Tang <geliang@kernel.org>
Signed-off-by: Quanye Yang <quanyeyang@proton.me>
---
net/mptcp/protocol.h | 6 ++++--
1 file changed, 4 insertions(+), 2 deletions(-)
diff --git a/net/mptcp/protocol.h b/net/mptcp/protocol.h
index b3121c8c766b..19ad2fe2036a 100644
--- a/net/mptcp/protocol.h
+++ b/net/mptcp/protocol.h
@@ -304,7 +304,9 @@ struct mptcp_sock {
* protection
*/
u64 bytes_acked;
- u64 snd_una;
+ u64 snd_una; /* updated under the msk data lock,
+ * lockless read
+ */
u64 wnd_end;
u32 last_data_sent;
u32 last_data_recv;
@@ -488,7 +490,7 @@ static inline struct mptcp_data_frag *mptcp_rtx_head(struct sock *sk)
{
struct mptcp_sock *msk = mptcp_sk(sk);
- if (msk->snd_una == msk->snd_nxt)
+ if (READ_ONCE(msk->snd_una) == msk->snd_nxt)
return NULL;
return list_first_entry_or_null(&msk->rtx_queue, struct mptcp_data_frag, list);
--
2.55.0
^ permalink raw reply related [flat|nested] 13+ messages in thread
* [PATCH mptcp-net v3 2/3] mptcp: fix data-race in mptcp_subflow_get_send / tcp_ack
2026-09-07 10:15 [PATCH mptcp-net v3 0/3] mptcp: fix reported data-races quanyeyang
@ 2026-09-07 10:15 ` Quanye Yang
2026-09-07 10:15 ` Quanye Yang
` (2 subsequent siblings)
3 siblings, 0 replies; 13+ messages in thread
From: Quanye Yang via B4 Relay @ 2026-09-07 10:15 UTC (permalink / raw)
To: MPTCP Linux; +Cc: geliang, matttbe, pabeni
From: Quanye Yang <quanyeyang@proton.me>
KCSAN reported a data race between the lockless read of
icsk->icsk_pending in mptcp_timeout_from_subflow() and the
smp_store_release() performed from tcp_ack() when clearing the
subflow retransmission timer.
The MPTCP socket lock held by the reader does not protect the TCP
subflow state. Use smp_load_acquire() to match the store-release
operations used by TCP and the other lockless readers of this field.
Fixes: 33d41c9cd74c ("mptcp: more accurate timeout")
Signed-off-by: Quanye Yang <quanyeyang@proton.me>
---
net/mptcp/protocol.c | 4 +++-
1 file changed, 3 insertions(+), 1 deletion(-)
diff --git a/net/mptcp/protocol.c b/net/mptcp/protocol.c
index 0b24e0afedfb..f4f63df9d787 100644
--- a/net/mptcp/protocol.c
+++ b/net/mptcp/protocol.c
@@ -606,7 +606,9 @@ static long mptcp_timeout_from_subflow(const struct mptcp_subflow_context *subfl
{
const struct sock *ssk = mptcp_subflow_tcp_sock(subflow);
- return inet_csk(ssk)->icsk_pending && !subflow->stale_count ?
+ /* Pair this lockless read with TCP's store-release updates. */
+ return smp_load_acquire(&inet_csk(ssk)->icsk_pending) &&
+ !subflow->stale_count ?
tcp_timeout_expires(ssk) - jiffies : 0;
}
--
2.55.0
^ permalink raw reply related [flat|nested] 13+ messages in thread
* [PATCH mptcp-net v3 2/3] mptcp: fix data-race in mptcp_subflow_get_send / tcp_ack
@ 2026-09-07 10:15 ` Quanye Yang
0 siblings, 0 replies; 13+ messages in thread
From: Quanye Yang @ 2026-09-07 10:15 UTC (permalink / raw)
To: MPTCP Linux; +Cc: geliang, matttbe, pabeni
KCSAN reported a data race between the lockless read of
icsk->icsk_pending in mptcp_timeout_from_subflow() and the
smp_store_release() performed from tcp_ack() when clearing the
subflow retransmission timer.
The MPTCP socket lock held by the reader does not protect the TCP
subflow state. Use smp_load_acquire() to match the store-release
operations used by TCP and the other lockless readers of this field.
Fixes: 33d41c9cd74c ("mptcp: more accurate timeout")
Signed-off-by: Quanye Yang <quanyeyang@proton.me>
---
net/mptcp/protocol.c | 4 +++-
1 file changed, 3 insertions(+), 1 deletion(-)
diff --git a/net/mptcp/protocol.c b/net/mptcp/protocol.c
index 0b24e0afedfb..f4f63df9d787 100644
--- a/net/mptcp/protocol.c
+++ b/net/mptcp/protocol.c
@@ -606,7 +606,9 @@ static long mptcp_timeout_from_subflow(const struct mptcp_subflow_context *subfl
{
const struct sock *ssk = mptcp_subflow_tcp_sock(subflow);
- return inet_csk(ssk)->icsk_pending && !subflow->stale_count ?
+ /* Pair this lockless read with TCP's store-release updates. */
+ return smp_load_acquire(&inet_csk(ssk)->icsk_pending) &&
+ !subflow->stale_count ?
tcp_timeout_expires(ssk) - jiffies : 0;
}
--
2.55.0
^ permalink raw reply related [flat|nested] 13+ messages in thread
* [PATCH mptcp-net v3 3/3] tcp: fix data-race in do_recvmmsg / mptcp_recvmsg
2026-09-07 10:15 [PATCH mptcp-net v3 0/3] mptcp: fix reported data-races quanyeyang
@ 2026-09-07 10:15 ` Quanye Yang
2026-09-07 10:15 ` Quanye Yang
` (2 subsequent siblings)
3 siblings, 0 replies; 13+ messages in thread
From: Quanye Yang via B4 Relay @ 2026-09-07 10:15 UTC (permalink / raw)
To: MPTCP Linux; +Cc: geliang, matttbe, pabeni
From: Quanye Yang <quanyeyang@proton.me>
KCSAN reported a data race between do_recvmmsg() and
mptcp_recvmsg() on sk->sk_err.
do_recvmmsg() calls sock_error() without holding the socket lock.
sock_error() atomically clears sk_err using xchg(), which can race
with the plain read in tcp_recv_should_stop(), even when its caller
holds the socket lock.
Use READ_ONCE() for the lockless read. No additional ordering is
required because the value is only used to decide whether receiving
should stop.
Fixes: 7a6a6cbc3e59 ("mptcp: recvmsg() can drain data from multiple subflows")
Signed-off-by: Quanye Yang <quanyeyang@proton.me>
---
include/net/tcp.h | 3 ++-
1 file changed, 2 insertions(+), 1 deletion(-)
diff --git a/include/net/tcp.h b/include/net/tcp.h
index 436495ff2271..c61d8678eafd 100644
--- a/include/net/tcp.h
+++ b/include/net/tcp.h
@@ -3082,7 +3082,8 @@ enum skb_drop_reason tcp_inbound_hash(struct sock *sk,
static inline int tcp_recv_should_stop(struct sock *sk)
{
- return sk->sk_err ||
+ /* sk_err can be cleared locklessly by sock_error(). */
+ return READ_ONCE(sk->sk_err) ||
sk->sk_state == TCP_CLOSE ||
(sk->sk_shutdown & RCV_SHUTDOWN) ||
signal_pending(current);
--
2.55.0
^ permalink raw reply related [flat|nested] 13+ messages in thread
* [PATCH mptcp-net v3 3/3] tcp: fix data-race in do_recvmmsg / mptcp_recvmsg
@ 2026-09-07 10:15 ` Quanye Yang
0 siblings, 0 replies; 13+ messages in thread
From: Quanye Yang @ 2026-09-07 10:15 UTC (permalink / raw)
To: MPTCP Linux; +Cc: geliang, matttbe, pabeni
KCSAN reported a data race between do_recvmmsg() and
mptcp_recvmsg() on sk->sk_err.
do_recvmmsg() calls sock_error() without holding the socket lock.
sock_error() atomically clears sk_err using xchg(), which can race
with the plain read in tcp_recv_should_stop(), even when its caller
holds the socket lock.
Use READ_ONCE() for the lockless read. No additional ordering is
required because the value is only used to decide whether receiving
should stop.
Fixes: 7a6a6cbc3e59 ("mptcp: recvmsg() can drain data from multiple subflows")
Signed-off-by: Quanye Yang <quanyeyang@proton.me>
---
include/net/tcp.h | 3 ++-
1 file changed, 2 insertions(+), 1 deletion(-)
diff --git a/include/net/tcp.h b/include/net/tcp.h
index 436495ff2271..c61d8678eafd 100644
--- a/include/net/tcp.h
+++ b/include/net/tcp.h
@@ -3082,7 +3082,8 @@ enum skb_drop_reason tcp_inbound_hash(struct sock *sk,
static inline int tcp_recv_should_stop(struct sock *sk)
{
- return sk->sk_err ||
+ /* sk_err can be cleared locklessly by sock_error(). */
+ return READ_ONCE(sk->sk_err) ||
sk->sk_state == TCP_CLOSE ||
(sk->sk_shutdown & RCV_SHUTDOWN) ||
signal_pending(current);
--
2.55.0
^ permalink raw reply related [flat|nested] 13+ messages in thread
* Re: [PATCH mptcp-net v3 3/3] tcp: fix data-race in do_recvmmsg / mptcp_recvmsg
2026-09-07 10:15 ` Quanye Yang
(?)
@ 2026-09-07 10:24 ` sashiko-bot
2026-09-07 13:23 ` quanyeyang
-1 siblings, 1 reply; 13+ messages in thread
From: sashiko-bot @ 2026-09-07 10:24 UTC (permalink / raw)
To: Quanye Yang; +Cc: mptcp
Thank you for your contribution! Sashiko AI review found 1 potential issue(s) to consider:
Pre-existing issues:
- [High] Incomplete fix for `sk->sk_err` data race leading to false EOFs. The patch fixes a lockless data race in `tcp_recv_should_stop()` by adding `READ_ONCE()`, but overlooks identical plain reads followed by `sock_error()` in `tcp_recvmsg_locked()`, `tcp_splice_read()`, `mptcp_recvmsg()`, and `tcp_bpf_sendmsg()`. If `sk->sk_err` is checked and evaluates to true, but is concurrently cleared to 0 by another thread calling `sock_error()` (e.g. via `do_recvmmsg` or `getsockopt(SO_ERROR)`), the subsequent assignment `copied = sock_error(sk)` will evaluate to 0. For receive and splice functions, this results in returning 0, which falsely signals a graceful connection closure (EOF) to the application rather than an error.
--
commit 0f561938f67b19245ef8889cb9ecc077fb82d7d0
Author: Quanye Yang <quanyeyang@proton.me>
tcp: fix data-race in do_recvmmsg / mptcp_recvmsg
This commit fixes a lockless data race in tcp_recv_should_stop() by using
READ_ONCE() for reading sk->sk_err. This prevents issues when sock_error()
atomically clears the error using xchg() without holding the socket lock.
> diff --git a/include/net/tcp.h b/include/net/tcp.h
> index 436495ff2271d..c61d8678eafd3 100644
> --- a/include/net/tcp.h
> +++ b/include/net/tcp.h
> @@ -3082,7 +3082,8 @@ enum skb_drop_reason tcp_inbound_hash(struct sock *sk,
>
> static inline int tcp_recv_should_stop(struct sock *sk)
> {
> - return sk->sk_err ||
> + /* sk_err can be cleared locklessly by sock_error(). */
> + return READ_ONCE(sk->sk_err) ||
[Severity: High]
This is a pre-existing issue, but does this fix also need to be applied to
other locations that read sk->sk_err locklessly?
If sk->sk_err is evaluated as non-zero, but the error is concurrently cleared
by another thread calling sock_error() (for example, via concurrent recvmmsg
or getsockopt calls), a subsequent call to sock_error(sk) will return 0. This
masks the socket error and falsely signals a graceful connection closure (EOF)
or 0 bytes sent.
This pattern appears to be present in several other functions:
net/ipv4/tcp.c:tcp_recvmsg_locked() {
...
if (sk->sk_err) {
copied = sock_error(sk);
break;
}
...
}
net/ipv4/tcp.c:tcp_splice_read() {
...
if (sk->sk_err) {
ret = sock_error(sk);
break;
}
...
}
net/mptcp/protocol.c:mptcp_recvmsg() {
...
if (sk->sk_err) {
copied = sock_error(sk);
break;
}
...
}
net/ipv4/tcp_bpf.c:tcp_bpf_sendmsg() {
...
if (sk->sk_err) {
err = -sk->sk_err;
goto out_err;
}
...
}
Would it be appropriate to update these instances as well to prevent
Time-Of-Check to Time-Of-Use races?
> sk->sk_state == TCP_CLOSE ||
> (sk->sk_shutdown & RCV_SHUTDOWN) ||
> signal_pending(current);
--
Sashiko AI review · https://sashiko.dev/#/patchset/20260907-mptcp-snd-una-race-v3-0-65ba0d5f0221@proton.me?part=3
^ permalink raw reply [flat|nested] 13+ messages in thread
* Re: [PATCH mptcp-net v3 0/3] mptcp: fix reported data-races
2026-09-07 10:15 [PATCH mptcp-net v3 0/3] mptcp: fix reported data-races quanyeyang
` (2 preceding siblings ...)
2026-09-07 10:15 ` Quanye Yang
@ 2026-09-07 11:28 ` MPTCP CI
3 siblings, 0 replies; 13+ messages in thread
From: MPTCP CI @ 2026-09-07 11:28 UTC (permalink / raw)
To: Quanye Yang; +Cc: mptcp
Hi Quanye,
Thank you for your modifications, that's great!
Our CI did some validations and here is its report:
- KVM Validation: normal (except selftest_mptcp_join): Success! ✅
- KVM Validation: normal (only selftest_mptcp_join): Success! ✅
- KVM Validation: debug (except selftest_mptcp_join): Success! ✅
- KVM Validation: debug (only selftest_mptcp_join): Success! ✅
- KVM Validation: btf-normal (only bpftest_all): Success! ✅
- KVM Validation: btf-debug (only bpftest_all): Success! ✅
- Perf:
- Task: https://github.com/multipath-tcp/mptcp_net-next/actions/runs/34112482705
Initiator: Patchew Applier
Commits: https://github.com/multipath-tcp/mptcp_net-next/commits/7d66f95769e2
Patchwork: https://patchwork.kernel.org/project/mptcp/list/?series=1159615
If there are some issues, you can reproduce them using the same environment as
the one used by the CI thanks to a docker image, e.g.:
$ cd [kernel source code]
$ docker run -v "${PWD}:${PWD}:rw" -w "${PWD}" --privileged --rm -it \
--pull always mptcp/mptcp-upstream-virtme-docker:latest \
auto-normal
For more details:
https://github.com/multipath-tcp/mptcp-upstream-virtme-docker
Please note that despite all the efforts that have been already done to have a
stable tests suite when executed on a public CI like here, it is possible some
reported issues are not due to your modifications. Still, do not hesitate to
help us improve that ;-)
Cheers,
MPTCP GH Action bot
Bot operated by Matthieu Baerts (NGI0 Core)
^ permalink raw reply [flat|nested] 13+ messages in thread
* Re: [PATCH mptcp-net v3 3/3] tcp: fix data-race in do_recvmmsg / mptcp_recvmsg
2026-09-07 10:24 ` sashiko-bot
@ 2026-09-07 13:23 ` quanyeyang
0 siblings, 0 replies; 13+ messages in thread
From: quanyeyang @ 2026-09-07 13:23 UTC (permalink / raw)
To: sashiko-reviews; +Cc: mptcp
On Monday, September 7th, 2026 at AM 3:24, sashiko-bot@kernel.org <sashiko-bot@kernel.org> wrote:
> Thank you for your contribution! Sashiko AI review found 1 potential issue(s) to consider:
>
> Pre-existing issues:
> - [High] Incomplete fix for `sk->sk_err` data race leading to false EOFs. The patch fixes a lockless data race in `tcp_recv_should_stop()` by adding `READ_ONCE()`, but overlooks identical plain reads followed by `sock_error()` in `tcp_recvmsg_locked()`, `tcp_splice_read()`, `mptcp_recvmsg()`, and `tcp_bpf_sendmsg()`. If `sk->sk_err` is checked and evaluates to true, but is concurrently cleared to 0 by another thread calling `sock_error()` (e.g. via `do_recvmmsg` or `getsockopt(SO_ERROR)`), the subsequent assignment `copied = sock_error(sk)` will evaluate to 0. For receive and splice functions, this results in returning 0, which falsely signals a graceful connection closure (EOF) to the application rather than an error.
> --
>
> commit 0f561938f67b19245ef8889cb9ecc077fb82d7d0
> Author: Quanye Yang <quanyeyang@proton.me>
>
> tcp: fix data-race in do_recvmmsg / mptcp_recvmsg
>
> This commit fixes a lockless data race in tcp_recv_should_stop() by using
> READ_ONCE() for reading sk->sk_err. This prevents issues when sock_error()
> atomically clears the error using xchg() without holding the socket lock.
>
> > diff --git a/include/net/tcp.h b/include/net/tcp.h
> > index 436495ff2271d..c61d8678eafd3 100644
> > --- a/include/net/tcp.h
> > +++ b/include/net/tcp.h
> > @@ -3082,7 +3082,8 @@ enum skb_drop_reason tcp_inbound_hash(struct sock *sk,
> >
> > static inline int tcp_recv_should_stop(struct sock *sk)
> > {
> > - return sk->sk_err ||
> > + /* sk_err can be cleared locklessly by sock_error(). */
> > + return READ_ONCE(sk->sk_err) ||
>
> [Severity: High]
> This is a pre-existing issue, but does this fix also need to be applied to
> other locations that read sk->sk_err locklessly?
>
> If sk->sk_err is evaluated as non-zero, but the error is concurrently cleared
> by another thread calling sock_error() (for example, via concurrent recvmmsg
> or getsockopt calls), a subsequent call to sock_error(sk) will return 0. This
> masks the socket error and falsely signals a graceful connection closure (EOF)
> or 0 bytes sent.
>
> This pattern appears to be present in several other functions:
>
> net/ipv4/tcp.c:tcp_recvmsg_locked() {
> ...
> if (sk->sk_err) {
> copied = sock_error(sk);
> break;
> }
> ...
> }
>
> net/ipv4/tcp.c:tcp_splice_read() {
> ...
> if (sk->sk_err) {
> ret = sock_error(sk);
> break;
> }
> ...
> }
>
> net/mptcp/protocol.c:mptcp_recvmsg() {
> ...
> if (sk->sk_err) {
> copied = sock_error(sk);
> break;
> }
> ...
> }
>
> net/ipv4/tcp_bpf.c:tcp_bpf_sendmsg() {
> ...
> if (sk->sk_err) {
> err = -sk->sk_err;
> goto out_err;
> }
> ...
> }
>
> Would it be appropriate to update these instances as well to prevent
> Time-Of-Check to Time-Of-Use races?
>
> > sk->sk_state == TCP_CLOSE ||
> > (sk->sk_shutdown & RCV_SHUTDOWN) ||
> > signal_pending(current);
>
> --
> Sashiko AI review · https://sashiko.dev/#/patchset/20260907-mptcp-snd-una-race-v3-0-65ba0d5f0221@proton.me?part=3
>
Thanks for the review.
The TOCTOU issue is valid, but it is separate from the reported KCSAN
race fixed by this patch. tcp_recv_should_stop() only peeks at sk_err
after data has already been copied, so READ_ONCE() is appropriate there
and the error must not be consumed.
For the no-data paths, the correct approach is to call sock_error()
once and only stop when it returns a non-zero error. This avoids
returning a false EOF if another thread consumes the error between the
check and sock_error().
The BPF function is tcp_bpf_recvmsg_parser(), not tcp_bpf_sendmsg().
There are also additional instances, including mptcp_splice_read() and
the LLC receive path, so updating only the locations listed above would
still be incomplete.
I will audit these check-then-sock_error() patterns and handle the
pre-existing TOCTOU issue separately, unless the maintainers prefer it
to be included as an additional patch in v4.
Thanks,
Quanye
^ permalink raw reply [flat|nested] 13+ messages in thread
* Re: [PATCH mptcp-net v3 1/3] mptcp: fix data-race in __mptcp_retrans / mptcp_incoming_options
2026-09-07 10:15 ` [PATCH mptcp-net v3 1/3] mptcp: fix data-race in __mptcp_retrans / mptcp_incoming_options quanyeyang
@ 2026-09-07 13:38 ` Paolo Abeni
2026-09-07 13:55 ` Matthieu Baerts
0 siblings, 1 reply; 13+ messages in thread
From: Paolo Abeni @ 2026-09-07 13:38 UTC (permalink / raw)
To: quanyeyang, MPTCP Linux; +Cc: Geliang Tang, Matthieu Baerts (NGI0)
On 9/7/26 12:15 PM, quanyeyang@proton.me wrote:
> From: "Matthieu Baerts (NGI0)" <matttbe@kernel.org>
>
> SyzKaller found this data-race:
>
> BUG: KCSAN: data-race in __mptcp_retrans / mptcp_incoming_options
>
> write (marked) to 0xffff888015e8e5f0 of 8 bytes by interrupt on cpu 0:
> __mptcp_snd_una_update net/mptcp/options.c:1055 [inline]
> mptcp_incoming_options+0x6a3/0x1ac0 net/mptcp/options.c:1183
> tcp_data_queue+0x101b/0x2440 net/ipv4/tcp_input.c:5583
> tcp_rcv_established+0x684/0x1fc0 net/ipv4/tcp_input.c:6654
> tcp_v4_do_rcv+0x35c/0x690 net/ipv4/tcp_ipv4.c:1866
> tcp_v4_rcv+0x1d91/0x25a0 net/ipv4/tcp_ipv4.c:2263
> ip_protocol_deliver_rcu+0x46/0x280 net/ipv4/ip_input.c:207
> ip_local_deliver_finish+0x190/0x270 net/ipv4/ip_input.c:241
> NF_HOOK include/linux/netfilter.h:318 [inline]
> NF_HOOK include/linux/netfilter.h:312 [inline]
> ip_local_deliver+0xe3/0x210 net/ipv4/ip_input.c:262
> dst_input include/net/dst.h:480 [inline]
> ip_rcv_finish net/ipv4/ip_input.c:492 [inline]
> NF_HOOK include/linux/netfilter.h:318 [inline]
> NF_HOOK include/linux/netfilter.h:312 [inline]
> ip_rcv+0x200/0x220 net/ipv4/ip_input.c:612
> __netif_receive_skb_one_core+0xeb/0x110 net/core/dev.c:6178
> __netif_receive_skb+0x1f/0xc0 net/core/dev.c:6291
> process_backlog+0x168/0x360 net/core/dev.c:6642
> __napi_poll+0x71/0x460 net/core/dev.c:7706
> napi_poll net/core/dev.c:7769 [inline]
> net_rx_action+0x6f8/0x810 net/core/dev.c:7926
> handle_softirqs+0xc9/0x2e0 kernel/softirq.c:622
> run_ksoftirqd kernel/softirq.c:1063 [inline]
> run_ksoftirqd+0x20/0x30 kernel/softirq.c:1055
> smpboot_thread_fn+0x287/0x520 kernel/smpboot.c:160
> kthread+0x1f2/0x240 kernel/kthread.c:436
> ret_from_fork+0x321/0x440 arch/x86/kernel/process.c:158
> ret_from_fork_asm+0x1a/0x30 arch/x86/entry/entry_64.S:245
>
> read to 0xffff888015e8e5f0 of 8 bytes by task 24 on cpu 1:
> mptcp_rtx_head net/mptcp/protocol.h:487 [inline]
> __mptcp_retrans+0x169/0x8f0 net/mptcp/protocol.c:2759
> mptcp_worker+0x6a6/0xb30 net/mptcp/protocol.c:2980
> process_one_work+0x3ee/0x970 kernel/workqueue.c:3275
> process_scheduled_works kernel/workqueue.c:3358 [inline]
> worker_thread+0x3c3/0x730 kernel/workqueue.c:3439
> kthread+0x1f2/0x240 kernel/kthread.c:436
> ret_from_fork+0x321/0x440 arch/x86/kernel/process.c:158
> ret_from_fork_asm+0x1a/0x30 arch/x86/entry/entry_64.S:245
I think this race is not present in the current tree, after commit
96d846e3e2a7 ("mptcp: let the retrans scheduler do its job").
/P
^ permalink raw reply [flat|nested] 13+ messages in thread
* Re: [PATCH mptcp-net v3 3/3] tcp: fix data-race in do_recvmmsg / mptcp_recvmsg
2026-09-07 10:15 ` Quanye Yang
(?)
(?)
@ 2026-09-07 13:46 ` Paolo Abeni
2026-09-07 14:38 ` quanyeyang
-1 siblings, 1 reply; 13+ messages in thread
From: Paolo Abeni @ 2026-09-07 13:46 UTC (permalink / raw)
To: quanyeyang, MPTCP Linux; +Cc: geliang, matttbe
On 9/7/26 12:15 PM, Quanye Yang via B4 Relay wrote:
> From: Quanye Yang <quanyeyang@proton.me>
>
> KCSAN reported a data race between do_recvmmsg() and
> mptcp_recvmsg() on sk->sk_err.
>
> do_recvmmsg() calls sock_error() without holding the socket lock.
> sock_error() atomically clears sk_err using xchg(), which can race
> with the plain read in tcp_recv_should_stop(), even when its caller
> holds the socket lock.
It looks like the same data-race is present for plain TCP betweem i.e.
multiple concurrent recvmmsg() reader? If so, I think this patch
could/should go separately directly into the net tree.
Also please include the full KCSAN splat.
/P
^ permalink raw reply [flat|nested] 13+ messages in thread
* Re: [PATCH mptcp-net v3 1/3] mptcp: fix data-race in __mptcp_retrans / mptcp_incoming_options
2026-09-07 13:38 ` Paolo Abeni
@ 2026-09-07 13:55 ` Matthieu Baerts
0 siblings, 0 replies; 13+ messages in thread
From: Matthieu Baerts @ 2026-09-07 13:55 UTC (permalink / raw)
To: Paolo Abeni; +Cc: Geliang Tang, quanyeyang, MPTCP Linux
Hi Paolo,
On 07/09/2026 15:38, Paolo Abeni wrote:
> On 9/7/26 12:15 PM, quanyeyang@proton.me wrote:
>> From: "Matthieu Baerts (NGI0)" <matttbe@kernel.org>
>>
>> SyzKaller found this data-race:
>>
>> BUG: KCSAN: data-race in __mptcp_retrans / mptcp_incoming_options
>>
>> write (marked) to 0xffff888015e8e5f0 of 8 bytes by interrupt on cpu 0:
>> __mptcp_snd_una_update net/mptcp/options.c:1055 [inline]
>> mptcp_incoming_options+0x6a3/0x1ac0 net/mptcp/options.c:1183
>> tcp_data_queue+0x101b/0x2440 net/ipv4/tcp_input.c:5583
>> tcp_rcv_established+0x684/0x1fc0 net/ipv4/tcp_input.c:6654
>> tcp_v4_do_rcv+0x35c/0x690 net/ipv4/tcp_ipv4.c:1866
>> tcp_v4_rcv+0x1d91/0x25a0 net/ipv4/tcp_ipv4.c:2263
>> ip_protocol_deliver_rcu+0x46/0x280 net/ipv4/ip_input.c:207
>> ip_local_deliver_finish+0x190/0x270 net/ipv4/ip_input.c:241
>> NF_HOOK include/linux/netfilter.h:318 [inline]
>> NF_HOOK include/linux/netfilter.h:312 [inline]
>> ip_local_deliver+0xe3/0x210 net/ipv4/ip_input.c:262
>> dst_input include/net/dst.h:480 [inline]
>> ip_rcv_finish net/ipv4/ip_input.c:492 [inline]
>> NF_HOOK include/linux/netfilter.h:318 [inline]
>> NF_HOOK include/linux/netfilter.h:312 [inline]
>> ip_rcv+0x200/0x220 net/ipv4/ip_input.c:612
>> __netif_receive_skb_one_core+0xeb/0x110 net/core/dev.c:6178
>> __netif_receive_skb+0x1f/0xc0 net/core/dev.c:6291
>> process_backlog+0x168/0x360 net/core/dev.c:6642
>> __napi_poll+0x71/0x460 net/core/dev.c:7706
>> napi_poll net/core/dev.c:7769 [inline]
>> net_rx_action+0x6f8/0x810 net/core/dev.c:7926
>> handle_softirqs+0xc9/0x2e0 kernel/softirq.c:622
>> run_ksoftirqd kernel/softirq.c:1063 [inline]
>> run_ksoftirqd+0x20/0x30 kernel/softirq.c:1055
>> smpboot_thread_fn+0x287/0x520 kernel/smpboot.c:160
>> kthread+0x1f2/0x240 kernel/kthread.c:436
>> ret_from_fork+0x321/0x440 arch/x86/kernel/process.c:158
>> ret_from_fork_asm+0x1a/0x30 arch/x86/entry/entry_64.S:245
>>
>> read to 0xffff888015e8e5f0 of 8 bytes by task 24 on cpu 1:
>> mptcp_rtx_head net/mptcp/protocol.h:487 [inline]
>> __mptcp_retrans+0x169/0x8f0 net/mptcp/protocol.c:2759
>> mptcp_worker+0x6a6/0xb30 net/mptcp/protocol.c:2980
>> process_one_work+0x3ee/0x970 kernel/workqueue.c:3275
>> process_scheduled_works kernel/workqueue.c:3358 [inline]
>> worker_thread+0x3c3/0x730 kernel/workqueue.c:3439
>> kthread+0x1f2/0x240 kernel/kthread.c:436
>> ret_from_fork+0x321/0x440 arch/x86/kernel/process.c:158
>> ret_from_fork_asm+0x1a/0x30 arch/x86/entry/entry_64.S:245
>
> I think this race is not present in the current tree, after commit
> 96d846e3e2a7 ("mptcp: let the retrans scheduler do its job").
Thank you for having checked. Good point, this old patch is maybe
outdated, I didn't check.
Note that on syzkaller side, the last occurrence I had for this issue
was on the 13th of May, and your patch was in on tree on the 3rd of
June. I guess it is indeed not needed then. (I don't know why I didn't
see it after the 13th of May, but there were no reproducers.)
Cheers,
Matt
--
Sponsored by the NGI0 Core fund.
^ permalink raw reply [flat|nested] 13+ messages in thread
* Re: [PATCH mptcp-net v3 3/3] tcp: fix data-race in do_recvmmsg / mptcp_recvmsg
2026-09-07 13:46 ` Paolo Abeni
@ 2026-09-07 14:38 ` quanyeyang
0 siblings, 0 replies; 13+ messages in thread
From: quanyeyang @ 2026-09-07 14:38 UTC (permalink / raw)
To: Paolo Abeni; +Cc: MPTCP Linux, geliang, matttbe
On Monday, September 7th, 2026 at AM 6:47, Paolo Abeni <pabeni@redhat.com> wrote:
> On 9/7/26 12:15 PM, Quanye Yang via B4 Relay wrote:
> > From: Quanye Yang <quanyeyang@proton.me>
> >
> > KCSAN reported a data race between do_recvmmsg() and
> > mptcp_recvmsg() on sk->sk_err.
> >
> > do_recvmmsg() calls sock_error() without holding the socket lock.
> > sock_error() atomically clears sk_err using xchg(), which can race
> > with the plain read in tcp_recv_should_stop(), even when its caller
> > holds the socket lock.
>
> It looks like the same data-race is present for plain TCP betweem i.e.
> multiple concurrent recvmmsg() reader? If so, I think this patch
> could/should go separately directly into the net tree.
>
> Also please include the full KCSAN splat.
>
> /P
>
>
Thanks for the feedback.
Patch 2/3 is independent from the other two patches and can be applied
on its own. Patch 1/3 can be dropped, and I will handle Patch 3/3
separately as a direct net submission.
Please let me know if you would still prefer a single-patch v4 for
Patch 2/3.
Thanks,
Quanye
^ permalink raw reply [flat|nested] 13+ messages in thread
end of thread, other threads:[~2026-09-07 14:38 UTC | newest]
Thread overview: 13+ messages (download: mbox.gz follow: Atom feed
-- links below jump to the message on this page --
2026-09-07 10:15 [PATCH mptcp-net v3 0/3] mptcp: fix reported data-races quanyeyang
2026-09-07 10:15 ` [PATCH mptcp-net v3 1/3] mptcp: fix data-race in __mptcp_retrans / mptcp_incoming_options quanyeyang
2026-09-07 13:38 ` Paolo Abeni
2026-09-07 13:55 ` Matthieu Baerts
2026-09-07 10:15 ` [PATCH mptcp-net v3 2/3] mptcp: fix data-race in mptcp_subflow_get_send / tcp_ack Quanye Yang via B4 Relay
2026-09-07 10:15 ` Quanye Yang
2026-09-07 10:15 ` [PATCH mptcp-net v3 3/3] tcp: fix data-race in do_recvmmsg / mptcp_recvmsg Quanye Yang via B4 Relay
2026-09-07 10:15 ` Quanye Yang
2026-09-07 10:24 ` sashiko-bot
2026-09-07 13:23 ` quanyeyang
2026-09-07 13:46 ` Paolo Abeni
2026-09-07 14:38 ` quanyeyang
2026-09-07 11:28 ` [PATCH mptcp-net v3 0/3] mptcp: fix reported data-races MPTCP CI
This is an external index of several public inboxes,
see mirroring instructions on how to clone and mirror
all data and code used by this external index.