* [PATCH] mptcp: push queued data on passive TFO subflows becoming established
@ 2026-10-08 9:52 T S Rameshkumar
2026-10-08 9:59 ` Petar Sakic
0 siblings, 1 reply; 3+ messages in thread
From: T S Rameshkumar @ 2026-10-08 9:52 UTC (permalink / raw)
To: Matthieu Baerts, Mat Martineau, Geliang Tang
Cc: netdev, mptcp, linux-kernel, Petar Sakic, T S Rameshkumar
With TCP Fast Open on an MPTCP listener, if the server application
writes data while the passive subflow is still in SYN_RECV (after
consuming the client's SYN data but before the MP_CAPABLE third ACK
arrives), __mptcp_subflow_active() refuses transmission and the data
is queued into the msk write queue.
When the MPC third ACK arrives, check_fully_established() marks the
subflow established, but because the third ACK carries no DSS data,
the queued bytes remain stranded until the peer sends more data.
Fix this by:
1. Invoking __mptcp_check_push() in check_fully_established() when
subflow->is_mptfo is set.
2. Setting MPTCP_PUSH_PENDING and scheduling the MPTCP worker in
subflow_state_change() so that once the underlying subflow transitions
to TCP_ESTABLISHED, pending queued bytes are immediately flushed.
Reported-by: Petar Sakic <petar.sakic@ink.fish>
Closes: https://lore.kernel.org/netdev/CAFPPu1gU2Y-D+d4i3F0MoNkYK+e1U+=X3qf6QycjfKBw+8snPg@mail.gmail.com/
Fixes: e00b63056fb4 ("fastopen: only mark MPTFO subflows with SYN data")
Signed-off-by: T S Rameshkumar <rameshkumar.t@phytecembedded.in>
---
net/mptcp/options.c | 7 +++++++
net/mptcp/subflow.c | 5 +++++
2 files changed, 12 insertions(+)
diff --git a/net/mptcp/options.c b/net/mptcp/options.c
index ce0de02f5..d5238fa11 100644
--- a/net/mptcp/options.c
+++ b/net/mptcp/options.c
@@ -1042,6 +1042,13 @@ static bool check_fully_established(struct mptcp_sock *msk, struct sock *ssk,
mptcp_data_lock((struct sock *)msk);
__mptcp_subflow_fully_established(msk, subflow, mp_opt);
+ /* Passive TFO: the application may have written data while the
+ * subflow was still in SYN_RECV; __mptcp_subflow_active() refused
+ * it then and nothing else spools the msk write queue when the
+ * MPC third ack (no DSS) arrives. Push it now.
+ */
+ if (subflow->is_mptfo)
+ __mptcp_check_push((struct sock *)msk, ssk);
mptcp_data_unlock((struct sock *)msk);
check_notify:
diff --git a/net/mptcp/subflow.c b/net/mptcp/subflow.c
index f0a6725d2..f499073a6 100644
--- a/net/mptcp/subflow.c
+++ b/net/mptcp/subflow.c
@@ -1894,6 +1894,11 @@ static void subflow_state_change(struct sock *sk)
if (subflow->resetting)
return;
+ if (subflow->is_mptfo) {
+ set_bit(MPTCP_PUSH_PENDING, &mptcp_sk(parent)->cb_flags);
+ mptcp_schedule_work(parent);
+ }
+
/* as recvmsg() does not acquire the subflow socket for ssk selection
* a fin packet carrying a DSS can be unnoticed if we don't trigger
* the data available machinery here.
--
2.34.1
^ permalink raw reply related [flat|nested] 3+ messages in thread* Re: [PATCH] mptcp: push queued data on passive TFO subflows becoming established
2026-10-08 9:52 [PATCH] mptcp: push queued data on passive TFO subflows becoming established T S Rameshkumar
@ 2026-10-08 9:59 ` Petar Sakic
0 siblings, 0 replies; 3+ messages in thread
From: Petar Sakic @ 2026-10-08 9:59 UTC (permalink / raw)
To: T S Rameshkumar
Cc: Matthieu Baerts, Mat Martineau, Geliang Tang, netdev, mptcp,
linux-kernel, T S Rameshkumar
Hi T S,
Thanks for picking this up. One note on the Fixes tag: we reproduced
the hang on 6.8 as well as 6.12.111, 6.18.15 and 7.2.6, so it predates
e00b63056fb4 (Aug 2026). That commit is related but not the origin.
Cheers
- Petar
On Thu, Oct 8, 2026 at 11:53 AM T S Rameshkumar <rameshsv06@gmail.com> wrote:
>
> With TCP Fast Open on an MPTCP listener, if the server application
> writes data while the passive subflow is still in SYN_RECV (after
> consuming the client's SYN data but before the MP_CAPABLE third ACK
> arrives), __mptcp_subflow_active() refuses transmission and the data
> is queued into the msk write queue.
>
> When the MPC third ACK arrives, check_fully_established() marks the
> subflow established, but because the third ACK carries no DSS data,
> the queued bytes remain stranded until the peer sends more data.
>
> Fix this by:
> 1. Invoking __mptcp_check_push() in check_fully_established() when
> subflow->is_mptfo is set.
> 2. Setting MPTCP_PUSH_PENDING and scheduling the MPTCP worker in
> subflow_state_change() so that once the underlying subflow transitions
> to TCP_ESTABLISHED, pending queued bytes are immediately flushed.
>
> Reported-by: Petar Sakic <petar.sakic@ink.fish>
> Closes: https://lore.kernel.org/netdev/CAFPPu1gU2Y-D+d4i3F0MoNkYK+e1U+=X3qf6QycjfKBw+8snPg@mail.gmail.com/
> Fixes: e00b63056fb4 ("fastopen: only mark MPTFO subflows with SYN data")
> Signed-off-by: T S Rameshkumar <rameshkumar.t@phytecembedded.in>
> ---
> net/mptcp/options.c | 7 +++++++
> net/mptcp/subflow.c | 5 +++++
> 2 files changed, 12 insertions(+)
>
> diff --git a/net/mptcp/options.c b/net/mptcp/options.c
> index ce0de02f5..d5238fa11 100644
> --- a/net/mptcp/options.c
> +++ b/net/mptcp/options.c
> @@ -1042,6 +1042,13 @@ static bool check_fully_established(struct mptcp_sock *msk, struct sock *ssk,
>
> mptcp_data_lock((struct sock *)msk);
> __mptcp_subflow_fully_established(msk, subflow, mp_opt);
> + /* Passive TFO: the application may have written data while the
> + * subflow was still in SYN_RECV; __mptcp_subflow_active() refused
> + * it then and nothing else spools the msk write queue when the
> + * MPC third ack (no DSS) arrives. Push it now.
> + */
> + if (subflow->is_mptfo)
> + __mptcp_check_push((struct sock *)msk, ssk);
> mptcp_data_unlock((struct sock *)msk);
>
> check_notify:
> diff --git a/net/mptcp/subflow.c b/net/mptcp/subflow.c
> index f0a6725d2..f499073a6 100644
> --- a/net/mptcp/subflow.c
> +++ b/net/mptcp/subflow.c
> @@ -1894,6 +1894,11 @@ static void subflow_state_change(struct sock *sk)
> if (subflow->resetting)
> return;
>
> + if (subflow->is_mptfo) {
> + set_bit(MPTCP_PUSH_PENDING, &mptcp_sk(parent)->cb_flags);
> + mptcp_schedule_work(parent);
> + }
> +
> /* as recvmsg() does not acquire the subflow socket for ssk selection
> * a fin packet carrying a DSS can be unnoticed if we don't trigger
> * the data available machinery here.
> --
> 2.34.1
>
^ permalink raw reply [flat|nested] 3+ messages in thread
* Bug report mptcp: passive TFO: data written by the server before the MPC third ACK is never sent
@ 2026-10-08 8:59 Petar Sakic
2026-10-08 10:05 ` [PATCH] mptcp: push queued data on passive TFO subflows becoming established T S Rameshkumar
0 siblings, 1 reply; 3+ messages in thread
From: Petar Sakic @ 2026-10-08 8:59 UTC (permalink / raw)
To: mptcp; +Cc: netdev, matttbe, martineau, geliang
Hi linux dev team,
There is bug I'd like to report:
Summary:
With TCP Fast Open on an MPTCP listener, if the server application
writes data while the passive subflow is still in SYN_RECV (i.e. after
it has consumed the client's SYN data but before the client's
MP_CAPABLE third ACK arrives), that data stays in the msk write queue
and is never transmitted. The connection hangs until the client gives
up and closes. Plain TCP with TFO is not affected.
Reproduced on:
Server kernels 6.8, 6.12.111, 6.18.15 and 7.2.6 (Debian builds),
client 6.12. Source read at 7.3-rc6 (602042bf, 7 Oct 2026). Lab:
client, NAT hop and server in separate network namespaces, ~120 ms
RTT.
Steps:
1. Server: `net.ipv4.tcp_fastopen=3`, an MPTCP listener with
`TCP_FASTOPEN`, an application that reads the request and replies
immediately (we used a proxy whose upstream answers in < 1 ms).
2. Client: `net.ipv4.tcp_fastopen=3`, MPTCP socket with
`TCP_FASTOPEN_CONNECT`; prime the cookie with one connection.
3. Second connection: the request fits in the SYN. Server replies
before the client's third ACK arrives.
Expected: the reply is sent once the subflow is fully established.
Actual: server `ss -M` shows FIN-WAIT-1 with the reply bytes queued
and `bytes_sent` unchanged; the wire shows the reply leave only after
the client closes. If the server's reply comes later than one RTT
(slow upstream), everything works.
Analysis:
- `__mptcp_subflow_active()` / `__tcp_can_send()` skip a subflow that
is not yet established, so the server's early write is queued on the
msk.
- When the MPC third ACK arrives, `check_fully_established()`
(options.c) marks the subflow fully established but nothing pushes the
msk write queue. The third ACK carries no DSS, so no later event
triggers a push until the peer sends data, and the peer is waiting for
the reply.
Candidate fix: (tested only in the lab: 30/30 requests succeed with
TFO in use on a patched 6.12.111)
```diff
--- a/net/mptcp/options.c
+++ b/net/mptcp/options.c
@@ -1058,6 +1058,13 @@ set_fully_established:
mptcp_data_lock((struct sock *)msk);
__mptcp_subflow_fully_established(msk, subflow, mp_opt);
+ /* Passive TFO: the application may have written data while the
+ * subflow was still in SYN_RECV; __mptcp_subflow_active() refused
+ * it then and nothing else spools the msk write queue when the
+ * MPC third ack (no DSS) arrives. Push it now.
+ */
+ if (subflow->is_mptfo)
+ __mptcp_check_push((struct sock *)msk, ssk);
mptcp_data_unlock((struct sock *)msk);
check_notify:
```
We are not kernel developers; the maintainers will know whether this
belongs here or in the fully-established path. Related but different:
e00b63056fb4 ("fastopen: only mark MPTFO subflows with SYN data", Aug
2026).
Impact we saw in production-like use
Shadowsocks-2022 proxies over MPTCP (xray) put the whole request in
the first write, so small requests ride the SYN; requests to targets
near the server hang. We have turned TFO off.
Thank you
Kind regards
--
Petar Sakic
ETO/AV-IT Engineer
Inkfish | www.ink.fish
Deeply curious
This e-mail is for the sole use of the intended recipient(s). It
contains information that may be confidential. If you believe that it
has been sent to you in error, please notify the sender immediately by
reply e-mail and destroy all copies of this message. Any disclosure,
copying, distribution, or use of this information by anyone other than
the intended recipient is prohibited.
^ permalink raw reply [flat|nested] 3+ messages in thread* Re: [PATCH] mptcp: push queued data on passive TFO subflows becoming established
2026-10-08 8:59 Bug report mptcp: passive TFO: data written by the server before the MPC third ACK is never sent Petar Sakic
@ 2026-10-08 10:05 ` T S Rameshkumar
0 siblings, 0 replies; 3+ messages in thread
From: T S Rameshkumar @ 2026-10-08 10:05 UTC (permalink / raw)
To: Petar Sakic
Cc: Matthieu Baerts, Mat Martineau, Geliang Tang, netdev, mptcp,
linux-kernel
Hi Petar, Matthieu, Mat,
Thanks for the feedback and confirmation!
Indeed, since the bug predates e00b63056fb4 and exists back to when server-side
Fast Open was originally introduced, the appropriate Fixes tag would be:
Fixes: fb7084501a61 ("mptcp: add support for TCP_FASTOPEN sockopt")
During our testing with simulated network latency in namespaces, we also found
that when the 3rd ACK arrives:
1. check_fully_established() marks subflow->fully_established = true.
2. But within tcp_ack(), the underlying TCP state is still TCP_SYN_RECV until
tcp_rcv_state_process() sets it to TCP_ESTABLISHED.
3. Therefore, invoking the push in subflow_state_change() ensures the subflow
is recognized as active by __tcp_can_send() and immediately flushes the
queued write buffer.
I can send a [PATCH net v2] with the updated Fixes tag once maintainers review.
Thanks,
T S Rameshkumar
^ permalink raw reply [flat|nested] 3+ messages in thread
end of thread, other threads:[~2026-10-08 10:06 UTC | newest]
Thread overview: 3+ messages (download: mbox.gz follow: Atom feed
-- links below jump to the message on this page --
2026-10-08 9:52 [PATCH] mptcp: push queued data on passive TFO subflows becoming established T S Rameshkumar
2026-10-08 9:59 ` Petar Sakic
-- strict thread matches above, loose matches on Subject: below --
2026-10-08 8:59 Bug report mptcp: passive TFO: data written by the server before the MPC third ACK is never sent Petar Sakic
2026-10-08 10:05 ` [PATCH] mptcp: push queued data on passive TFO subflows becoming established T S Rameshkumar
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox