* [PATCH net] tcp: prevent collapsing skbs across boundary in rtx queue
@ 2026-09-24 15:44 Willem de Bruijn
2026-09-24 16:37 ` Eric Dumazet
` (2 more replies)
0 siblings, 3 replies; 4+ messages in thread
From: Willem de Bruijn @ 2026-09-24 15:44 UTC (permalink / raw)
To: netdev
Cc: davem, kuba, edumazet, pabeni, horms, andrew+netdev, dzahka,
john.fastabend, sd, Willem de Bruijn, stable
From: Willem de Bruijn <willemb@google.com>
tcp_write_collapse_fence() sets TCP_SKB_CB(skb)->eor = 1 on
tcp_write_queue_tail(sk) to prevent skbs queued after a switch to
device encryption from being collapsed into earlier skbs.
The fence is a no-op if all earlier data has already been transmitted
when the switch happens: sk->sk_write_queue is empty. The not yet
acknowledged earlier skbs wait in sk->tcp_rtx_queue with eor 0.
On a subsequent retransmit or SACK shift, tcp_retrans_try_collapse() or
tcp_shift_skb_data() can then merge an skb queued after the switch into
one queued before it.
Both users of the fence are affected:
- psp: devices only encrypt skbs with skb->decrypted set. The merged skb
keeps decrypted = 0 from the earlier skb, so merged data sent after
psp_sock_assoc_set_tx() is retransmitted in cleartext.
- tls device offload: the merged skb straddles the start marker set in
tls_set_device_offload(). The software fallback (fill_sg_in() returns
-EINVAL) and the mlx5, nfp and funeth drivers cannot handle such an
skb and drop it. Every retransmit rebuilds the same skb, so the
connection stalls.
Fix this in two places, for defense in depth:
1. Fall back to tcp_rtx_queue_tail(sk) in tcp_write_collapse_fence()
when tcp_write_queue_tail(sk) is NULL.
2. Check !skb_cmp_decrypted(to, from) in tcp_skb_can_collapse(), as
tcp_skb_can_collapse_rx() does on receive. skb_shift(), which both
collapse paths call, already has a DEBUG_NET_WARN_ON_ONCE() for this
condition.
Fixes: e8f69799810c ("net/tls: Add generic NIC offload infrastructure")
Cc: stable@vger.kernel.org
Signed-off-by: Willem de Bruijn <willemb@google.com>
---
include/net/tcp.h | 4 ++--
1 file changed, 2 insertions(+), 2 deletions(-)
diff --git a/include/net/tcp.h b/include/net/tcp.h
index 436495ff2271..4416cdf9bf30 100644
--- a/include/net/tcp.h
+++ b/include/net/tcp.h
@@ -1232,9 +1232,9 @@ static inline bool tcp_skb_can_collapse_to(const struct sk_buff *skb)
static inline bool tcp_skb_can_collapse(const struct sk_buff *to,
const struct sk_buff *from)
{
- /* skb_cmp_decrypted() not needed, use tcp_write_collapse_fence() */
return likely(tcp_skb_can_collapse_to(to) &&
mptcp_skb_can_collapse(to, from) &&
+ !skb_cmp_decrypted(to, from) &&
skb_pure_zcopy_same(to, from) &&
skb_frags_readable(to) == skb_frags_readable(from));
}
@@ -2327,7 +2327,7 @@ static inline void tcp_rtx_queue_unlink_and_free(struct sk_buff *skb, struct soc
static inline void tcp_write_collapse_fence(struct sock *sk)
{
- struct sk_buff *skb = tcp_write_queue_tail(sk);
+ struct sk_buff *skb = tcp_write_queue_tail(sk) ?: tcp_rtx_queue_tail(sk);
if (skb)
TCP_SKB_CB(skb)->eor = 1;
--
2.56.0.rc1.310.g51773c2048-goog
^ permalink raw reply related [flat|nested] 4+ messages in thread
* Re: [PATCH net] tcp: prevent collapsing skbs across boundary in rtx queue
2026-09-24 15:44 [PATCH net] tcp: prevent collapsing skbs across boundary in rtx queue Willem de Bruijn
@ 2026-09-24 16:37 ` Eric Dumazet
2026-09-24 16:53 ` Daniel Zahka
2026-09-24 18:10 ` patchwork-bot+netdevbpf
2 siblings, 0 replies; 4+ messages in thread
From: Eric Dumazet @ 2026-09-24 16:37 UTC (permalink / raw)
To: Willem de Bruijn
Cc: netdev, davem, kuba, pabeni, horms, andrew+netdev, dzahka,
john.fastabend, sd, Willem de Bruijn, stable
On Thu, Sep 24, 2026 at 5:44 PM Willem de Bruijn
<willemdebruijn.kernel@gmail.com> wrote:
>
> From: Willem de Bruijn <willemb@google.com>
>
> tcp_write_collapse_fence() sets TCP_SKB_CB(skb)->eor = 1 on
> tcp_write_queue_tail(sk) to prevent skbs queued after a switch to
> device encryption from being collapsed into earlier skbs.
>
> The fence is a no-op if all earlier data has already been transmitted
> when the switch happens: sk->sk_write_queue is empty. The not yet
> acknowledged earlier skbs wait in sk->tcp_rtx_queue with eor 0.
>
> On a subsequent retransmit or SACK shift, tcp_retrans_try_collapse() or
> tcp_shift_skb_data() can then merge an skb queued after the switch into
> one queued before it.
>
> Both users of the fence are affected:
>
> - psp: devices only encrypt skbs with skb->decrypted set. The merged skb
> keeps decrypted = 0 from the earlier skb, so merged data sent after
> psp_sock_assoc_set_tx() is retransmitted in cleartext.
>
> - tls device offload: the merged skb straddles the start marker set in
> tls_set_device_offload(). The software fallback (fill_sg_in() returns
> -EINVAL) and the mlx5, nfp and funeth drivers cannot handle such an
> skb and drop it. Every retransmit rebuilds the same skb, so the
> connection stalls.
>
> Fix this in two places, for defense in depth:
>
> 1. Fall back to tcp_rtx_queue_tail(sk) in tcp_write_collapse_fence()
> when tcp_write_queue_tail(sk) is NULL.
>
> 2. Check !skb_cmp_decrypted(to, from) in tcp_skb_can_collapse(), as
> tcp_skb_can_collapse_rx() does on receive. skb_shift(), which both
> collapse paths call, already has a DEBUG_NET_WARN_ON_ONCE() for this
> condition.
>
> Fixes: e8f69799810c ("net/tls: Add generic NIC offload infrastructure")
> Cc: stable@vger.kernel.org
> Signed-off-by: Willem de Bruijn <willemb@google.com>
Reviewed-by: Eric Dumazet <edumazet@google.com>
Thanks!
^ permalink raw reply [flat|nested] 4+ messages in thread
* Re: [PATCH net] tcp: prevent collapsing skbs across boundary in rtx queue
2026-09-24 15:44 [PATCH net] tcp: prevent collapsing skbs across boundary in rtx queue Willem de Bruijn
2026-09-24 16:37 ` Eric Dumazet
@ 2026-09-24 16:53 ` Daniel Zahka
2026-09-24 18:10 ` patchwork-bot+netdevbpf
2 siblings, 0 replies; 4+ messages in thread
From: Daniel Zahka @ 2026-09-24 16:53 UTC (permalink / raw)
To: Willem de Bruijn, netdev
Cc: davem, kuba, edumazet, pabeni, horms, andrew+netdev, dzahka,
john.fastabend, sd, Willem de Bruijn, stable
On Thu Sep 24, 2026 at 11:44 AM EDT, Willem de Bruijn wrote:
> From: Willem de Bruijn <willemb@google.com>
>
> tcp_write_collapse_fence() sets TCP_SKB_CB(skb)->eor = 1 on
> tcp_write_queue_tail(sk) to prevent skbs queued after a switch to
> device encryption from being collapsed into earlier skbs.
>
> The fence is a no-op if all earlier data has already been transmitted
> when the switch happens: sk->sk_write_queue is empty. The not yet
> acknowledged earlier skbs wait in sk->tcp_rtx_queue with eor 0.
>
> On a subsequent retransmit or SACK shift, tcp_retrans_try_collapse() or
> tcp_shift_skb_data() can then merge an skb queued after the switch into
> one queued before it.
>
> Both users of the fence are affected:
>
> - psp: devices only encrypt skbs with skb->decrypted set. The merged skb
> keeps decrypted = 0 from the earlier skb, so merged data sent after
> psp_sock_assoc_set_tx() is retransmitted in cleartext.
>
> - tls device offload: the merged skb straddles the start marker set in
> tls_set_device_offload(). The software fallback (fill_sg_in() returns
> -EINVAL) and the mlx5, nfp and funeth drivers cannot handle such an
> skb and drop it. Every retransmit rebuilds the same skb, so the
> connection stalls.
>
> Fix this in two places, for defense in depth:
>
> 1. Fall back to tcp_rtx_queue_tail(sk) in tcp_write_collapse_fence()
> when tcp_write_queue_tail(sk) is NULL.
>
> 2. Check !skb_cmp_decrypted(to, from) in tcp_skb_can_collapse(), as
> tcp_skb_can_collapse_rx() does on receive. skb_shift(), which both
> collapse paths call, already has a DEBUG_NET_WARN_ON_ONCE() for this
> condition.
>
> Fixes: e8f69799810c ("net/tls: Add generic NIC offload infrastructure")
> Cc: stable@vger.kernel.org
> Signed-off-by: Willem de Bruijn <willemb@google.com>
> ---
Reviewed-by: Daniel Zahka <daniel.zahka@gmail.com>
^ permalink raw reply [flat|nested] 4+ messages in thread
* Re: [PATCH net] tcp: prevent collapsing skbs across boundary in rtx queue
2026-09-24 15:44 [PATCH net] tcp: prevent collapsing skbs across boundary in rtx queue Willem de Bruijn
2026-09-24 16:37 ` Eric Dumazet
2026-09-24 16:53 ` Daniel Zahka
@ 2026-09-24 18:10 ` patchwork-bot+netdevbpf
2 siblings, 0 replies; 4+ messages in thread
From: patchwork-bot+netdevbpf @ 2026-09-24 18:10 UTC (permalink / raw)
To: Willem de Bruijn
Cc: netdev, davem, kuba, edumazet, pabeni, horms, andrew+netdev,
dzahka, john.fastabend, sd, willemb, stable
Hello:
This patch was applied to netdev/net.git (main)
by Jakub Kicinski <kuba@kernel.org>:
On Thu, 24 Sep 2026 11:44:12 -0400 you wrote:
> From: Willem de Bruijn <willemb@google.com>
>
> tcp_write_collapse_fence() sets TCP_SKB_CB(skb)->eor = 1 on
> tcp_write_queue_tail(sk) to prevent skbs queued after a switch to
> device encryption from being collapsed into earlier skbs.
>
> The fence is a no-op if all earlier data has already been transmitted
> when the switch happens: sk->sk_write_queue is empty. The not yet
> acknowledged earlier skbs wait in sk->tcp_rtx_queue with eor 0.
>
> [...]
Here is the summary with links:
- [net] tcp: prevent collapsing skbs across boundary in rtx queue
https://git.kernel.org/netdev/net/c/fc6d80eb5044
You are awesome, thank you!
--
Deet-doot-dot, I am a bot.
https://korg.docs.kernel.org/patchwork/pwbot.html
^ permalink raw reply [flat|nested] 4+ messages in thread
end of thread, other threads:[~2026-09-24 18:11 UTC | newest]
Thread overview: 4+ messages (download: mbox.gz follow: Atom feed
-- links below jump to the message on this page --
2026-09-24 15:44 [PATCH net] tcp: prevent collapsing skbs across boundary in rtx queue Willem de Bruijn
2026-09-24 16:37 ` Eric Dumazet
2026-09-24 16:53 ` Daniel Zahka
2026-09-24 18:10 ` patchwork-bot+netdevbpf
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox