All of lore.kernel.org
 help / color / mirror / Atom feed
From: Glenn Judd <gmj@meta.com>
To: "David S. Miller" <davem@davemloft.net>,
	Eric Dumazet <edumazet@google.com>,
	Jakub Kicinski <kuba@kernel.org>, Paolo Abeni <pabeni@redhat.com>,
	<netdev@vger.kernel.org>
Cc: Simon Horman <horms@kernel.org>,
	Willem de Bruijn <willemb@google.com>,
	Kuniyuki Iwashima <kuniyu@google.com>,
	Richard Gobert <richardbgobert@gmail.com>,
	Kees Cook <kees@kernel.org>,
	Jiayuan Chen <jiayuan.chen@linux.dev>,
	<linux-kernel@vger.kernel.org>, Glenn Judd <gmj@meta.com>
Subject: [RFC PATCH net-next] net: gro: coalesce padded small IPv4 TCP segments
Date: Fri, 31 Jul 2026 11:54:31 -0700	[thread overview]
Message-ID: <20260731185431.2777685-1-gmj@meta.com> (raw)

Software GRO fails to coalesce small IPv4/TCP segment that was
padded up to the 60-byte minimum Ethernet frame.

The selftest tools/testing/selftests/drivers/net/hw/gro.py subtest
sw_ipv4_data_lrg_1byte sends {100, 1} expecting to receive {101}.
In current code, it receives {100, 1} (no coalescing) instead.

Cause: inet_gro_receive() computes its flush term from
tot_len ^ skb_gro_len() before skb_gro_pull(), while skb_gro_len()
still includes trailing Ethernet padding. A small IPv4/TCP segment
padded up to the 60-byte minimum frame has tot_len != skb_gro_len(),
so flush is set and the runt never coalesces.

Assisted-by: Claude:claude-opus-4-8
Assisted-by: Codex:gpt-5.6
Assisted-by: Meta:internal-AI-tooling
Signed-off-by: Glenn Judd <gmj@meta.com>
---

Notes:
    RFC notes
    ---------
    Per Jakub Kicinski, the open question is fast-path cost: this adds two
    operations to the common IPv4 GRO path for every packet -- reading
    iph->tot_len and the skb_gro_len() comparison.  Everything expensive
    (linear check, trim, pointer refresh, csum recompute) is behind unlikely()
    on the slow path.  Is that per-packet cost worth the coalescing win for
    padded runts?
    
    Testing: netdevsim cannot reproduce this -- it never pads short frames to
    ETH_ZLEN -- so sw_ipv4_data_lrg_1byte passes trivially there.  Reproduced
    and fixed on a real NIC (cx7): baseline FAIL -> patched PASS.  Also
    validated locally under KASAN + CONFIG_FAIL_SKB_REALLOC (no UAF).

 net/ipv4/af_inet.c | 20 ++++++++++++++++++++
 1 file changed, 20 insertions(+)

diff --git a/net/ipv4/af_inet.c b/net/ipv4/af_inet.c
index 32d006c1a8ee..998ff77fd7b9 100644
--- a/net/ipv4/af_inet.c
+++ b/net/ipv4/af_inet.c
@@ -1470,6 +1470,7 @@ struct sk_buff *inet_gro_receive(struct list_head *head, struct sk_buff *skb)
 	const struct net_offload *ops;
 	struct sk_buff *pp = NULL;
 	const struct iphdr *iph;
+	unsigned int tot_len;
 	struct sk_buff *p;
 	unsigned int hlen;
 	unsigned int off;
@@ -1498,6 +1499,25 @@ struct sk_buff *inet_gro_receive(struct list_head *head, struct sk_buff *skb)
 		goto out;
 
 	NAPI_GRO_CB(skb)->proto = proto;
+
+	tot_len = ntohs(iph->tot_len);
+	if (unlikely(skb_gro_len(skb) > tot_len)) {
+		if (!skb_is_nonlinear(skb)) {
+			if (tot_len < sizeof(*iph) ||
+			    pskb_trim_rcsum(skb, off + tot_len))
+				goto out;
+
+			NAPI_GRO_CB(skb)->frag0 = skb->data;
+			NAPI_GRO_CB(skb)->frag0_len = skb->len;
+			iph = skb_gro_header(skb, hlen, off);
+			if (unlikely(!iph))
+				goto out;
+			if (skb->ip_summed == CHECKSUM_COMPLETE)
+				NAPI_GRO_CB(skb)->csum =
+					skb_checksum(skb, off, tot_len, 0);
+		}
+	}
+
 	flush = (u16)((ntohl(*(__be32 *)iph) ^ skb_gro_len(skb)) | (ntohl(*(__be32 *)&iph->id) & ~IP_DF));
 
 	list_for_each_entry(p, head, list) {

base-commit: 2fbade66245059c78daeaccfce13ecf499fffb51
-- 
2.53.0-Meta


             reply	other threads:[~2026-07-31 18:54 UTC|newest]

Thread overview: 6+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-07-31 18:54 Glenn Judd [this message]
2026-08-10 12:11 ` [RFC PATCH net-next] net: gro: coalesce padded small IPv4 TCP segments Richard Gobert
2026-08-13 19:14   ` Glenn Judd
2026-08-13 19:54 ` Eric Dumazet
2026-08-13 20:00   ` Eric Dumazet
2026-08-16  1:44     ` Glenn Judd

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=20260731185431.2777685-1-gmj@meta.com \
    --to=gmj@meta.com \
    --cc=davem@davemloft.net \
    --cc=edumazet@google.com \
    --cc=horms@kernel.org \
    --cc=jiayuan.chen@linux.dev \
    --cc=kees@kernel.org \
    --cc=kuba@kernel.org \
    --cc=kuniyu@google.com \
    --cc=linux-kernel@vger.kernel.org \
    --cc=netdev@vger.kernel.org \
    --cc=pabeni@redhat.com \
    --cc=richardbgobert@gmail.com \
    --cc=willemb@google.com \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is an external index of several public inboxes,
see mirroring instructions on how to clone and mirror
all data and code used by this external index.