All of lore.kernel.org
 help / color / mirror / Atom feed
From: Ido Schimmel <idosch@nvidia.com>
To: Chengfeng Ye <nicoyip.dev@gmail.com>
Cc: David Ahern <dsahern@kernel.org>,
	netdev@vger.kernel.org, "David S . Miller" <davem@davemloft.net>,
	Eric Dumazet <edumazet@google.com>,
	Jakub Kicinski <kuba@kernel.org>, Paolo Abeni <pabeni@redhat.com>,
	Simon Horman <horms@kernel.org>, Xin Long <lucien.xin@gmail.com>,
	William Tu <u9012063@gmail.com>,
	linux-kernel@vger.kernel.org, stable@vger.kernel.org
Subject: Re: [PATCH net v4] ip_tunnel: reserve FOU/GUE headroom before encapsulation
Date: Sun, 6 Sep 2026 18:33:54 +0300	[thread overview]
Message-ID: <20260906153354.GB323313@shredder> (raw)
In-Reply-To: <20260904165544.1362052-1-nicoyip.dev@gmail.com>

On Sat, Sep 05, 2026 at 12:55:44AM +0800, Chengfeng Ye wrote:
> ip_tunnel_encap() expects its callers to reserve headroom based on
> ip_encap_hlen(). ip_tunnel_xmit() currently pushes FOU and GUE headers
> before it grows the skb headroom.
> 
> That becomes visible when ipgre_changelink() publishes UDP
> encapsulation before it updates the device headroom. The transmit path
> does not serialize with RTNL, so it can interleave as follows:
> 
>   CPU 0 (ipgre_changelink)        CPU 1 (ipgre_xmit)
>   install GUE encapsulation
>                                   reserve the old needed_headroom
>   publish larger GRE flags
>   update tunnel->tun_hlen
>                                   push the larger GRE header
>                                   push the GUE and UDP headers
>   update dev->needed_headroom
> 
> With REMCSUM, the new layout can push 16 bytes of GRE and 20 bytes of
> GUE/UDP headers into an skb with only 32 bytes of actual headroom. The
> final UDP push writes four bytes before skb->head.
> 
> With the update window widened, the kernel reported:
> 
>   skbuff: skb_under_panic: ... len:128 put:8 ... dev:gre0poc
>   kernel BUG at net/core/skbuff.c:214!
>   Oops: invalid opcode: 0000 [#1] SMP KASAN NOPTI
>   Call Trace:
>    skb_push
>    fou_build_udp
>    gue_build_header
>    ip_tunnel_xmit
>    __gre_xmit
>    ipgre_xmit
> 
> Snapshot tunnel->encap, route and perform PMTU handling first, then
> reserve the final headroom before ip_tunnel_encap() builds the UDP
> tunnel headers. Because encapsulation is now delayed, pass the
> snapshotted encap length into tnl_update_pmtu() so the inner packet
> size stays skb->len + encap_hlen - tunnel_hlen.
> 
> ip_md_tunnel_xmit() is left unchanged. It takes FOU/GUE parameters
> from the skb metadata dst, not from the device configuration, so it
> is not exposed to this race. tnl_update_pmtu() still runs after
> encapsulation there, so that path passes 0 for the extra length.
> 
> The IPv6 analogue of this headroom reservation is still work in
> progress.

Please wait for the IPv6 part to be accepted (some changes are needed
there and here as well) and then mention its sha1 in the commit message:

https://lore.kernel.org/netdev/2ce8f3e4bbed6060afb0aee33dc4e1d9ae021280.1788262122.git.petalzu987@gmail.com/

> 
> Link: https://lore.kernel.org/netdev/b58876297f7d45de008f2e94b6ecab8b2ed84d21.1786088695.git.petalzu987@gmail.com/
> Fixes: dd9d598c6657 ("ip_gre: add the support for i/o_flags update via netlink")

Looks like a more suitable Fixes tag:

Fixes: 56328486539d ("net: Changes to ip_tunnel to support foo-over-udp encapsulation")

> Cc: stable@vger.kernel.org
> Signed-off-by: Chengfeng Ye <nicoyip.dev@gmail.com>
> ---
> Changes in v4:
> - Not restructure ip_md_tunnel_xmit(); only pass 0 into the new
>   tnl_update_pmtu() argument.
> - Snapshot tunnel->encap with data_race() and reuse the copy for both
>   ip_encap_hlen() and ip_tunnel_encap().
> - Fix tnl_update_pmtu() inner packet size after delaying the encap
>   push: pkt_size = skb->len + encap_hlen - tunnel_hlen.
> - Keep reverse xmas tree for the new locals.
> - Note that the IPv6 FOU/GUE headroom fix is still WIP.
> 
> Changes in v3:
> - Move the headroom reservation into ip_tunnel_xmit() instead of
>   growing the skb inside the FOU/GUE builders.
> - Use ip_encap_hlen() to reserve the final caller-side headroom before
>   ip_tunnel_encap().
> - Drop the IPv4 raw-pointer refreshes that were only needed when
>   skb_cow_head() could run inside the encapsulation builders.
> 
> Link: https://lore.kernel.org/netdev/20260824111944.187200-1-nicoyip.dev@gmail.com/ [v3]
> Link: https://lore.kernel.org/netdev/20260808005956.3761487-1-nicoyip.dev@gmail.com/ [v2]
> Link: https://lore.kernel.org/netdev/20260801060115.3538849-1-nicoyip.dev@gmail.com/ [v1]
> ---
>  net/ipv4/ip_tunnel.c | 24 ++++++++++++++++++------
>  1 file changed, 18 insertions(+), 6 deletions(-)

ipip6_tunnel_xmit() also calls ip_tunnel_encap() and seems to have the
same problem as ip6_tnl_xmit().

  reply	other threads:[~2026-09-06 15:34 UTC|newest]

Thread overview: 3+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-09-04 16:55 [PATCH net v4] ip_tunnel: reserve FOU/GUE headroom before encapsulation Chengfeng Ye
2026-09-06 15:33 ` Ido Schimmel [this message]
2026-09-09  4:57 ` netdev-bot+sashiko

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=20260906153354.GB323313@shredder \
    --to=idosch@nvidia.com \
    --cc=davem@davemloft.net \
    --cc=dsahern@kernel.org \
    --cc=edumazet@google.com \
    --cc=horms@kernel.org \
    --cc=kuba@kernel.org \
    --cc=linux-kernel@vger.kernel.org \
    --cc=lucien.xin@gmail.com \
    --cc=netdev@vger.kernel.org \
    --cc=nicoyip.dev@gmail.com \
    --cc=pabeni@redhat.com \
    --cc=stable@vger.kernel.org \
    --cc=u9012063@gmail.com \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is an external index of several public inboxes,
see mirroring instructions on how to clone and mirror
all data and code used by this external index.