Netdev List
 help / color / mirror / Atom feed
* [PATCH net] net/sched: sch_frag: reject a negative L2 length in sch_fragment()
@ 2026-08-27  9:23 Fourie Zhang
  2026-08-28 10:43 ` Jamal Hadi Salim
  0 siblings, 1 reply; 4+ messages in thread
From: Fourie Zhang @ 2026-08-27  9:23 UTC (permalink / raw)
  To: Jamal Hadi Salim, Jiri Pirko, David S . Miller, Eric Dumazet,
	Jakub Kicinski, Paolo Abeni, Simon Horman, Cong Wang, wenxu
  Cc: netdev, Fourie Zhang, stable, TencentOS Corvus AI

sch_fragment() bounds the L2 header length only from above:

	if (skb_network_offset(skb) > VLAN_ETH_HLEN)

skb_network_offset() returns int and the comparison is signed, so a
negative offset passes. sch_frag_prepare_frag() then stores it in an
unsigned int and uses it as a memcpy() length into the 18-byte per-CPU
l2_data buffer:

	unsigned int hlen = skb_network_offset(skb);
	memcpy(&data->l2_data, skb->data, hlen);

The offset is negative whenever the network header sits behind
skb->data. ipv6_srh_rcv() runs in that state, because
ip6_protocol_deliver_rcu() pulls each extension header as it walks the
chain and the segments_left > 0 branch pushes back only
sizeof(struct ipv6hdr). act_ct enables the static key that gates
sch_frag_xmit_hook() and sets tc_skb_cb(skb)->mru, which survives
receive and forward, so an act_mirred redirect from a root qdisc can
reach sch_fragment() with the offset still negative.

A negative L2 header length is invalid, so reject it through the existing
"L2 header too long to fragment" path. A non-negative offset behaves
exactly as before.

Reproduced on v7.2 by an unprivileged user with only
unshare(CLONE_NEWUSER|CLONE_NEWNET). RDX is the truncated length,
(unsigned int)(-42):

  RIP: 0010:memcpy+0x8/0x20
  RDX: 00000000ffffffd6
   sch_frag_prepare_frag+0x364/0x440   net/sched/sch_frag.c:74
   sch_fragment+0x1c7/0x7a0            net/sched/sch_frag.c:122
   tcf_mirred_to_dev+0x7a9/0xf10
   ip6_forward+0x114b/0x31b0
   ipv6_rthdr_rcv+0x40d7/0x5e10

With the patch applied the drop path fires and the packet is discarded.

Fixes: c129412f74e9 ("net/sched: sch_frag: add generic packet fragment support.")
Cc: stable@kernel.org
Reported-by: TencentOS Corvus AI <corvus@tencent.com>
Assisted-by: tencentos-corvus-ai:kimi-k3
Signed-off-by: Fourie Zhang <fouriezhang@tencent.com>
---
 net/sched/sch_frag.c | 3 ++-
 1 file changed, 2 insertions(+), 1 deletion(-)

diff --git a/net/sched/sch_frag.c b/net/sched/sch_frag.c
index 75ee52750919..b68a198b1646 100644
--- a/net/sched/sch_frag.c
+++ b/net/sched/sch_frag.c
@@ -89,9 +89,10 @@ static struct dst_ops sch_frag_dst_ops = {
 static int sch_fragment(struct net *net, struct sk_buff *skb,
 			u16 mru, int (*xmit)(struct sk_buff *skb))
 {
+	int hlen = skb_network_offset(skb);
 	int ret = -1;
 
-	if (skb_network_offset(skb) > VLAN_ETH_HLEN) {
+	if (hlen < 0 || hlen > VLAN_ETH_HLEN) {
 		net_warn_ratelimited("L2 header too long to fragment\n");
 		goto err;
 	}
-- 
2.43.7


^ permalink raw reply related	[flat|nested] 4+ messages in thread

* Re: [PATCH net] net/sched: sch_frag: reject a negative L2 length in sch_fragment()
  2026-08-27  9:23 [PATCH net] net/sched: sch_frag: reject a negative L2 length in sch_fragment() Fourie Zhang
@ 2026-08-28 10:43 ` Jamal Hadi Salim
  2026-08-28 12:22   ` Eric Dumazet
  0 siblings, 1 reply; 4+ messages in thread
From: Jamal Hadi Salim @ 2026-08-28 10:43 UTC (permalink / raw)
  To: Fourie Zhang
  Cc: Jiri Pirko, David S . Miller, Eric Dumazet, Jakub Kicinski,
	Paolo Abeni, Simon Horman, Cong Wang, wenxu, netdev, Fourie Zhang,
	stable, TencentOS Corvus AI

On Thu, Aug 27, 2026 at 5:24 AM Fourie Zhang <littleddfu@gmail.com> wrote:
>
> sch_fragment() bounds the L2 header length only from above:
>
>         if (skb_network_offset(skb) > VLAN_ETH_HLEN)
>
> skb_network_offset() returns int and the comparison is signed, so a
> negative offset passes. sch_frag_prepare_frag() then stores it in an
> unsigned int and uses it as a memcpy() length into the 18-byte per-CPU
> l2_data buffer:

The patch looks good. Comments:
1. Cc: stable@kernel.org should be Cc: stable@vger.kernel.org.
2. Looking at the current sashiko review, the two issues it mentions
certainly do not require a v2. A followup is needed.
I am recording the followups - do you want to send thos patches?
Otherwise we will.
3. In the future, if you send a patch make sure you include a tdc test
that would fail without your patch and pass with your patch.We test
everything tc related. Can you provide a tdc test or a standalone
reproducer script? If it is sensitive just send directly to me.

Handwaving: you dont need a v2 except for #1 above. Unfortunately, it
seems the nipa sashiko now looks only 72 hours later, so you are not
off the hook yet.  Lets wait to see what the overlord says before
deciding  if there's a need for v2.

cheers,
jamal

>
>         unsigned int hlen = skb_network_offset(skb);
>         memcpy(&data->l2_data, skb->data, hlen);
>
> The offset is negative whenever the network header sits behind
> skb->data. ipv6_srh_rcv() runs in that state, because
> ip6_protocol_deliver_rcu() pulls each extension header as it walks the
> chain and the segments_left > 0 branch pushes back only
> sizeof(struct ipv6hdr). act_ct enables the static key that gates
> sch_frag_xmit_hook() and sets tc_skb_cb(skb)->mru, which survives
> receive and forward, so an act_mirred redirect from a root qdisc can
> reach sch_fragment() with the offset still negative.
>
> A negative L2 header length is invalid, so reject it through the existing
> "L2 header too long to fragment" path. A non-negative offset behaves
> exactly as before.
>
> Reproduced on v7.2 by an unprivileged user with only
> unshare(CLONE_NEWUSER|CLONE_NEWNET). RDX is the truncated length,
> (unsigned int)(-42):
>
>   RIP: 0010:memcpy+0x8/0x20
>   RDX: 00000000ffffffd6
>    sch_frag_prepare_frag+0x364/0x440   net/sched/sch_frag.c:74
>    sch_fragment+0x1c7/0x7a0            net/sched/sch_frag.c:122
>    tcf_mirred_to_dev+0x7a9/0xf10
>    ip6_forward+0x114b/0x31b0
>    ipv6_rthdr_rcv+0x40d7/0x5e10
>
> With the patch applied the drop path fires and the packet is discarded.
>
> Fixes: c129412f74e9 ("net/sched: sch_frag: add generic packet fragment support.")
> Cc: stable@kernel.org
> Reported-by: TencentOS Corvus AI <corvus@tencent.com>
> Assisted-by: tencentos-corvus-ai:kimi-k3
> Signed-off-by: Fourie Zhang <fouriezhang@tencent.com>
> ---
>  net/sched/sch_frag.c | 3 ++-
>  1 file changed, 2 insertions(+), 1 deletion(-)
>
> diff --git a/net/sched/sch_frag.c b/net/sched/sch_frag.c
> index 75ee52750919..b68a198b1646 100644
> --- a/net/sched/sch_frag.c
> +++ b/net/sched/sch_frag.c
> @@ -89,9 +89,10 @@ static struct dst_ops sch_frag_dst_ops = {
>  static int sch_fragment(struct net *net, struct sk_buff *skb,
>                         u16 mru, int (*xmit)(struct sk_buff *skb))
>  {
> +       int hlen = skb_network_offset(skb);
>         int ret = -1;
>
> -       if (skb_network_offset(skb) > VLAN_ETH_HLEN) {
> +       if (hlen < 0 || hlen > VLAN_ETH_HLEN) {
>                 net_warn_ratelimited("L2 header too long to fragment\n");
>                 goto err;
>         }
> --
> 2.43.7
>

^ permalink raw reply	[flat|nested] 4+ messages in thread

* Re: [PATCH net] net/sched: sch_frag: reject a negative L2 length in sch_fragment()
  2026-08-28 10:43 ` Jamal Hadi Salim
@ 2026-08-28 12:22   ` Eric Dumazet
  2026-08-28 13:36     ` Eric Dumazet
  0 siblings, 1 reply; 4+ messages in thread
From: Eric Dumazet @ 2026-08-28 12:22 UTC (permalink / raw)
  To: Jamal Hadi Salim
  Cc: Fourie Zhang, Jiri Pirko, David S . Miller, Jakub Kicinski,
	Paolo Abeni, Simon Horman, Cong Wang, wenxu, netdev, Fourie Zhang,
	stable, TencentOS Corvus AI

On Fri, Aug 28, 2026 at 12:43 PM Jamal Hadi Salim <jhs@mojatatu.com> wrote:
>
> On Thu, Aug 27, 2026 at 5:24 AM Fourie Zhang <littleddfu@gmail.com> wrote:
> >
> > sch_fragment() bounds the L2 header length only from above:
> >
> >         if (skb_network_offset(skb) > VLAN_ETH_HLEN)
> >
> > skb_network_offset() returns int and the comparison is signed, so a
> > negative offset passes. sch_frag_prepare_frag() then stores it in an
> > unsigned int and uses it as a memcpy() length into the 18-byte per-CPU
> > l2_data buffer:
>
> The patch looks good. Comments:
> 1. Cc: stable@kernel.org should be Cc: stable@vger.kernel.org.
> 2. Looking at the current sashiko review, the two issues it mentions
> certainly do not require a v2. A followup is needed.
> I am recording the followups - do you want to send thos patches?
> Otherwise we will.
> 3. In the future, if you send a patch make sure you include a tdc test
> that would fail without your patch and pass with your patch.We test
> everything tc related. Can you provide a tdc test or a standalone
> reproducer script? If it is sensitive just send directly to me.
>
> Handwaving: you dont need a v2 except for #1 above. Unfortunately, it
> seems the nipa sashiko now looks only 72 hours later, so you are not
> off the hook yet.  Lets wait to see what the overlord says before
> deciding  if there's a need for v2.
>
> cheers,
> jamal
>
> >
> >         unsigned int hlen = skb_network_offset(skb);
> >         memcpy(&data->l2_data, skb->data, hlen);
> >
> > The offset is negative whenever the network header sits behind
> > skb->data. ipv6_srh_rcv() runs in that state, because
> > ip6_protocol_deliver_rcu() pulls each extension header as it walks the
> > chain and the segments_left > 0 branch pushes back only
> > sizeof(struct ipv6hdr). act_ct enables the static key that gates
> > sch_frag_xmit_hook() and sets tc_skb_cb(skb)->mru, which survives
> > receive and forward, so an act_mirred redirect from a root qdisc can
> > reach sch_fragment() with the offset still negative.
> >
> > A negative L2 header length is invalid, so reject it through the existing
> > "L2 header too long to fragment" path. A non-negative offset behaves
> > exactly as before.
> >
> > Reproduced on v7.2 by an unprivileged user with only
> > unshare(CLONE_NEWUSER|CLONE_NEWNET). RDX is the truncated length,
> > (unsigned int)(-42):
> >
> >   RIP: 0010:memcpy+0x8/0x20
> >   RDX: 00000000ffffffd6
> >    sch_frag_prepare_frag+0x364/0x440   net/sched/sch_frag.c:74
> >    sch_fragment+0x1c7/0x7a0            net/sched/sch_frag.c:122
> >    tcf_mirred_to_dev+0x7a9/0xf10
> >    ip6_forward+0x114b/0x31b0
> >    ipv6_rthdr_rcv+0x40d7/0x5e10
> >
> > With the patch applied the drop path fires and the packet is discarded.
> >
> > Fixes: c129412f74e9 ("net/sched: sch_frag: add generic packet fragment support.")
> > Cc: stable@kernel.org
> > Reported-by: TencentOS Corvus AI <corvus@tencent.com>
> > Assisted-by: tencentos-corvus-ai:kimi-k3
> > Signed-off-by: Fourie Zhang <fouriezhang@tencent.com>
> > ---
> >  net/sched/sch_frag.c | 3 ++-
> >  1 file changed, 2 insertions(+), 1 deletion(-)
> >
> > diff --git a/net/sched/sch_frag.c b/net/sched/sch_frag.c
> > index 75ee52750919..b68a198b1646 100644
> > --- a/net/sched/sch_frag.c
> > +++ b/net/sched/sch_frag.c
> > @@ -89,9 +89,10 @@ static struct dst_ops sch_frag_dst_ops = {
> >  static int sch_fragment(struct net *net, struct sk_buff *skb,
> >                         u16 mru, int (*xmit)(struct sk_buff *skb))
> >  {
> > +       int hlen = skb_network_offset(skb);
> >         int ret = -1;
> >
> > -       if (skb_network_offset(skb) > VLAN_ETH_HLEN) {
> > +       if (hlen < 0 || hlen > VLAN_ETH_HLEN) {
> >                 net_warn_ratelimited("L2 header too long to fragment\n");
> >                 goto err;
> >         }

I am only speaking for mysef, but

pw-bot: cr

IMO this is defensive programming working around the real bug in ipv6_srh_rcv().

Let's fix the root cause.

I saw a recent patch which had some issues:
https://sashiko.dev/#/patchset/20260817104128.22681-1-juny24602@gmail.com?part=1

It's funny that both patches were assisted by tencentos-corvus-ai:kimi-k3 :/

I think something like this could fix both issues:

diff --git a/net/ipv6/exthdrs.c b/net/ipv6/exthdrs.c
index 51941ad656a36e739388c7da2fcd639ab0e453b5..09a4552f7f08aa8208eb981c0536c58a3561ecbd
100644
--- a/net/ipv6/exthdrs.c
+++ b/net/ipv6/exthdrs.c
@@ -445,7 +445,7 @@ static int ipv6_srh_rcv(struct sk_buff *skb,
struct inet6_dev *idev)
        hdr->segments_left--;
        addr = hdr->segments + hdr->segments_left;

-       skb_push(skb, sizeof(struct ipv6hdr));
+       skb_push(skb, -skb_network_offset(skb));

        if (skb->ip_summed == CHECKSUM_COMPLETE)
                seg6_update_csum(skb);
@@ -469,7 +469,7 @@ static int ipv6_srh_rcv(struct sk_buff *skb,
struct inet6_dev *idev)
                }
                ipv6_hdr(skb)->hop_limit--;

-               skb_pull(skb, sizeof(struct ipv6hdr));
+               skb_pull(skb, skb_transport_offset(skb));
                goto looped_back;
        }

^ permalink raw reply	[flat|nested] 4+ messages in thread

* Re: [PATCH net] net/sched: sch_frag: reject a negative L2 length in sch_fragment()
  2026-08-28 12:22   ` Eric Dumazet
@ 2026-08-28 13:36     ` Eric Dumazet
  0 siblings, 0 replies; 4+ messages in thread
From: Eric Dumazet @ 2026-08-28 13:36 UTC (permalink / raw)
  To: Jamal Hadi Salim
  Cc: Fourie Zhang, Jiri Pirko, David S . Miller, Jakub Kicinski,
	Paolo Abeni, Simon Horman, netdev, Fourie Zhang, stable,
	TencentOS Corvus AI

On Fri, Aug 28, 2026 at 2:22 PM Eric Dumazet <edumazet@google.com> wrote:
>

> IMO this is defensive programming working around the real bug in ipv6_srh_rcv().
>
> Let's fix the root cause.
>
> I saw a recent patch which had some issues:
> https://sashiko.dev/#/patchset/20260817104128.22681-1-juny24602@gmail.com?part=1
>
> It's funny that both patches were assisted by tencentos-corvus-ai:kimi-k3 :/
>
> I think something like this could fix both issues:
>
> diff --git a/net/ipv6/exthdrs.c b/net/ipv6/exthdrs.c
> index 51941ad656a36e739388c7da2fcd639ab0e453b5..09a4552f7f08aa8208eb981c0536c58a3561ecbd
> 100644
> --- a/net/ipv6/exthdrs.c
> +++ b/net/ipv6/exthdrs.c
> @@ -445,7 +445,7 @@ static int ipv6_srh_rcv(struct sk_buff *skb,
> struct inet6_dev *idev)
>         hdr->segments_left--;
>         addr = hdr->segments + hdr->segments_left;
>
> -       skb_push(skb, sizeof(struct ipv6hdr));
> +       skb_push(skb, -skb_network_offset(skb));
>
>         if (skb->ip_summed == CHECKSUM_COMPLETE)
>                 seg6_update_csum(skb);
> @@ -469,7 +469,7 @@ static int ipv6_srh_rcv(struct sk_buff *skb,
> struct inet6_dev *idev)
>                 }
>                 ipv6_hdr(skb)->hop_limit--;
>
> -               skb_pull(skb, sizeof(struct ipv6hdr));
> +               skb_pull(skb, skb_transport_offset(skb));
>                 goto looped_back;
>         }

I am running our net selftests, and if all green will send this patch
with a changelog like:

ipv6: sr: restore network header before routing and forwarding

ipv6_srh_rcv() runs with skb->data at the Segment Routing Header (SRH)
while skb_network_header() points at the IPv6 header.

When segments_left > 0, ipv6_srh_rcv() previously restored the skb->data
position by pushing sizeof(struct ipv6hdr), assuming the SRH immediately
followed the fixed IPv6 header. If another extension header (such as a
Hop-by-Hop options header) precedes the SRH, skb_network_offset()
remained negative.

This led to two problems:
1. During ip6_route_input(), fib6_rules_early_flow_dissect() invokes
   __skb_flow_dissect() which passes the negative skb_network_offset()
   to flow dissection, breaking BPF and C flow dissector logic.
2. If forwarded via ip6_forward() or redirected via act_mirred, downstream
   handlers (like sch_fragment() or neighbour output) pass the negative
   offset as an unsigned length, triggering OOB memcpy or buffer overflows.

Fix this by pushing -skb_network_offset(skb) before routing, ensuring
skb_network_offset(skb) is 0 for route lookup / flow dissection as well as
downstream forwarding. On the loopback path, pull skb_transport_offset(skb)
to restore skb->data to the SRH before looping back.

Fixes: 1ababeba4a21 ("ipv6: implement dataplane support for rthdr type
4 (Segment Routing Header)")
Reported-by: TencentOS Corvus AI <corvus@tencent.com>
Reported-by: Jun Yang <junvyyang@tencent.com>
Reported-by: Fourie Zhang <fouriezhang@tencent.com>
Closes: https://lore.kernel.org/netdev/20260817104128.22681-1-juny24602@gmail.com/
Closes: https://lore.kernel.org/netdev/20260827092345.2301937-1-fouriezhang@tencent.com/
Signed-off-by: Eric Dumazet <edumazet@google.com>

^ permalink raw reply	[flat|nested] 4+ messages in thread

end of thread, other threads:[~2026-08-28 13:37 UTC | newest]

Thread overview: 4+ messages (download: mbox.gz follow: Atom feed
-- links below jump to the message on this page --
2026-08-27  9:23 [PATCH net] net/sched: sch_frag: reject a negative L2 length in sch_fragment() Fourie Zhang
2026-08-28 10:43 ` Jamal Hadi Salim
2026-08-28 12:22   ` Eric Dumazet
2026-08-28 13:36     ` Eric Dumazet

This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox