Netdev List
 help / color / mirror / Atom feed
From: Andrea Mayer <andrea.mayer@uniroma2.it>
To: Yuya Kusakabe <yuya.kusakabe@gmail.com>
Cc: "David S. Miller" <davem@davemloft.net>,
	Eric Dumazet <edumazet@kernel.org>,
	Jakub Kicinski <kuba@kernel.org>, Paolo Abeni <pabeni@redhat.com>,
	Simon Horman <horms@kernel.org>,
	netdev@vger.kernel.org, linux-kernel@vger.kernel.org,
	stefano.salsano@uniroma2.it,
	Andrea Mayer <andrea.mayer@uniroma2.it>
Subject: Re: [PATCH net-next 1/2] seg6: skip the route refcount in seg6_lookup_any_nexthop()
Date: Fri, 9 Oct 2026 02:43:41 +0200	[thread overview]
Message-ID: <20261009024341.687e7226cd6abca80e779fda@uniroma2.it> (raw)
In-Reply-To: <20261005-seg6-lookup-noref-v1-1-092e84774e80@gmail.com>

On Mon, 05 Oct 2026 14:31:53 +0900
Yuya Kusakabe <yuya.kusakabe@gmail.com> wrote:

> ip6_route_input() attaches the input route to the skb without taking a
> reference, relying on the RCU read-side section of the receive path.
> seg6_lookup_any_nexthop() is called from that same section, either
> directly from the seg6local input handler or from a netfilter okfn,
> but still takes and drops a route reference for every packet.

Hi Yuya,

Thanks for the patch and the throughput numbers. The Sashiko review asked
whether this also holds for the BPF callers in filter.c. Under
BPF_PROG_TEST_RUN the call does not come from the receive path: LWT_IN
programs reach seg6_lookup_nexthop() through bpf_push_seg6_encap(),
bpf_test_run() runs them many times on the same skb, and the RCU section
can end between runs. With the patch, on 32-bit ARM
(CONFIG_PREEMPT_VOLUNTARY, KASAN):

  BUG: KASAN: slab-use-after-free in __seg6_do_srh_encap+0x3c/0x52c
  Read of size 4 at addr c46b8100 by task bpftool/119
  CPU: 0 UID: 0 PID: 119 Comm: bpftool [snip] 7.3.0-rc5 #1 VOLUNTARY
  Tainted: [O]=OOT_MODULE
  Hardware name: Generic DT based system
  Call trace:
    [snip]
    __seg6_do_srh_encap from seg6_do_srh_encap+0x1c/0x24
    seg6_do_srh_encap from bpf_push_seg6_encap+0x184/0x194
    bpf_push_seg6_encap from bpf_lwt_in_push_encap+0x48/0x5c
    bpf_lwt_in_push_encap from ___bpf_prog_run+0x1aa0/0x4334
    ___bpf_prog_run from __bpf_prog_run32+0xd4/0x114
    __bpf_prog_run32 from bpf_test_run+0x2bc/0x57c
    bpf_test_run from bpf_prog_test_run_skb+0x948/0x1100
    bpf_prog_test_run_skb from __sys_bpf+0x10a8/0x2eac
    __sys_bpf from sys_bpf+0x50/0x64
    sys_bpf from ret_fast_syscall+0x0/0x54
    [snip]
  Allocated by task 119:
    [snip]
    dst_alloc+0x50/0xac
    ip6_dst_alloc+0x24/0x7c
    ip6_pol_route+0x468/0x780
    ip6_pol_route_input+0x48/0x50
    fib6_rule_action+0x16c/0x2e8
    fib_rules_lookup+0x2b8/0x3ec
    fib6_rule_lookup+0x228/0x308
    seg6_lookup_any_nexthop+0x3e8/0x4d0
    seg6_lookup_nexthop+0x1c/0x24
    bpf_lwt_in_push_encap+0x48/0x5c
    ___bpf_prog_run+0x1aa0/0x4334
    __bpf_prog_run32+0xd4/0x114
    bpf_test_run+0x2bc/0x57c
    bpf_prog_test_run_skb+0x948/0x1100
    __sys_bpf+0x10a8/0x2eac
    sys_bpf+0x50/0x64
    ret_fast_syscall+0x0/0x54
  Freed by task 123:
    [snip]
    kmem_cache_free+0xd8/0x3cc
    rcu_core+0x434/0x1154
    handle_softirqs+0x1d4/0x660
    __irq_exit_rcu+0xe8/0x198
    irq_exit+0x10/0x18
    call_with_stack+0x18/0x20
    __irq_svc+0x98/0xb0
    finish_task_switch+0x158/0x4f4
    __schedule+0x6c8/0x1710
    schedule+0x38/0x124
    schedule_hrtimeout_range_clock+0x170/0x1e0
    usleep_range_state+0xf4/0x178
    trigger_fn+0x1dc/0x238 [seg6_noref_trigger]
    kthread+0x1c4/0x200
    ret_from_fork+0x14/0x28
  Last potentially related work creation:
    [snip]
    __call_rcu_common.constprop.0+0x4c/0x608
    ip6_pol_route+0x348/0x780
    ip6_pol_route_output+0x48/0x50
    fib6_rule_action+0x16c/0x2e8
    fib_rules_lookup+0x2b8/0x3ec
    fib6_rule_lookup+0x228/0x308
    ip6_route_output_flags+0x100/0x220
    trigger_fn+0x20c/0x238 [seg6_noref_trigger]
    kthread+0x1c4/0x200
    ret_from_fork+0x14/0x28
  The buggy address belongs to the object at c46b8100
   which belongs to the cache ip6_dst_cache of size 156
  The buggy address is located 0 bytes inside of
   freed 156-byte region [c46b8100, c46b819c)
  [snip]

The module in the trace is a small test helper. In a loop it calls
rt_genid_bump_ipv6(), as "ip -6 rule add" or a nexthop replace does, and
then does an output lookup on a route that shares the nexthop. I used the
module to make the window easier to hit.

Before the patch seg6_lookup_nexthop() used skb_dst_set(), so the skb held
a reference on its dst between runs.

> Look the route up with RT6_LOOKUP_F_DST_NOREF, as ip6_route_input()
> does, for the input and table lookups.
> [snip]

This also changes the contract of seg6_lookup_nexthop(). It does not
always take a reference now, and when it does not, the dst on the skb is
valid only inside the RCU section of the caller. seg6_lookup_nexthop() is
shared: it is declared in include/net/seg6.h and seg6_local.h, filter.c
uses it too, and neither the call sites nor the declarations say that the
lifetime changed.

One possibility would be to split it into two functions, conceptually
similar to ip6_route_output_flags() and ip6_route_output_flags_noref(): a
noref primitive for the callers in the seg6local input path, and a
refcounted wrapper around the primitive, which bpf_push_seg6_encap() would
call. The split would put the contract in the name and fix the
use-after-free above.

Ciao,
Andrea

  parent reply	other threads:[~2026-10-09  0:44 UTC|newest]

Thread overview: 6+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-10-05  5:31 [PATCH net-next 0/2] seg6: skip route refcounts in seg6local lookups Yuya Kusakabe
2026-10-05  5:31 ` [PATCH net-next 1/2] seg6: skip the route refcount in seg6_lookup_any_nexthop() Yuya Kusakabe
2026-10-08  5:32   ` netdev-bot+sashiko
2026-10-09  0:43   ` Andrea Mayer [this message]
2026-10-09  5:39     ` Yuya Kusakabe
2026-10-05  5:31 ` [PATCH net-next 2/2] seg6: skip the route refcount in seg6local IPv4 lookups Yuya Kusakabe

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=20261009024341.687e7226cd6abca80e779fda@uniroma2.it \
    --to=andrea.mayer@uniroma2.it \
    --cc=davem@davemloft.net \
    --cc=edumazet@kernel.org \
    --cc=horms@kernel.org \
    --cc=kuba@kernel.org \
    --cc=linux-kernel@vger.kernel.org \
    --cc=netdev@vger.kernel.org \
    --cc=pabeni@redhat.com \
    --cc=stefano.salsano@uniroma2.it \
    --cc=yuya.kusakabe@gmail.com \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox