* [PATCH net-next 0/2] seg6: skip route refcounts in seg6local lookups
@ 2026-10-05 5:31 Yuya Kusakabe
2026-10-05 5:31 ` [PATCH net-next 1/2] seg6: skip the route refcount in seg6_lookup_any_nexthop() Yuya Kusakabe
2026-10-05 5:31 ` [PATCH net-next 2/2] seg6: skip the route refcount in seg6local IPv4 lookups Yuya Kusakabe
0 siblings, 2 replies; 6+ messages in thread
From: Yuya Kusakabe @ 2026-10-05 5:31 UTC (permalink / raw)
To: Andrea Mayer, David S. Miller, Eric Dumazet, Jakub Kicinski,
Paolo Abeni, Simon Horman
Cc: netdev, linux-kernel, Yuya Kusakabe
seg6local looks up the nexthop of its IPv6 and IPv4 behaviors from the
RCU read-side section of the receive path, yet takes and drops a route
reference for every packet. Skip the reference, as ip6_route_input()
and ip_rcv_finish() do.
---
Yuya Kusakabe (2):
seg6: skip the route refcount in seg6_lookup_any_nexthop()
seg6: skip the route refcount in seg6local IPv4 lookups
net/ipv6/seg6_local.c | 14 ++++++++++----
1 file changed, 10 insertions(+), 4 deletions(-)
---
base-commit: cfb7793d1bc0f7d90571611979654cf1b3886b29
change-id: 20261004-seg6-lookup-noref-e016495b60ed
Best regards,
--
Yuya Kusakabe <yuya.kusakabe@gmail.com>
^ permalink raw reply [flat|nested] 6+ messages in thread* [PATCH net-next 1/2] seg6: skip the route refcount in seg6_lookup_any_nexthop() 2026-10-05 5:31 [PATCH net-next 0/2] seg6: skip route refcounts in seg6local lookups Yuya Kusakabe @ 2026-10-05 5:31 ` Yuya Kusakabe 2026-10-08 5:32 ` netdev-bot+sashiko 2026-10-09 0:43 ` Andrea Mayer 2026-10-05 5:31 ` [PATCH net-next 2/2] seg6: skip the route refcount in seg6local IPv4 lookups Yuya Kusakabe 1 sibling, 2 replies; 6+ messages in thread From: Yuya Kusakabe @ 2026-10-05 5:31 UTC (permalink / raw) To: Andrea Mayer, David S. Miller, Eric Dumazet, Jakub Kicinski, Paolo Abeni, Simon Horman Cc: netdev, linux-kernel, Yuya Kusakabe ip6_route_input() attaches the input route to the skb without taking a reference, relying on the RCU read-side section of the receive path. seg6_lookup_any_nexthop() is called from that same section, either directly from the seg6local input handler or from a netfilter okfn, but still takes and drops a route reference for every packet. Look the route up with RT6_LOOKUP_F_DST_NOREF, as ip6_route_input() does, for the input and table lookups. Throughput at 0.5% packet loss of 94-byte End packets (two-segment SRH, no payload), spread by RSS on an ixgbe 82599ES across two 2.30 GHz Xeon E5-2650 v3 with 10 cores each, offered by TRex and binary-searched over 10 runs of 10 s, in Mpps: queues cores before after 1 1 0.735 0.770 (+4.8%) 10 10 on the NIC's socket 6.468 6.976 (+7.9%) 16 10 + 6 on two sockets 6.931 10.128 (+46.1%) With the default rules, fib6_rule_lookup() tries the local table first, and the miss takes and drops a reference on ip6_null_entry, one route shared by every CPU. At 16 queues that cache line bounces between the sockets, and ip6_hold_safe() and dst_release() together take more than a third of the cycles. Assisted-by: LLM Signed-off-by: Yuya Kusakabe <yuya.kusakabe@gmail.com> --- net/ipv6/seg6_local.c | 10 ++++++++-- 1 file changed, 8 insertions(+), 2 deletions(-) diff --git a/net/ipv6/seg6_local.c b/net/ipv6/seg6_local.c index 0c3d1be7f129..4bde3669d0dd 100644 --- a/net/ipv6/seg6_local.c +++ b/net/ipv6/seg6_local.c @@ -342,6 +342,7 @@ seg6_lookup_any_nexthop(struct sk_buff *skb, struct in6_addr *nhaddr, fl6.flowi6_flags = FLOWI_FLAG_KNOWN_NH; if (!tbl_id && !oif) { + flags |= RT6_LOOKUP_F_DST_NOREF; dst = ip6_route_input_lookup(net, skb->dev, &fl6, skb, flags); } else if (tbl_id) { struct fib6_table *table; @@ -350,6 +351,7 @@ seg6_lookup_any_nexthop(struct sk_buff *skb, struct in6_addr *nhaddr, if (!table) goto out; + flags |= RT6_LOOKUP_F_DST_NOREF; rt = ip6_pol_route(net, table, oif, &fl6, skb, flags); dst = &rt->dst; } else { @@ -363,7 +365,7 @@ seg6_lookup_any_nexthop(struct sk_buff *skb, struct in6_addr *nhaddr, dev_flags |= IFF_LOOPBACK; if (dst && (dst_dev(dst)->flags & dev_flags) && !dst->error) { - dst_release(dst); + ip6_rt_put_flags(dst_rt6_info(dst), flags); dst = NULL; } @@ -372,10 +374,14 @@ seg6_lookup_any_nexthop(struct sk_buff *skb, struct in6_addr *nhaddr, rt = net->ipv6.ip6_blk_hole_entry; dst = &rt->dst; dst_hold(dst); + flags &= ~RT6_LOOKUP_F_DST_NOREF; } skb_dst_drop(skb); - skb_dst_set(skb, dst); + if ((flags & RT6_LOOKUP_F_DST_NOREF) && !dst->rt_uncached_list) + skb_dst_set_noref(skb, dst); + else + skb_dst_set(skb, dst); return dst->error; } -- 2.50.1 ^ permalink raw reply related [flat|nested] 6+ messages in thread
* Re: [PATCH net-next 1/2] seg6: skip the route refcount in seg6_lookup_any_nexthop() 2026-10-05 5:31 ` [PATCH net-next 1/2] seg6: skip the route refcount in seg6_lookup_any_nexthop() Yuya Kusakabe @ 2026-10-08 5:32 ` netdev-bot+sashiko 2026-10-09 0:43 ` Andrea Mayer 1 sibling, 0 replies; 6+ messages in thread From: netdev-bot+sashiko @ 2026-10-08 5:32 UTC (permalink / raw) To: yuya.kusakabe Cc: andrea.mayer, davem, edumazet, kuba, pabeni, horms, netdev, linux-kernel Thank you for your contribution! Sashiko AI review found 1 potential issue(s) to consider. Critical: 0 · High: 1 · Medium: 0 · Low: 0 If you believe none of the issues are legitimate, please reply to each issue of Medium severity or higher and explain why the AI is wrong. If any of them are legitimate and you plan to respin, please let patchwork know by including "pw-bot: cr" as a separate line at the end of your reply (one such reply per series is enough). - [High] After this patch, the exported helper seg6_lookup_nexthop()/seg6_lookup_any_nexthop()… -- seg6: skip the route refcount in seg6_lookup_any_nexthop() This patch makes seg6_lookup_any_nexthop() look up input and table routes with RT6_LOOKUP_F_DST_NOREF. When the route is not on the uncached list, it attaches the result with skb_dst_set_noref(). The commit message assumes every caller runs inside the receive path's RCU read-side section. > diff --git a/net/ipv6/seg6_local.c b/net/ipv6/seg6_local.c > index 0c3d1be7f1299..4bde3669d0dda 100644 > --- a/net/ipv6/seg6_local.c > +++ b/net/ipv6/seg6_local.c > @@ -342,6 +342,7 @@ seg6_lookup_any_nexthop(struct sk_buff *skb, struct in6_addr *nhaddr, > fl6.flowi6_flags = FLOWI_FLAG_KNOWN_NH; > > if (!tbl_id && !oif) { > + flags |= RT6_LOOKUP_F_DST_NOREF; > dst = ip6_route_input_lookup(net, skb->dev, &fl6, skb, flags); > } else if (tbl_id) { > struct fib6_table *table; [ ... ] > @@ -372,10 +374,14 @@ seg6_lookup_any_nexthop(struct sk_buff *skb, struct in6_addr *nhaddr, > rt = net->ipv6.ip6_blk_hole_entry; > dst = &rt->dst; > dst_hold(dst); > + flags &= ~RT6_LOOKUP_F_DST_NOREF; > } > > skb_dst_drop(skb); > - skb_dst_set(skb, dst); > + if ((flags & RT6_LOOKUP_F_DST_NOREF) && !dst->rt_uncached_list) > + skb_dst_set_noref(skb, dst); > + else > + skb_dst_set(skb, dst); > return dst->error; > } [Severity: High] The commit message says the helper is called "either directly from the seg6local input handler or from a netfilter okfn". Does that also hold for the BPF callers in net/core/filter.c? bpf_push_seg6_encap() ends with: net/core/filter.c:bpf_push_seg6_encap() { ... return seg6_lookup_nexthop(skb, NULL, 0); } LWT_IN programs get there through bpf_lwt_push_encap() with BPF_LWT_ENCAP_SEG6 or BPF_LWT_ENCAP_SEG6_INLINE. lwt_in_prog_ops also sets .test_run = bpf_prog_test_run_skb. In that path, bpf_test_run() runs the program on the same skb for every iteration. Between iterations it can drop RCU and reschedule: net/bpf/test_run.c:bpf_test_timer_continue() { ... if (need_resched()) { /* During iteration: we need to reschedule between runs. */ t->time_spent += ktime_get_ns() - t->time_start; bpf_test_timer_leave(t); cond_resched(); bpf_test_timer_enter(t); } ... } With RT6_LOOKUP_F_DST_NOREF, ip6_pol_route() no longer takes a reference on pcpu or exception routes. Either kind can now stay on the skb as a noref dst through that window. Suppose the route is released during that window, either by a route delete or replace (fib6_nh_release_dsts() -> dst_release() -> call_rcu) or by exception aging. Can the next iteration then read the freed dst before a new lookup replaces it? A few examples: bpf_get_route_realm() dst_tclassid(skb) reads skb_dst(skb)->tclassid bpf_lwt_push_encap(BPF_LWT_ENCAP_SEG6) bpf_push_seg6_encap() seg6_do_srh_encap() __seg6_do_srh_encap() struct dst_entry *dst = skb_dst(skb); struct net_device *dev = dst_dev(dst); struct net *net = dev_net(dev); bpf_lwt_push_encap(BPF_LWT_ENCAP_SEG6_INLINE) with an HMAC TLV skb_dst_dev_net(skb) Before this patch, skb_dst_set() always attached a refcounted dst, so the dst stayed valid across iterations. The follow-up patch in the series ("seg6: skip the route refcount in seg6local IPv4 lookups") only changes input_action_end_dx4_finish() and input_action_end_dt4(). It does not touch filter.c, so the problem is still there at the end of the series. Would it make sense to keep a refcounted dst for the BPF callers? One option is skb_dst_force() in bpf_push_seg6_encap(). Another is to limit the NOREF lookup to seg6local's own input actions and leave the shared exported helper alone. The declarations of seg6_lookup_nexthop() in include/net/seg6.h and seg6_lookup_any_nexthop() in include/net/seg6_local.h also don't mention that the attached dst is now valid only inside the caller's RCU read-side section. -- Sashiko AI review · https://netdev-ai.bots.linux.dev/sashiko/#/patchset/20261005-seg6-lookup-noref-v1-0-092e84774e80%40gmail.com ^ permalink raw reply [flat|nested] 6+ messages in thread
* Re: [PATCH net-next 1/2] seg6: skip the route refcount in seg6_lookup_any_nexthop() 2026-10-05 5:31 ` [PATCH net-next 1/2] seg6: skip the route refcount in seg6_lookup_any_nexthop() Yuya Kusakabe 2026-10-08 5:32 ` netdev-bot+sashiko @ 2026-10-09 0:43 ` Andrea Mayer 2026-10-09 5:39 ` Yuya Kusakabe 1 sibling, 1 reply; 6+ messages in thread From: Andrea Mayer @ 2026-10-09 0:43 UTC (permalink / raw) To: Yuya Kusakabe Cc: David S. Miller, Eric Dumazet, Jakub Kicinski, Paolo Abeni, Simon Horman, netdev, linux-kernel, stefano.salsano, Andrea Mayer On Mon, 05 Oct 2026 14:31:53 +0900 Yuya Kusakabe <yuya.kusakabe@gmail.com> wrote: > ip6_route_input() attaches the input route to the skb without taking a > reference, relying on the RCU read-side section of the receive path. > seg6_lookup_any_nexthop() is called from that same section, either > directly from the seg6local input handler or from a netfilter okfn, > but still takes and drops a route reference for every packet. Hi Yuya, Thanks for the patch and the throughput numbers. The Sashiko review asked whether this also holds for the BPF callers in filter.c. Under BPF_PROG_TEST_RUN the call does not come from the receive path: LWT_IN programs reach seg6_lookup_nexthop() through bpf_push_seg6_encap(), bpf_test_run() runs them many times on the same skb, and the RCU section can end between runs. With the patch, on 32-bit ARM (CONFIG_PREEMPT_VOLUNTARY, KASAN): BUG: KASAN: slab-use-after-free in __seg6_do_srh_encap+0x3c/0x52c Read of size 4 at addr c46b8100 by task bpftool/119 CPU: 0 UID: 0 PID: 119 Comm: bpftool [snip] 7.3.0-rc5 #1 VOLUNTARY Tainted: [O]=OOT_MODULE Hardware name: Generic DT based system Call trace: [snip] __seg6_do_srh_encap from seg6_do_srh_encap+0x1c/0x24 seg6_do_srh_encap from bpf_push_seg6_encap+0x184/0x194 bpf_push_seg6_encap from bpf_lwt_in_push_encap+0x48/0x5c bpf_lwt_in_push_encap from ___bpf_prog_run+0x1aa0/0x4334 ___bpf_prog_run from __bpf_prog_run32+0xd4/0x114 __bpf_prog_run32 from bpf_test_run+0x2bc/0x57c bpf_test_run from bpf_prog_test_run_skb+0x948/0x1100 bpf_prog_test_run_skb from __sys_bpf+0x10a8/0x2eac __sys_bpf from sys_bpf+0x50/0x64 sys_bpf from ret_fast_syscall+0x0/0x54 [snip] Allocated by task 119: [snip] dst_alloc+0x50/0xac ip6_dst_alloc+0x24/0x7c ip6_pol_route+0x468/0x780 ip6_pol_route_input+0x48/0x50 fib6_rule_action+0x16c/0x2e8 fib_rules_lookup+0x2b8/0x3ec fib6_rule_lookup+0x228/0x308 seg6_lookup_any_nexthop+0x3e8/0x4d0 seg6_lookup_nexthop+0x1c/0x24 bpf_lwt_in_push_encap+0x48/0x5c ___bpf_prog_run+0x1aa0/0x4334 __bpf_prog_run32+0xd4/0x114 bpf_test_run+0x2bc/0x57c bpf_prog_test_run_skb+0x948/0x1100 __sys_bpf+0x10a8/0x2eac sys_bpf+0x50/0x64 ret_fast_syscall+0x0/0x54 Freed by task 123: [snip] kmem_cache_free+0xd8/0x3cc rcu_core+0x434/0x1154 handle_softirqs+0x1d4/0x660 __irq_exit_rcu+0xe8/0x198 irq_exit+0x10/0x18 call_with_stack+0x18/0x20 __irq_svc+0x98/0xb0 finish_task_switch+0x158/0x4f4 __schedule+0x6c8/0x1710 schedule+0x38/0x124 schedule_hrtimeout_range_clock+0x170/0x1e0 usleep_range_state+0xf4/0x178 trigger_fn+0x1dc/0x238 [seg6_noref_trigger] kthread+0x1c4/0x200 ret_from_fork+0x14/0x28 Last potentially related work creation: [snip] __call_rcu_common.constprop.0+0x4c/0x608 ip6_pol_route+0x348/0x780 ip6_pol_route_output+0x48/0x50 fib6_rule_action+0x16c/0x2e8 fib_rules_lookup+0x2b8/0x3ec fib6_rule_lookup+0x228/0x308 ip6_route_output_flags+0x100/0x220 trigger_fn+0x20c/0x238 [seg6_noref_trigger] kthread+0x1c4/0x200 ret_from_fork+0x14/0x28 The buggy address belongs to the object at c46b8100 which belongs to the cache ip6_dst_cache of size 156 The buggy address is located 0 bytes inside of freed 156-byte region [c46b8100, c46b819c) [snip] The module in the trace is a small test helper. In a loop it calls rt_genid_bump_ipv6(), as "ip -6 rule add" or a nexthop replace does, and then does an output lookup on a route that shares the nexthop. I used the module to make the window easier to hit. Before the patch seg6_lookup_nexthop() used skb_dst_set(), so the skb held a reference on its dst between runs. > Look the route up with RT6_LOOKUP_F_DST_NOREF, as ip6_route_input() > does, for the input and table lookups. > [snip] This also changes the contract of seg6_lookup_nexthop(). It does not always take a reference now, and when it does not, the dst on the skb is valid only inside the RCU section of the caller. seg6_lookup_nexthop() is shared: it is declared in include/net/seg6.h and seg6_local.h, filter.c uses it too, and neither the call sites nor the declarations say that the lifetime changed. One possibility would be to split it into two functions, conceptually similar to ip6_route_output_flags() and ip6_route_output_flags_noref(): a noref primitive for the callers in the seg6local input path, and a refcounted wrapper around the primitive, which bpf_push_seg6_encap() would call. The split would put the contract in the name and fix the use-after-free above. Ciao, Andrea ^ permalink raw reply [flat|nested] 6+ messages in thread
* Re: [PATCH net-next 1/2] seg6: skip the route refcount in seg6_lookup_any_nexthop() 2026-10-09 0:43 ` Andrea Mayer @ 2026-10-09 5:39 ` Yuya Kusakabe 0 siblings, 0 replies; 6+ messages in thread From: Yuya Kusakabe @ 2026-10-09 5:39 UTC (permalink / raw) To: Andrea Mayer Cc: David S. Miller, Eric Dumazet, Jakub Kicinski, Paolo Abeni, Simon Horman, netdev, linux-kernel, stefano.salsano On Fri, Oct 09, 2026 at 02:43:41AM +0200, Andrea Mayer wrote: > One possibility would be to split it into two functions, conceptually > similar to ip6_route_output_flags() and ip6_route_output_flags_noref(): a > noref primitive for the callers in the seg6local input path, and a > refcounted wrapper around the primitive, which bpf_push_seg6_encap() would > call. The split would put the contract in the name and fix the > use-after-free above. Thanks for reproducing it with KASAN. I agree, and v2 will do that split: seg6_lookup_nexthop() will still attach a refcounted dst to the skb for filter.c, and only the seg6local input actions will use the noref variant. The lifetime of the dst from the exported helper then does not change, so its declarations need no update. The commit message of patch 1/2 will also drop the claim that every caller runs in the receive path. pw-bot: cr ^ permalink raw reply [flat|nested] 6+ messages in thread
* [PATCH net-next 2/2] seg6: skip the route refcount in seg6local IPv4 lookups 2026-10-05 5:31 [PATCH net-next 0/2] seg6: skip route refcounts in seg6local lookups Yuya Kusakabe 2026-10-05 5:31 ` [PATCH net-next 1/2] seg6: skip the route refcount in seg6_lookup_any_nexthop() Yuya Kusakabe @ 2026-10-05 5:31 ` Yuya Kusakabe 1 sibling, 0 replies; 6+ messages in thread From: Yuya Kusakabe @ 2026-10-05 5:31 UTC (permalink / raw) To: Andrea Mayer, David S. Miller, Eric Dumazet, Jakub Kicinski, Paolo Abeni, Simon Horman Cc: netdev, linux-kernel, Yuya Kusakabe ip_route_input() looks the input route up with ip_route_input_noref() and then forces a reference on it, which is dropped again when the packet is transmitted. End.DX4 and End.DT4 call it from the RCU read-side section of the receive path, either directly from the seg6local input handler or from a netfilter okfn, and hand the skb to dst_input() before leaving that section. Use ip_route_input_noref(), as ip_rcv_finish() does. Throughput at 0.5% packet loss of 82-byte IPv4-in-IPv6 UDP packets (no SRH), spread by RSS on an ixgbe 82599ES across two 10-core Xeon E5-2650 v3, offered by TRex over 10 runs of 10 s, in Mpps: End.DX4 queues cores before after 1 1 0.816 0.837 (+2.6%) 10 10 on the NIC's socket 6.534 7.621 (+16.6%) 16 10 + 6 on two sockets 3.369 10.208 (+203.0%) End.DT4 queues cores before after 1 1 0.497 0.501 (+0.8%) 10 10 on the NIC's socket 4.410 4.679 (+6.1%) 16 10 + 6 on two sockets 3.259 7.431 (+128.0%) Every CPU takes and drops a reference on the same input route, the one cached on the nexthop. At 16 queues that cache line bounces between the sockets, and skb_dst_force() and dst_release() together take about 60% of the cycles for End.DX4 and over 40% for End.DT4. Assisted-by: LLM Signed-off-by: Yuya Kusakabe <yuya.kusakabe@gmail.com> --- net/ipv6/seg6_local.c | 4 ++-- 1 file changed, 2 insertions(+), 2 deletions(-) diff --git a/net/ipv6/seg6_local.c b/net/ipv6/seg6_local.c index 4bde3669d0dd..1dfe20b365b7 100644 --- a/net/ipv6/seg6_local.c +++ b/net/ipv6/seg6_local.c @@ -1079,7 +1079,7 @@ static int input_action_end_dx4_finish(struct net *net, struct sock *sk, skb_dst_drop(skb); - reason = ip_route_input(skb, nhaddr, iph->saddr, 0, skb->dev); + reason = ip_route_input_noref(skb, nhaddr, iph->saddr, 0, skb->dev); if (reason) { kfree_skb_reason(skb, reason); return -EINVAL; @@ -1305,7 +1305,7 @@ static int input_action_end_dt4(struct sk_buff *skb, iph = ip_hdr(skb); - reason = ip_route_input(skb, iph->daddr, iph->saddr, 0, skb->dev); + reason = ip_route_input_noref(skb, iph->daddr, iph->saddr, 0, skb->dev); if (unlikely(reason)) goto drop; -- 2.50.1 ^ permalink raw reply related [flat|nested] 6+ messages in thread
end of thread, other threads:[~2026-10-09 5:39 UTC | newest] Thread overview: 6+ messages (download: mbox.gz follow: Atom feed -- links below jump to the message on this page -- 2026-10-05 5:31 [PATCH net-next 0/2] seg6: skip route refcounts in seg6local lookups Yuya Kusakabe 2026-10-05 5:31 ` [PATCH net-next 1/2] seg6: skip the route refcount in seg6_lookup_any_nexthop() Yuya Kusakabe 2026-10-08 5:32 ` netdev-bot+sashiko 2026-10-09 0:43 ` Andrea Mayer 2026-10-09 5:39 ` Yuya Kusakabe 2026-10-05 5:31 ` [PATCH net-next 2/2] seg6: skip the route refcount in seg6local IPv4 lookups Yuya Kusakabe
This is a public inbox, see mirroring instructions for how to clone and mirror all data and code used for this inbox