Netdev List
 help / color / mirror / Atom feed
From: netdev-bot+sashiko@kernel.org
To: kuniyu@google.com
Cc: dsahern@kernel.org, idosch@nvidia.com, davem@davemloft.net,
	edumazet@kernel.org, kuba@kernel.org, pabeni@redhat.com,
	horms@kernel.org, carges@cloudflare.com, kuni1840@gmail.com,
	netdev@vger.kernel.org
Subject: Re: [PATCH v2 net-next 2/2] ipv6: Batch rt6_uncached_list_flush_dev() in netdev_run_todo().
Date: Thu, 01 Oct 2026 02:15:10 +0000	[thread overview]
Message-ID: <179082091084.434549.1540561487678718082@kernel.org> (raw)
In-Reply-To: <20260929201546.1533268-3-kuniyu@google.com>

Thank you for your contribution! Sashiko AI review found 4 potential
issue(s) to consider.

Critical: 0 · High: 1 · Medium: 2 · Low: 1

If you believe none of the issues are legitimate, please reply to each
issue of Medium severity or higher and explain why the AI is wrong.
If any of them are legitimate and you plan to respin, please let patchwork
know by including "pw-bot: cr" as a separate line at the end of your reply
(one such reply per series is enough).

- [High] Moving the IPv6 uncached-route flush into netdev_run_todo()
  removes the RCU grace period that used to follow the device-pointer…
- [Medium] Build failure with CONFIG_IPV6=n.
- [Medium] The batched flush reads dev->reg_state twice per route with no
  consistent snapshot, and another unregister batch can change reg_state…
- [Low] The new initialization guard `if (unlikely(!ul->head.next))
  return;` in rt6_uncached_list_flush_dev() (net/ipv6/route.c:172) is a…

--

ipv6: Batch rt6_uncached_list_flush_dev() in netdev_run_todo().

This patch moves the IPv6 uncached route flush out of the per-device
NETDEV_UNREGISTER path. It now runs once per batched unregistration from
netdev_run_todo(), without RTNL. With dev == NULL, a route's device is
matched by reg_state == NETREG_UNREGISTERED.

Closes: https://lore.kernel.org/netdev/20260917-hash-bucket-route-lists-v3-0-30493a37b6eb@cloudflare.com/

> diff --git a/net/core/dev.c b/net/core/dev.c
> index 9cedc17b5bdc2..0b3213ff56e70 100644
> --- a/net/core/dev.c
> +++ b/net/core/dev.c
> @@ -11838,8 +11838,10 @@ void netdev_run_todo(void)
>  		linkwatch_sync_dev(dev);
>  	}
>  
> -	if (!list_empty(&list))
> +	if (!list_empty(&list)) {
>  		rt_flush_dev(NULL);
> +		rt6_uncached_list_flush_dev(NULL);

[Severity: Medium]
With CONFIG_IPV6=n, is the new stub for rt6_uncached_list_flush_dev()
visible here?

net/core/dev.c doesn't include <net/ip6_route.h> directly. The only path
seems to be net/udp_tunnel.h -> net/ip_tunnels.h, and that include is
compiled out when IPv6 is off:

include/net/ip_tunnels.h:
#if IS_ENABLED(CONFIG_IPV6)
#include <net/ipv6.h>
#include <net/ip6_fib.h>
#include <net/ip6_route.h>
#endif

So when IPv6 is disabled, dev.c would see neither the prototype nor the
static inline stub added to include/net/ip6_route.h. Would this call then
fail with -Werror=implicit-function-declaration? It would also fail to
link, since net/ipv6/route.c isn't built.

rt_flush_dev() doesn't hit this, because net/route.h always reaches
dev.c through net/ip.h.

Should dev.c include <net/ip6_route.h> directly, or should the
declaration and stub go in a header that dev.c always includes?

> +	}
>  
>  	cnt = 0;
>  	while (!list_empty(&list)) {
> diff --git a/net/ipv6/route.c b/net/ipv6/route.c
> index a76869ff87cd6..38261775e7137 100644
> --- a/net/ipv6/route.c
> +++ b/net/ipv6/route.c
> @@ -158,14 +158,20 @@ void rt6_uncached_list_del(struct rt6_info *rt)
>  	}
>  }
>  
> -static void rt6_uncached_list_flush_dev(struct net_device *dev)
> +void rt6_uncached_list_flush_dev(struct net_device *dev)
>  {
>  	int cpu;
>  
> +	if (dev && dev->dismantle)
> +		return;
> +
>  	for_each_possible_cpu(cpu) {
>  		struct uncached_list *ul = per_cpu_ptr(&rt6_uncached_list, cpu);
>  		struct rt6_info *rt, *safe;
>  
> +		if (unlikely(!ul->head.next))
> +			return;

[Severity: Low]
This isn't a functional bug, but ul->head.next is read here without
ul->lock. Other CPUs can update it concurrently under the lock:
list_add_tail() in rt6_uncached_list_add() when the list is empty, and
list_del_init() in rt6_uncached_list_del().

The list_empty() check just below already uses READ_ONCE(). Should this
be READ_ONCE(ul->head.next) so that KCSAN doesn't report it?

> +
>  		if (list_empty(&ul->head))
>  			continue;
>  
> @@ -175,19 +181,23 @@ static void rt6_uncached_list_flush_dev(struct net_device *dev)
>  			struct net_device *rt_dev = rt->dst.dev;
>  			bool handled = false;
>  
> -			if (rt_idev && rt_idev->dev == dev) {
> +			if (rt_idev &&
> +			    (dev ? rt_idev->dev == dev :
> +			     READ_ONCE(rt_idev->dev->reg_state) == NETREG_UNREGISTERED)) {
>  				rt->rt6i_idev = in6_dev_get(blackhole_netdev);
>  				in6_dev_put(rt_idev);
>  				handled = true;
>  			}
>  
> -			if (rt_dev == dev) {
> -				rt->dst.dev = blackhole_netdev;
> +			if (dev ? rt_dev == dev :
> +			    READ_ONCE(rt_dev->reg_state) == NETREG_UNREGISTERED) {

[Severity: Medium]
If two netdev_run_todo() instances overlap, can these two reg_state reads
see different values for the same device?

netdev_run_todo() calls __rtnl_unlock() before it sets each device to
NETREG_UNREGISTERED, and that write is done under netdev_lock(), not
ul->lock. So batch A can be scanning while batch B updates a device D
that is on B's todo list. Take a route whose rt6i_idev->dev and dst.dev
are both D:

CPU1 (batch A, flush(NULL))          CPU2 (batch B, owns D)
READ_ONCE(rt_idev->dev->reg_state)
  sees NETREG_UNREGISTERING,
  rt6i_idev is not swapped
                                     WRITE_ONCE(D->reg_state,
                                                NETREG_UNREGISTERED)
READ_ONCE(rt_dev->reg_state)
  sees NETREG_UNREGISTERED,
  dst.dev swapped, handled = true
list_del_init(&rt->dst.rt_uncached)

By the time batch B's own flush runs, the route is off the list. Its
rt6i_idev still points at D's inet6_dev, which holds a reference on D.

The NETDEV_UNREGISTER rebroadcasts can't fix this either.
rt6_disable_ip()->rt6_uncached_list_flush_dev(dev) now returns early on
dev->dismantle, and the route is no longer listed.

Could batch B then sit in netdev_wait_allrefs_any() printing "waiting for
D to become free" for as long as something holds the uncached route? The
window is narrow, but the old code compared pointers under RTNL, so it
could not race this way.

> +				rcu_assign_pointer(rt->dst.dev_rcu, blackhole_netdev);
>  				netdev_ref_replace(rt_dev, blackhole_netdev,
>  						   &rt->dst.dev_tracker,
>  						   GFP_ATOMIC);

[Severity: High]
Is there still an RCU grace period between this pointer swap and the
old device being freed?

Before this patch, the flush ran from the NETDEV_UNREGISTER notifier in
unregister_netdevice_many_notify(). That is followed by synchronize_net(),
and later by the rcu_barrier() at the top of netdev_run_todo().

Now the flush runs after that rcu_barrier():

netdev_run_todo()
    __rtnl_unlock();
    rcu_barrier();
    ...
    rt_flush_dev(NULL);
    rt6_uncached_list_flush_dev(NULL);
        rcu_assign_pointer(rt->dst.dev_rcu, blackhole_netdev);
        netdev_ref_replace(rt_dev, ...);  /* drops ref on old dev */
    netdev_wait_allrefs_any(&list);       /* returns if refcnt == 1 */
    ...
    free_netdev() / kobject_put()
        -> netdev_release()
            kvfree(dev);

netdev_release() frees the device without waiting for readers:

	/* no need to wait for rcu grace period:
	 * device is dead and about to be freed.
	 */

ip6_output() is one reader of this pointer:

	rcu_read_lock();
	dev = dst_dev_rcu(dst);

It then passes dev on to NF_HOOK_COND(), ip6_finish_output() and
dev_queue_xmit(). Suppose a reader loads the old pointer after the
rcu_barrier() but before this swap. Can it still be using dev, and its
tx queues, after they are freed?

The in6_dev_put(rt_idev) branch above has the same problem.
in6_dev_finish_destroy() drops the netdev reference before its call_rcu().

The previous patch in this series, "ipv4: Batch rt_flush_dev() in
netdev_run_todo().", uses the same ordering for IPv4.

>  				handled = true;
>  			}
> +
>  			if (handled)
>  				list_del_init(&rt->dst.rt_uncached);
>  		}

-- 
Sashiko AI review · https://netdev-ai.bots.linux.dev/sashiko/#/patchset/20260929201546.1533268-1-kuniyu%40google.com

  reply	other threads:[~2026-10-01  2:15 UTC|newest]

Thread overview: 8+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-09-29 20:15 [PATCH v2 net-next 0/2] ip: Batch flushing uncached routes per batched device unregistration Kuniyuki Iwashima
2026-09-29 20:15 ` [PATCH v2 net-next 1/2] ipv4: Batch rt_flush_dev() in netdev_run_todo() Kuniyuki Iwashima
2026-10-01  2:15   ` netdev-bot+sashiko
2026-10-01  5:27     ` Kuniyuki Iwashima
2026-09-29 20:15 ` [PATCH v2 net-next 2/2] ipv6: Batch rt6_uncached_list_flush_dev() " Kuniyuki Iwashima
2026-10-01  2:15   ` netdev-bot+sashiko [this message]
2026-10-01  5:36     ` Kuniyuki Iwashima
2026-09-30  2:33 ` [PATCH v2 net-next 0/2] ip: Batch flushing uncached routes per batched device unregistration Chris Arges

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=179082091084.434549.1540561487678718082@kernel.org \
    --to=netdev-bot+sashiko@kernel.org \
    --cc=carges@cloudflare.com \
    --cc=davem@davemloft.net \
    --cc=dsahern@kernel.org \
    --cc=edumazet@kernel.org \
    --cc=horms@kernel.org \
    --cc=idosch@nvidia.com \
    --cc=kuba@kernel.org \
    --cc=kuni1840@gmail.com \
    --cc=kuniyu@google.com \
    --cc=netdev@vger.kernel.org \
    --cc=pabeni@redhat.com \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox