From: Kuniyuki Iwashima <kuniyu@google.com>
To: David Ahern <dsahern@kernel.org>,
Ido Schimmel <idosch@nvidia.com>,
Andrew Lunn <andrew+netdev@lunn.ch>,
"David S . Miller" <davem@davemloft.net>,
Eric Dumazet <edumazet@kernel.org>,
Jakub Kicinski <kuba@kernel.org>, Paolo Abeni <pabeni@redhat.com>
Cc: Simon Horman <horms@kernel.org>,
Chris J Arges <carges@cloudflare.com>,
Kuniyuki Iwashima <kuniyu@google.com>,
Kuniyuki Iwashima <kuni1840@gmail.com>,
netdev@vger.kernel.org
Subject: [PATCH v3 net-next 5/6] ipv4: Batch rt_flush_dev() in netdev_run_todo().
Date: Thu, 1 Oct 2026 20:47:17 +0000 [thread overview]
Message-ID: <20261001204752.2572265-6-kuniyu@google.com> (raw)
In-Reply-To: <20261001204752.2572265-1-kuniyu@google.com>
IPv4 uncached routes are linked to the global per-cpu lists,
rt_uncached_list.
When unregistering a netdev, rt_flush_dev() iterates over the
potentially long lists to find uncached routes tied to the device
and swap it with blackhole_netdev.
Since it is called for every device in dying netns under RTNL,
it adds O(N_dev x (N_cpu + N_route)) costs to any batched device
unregistration.
Let's call it once per batched device unregistration without RTNL.
Note that rt_flush_dev() must be called after setting dev->reg_state
to NETREG_UNREGISTERED. Otherwise, because rt_flush_dev(NULL) runs
without RTNL, it could race with unregister_netdevice_many_notify()
and prematurely purge routes for NETREG_UNREGISTERING dev, for
which flush_all_backlogs() has not been called yet.
Reported-by: Chris J Arges <carges@cloudflare.com>
Closes: https://lore.kernel.org/netdev/20260917-hash-bucket-route-lists-v3-0-30493a37b6eb@cloudflare.com/
Signed-off-by: Kuniyuki Iwashima <kuniyu@google.com>
---
include/net/route.h | 6 ++++++
net/core/dev.c | 4 ++++
net/ipv4/route.c | 11 +++++++++--
3 files changed, 19 insertions(+), 2 deletions(-)
diff --git a/include/net/route.h b/include/net/route.h
index 6b55de2e4df8..3fccd31eb74d 100644
--- a/include/net/route.h
+++ b/include/net/route.h
@@ -129,7 +129,13 @@ struct in_device;
int ip_rt_init(void);
void rt_cache_flush(struct net *net);
+#ifdef CONFIG_INET
void rt_flush_dev(struct net_device *dev);
+#else
+static inline void rt_flush_dev(struct net_device *dev)
+{
+}
+#endif
static inline void inet_sk_init_flowi4(const struct inet_sock *inet,
struct flowi4 *fl4)
diff --git a/net/core/dev.c b/net/core/dev.c
index a8eb382f40ca..7dba0292f052 100644
--- a/net/core/dev.c
+++ b/net/core/dev.c
@@ -129,6 +129,7 @@
#include <linux/ip.h>
#include <net/ip.h>
#include <net/mpls.h>
+#include <net/route.h>
#include <linux/ipv6.h>
#include <linux/in.h>
#include <linux/jhash.h>
@@ -11838,6 +11839,9 @@ void netdev_run_todo(void)
linkwatch_sync_dev(dev);
}
+ if (!list_empty(&list))
+ rt_flush_dev(NULL);
+
cnt = 0;
while (!list_empty(&list)) {
dev = netdev_wait_allrefs_any(&list);
diff --git a/net/ipv4/route.c b/net/ipv4/route.c
index d7da2f1acbb5..1b641901f10d 100644
--- a/net/ipv4/route.c
+++ b/net/ipv4/route.c
@@ -1587,6 +1587,9 @@ void rt_flush_dev(struct net_device *dev)
struct rtable *rt, *safe;
int cpu;
+ if (dev && dev->dismantle)
+ return;
+
for_each_possible_cpu(cpu) {
struct uncached_list *ul = &per_cpu(rt_uncached_list, cpu);
@@ -1595,10 +1598,14 @@ void rt_flush_dev(struct net_device *dev)
spin_lock_bh(&ul->lock);
list_for_each_entry_safe(rt, safe, &ul->head, dst.rt_uncached) {
- if (rt->dst.dev != dev)
+ struct net_device *rt_dev = rt->dst.dev;
+
+ if (dev ? rt_dev != dev :
+ READ_ONCE(rt_dev->reg_state) != NETREG_UNREGISTERED)
continue;
+
rcu_assign_pointer(rt->dst.dev_rcu, blackhole_netdev);
- netdev_ref_replace(dev, blackhole_netdev,
+ netdev_ref_replace(rt_dev, blackhole_netdev,
&rt->dst.dev_tracker, GFP_ATOMIC);
list_del_init(&rt->dst.rt_uncached);
}
--
2.56.0.rc1.315.gc6ed9934b7-goog
next prev parent reply other threads:[~2026-10-01 20:48 UTC|newest]
Thread overview: 17+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-10-01 20:47 [PATCH v3 net-next 0/6] ip: Batch flushing uncached routes per batched device unregistration Kuniyuki Iwashima
2026-10-01 20:47 ` [PATCH v3 net-next 1/6] net: Order blackhole_netdev_init(), inet_init(), and inet6_init() Kuniyuki Iwashima
2026-10-05 12:18 ` Fernando Fernandez Mancera
2026-10-01 20:47 ` [PATCH v3 net-next 2/6] ipv4: Inline inet_blackhole_dev_init() to devinet_init() Kuniyuki Iwashima
2026-10-04 23:02 ` netdev-bot+sashiko
2026-10-01 20:47 ` [PATCH v3 net-next 3/6] net: Rename dev_isalive() to netif_is_alive() Kuniyuki Iwashima
2026-10-01 20:47 ` [PATCH v3 net-next 4/6] xfrm: Check netif_is_alive() in xfrm_bundle_create() and xfrm_create_dummy_bundle() Kuniyuki Iwashima
2026-10-01 20:47 ` Kuniyuki Iwashima [this message]
2026-10-04 23:02 ` [PATCH v3 net-next 5/6] ipv4: Batch rt_flush_dev() in netdev_run_todo() netdev-bot+sashiko
2026-10-05 0:51 ` Kuniyuki Iwashima
2026-10-01 20:47 ` [PATCH v3 net-next 6/6] ipv6: Batch rt6_uncached_list_flush_dev() " Kuniyuki Iwashima
2026-10-04 23:02 ` netdev-bot+sashiko
2026-10-05 0:49 ` Kuniyuki Iwashima
2026-10-06 8:12 ` [PATCH v3 net-next 0/6] ip: Batch flushing uncached routes per batched device unregistration Ido Schimmel
2026-10-06 17:45 ` Kuniyuki Iwashima
2026-10-06 20:07 ` Chris Arges
2026-10-06 23:10 ` patchwork-bot+netdevbpf
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=20261001204752.2572265-6-kuniyu@google.com \
--to=kuniyu@google.com \
--cc=andrew+netdev@lunn.ch \
--cc=carges@cloudflare.com \
--cc=davem@davemloft.net \
--cc=dsahern@kernel.org \
--cc=edumazet@kernel.org \
--cc=horms@kernel.org \
--cc=idosch@nvidia.com \
--cc=kuba@kernel.org \
--cc=kuni1840@gmail.com \
--cc=netdev@vger.kernel.org \
--cc=pabeni@redhat.com \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is an external index of several public inboxes,
see mirroring instructions on how to clone and mirror
all data and code used by this external index.