From: Kuniyuki Iwashima <kuniyu@google.com>
To: David Ahern <dsahern@kernel.org>,
Ido Schimmel <idosch@nvidia.com>,
"David S . Miller" <davem@davemloft.net>,
Eric Dumazet <edumazet@kernel.org>,
Jakub Kicinski <kuba@kernel.org>,
Paolo Abeni <pabeni@redhat.com>
Cc: Simon Horman <horms@kernel.org>,
Chris J Arges <carges@cloudflare.com>,
Kuniyuki Iwashima <kuniyu@google.com>,
Kuniyuki Iwashima <kuni1840@gmail.com>,
netdev@vger.kernel.org
Subject: [PATCH v1 net-next 4/5] ipv4: Batch rt_flush_dev() for dying netns.
Date: Sun, 27 Sep 2026 20:23:41 +0000 [thread overview]
Message-ID: <20260927202429.2452589-5-kuniyu@google.com> (raw)
In-Reply-To: <20260927202429.2452589-1-kuniyu@google.com>
IPv4 uncached routes are linked to the global per-cpu lists,
rt_uncached_list.
When unregistering a netdev, rt_flush_dev() iterates over the
potentially long lists to find uncached routes tied to the device
and swap it with blackhole_netdev.
Since it is called for every device in dying netns under RTNL,
it adds O(N_dev x (N_cpu + N_route)) costs to netns dismantle.
Let's call it (almost) once per cleanup_net().
When rt_flush_dev() is called with NULL from ->pre_exit_batch(),
it purges every entry in dying netns, reducing the cost to
O(N_cpu + N_route).
Since ->pre_exit_batch() is called before synchronize_rcu(),
we must prevent adding a new route for dying netns, so now
rt_add_uncached_list() checks !check_net() and swaps the device
with blackhole_netdev.
When rt_flush_dev() is later called again from fib_netdev_event()
via NETDEV_UNREGISTER, it just returns.
Note that net_pre_exit_done() cannot be replaced with !check_net()
because:
1. some ->pre_exit() call unregister_netdevice() before
fib_net_ops (e.g. ovs_pre_exit_net(), l2tp_pre_exit_net()).
2. ->dellink() could call unregister_netdevice() for another
netdev in a dying netns queued for the next cleanup_net()
batch, for which ->pre_exit_batch() has not been called
yet (e.g. veth).
Signed-off-by: Kuniyuki Iwashima <kuniyu@google.com>
---
net/ipv4/fib_frontend.c | 6 ++++++
net/ipv4/route.c | 28 +++++++++++++++++++++++-----
2 files changed, 29 insertions(+), 5 deletions(-)
diff --git a/net/ipv4/fib_frontend.c b/net/ipv4/fib_frontend.c
index 8a3dc04e8cac..b8d76b6279e1 100644
--- a/net/ipv4/fib_frontend.c
+++ b/net/ipv4/fib_frontend.c
@@ -1685,6 +1685,11 @@ static void __net_exit fib_net_pre_exit(struct net *net)
nl_fib_lookup_exit(net);
}
+static void __net_exit fib_net_pre_exit_batch(struct list_head *net_exit_list)
+{
+ rt_flush_dev(NULL);
+}
+
static void __net_exit fib_net_exit_rtnl(struct net *net,
struct list_head *dev_kill_list)
{
@@ -1704,6 +1709,7 @@ static void __net_exit fib_net_exit(struct net *net)
static struct pernet_operations fib_net_ops = {
.init = fib_net_init,
.pre_exit = fib_net_pre_exit,
+ .pre_exit_batch = fib_net_pre_exit_batch,
.exit_rtnl = fib_net_exit_rtnl,
.exit = fib_net_exit,
};
diff --git a/net/ipv4/route.c b/net/ipv4/route.c
index d7da2f1acbb5..cbe328b3f254 100644
--- a/net/ipv4/route.c
+++ b/net/ipv4/route.c
@@ -1554,14 +1554,29 @@ struct uncached_list {
static DEFINE_PER_CPU_ALIGNED(struct uncached_list, rt_uncached_list);
+static void rt_replace_uncached_list(struct rtable *rt)
+{
+ struct net_device *dev = dst_dev(&rt->dst);
+
+ rcu_assign_pointer(rt->dst.dev_rcu, blackhole_netdev);
+ netdev_ref_replace(dev, blackhole_netdev,
+ &rt->dst.dev_tracker, GFP_ATOMIC);
+}
+
void rt_add_uncached_list(struct rtable *rt)
{
struct uncached_list *ul = raw_cpu_ptr(&rt_uncached_list);
+ /* Set once and never cleared: non-NULL marks an uncached route. */
rt->dst.rt_uncached_list = ul;
spin_lock_bh(&ul->lock);
- list_add_tail(&rt->dst.rt_uncached, &ul->head);
+
+ if (check_net(dst_dev_net_rcu(&rt->dst)))
+ list_add_tail(&rt->dst.rt_uncached, &ul->head);
+ else
+ rt_replace_uncached_list(rt);
+
spin_unlock_bh(&ul->lock);
}
@@ -1587,6 +1602,9 @@ void rt_flush_dev(struct net_device *dev)
struct rtable *rt, *safe;
int cpu;
+ if (dev && net_pre_exit_done(dev_net(dev)))
+ return;
+
for_each_possible_cpu(cpu) {
struct uncached_list *ul = &per_cpu(rt_uncached_list, cpu);
@@ -1595,11 +1613,11 @@ void rt_flush_dev(struct net_device *dev)
spin_lock_bh(&ul->lock);
list_for_each_entry_safe(rt, safe, &ul->head, dst.rt_uncached) {
- if (rt->dst.dev != dev)
+ if (rt->dst.dev != dev &&
+ (dev || check_net(dev_net(rt->dst.dev))))
continue;
- rcu_assign_pointer(rt->dst.dev_rcu, blackhole_netdev);
- netdev_ref_replace(dev, blackhole_netdev,
- &rt->dst.dev_tracker, GFP_ATOMIC);
+
+ rt_replace_uncached_list(rt);
list_del_init(&rt->dst.rt_uncached);
}
spin_unlock_bh(&ul->lock);
--
2.56.0.rc1.315.gc6ed9934b7-goog
next prev parent reply other threads:[~2026-09-27 20:24 UTC|newest]
Thread overview: 12+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-09-27 20:23 [PATCH v1 net-next 0/5] ip: Batch flushing uncached routes for dying netns Kuniyuki Iwashima
2026-09-27 20:23 ` [PATCH v1 net-next 1/5] net: Remove net->is_dying Kuniyuki Iwashima
2026-09-29 6:26 ` netdev-bot+sashiko
2026-09-27 20:23 ` [PATCH v1 net-next 2/5] net: Add ->pre_exit_batch() to struct pernet_operations Kuniyuki Iwashima
2026-09-29 6:26 ` netdev-bot+sashiko
2026-09-27 20:23 ` [PATCH v1 net-next 3/5] net: Track state in ops_undo_list() Kuniyuki Iwashima
2026-09-27 20:23 ` Kuniyuki Iwashima [this message]
2026-09-27 22:51 ` [PATCH v1 net-next 4/5] ipv4: Batch rt_flush_dev() for dying netns Eric Dumazet
2026-09-28 16:33 ` Kuniyuki Iwashima
2026-09-29 6:26 ` netdev-bot+sashiko
2026-09-27 20:23 ` [PATCH v1 net-next 5/5] ipv6: Batch rt6_uncached_list_flush_dev() " Kuniyuki Iwashima
2026-09-29 6:26 ` netdev-bot+sashiko
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=20260927202429.2452589-5-kuniyu@google.com \
--to=kuniyu@google.com \
--cc=carges@cloudflare.com \
--cc=davem@davemloft.net \
--cc=dsahern@kernel.org \
--cc=edumazet@kernel.org \
--cc=horms@kernel.org \
--cc=idosch@nvidia.com \
--cc=kuba@kernel.org \
--cc=kuni1840@gmail.com \
--cc=netdev@vger.kernel.org \
--cc=pabeni@redhat.com \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox