All of lore.kernel.org
 help / color / mirror / Atom feed
From: Kuniyuki Iwashima <kuniyu@google.com>
To: David Ahern <dsahern@kernel.org>,
	Ido Schimmel <idosch@nvidia.com>,
	 "David S . Miller" <davem@davemloft.net>,
	Eric Dumazet <edumazet@kernel.org>,
	 Jakub Kicinski <kuba@kernel.org>,
	Paolo Abeni <pabeni@redhat.com>
Cc: Simon Horman <horms@kernel.org>,
	Chris J Arges <carges@cloudflare.com>,
	 Kuniyuki Iwashima <kuniyu@google.com>,
	Kuniyuki Iwashima <kuni1840@gmail.com>,
	netdev@vger.kernel.org
Subject: [PATCH v1 net-next 4/5] ipv4: Batch rt_flush_dev() for dying netns.
Date: Sun, 27 Sep 2026 20:23:41 +0000	[thread overview]
Message-ID: <20260927202429.2452589-5-kuniyu@google.com> (raw)
In-Reply-To: <20260927202429.2452589-1-kuniyu@google.com>

IPv4 uncached routes are linked to the global per-cpu lists,
rt_uncached_list.

When unregistering a netdev, rt_flush_dev() iterates over the
potentially long lists to find uncached routes tied to the device
and swap it with blackhole_netdev.

Since it is called for every device in dying netns under RTNL,
it adds O(N_dev x (N_cpu + N_route)) costs to netns dismantle.

Let's call it (almost) once per cleanup_net().

When rt_flush_dev() is called with NULL from ->pre_exit_batch(),
it purges every entry in dying netns, reducing the cost to
O(N_cpu + N_route).

Since ->pre_exit_batch() is called before synchronize_rcu(),
we must prevent adding a new route for dying netns, so now
rt_add_uncached_list() checks !check_net() and swaps the device
with blackhole_netdev.

When rt_flush_dev() is later called again from fib_netdev_event()
via NETDEV_UNREGISTER, it just returns.

Note that net_pre_exit_done() cannot be replaced with !check_net()
because:

  1. some ->pre_exit() call unregister_netdevice() before
     fib_net_ops (e.g. ovs_pre_exit_net(), l2tp_pre_exit_net()).

  2. ->dellink() could call unregister_netdevice() for another
     netdev in a dying netns queued for the next cleanup_net()
     batch, for which ->pre_exit_batch() has not been called
     yet (e.g. veth).

Signed-off-by: Kuniyuki Iwashima <kuniyu@google.com>
---
 net/ipv4/fib_frontend.c |  6 ++++++
 net/ipv4/route.c        | 28 +++++++++++++++++++++++-----
 2 files changed, 29 insertions(+), 5 deletions(-)

diff --git a/net/ipv4/fib_frontend.c b/net/ipv4/fib_frontend.c
index 8a3dc04e8cac..b8d76b6279e1 100644
--- a/net/ipv4/fib_frontend.c
+++ b/net/ipv4/fib_frontend.c
@@ -1685,6 +1685,11 @@ static void __net_exit fib_net_pre_exit(struct net *net)
 	nl_fib_lookup_exit(net);
 }
 
+static void __net_exit fib_net_pre_exit_batch(struct list_head *net_exit_list)
+{
+	rt_flush_dev(NULL);
+}
+
 static void __net_exit fib_net_exit_rtnl(struct net *net,
 					 struct list_head *dev_kill_list)
 {
@@ -1704,6 +1709,7 @@ static void __net_exit fib_net_exit(struct net *net)
 static struct pernet_operations fib_net_ops = {
 	.init = fib_net_init,
 	.pre_exit = fib_net_pre_exit,
+	.pre_exit_batch = fib_net_pre_exit_batch,
 	.exit_rtnl = fib_net_exit_rtnl,
 	.exit = fib_net_exit,
 };
diff --git a/net/ipv4/route.c b/net/ipv4/route.c
index d7da2f1acbb5..cbe328b3f254 100644
--- a/net/ipv4/route.c
+++ b/net/ipv4/route.c
@@ -1554,14 +1554,29 @@ struct uncached_list {
 
 static DEFINE_PER_CPU_ALIGNED(struct uncached_list, rt_uncached_list);
 
+static void rt_replace_uncached_list(struct rtable *rt)
+{
+	struct net_device *dev = dst_dev(&rt->dst);
+
+	rcu_assign_pointer(rt->dst.dev_rcu, blackhole_netdev);
+	netdev_ref_replace(dev, blackhole_netdev,
+			   &rt->dst.dev_tracker, GFP_ATOMIC);
+}
+
 void rt_add_uncached_list(struct rtable *rt)
 {
 	struct uncached_list *ul = raw_cpu_ptr(&rt_uncached_list);
 
+	/* Set once and never cleared: non-NULL marks an uncached route. */
 	rt->dst.rt_uncached_list = ul;
 
 	spin_lock_bh(&ul->lock);
-	list_add_tail(&rt->dst.rt_uncached, &ul->head);
+
+	if (check_net(dst_dev_net_rcu(&rt->dst)))
+		list_add_tail(&rt->dst.rt_uncached, &ul->head);
+	else
+		rt_replace_uncached_list(rt);
+
 	spin_unlock_bh(&ul->lock);
 }
 
@@ -1587,6 +1602,9 @@ void rt_flush_dev(struct net_device *dev)
 	struct rtable *rt, *safe;
 	int cpu;
 
+	if (dev && net_pre_exit_done(dev_net(dev)))
+		return;
+
 	for_each_possible_cpu(cpu) {
 		struct uncached_list *ul = &per_cpu(rt_uncached_list, cpu);
 
@@ -1595,11 +1613,11 @@ void rt_flush_dev(struct net_device *dev)
 
 		spin_lock_bh(&ul->lock);
 		list_for_each_entry_safe(rt, safe, &ul->head, dst.rt_uncached) {
-			if (rt->dst.dev != dev)
+			if (rt->dst.dev != dev &&
+			    (dev || check_net(dev_net(rt->dst.dev))))
 				continue;
-			rcu_assign_pointer(rt->dst.dev_rcu, blackhole_netdev);
-			netdev_ref_replace(dev, blackhole_netdev,
-					   &rt->dst.dev_tracker, GFP_ATOMIC);
+
+			rt_replace_uncached_list(rt);
 			list_del_init(&rt->dst.rt_uncached);
 		}
 		spin_unlock_bh(&ul->lock);
-- 
2.56.0.rc1.315.gc6ed9934b7-goog


  parent reply	other threads:[~2026-09-27 20:24 UTC|newest]

Thread overview: 12+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-09-27 20:23 [PATCH v1 net-next 0/5] ip: Batch flushing uncached routes for dying netns Kuniyuki Iwashima
2026-09-27 20:23 ` [PATCH v1 net-next 1/5] net: Remove net->is_dying Kuniyuki Iwashima
2026-09-29  6:26   ` netdev-bot+sashiko
2026-09-27 20:23 ` [PATCH v1 net-next 2/5] net: Add ->pre_exit_batch() to struct pernet_operations Kuniyuki Iwashima
2026-09-29  6:26   ` netdev-bot+sashiko
2026-09-27 20:23 ` [PATCH v1 net-next 3/5] net: Track state in ops_undo_list() Kuniyuki Iwashima
2026-09-27 20:23 ` Kuniyuki Iwashima [this message]
2026-09-27 22:51   ` [PATCH v1 net-next 4/5] ipv4: Batch rt_flush_dev() for dying netns Eric Dumazet
2026-09-28 16:33     ` Kuniyuki Iwashima
2026-09-29  6:26   ` netdev-bot+sashiko
2026-09-27 20:23 ` [PATCH v1 net-next 5/5] ipv6: Batch rt6_uncached_list_flush_dev() " Kuniyuki Iwashima
2026-09-29  6:26   ` netdev-bot+sashiko

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=20260927202429.2452589-5-kuniyu@google.com \
    --to=kuniyu@google.com \
    --cc=carges@cloudflare.com \
    --cc=davem@davemloft.net \
    --cc=dsahern@kernel.org \
    --cc=edumazet@kernel.org \
    --cc=horms@kernel.org \
    --cc=idosch@nvidia.com \
    --cc=kuba@kernel.org \
    --cc=kuni1840@gmail.com \
    --cc=netdev@vger.kernel.org \
    --cc=pabeni@redhat.com \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is an external index of several public inboxes,
see mirroring instructions on how to clone and mirror
all data and code used by this external index.