All of lore.kernel.org
 help / color / mirror / Atom feed
* [PATCH v1 net-next 0/5] ip: Batch flushing uncached routes for dying netns.
@ 2026-09-27 20:23 Kuniyuki Iwashima
  2026-09-27 20:23 ` [PATCH v1 net-next 1/5] net: Remove net->is_dying Kuniyuki Iwashima
                   ` (4 more replies)
  0 siblings, 5 replies; 12+ messages in thread
From: Kuniyuki Iwashima @ 2026-09-27 20:23 UTC (permalink / raw)
  To: David Ahern, Ido Schimmel, David S . Miller, Eric Dumazet,
	Jakub Kicinski, Paolo Abeni
  Cc: Simon Horman, Chris J Arges, Kuniyuki Iwashima, Kuniyuki Iwashima,
	netdev

Chris J Arges reported high RTNL contention during cleanup_net()
caused by rt_flush_dev() and rt6_uncached_list_flush_dev() iterating
over the global per-cpu uncached route lists for every netdev in
dying netns [0].

This series resolves the issue by introducing a new ->pre_exit_batch()
hook in struct pernet_operations and batching the uncached route
cleanup once per cleanup_net() without holding RTNL.

[0]: https://lore.kernel.org/netdev/20260917-hash-bucket-route-lists-v3-0-30493a37b6eb@cloudflare.com/


Kuniyuki Iwashima (5):
  net: Remove net->is_dying.
  net: Add ->pre_exit_batch() to struct pernet_operations.
  net: Track state in ops_undo_list().
  ipv4: Batch rt_flush_dev() for dying netns.
  ipv6: Batch rt6_uncached_list_flush_dev() for dying netns.

 include/net/net_namespace.h | 11 +++++++-
 net/core/net_namespace.c    | 13 +++++++--
 net/ipv4/fib_frontend.c     |  6 +++++
 net/ipv4/route.c            | 28 +++++++++++++++----
 net/ipv6/route.c            | 54 ++++++++++++++++++++++++++-----------
 5 files changed, 89 insertions(+), 23 deletions(-)

-- 
2.56.0.rc1.315.gc6ed9934b7-goog


^ permalink raw reply	[flat|nested] 12+ messages in thread

* [PATCH v1 net-next 1/5] net: Remove net->is_dying.
  2026-09-27 20:23 [PATCH v1 net-next 0/5] ip: Batch flushing uncached routes for dying netns Kuniyuki Iwashima
@ 2026-09-27 20:23 ` Kuniyuki Iwashima
  2026-09-29  6:26   ` netdev-bot+sashiko
  2026-09-27 20:23 ` [PATCH v1 net-next 2/5] net: Add ->pre_exit_batch() to struct pernet_operations Kuniyuki Iwashima
                   ` (3 subsequent siblings)
  4 siblings, 1 reply; 12+ messages in thread
From: Kuniyuki Iwashima @ 2026-09-27 20:23 UTC (permalink / raw)
  To: David Ahern, Ido Schimmel, David S . Miller, Eric Dumazet,
	Jakub Kicinski, Paolo Abeni
  Cc: Simon Horman, Chris J Arges, Kuniyuki Iwashima, Kuniyuki Iwashima,
	netdev

Commit 7acee67a6bce ("netns: optimize netns cleaning by batching
unhash_nsid calls") added the is_dying flag in struct net.

It is an alias of !check_net(net), and having two ways to represent
the same state is confusing.

Let's remove the flag.

Signed-off-by: Kuniyuki Iwashima <kuniyu@google.com>
---
 include/net/net_namespace.h | 1 -
 net/core/net_namespace.c    | 3 +--
 2 files changed, 1 insertion(+), 3 deletions(-)

diff --git a/include/net/net_namespace.h b/include/net/net_namespace.h
index 46b4c67e2966..b4247e767d13 100644
--- a/include/net/net_namespace.h
+++ b/include/net/net_namespace.h
@@ -125,7 +125,6 @@ struct net {
 	 * it is critical that it is on a read_mostly cache line.
 	 */
 	u32			hash_mix;
-	bool			is_dying;
 
 	struct net_device       *loopback_dev;          /* The loopback */
 
diff --git a/net/core/net_namespace.c b/net/core/net_namespace.c
index da5f881fbd3b..0e13de0cd36b 100644
--- a/net/core/net_namespace.c
+++ b/net/core/net_namespace.c
@@ -644,7 +644,7 @@ static void unhash_nsid(struct net *last)
 			int curr_id = id;
 
 			id++;
-			if (!peer->is_dying)
+			if (check_net(peer))
 				continue;
 
 			idr_remove(&tmp->netns_ids, curr_id);
@@ -681,7 +681,6 @@ static void cleanup_net(struct work_struct *work)
 	llist_for_each_entry(net, net_kill_list, cleanup_list) {
 		ns_tree_remove(net);
 		list_del_rcu(&net->list);
-		net->is_dying = true;
 	}
 	/* Cache last net. After we unlock rtnl, no one new net
 	 * added to net_namespace_list can assign nsid pointer
-- 
2.56.0.rc1.315.gc6ed9934b7-goog


^ permalink raw reply related	[flat|nested] 12+ messages in thread

* [PATCH v1 net-next 2/5] net: Add ->pre_exit_batch() to struct pernet_operations.
  2026-09-27 20:23 [PATCH v1 net-next 0/5] ip: Batch flushing uncached routes for dying netns Kuniyuki Iwashima
  2026-09-27 20:23 ` [PATCH v1 net-next 1/5] net: Remove net->is_dying Kuniyuki Iwashima
@ 2026-09-27 20:23 ` Kuniyuki Iwashima
  2026-09-29  6:26   ` netdev-bot+sashiko
  2026-09-27 20:23 ` [PATCH v1 net-next 3/5] net: Track state in ops_undo_list() Kuniyuki Iwashima
                   ` (2 subsequent siblings)
  4 siblings, 1 reply; 12+ messages in thread
From: Kuniyuki Iwashima @ 2026-09-27 20:23 UTC (permalink / raw)
  To: David Ahern, Ido Schimmel, David S . Miller, Eric Dumazet,
	Jakub Kicinski, Paolo Abeni
  Cc: Simon Horman, Chris J Arges, Kuniyuki Iwashima, Kuniyuki Iwashima,
	netdev

While unregistering a netdev, we must find and clean up all
resources holding refcounts of the device.

This operation is usually triggered for each component via
NETDEV_UNREGISTER and typically requires a full scan of hash
tables, etc.

For example, rt_flush_dev() and rt6_uncached_list_flush_dev()
are called for each netdev and iterate over the global per-cpu
lists to find uncached routes tied to the device, resulting
in O(N_dev x (N_cpu + N_route)) loops during netns dismantle.

Moreover, rt_flush_dev() and rt6_uncached_list_flush_dev()
do not even require RTNL.

To batch such cleanup across dying netns without holding RTNL
before ops_exit_rtnl_list(), let's add a new ->pre_exit_batch()
hook to pernet_operations.

Signed-off-by: Kuniyuki Iwashima <kuniyu@google.com>
---
 include/net/net_namespace.h | 1 +
 net/core/net_namespace.c    | 3 +++
 2 files changed, 4 insertions(+)

diff --git a/include/net/net_namespace.h b/include/net/net_namespace.h
index b4247e767d13..58b2601bb869 100644
--- a/include/net/net_namespace.h
+++ b/include/net/net_namespace.h
@@ -488,6 +488,7 @@ struct pernet_operations {
 	 */
 	int (*init)(struct net *net);
 	void (*pre_exit)(struct net *net);
+	void (*pre_exit_batch)(struct list_head *net_exit_list);
 	void (*exit)(struct net *net);
 	void (*exit_batch)(struct list_head *net_exit_list);
 	/* Following method is called with RTNL held. */
diff --git a/net/core/net_namespace.c b/net/core/net_namespace.c
index 0e13de0cd36b..476fbf913bad 100644
--- a/net/core/net_namespace.c
+++ b/net/core/net_namespace.c
@@ -160,6 +160,9 @@ static void ops_pre_exit_list(const struct pernet_operations *ops,
 		list_for_each_entry(net, net_exit_list, exit_list)
 			ops->pre_exit(net);
 	}
+
+	if (ops->pre_exit_batch)
+		ops->pre_exit_batch(net_exit_list);
 }
 
 static void ops_exit_rtnl_list(const struct list_head *ops_list,
-- 
2.56.0.rc1.315.gc6ed9934b7-goog


^ permalink raw reply related	[flat|nested] 12+ messages in thread

* [PATCH v1 net-next 3/5] net: Track state in ops_undo_list().
  2026-09-27 20:23 [PATCH v1 net-next 0/5] ip: Batch flushing uncached routes for dying netns Kuniyuki Iwashima
  2026-09-27 20:23 ` [PATCH v1 net-next 1/5] net: Remove net->is_dying Kuniyuki Iwashima
  2026-09-27 20:23 ` [PATCH v1 net-next 2/5] net: Add ->pre_exit_batch() to struct pernet_operations Kuniyuki Iwashima
@ 2026-09-27 20:23 ` Kuniyuki Iwashima
  2026-09-27 20:23 ` [PATCH v1 net-next 4/5] ipv4: Batch rt_flush_dev() for dying netns Kuniyuki Iwashima
  2026-09-27 20:23 ` [PATCH v1 net-next 5/5] ipv6: Batch rt6_uncached_list_flush_dev() " Kuniyuki Iwashima
  4 siblings, 0 replies; 12+ messages in thread
From: Kuniyuki Iwashima @ 2026-09-27 20:23 UTC (permalink / raw)
  To: David Ahern, Ido Schimmel, David S . Miller, Eric Dumazet,
	Jakub Kicinski, Paolo Abeni
  Cc: Simon Horman, Chris J Arges, Kuniyuki Iwashima, Kuniyuki Iwashima,
	netdev

We will call rt_flush_dev() and rt6_uncached_list_flush_dev()
from ->pre_exit_batch().

Then, we want them to return early when called again from
 ->exit_rtnl() or default_device_exit_batch().

However, we cannot simply return early when !check_net(net).

In the following cases, even if check_net(net) is false,
we cannot skip rt_flush_dev() / rt6_uncached_list_flush_dev():

  1. some ->pre_exit() call unregister_netdevice() before
     fib_net_ops (e.g. ovs_pre_exit_net(), l2tp_pre_exit_net()).

  2. ->dellink() could call unregister_netdevice() for another
     netdev in a dying netns queued for the next cleanup_net()
     batch, for which ->pre_exit_batch() has not been called
     yet (e.g. veth).

Thus, we need a clear flag to indicate that ->pre_exit_batch()
has already been called.

Let's add net->undo_state and update it only for dying netns.

Signed-off-by: Kuniyuki Iwashima <kuniyu@google.com>
---
 include/net/net_namespace.h | 9 +++++++++
 net/core/net_namespace.c    | 7 +++++++
 2 files changed, 16 insertions(+)

diff --git a/include/net/net_namespace.h b/include/net/net_namespace.h
index 58b2601bb869..6a817326fbb8 100644
--- a/include/net/net_namespace.h
+++ b/include/net/net_namespace.h
@@ -57,6 +57,9 @@ struct uevent_sock;
 struct netns_ipvs;
 struct bpf_prog;
 
+enum {
+	NET_PRE_EXIT_DONE = 1,
+};
 
 #define NETDEV_HASHBITS    8
 #define NETDEV_HASHENTRIES (1 << NETDEV_HASHBITS)
@@ -125,6 +128,7 @@ struct net {
 	 * it is critical that it is on a read_mostly cache line.
 	 */
 	u32			hash_mix;
+	u8			undo_state;
 
 	struct net_device       *loopback_dev;          /* The loopback */
 
@@ -364,6 +368,11 @@ static inline bool net_initialized(const struct net *net)
 	return READ_ONCE(net->list.next);
 }
 
+static inline bool net_pre_exit_done(const struct net *net)
+{
+	return READ_ONCE(net->undo_state) >= NET_PRE_EXIT_DONE;
+}
+
 static inline void __netns_tracker_alloc(struct net *net,
 					 netns_tracker *tracker,
 					 bool refcounted,
diff --git a/net/core/net_namespace.c b/net/core/net_namespace.c
index 476fbf913bad..9e2462bfa478 100644
--- a/net/core/net_namespace.c
+++ b/net/core/net_namespace.c
@@ -226,7 +226,9 @@ static void ops_undo_list(const struct list_head *ops_list,
 			  bool expedite_rcu)
 {
 	const struct pernet_operations *saved_ops;
+	bool dying = ops_list == &pernet_list;
 	bool hold_rtnl = false;
+	struct net *net;
 
 	if (!ops)
 		ops = list_entry(ops_list, typeof(*ops), list);
@@ -248,6 +250,11 @@ static void ops_undo_list(const struct list_head *ops_list,
 	else
 		synchronize_rcu();
 
+	if (dying) {
+		list_for_each_entry(net, net_exit_list, exit_list)
+			WRITE_ONCE(net->undo_state, NET_PRE_EXIT_DONE);
+	}
+
 	if (hold_rtnl)
 		ops_exit_rtnl_list(ops_list, saved_ops, net_exit_list);
 
-- 
2.56.0.rc1.315.gc6ed9934b7-goog


^ permalink raw reply related	[flat|nested] 12+ messages in thread

* [PATCH v1 net-next 4/5] ipv4: Batch rt_flush_dev() for dying netns.
  2026-09-27 20:23 [PATCH v1 net-next 0/5] ip: Batch flushing uncached routes for dying netns Kuniyuki Iwashima
                   ` (2 preceding siblings ...)
  2026-09-27 20:23 ` [PATCH v1 net-next 3/5] net: Track state in ops_undo_list() Kuniyuki Iwashima
@ 2026-09-27 20:23 ` Kuniyuki Iwashima
  2026-09-27 22:51   ` Eric Dumazet
  2026-09-29  6:26   ` netdev-bot+sashiko
  2026-09-27 20:23 ` [PATCH v1 net-next 5/5] ipv6: Batch rt6_uncached_list_flush_dev() " Kuniyuki Iwashima
  4 siblings, 2 replies; 12+ messages in thread
From: Kuniyuki Iwashima @ 2026-09-27 20:23 UTC (permalink / raw)
  To: David Ahern, Ido Schimmel, David S . Miller, Eric Dumazet,
	Jakub Kicinski, Paolo Abeni
  Cc: Simon Horman, Chris J Arges, Kuniyuki Iwashima, Kuniyuki Iwashima,
	netdev

IPv4 uncached routes are linked to the global per-cpu lists,
rt_uncached_list.

When unregistering a netdev, rt_flush_dev() iterates over the
potentially long lists to find uncached routes tied to the device
and swap it with blackhole_netdev.

Since it is called for every device in dying netns under RTNL,
it adds O(N_dev x (N_cpu + N_route)) costs to netns dismantle.

Let's call it (almost) once per cleanup_net().

When rt_flush_dev() is called with NULL from ->pre_exit_batch(),
it purges every entry in dying netns, reducing the cost to
O(N_cpu + N_route).

Since ->pre_exit_batch() is called before synchronize_rcu(),
we must prevent adding a new route for dying netns, so now
rt_add_uncached_list() checks !check_net() and swaps the device
with blackhole_netdev.

When rt_flush_dev() is later called again from fib_netdev_event()
via NETDEV_UNREGISTER, it just returns.

Note that net_pre_exit_done() cannot be replaced with !check_net()
because:

  1. some ->pre_exit() call unregister_netdevice() before
     fib_net_ops (e.g. ovs_pre_exit_net(), l2tp_pre_exit_net()).

  2. ->dellink() could call unregister_netdevice() for another
     netdev in a dying netns queued for the next cleanup_net()
     batch, for which ->pre_exit_batch() has not been called
     yet (e.g. veth).

Signed-off-by: Kuniyuki Iwashima <kuniyu@google.com>
---
 net/ipv4/fib_frontend.c |  6 ++++++
 net/ipv4/route.c        | 28 +++++++++++++++++++++++-----
 2 files changed, 29 insertions(+), 5 deletions(-)

diff --git a/net/ipv4/fib_frontend.c b/net/ipv4/fib_frontend.c
index 8a3dc04e8cac..b8d76b6279e1 100644
--- a/net/ipv4/fib_frontend.c
+++ b/net/ipv4/fib_frontend.c
@@ -1685,6 +1685,11 @@ static void __net_exit fib_net_pre_exit(struct net *net)
 	nl_fib_lookup_exit(net);
 }
 
+static void __net_exit fib_net_pre_exit_batch(struct list_head *net_exit_list)
+{
+	rt_flush_dev(NULL);
+}
+
 static void __net_exit fib_net_exit_rtnl(struct net *net,
 					 struct list_head *dev_kill_list)
 {
@@ -1704,6 +1709,7 @@ static void __net_exit fib_net_exit(struct net *net)
 static struct pernet_operations fib_net_ops = {
 	.init = fib_net_init,
 	.pre_exit = fib_net_pre_exit,
+	.pre_exit_batch = fib_net_pre_exit_batch,
 	.exit_rtnl = fib_net_exit_rtnl,
 	.exit = fib_net_exit,
 };
diff --git a/net/ipv4/route.c b/net/ipv4/route.c
index d7da2f1acbb5..cbe328b3f254 100644
--- a/net/ipv4/route.c
+++ b/net/ipv4/route.c
@@ -1554,14 +1554,29 @@ struct uncached_list {
 
 static DEFINE_PER_CPU_ALIGNED(struct uncached_list, rt_uncached_list);
 
+static void rt_replace_uncached_list(struct rtable *rt)
+{
+	struct net_device *dev = dst_dev(&rt->dst);
+
+	rcu_assign_pointer(rt->dst.dev_rcu, blackhole_netdev);
+	netdev_ref_replace(dev, blackhole_netdev,
+			   &rt->dst.dev_tracker, GFP_ATOMIC);
+}
+
 void rt_add_uncached_list(struct rtable *rt)
 {
 	struct uncached_list *ul = raw_cpu_ptr(&rt_uncached_list);
 
+	/* Set once and never cleared: non-NULL marks an uncached route. */
 	rt->dst.rt_uncached_list = ul;
 
 	spin_lock_bh(&ul->lock);
-	list_add_tail(&rt->dst.rt_uncached, &ul->head);
+
+	if (check_net(dst_dev_net_rcu(&rt->dst)))
+		list_add_tail(&rt->dst.rt_uncached, &ul->head);
+	else
+		rt_replace_uncached_list(rt);
+
 	spin_unlock_bh(&ul->lock);
 }
 
@@ -1587,6 +1602,9 @@ void rt_flush_dev(struct net_device *dev)
 	struct rtable *rt, *safe;
 	int cpu;
 
+	if (dev && net_pre_exit_done(dev_net(dev)))
+		return;
+
 	for_each_possible_cpu(cpu) {
 		struct uncached_list *ul = &per_cpu(rt_uncached_list, cpu);
 
@@ -1595,11 +1613,11 @@ void rt_flush_dev(struct net_device *dev)
 
 		spin_lock_bh(&ul->lock);
 		list_for_each_entry_safe(rt, safe, &ul->head, dst.rt_uncached) {
-			if (rt->dst.dev != dev)
+			if (rt->dst.dev != dev &&
+			    (dev || check_net(dev_net(rt->dst.dev))))
 				continue;
-			rcu_assign_pointer(rt->dst.dev_rcu, blackhole_netdev);
-			netdev_ref_replace(dev, blackhole_netdev,
-					   &rt->dst.dev_tracker, GFP_ATOMIC);
+
+			rt_replace_uncached_list(rt);
 			list_del_init(&rt->dst.rt_uncached);
 		}
 		spin_unlock_bh(&ul->lock);
-- 
2.56.0.rc1.315.gc6ed9934b7-goog


^ permalink raw reply related	[flat|nested] 12+ messages in thread

* [PATCH v1 net-next 5/5] ipv6: Batch rt6_uncached_list_flush_dev() for dying netns.
  2026-09-27 20:23 [PATCH v1 net-next 0/5] ip: Batch flushing uncached routes for dying netns Kuniyuki Iwashima
                   ` (3 preceding siblings ...)
  2026-09-27 20:23 ` [PATCH v1 net-next 4/5] ipv4: Batch rt_flush_dev() for dying netns Kuniyuki Iwashima
@ 2026-09-27 20:23 ` Kuniyuki Iwashima
  2026-09-29  6:26   ` netdev-bot+sashiko
  4 siblings, 1 reply; 12+ messages in thread
From: Kuniyuki Iwashima @ 2026-09-27 20:23 UTC (permalink / raw)
  To: David Ahern, Ido Schimmel, David S . Miller, Eric Dumazet,
	Jakub Kicinski, Paolo Abeni
  Cc: Simon Horman, Chris J Arges, Kuniyuki Iwashima, Kuniyuki Iwashima,
	netdev

IPv6 uncached routes are linked to the global per-cpu lists,
rt6_uncached_list.

When unregistering a netdev, rt6_uncached_list_flush_dev()
iterates over the potentially long lists to find uncached
routes tied to the device and swap it with blackhole_netdev.

Since it is called for every device in dying netns under RTNL,
it adds O(N_dev x (N_cpu + N_route)) costs to netns dismantle.

Similar to IPv4, let's call it (almost) once per cleanup_net().

Chris J Arges verified the changes improve RTNL wait time
during cleanup_net() by 75x, quoted from [0]:

  <...> I was able to confirm even greater reduction in contention
  as measured by how much latency an unrelated process takes when
  waiting for cleanup_net to complete. <...>

  Some rough average latency numbers with 36 devices, 160k routes,
  8 vCPUs:

  - main: 137ms
  <...>
  - your patchset: 1.8ms

Reported-by: Chris J Arges <carges@cloudflare.com>
Closes: https://lore.kernel.org/netdev/20260917-hash-bucket-route-lists-v3-0-30493a37b6eb@cloudflare.com/
Link: https://lore.kernel.org/netdev/aq2B8PSfjn-xau4V@20HS2G4/ #[0]
Signed-off-by: Kuniyuki Iwashima <kuniyu@google.com>
---
 net/ipv6/route.c | 54 ++++++++++++++++++++++++++++++++++--------------
 1 file changed, 39 insertions(+), 15 deletions(-)

diff --git a/net/ipv6/route.c b/net/ipv6/route.c
index 475ced827ec5..7747e4f20fee 100644
--- a/net/ipv6/route.c
+++ b/net/ipv6/route.c
@@ -135,6 +135,22 @@ struct uncached_list {
 
 static DEFINE_PER_CPU_ALIGNED(struct uncached_list, rt6_uncached_list);
 
+static void rt6_uncached_list_replace(struct rt6_info *rt)
+{
+	struct net_device *dev = dst_dev(&rt->dst);
+	struct inet6_dev *rt_idev = rt->rt6i_idev;
+
+	if (rt_idev) {
+		rt->rt6i_idev = in6_dev_get(blackhole_netdev);
+		in6_dev_put(rt_idev);
+	}
+
+	rcu_assign_pointer(rt->dst.dev_rcu, blackhole_netdev);
+	netdev_ref_replace(dev, blackhole_netdev,
+			   &rt->dst.dev_tracker,
+			   GFP_ATOMIC);
+}
+
 void rt6_uncached_list_add(struct rt6_info *rt)
 {
 	struct uncached_list *ul = raw_cpu_ptr(&rt6_uncached_list);
@@ -143,7 +159,12 @@ void rt6_uncached_list_add(struct rt6_info *rt)
 	rt->dst.rt_uncached_list = ul;
 
 	spin_lock_bh(&ul->lock);
-	list_add_tail(&rt->dst.rt_uncached, &ul->head);
+
+	if (check_net(dst_dev_net_rcu(&rt->dst)))
+		list_add_tail(&rt->dst.rt_uncached, &ul->head);
+	else
+		rt6_uncached_list_replace(rt);
+
 	spin_unlock_bh(&ul->lock);
 }
 
@@ -162,6 +183,9 @@ static void rt6_uncached_list_flush_dev(struct net_device *dev)
 {
 	int cpu;
 
+	if (dev && net_pre_exit_done(dev_net(dev)))
+		return;
+
 	for_each_possible_cpu(cpu) {
 		struct uncached_list *ul = per_cpu_ptr(&rt6_uncached_list, cpu);
 		struct rt6_info *rt, *safe;
@@ -173,23 +197,17 @@ static void rt6_uncached_list_flush_dev(struct net_device *dev)
 		list_for_each_entry_safe(rt, safe, &ul->head, dst.rt_uncached) {
 			struct inet6_dev *rt_idev = rt->rt6i_idev;
 			struct net_device *rt_dev = rt->dst.dev;
-			bool handled = false;
 
-			if (rt_idev && rt_idev->dev == dev) {
-				rt->rt6i_idev = in6_dev_get(blackhole_netdev);
-				in6_dev_put(rt_idev);
-				handled = true;
+			if (dev) {
+				if (rt_dev != dev &&
+				    (!rt_idev || rt_idev->dev != dev))
+					continue;
+			} else if (check_net(dev_net(rt_dev))) {
+				continue;
 			}
 
-			if (rt_dev == dev) {
-				rt->dst.dev = blackhole_netdev;
-				netdev_ref_replace(rt_dev, blackhole_netdev,
-						   &rt->dst.dev_tracker,
-						   GFP_ATOMIC);
-				handled = true;
-			}
-			if (handled)
-				list_del_init(&rt->dst.rt_uncached);
+			rt6_uncached_list_replace(rt);
+			list_del_init(&rt->dst.rt_uncached);
 		}
 		spin_unlock_bh(&ul->lock);
 	}
@@ -6801,6 +6819,11 @@ static int __net_init ip6_route_net_init(struct net *net)
 	goto out;
 }
 
+static void __net_exit ip6_route_net_pre_exit_batch(struct list_head *net_exit_list)
+{
+	rt6_uncached_list_flush_dev(NULL);
+}
+
 static void __net_exit ip6_route_net_exit(struct net *net)
 {
 	kfree(net->ipv6.fib6_null_entry);
@@ -6839,6 +6862,7 @@ static void __net_exit ip6_route_net_exit_late(struct net *net)
 
 static struct pernet_operations ip6_route_net_ops = {
 	.init = ip6_route_net_init,
+	.pre_exit_batch = ip6_route_net_pre_exit_batch,
 	.exit = ip6_route_net_exit,
 };
 
-- 
2.56.0.rc1.315.gc6ed9934b7-goog


^ permalink raw reply related	[flat|nested] 12+ messages in thread

* Re: [PATCH v1 net-next 4/5] ipv4: Batch rt_flush_dev() for dying netns.
  2026-09-27 20:23 ` [PATCH v1 net-next 4/5] ipv4: Batch rt_flush_dev() for dying netns Kuniyuki Iwashima
@ 2026-09-27 22:51   ` Eric Dumazet
  2026-09-28 16:33     ` Kuniyuki Iwashima
  2026-09-29  6:26   ` netdev-bot+sashiko
  1 sibling, 1 reply; 12+ messages in thread
From: Eric Dumazet @ 2026-09-27 22:51 UTC (permalink / raw)
  To: Kuniyuki Iwashima
  Cc: David Ahern, Ido Schimmel, David S . Miller, Jakub Kicinski,
	Paolo Abeni, Simon Horman, Chris J Arges, Kuniyuki Iwashima,
	netdev

On Sun, Sep 27, 2026 at 10:24 PM Kuniyuki Iwashima <kuniyu@google.com> wrote:
>
> IPv4 uncached routes are linked to the global per-cpu lists,
> rt_uncached_list.
>
> When unregistering a netdev, rt_flush_dev() iterates over the
> potentially long lists to find uncached routes tied to the device
> and swap it with blackhole_netdev.
>
> Since it is called for every device in dying netns under RTNL,
> it adds O(N_dev x (N_cpu + N_route)) costs to netns dismantle.
>
> Let's call it (almost) once per cleanup_net().
>
> When rt_flush_dev() is called with NULL from ->pre_exit_batch(),
> it purges every entry in dying netns, reducing the cost to
> O(N_cpu + N_route).
>
> Since ->pre_exit_batch() is called before synchronize_rcu(),
> we must prevent adding a new route for dying netns, so now
> rt_add_uncached_list() checks !check_net() and swaps the device
> with blackhole_netdev.
>
> When rt_flush_dev() is later called again from fib_netdev_event()
> via NETDEV_UNREGISTER, it just returns.
>
> Note that net_pre_exit_done() cannot be replaced with !check_net()
> because:
>
>   1. some ->pre_exit() call unregister_netdevice() before
>      fib_net_ops (e.g. ovs_pre_exit_net(), l2tp_pre_exit_net()).
>
>   2. ->dellink() could call unregister_netdevice() for another
>      netdev in a dying netns queued for the next cleanup_net()
>      batch, for which ->pre_exit_batch() has not been called
>      yet (e.g. veth).
>
> Signed-off-by: Kuniyuki Iwashima <kuniyu@google.com>
> ---
>  net/ipv4/fib_frontend.c |  6 ++++++
>  net/ipv4/route.c        | 28 +++++++++++++++++++++++-----
>  2 files changed, 29 insertions(+), 5 deletions(-)
>
> diff --git a/net/ipv4/fib_frontend.c b/net/ipv4/fib_frontend.c
> index 8a3dc04e8cac..b8d76b6279e1 100644
> --- a/net/ipv4/fib_frontend.c
> +++ b/net/ipv4/fib_frontend.c
> @@ -1685,6 +1685,11 @@ static void __net_exit fib_net_pre_exit(struct net *net)
>         nl_fib_lookup_exit(net);
>  }
>
> +static void __net_exit fib_net_pre_exit_batch(struct list_head *net_exit_list)
> +{
> +       rt_flush_dev(NULL);
> +}
> +
>  static void __net_exit fib_net_exit_rtnl(struct net *net,
>                                          struct list_head *dev_kill_list)
>  {
> @@ -1704,6 +1709,7 @@ static void __net_exit fib_net_exit(struct net *net)
>  static struct pernet_operations fib_net_ops = {
>         .init = fib_net_init,
>         .pre_exit = fib_net_pre_exit,
> +       .pre_exit_batch = fib_net_pre_exit_batch,
>         .exit_rtnl = fib_net_exit_rtnl,
>         .exit = fib_net_exit,
>  };
> diff --git a/net/ipv4/route.c b/net/ipv4/route.c
> index d7da2f1acbb5..cbe328b3f254 100644
> --- a/net/ipv4/route.c
> +++ b/net/ipv4/route.c
> @@ -1554,14 +1554,29 @@ struct uncached_list {
>
>  static DEFINE_PER_CPU_ALIGNED(struct uncached_list, rt_uncached_list);
>
> +static void rt_replace_uncached_list(struct rtable *rt)
> +{
> +       struct net_device *dev = dst_dev(&rt->dst);
> +
> +       rcu_assign_pointer(rt->dst.dev_rcu, blackhole_netdev);
> +       netdev_ref_replace(dev, blackhole_netdev,
> +                          &rt->dst.dev_tracker, GFP_ATOMIC);
> +}
> +
>  void rt_add_uncached_list(struct rtable *rt)
>  {
>         struct uncached_list *ul = raw_cpu_ptr(&rt_uncached_list);
>
> +       /* Set once and never cleared: non-NULL marks an uncached route. */
>         rt->dst.rt_uncached_list = ul;
>
>         spin_lock_bh(&ul->lock);
> -       list_add_tail(&rt->dst.rt_uncached, &ul->head);
> +
> +       if (check_net(dst_dev_net_rcu(&rt->dst)))
> +               list_add_tail(&rt->dst.rt_uncached, &ul->head);
> +       else
> +               rt_replace_uncached_list(rt);
> +
>         spin_unlock_bh(&ul->lock);
>  }

Hi Kuniyuki,

Doing this in ->pre_exit_batch() and rt_add_uncached_list() /
rt6_uncached_list_add() is too early.

1) check_net(net) becomes false as soon as the last put_net() drops
   __ns_ref to 0, while devices in the netns are still UP and sockets /
   packets are still active (until default_device_exit_batch()).
   Swapping rt->dst.dev_rcu to blackhole_netdev on live routes (while
   rt->dst.input and rt->dst.output are not dst_discard) changes
   skb_dst_dev_net_rcu(skb) / dst_dev_net_rcu(&rt->dst) to &init_net.
   RX paths (__inet_lookup_skb, __inet6_lookup_skb, tcp_v4_rcv,
   icmp_rcv, ip_error, ...) will then look up sockets in init_net or
   send RST/ICMP from init_net, and TX paths (ip_route_output_flow,
   icmp6_dst_alloc, ip6_output) will see blackhole_netdev / init_net.

2) In patch 3/5, ops_undo_list() is also called from setup_net()
   (out_undo) when an ->init() callback fails. At that point
   check_net(net) is still true (__ns_ref == 1), so ->pre_exit_batch()
   will skip net (or not be called at all if setup_net() failed before
   fib_net_ops), yet ops_undo_list() sets NET_PRE_EXIT_DONE because
   ops_list == &pernet_list, causing default_device_exit_batch() to
   skip rt_flush_dev(dev).

3) In rt_flush_dev() / rt6_uncached_list_flush_dev(), the lockless
   if (list_empty(&ul->head)) check can race with a concurrent
   rt_add_uncached_list() that observed check_net(net) == true just
   before the last put_net() but has not yet executed list_add_tail().
   Since NET_PRE_EXIT_DONE disables all subsequent rt_flush_dev(dev)
   calls (including the retry in netdev_wait_allrefs_any()), that dev
   reference would never be released.

Instead of ->pre_exit_batch(), could we batch this at the
unregister_netdevice_many_notify() stage (after the NETDEV_UNREGISTER
loop, when all devices in the batch are already down, unlisted, and
marked NETREG_UNREGISTERING)? That would avoid ->pre_exit_batch(),
net->undo_state, and touching rt_add_uncached_list(), while also
speeding up any batched device unregistration.

^ permalink raw reply	[flat|nested] 12+ messages in thread

* Re: [PATCH v1 net-next 4/5] ipv4: Batch rt_flush_dev() for dying netns.
  2026-09-27 22:51   ` Eric Dumazet
@ 2026-09-28 16:33     ` Kuniyuki Iwashima
  0 siblings, 0 replies; 12+ messages in thread
From: Kuniyuki Iwashima @ 2026-09-28 16:33 UTC (permalink / raw)
  To: Eric Dumazet
  Cc: David Ahern, Ido Schimmel, David S . Miller, Jakub Kicinski,
	Paolo Abeni, Simon Horman, Chris J Arges, Kuniyuki Iwashima,
	netdev

On Sun, Sep 27, 2026 at 3:52 PM Eric Dumazet <edumazet@kernel.org> wrote:
>
> On Sun, Sep 27, 2026 at 10:24 PM Kuniyuki Iwashima <kuniyu@google.com> wrote:
> >
> > IPv4 uncached routes are linked to the global per-cpu lists,
> > rt_uncached_list.
> >
> > When unregistering a netdev, rt_flush_dev() iterates over the
> > potentially long lists to find uncached routes tied to the device
> > and swap it with blackhole_netdev.
> >
> > Since it is called for every device in dying netns under RTNL,
> > it adds O(N_dev x (N_cpu + N_route)) costs to netns dismantle.
> >
> > Let's call it (almost) once per cleanup_net().
> >
> > When rt_flush_dev() is called with NULL from ->pre_exit_batch(),
> > it purges every entry in dying netns, reducing the cost to
> > O(N_cpu + N_route).
> >
> > Since ->pre_exit_batch() is called before synchronize_rcu(),
> > we must prevent adding a new route for dying netns, so now
> > rt_add_uncached_list() checks !check_net() and swaps the device
> > with blackhole_netdev.
> >
> > When rt_flush_dev() is later called again from fib_netdev_event()
> > via NETDEV_UNREGISTER, it just returns.
> >
> > Note that net_pre_exit_done() cannot be replaced with !check_net()
> > because:
> >
> >   1. some ->pre_exit() call unregister_netdevice() before
> >      fib_net_ops (e.g. ovs_pre_exit_net(), l2tp_pre_exit_net()).
> >
> >   2. ->dellink() could call unregister_netdevice() for another
> >      netdev in a dying netns queued for the next cleanup_net()
> >      batch, for which ->pre_exit_batch() has not been called
> >      yet (e.g. veth).
> >
> > Signed-off-by: Kuniyuki Iwashima <kuniyu@google.com>
> > ---
> >  net/ipv4/fib_frontend.c |  6 ++++++
> >  net/ipv4/route.c        | 28 +++++++++++++++++++++++-----
> >  2 files changed, 29 insertions(+), 5 deletions(-)
> >
> > diff --git a/net/ipv4/fib_frontend.c b/net/ipv4/fib_frontend.c
> > index 8a3dc04e8cac..b8d76b6279e1 100644
> > --- a/net/ipv4/fib_frontend.c
> > +++ b/net/ipv4/fib_frontend.c
> > @@ -1685,6 +1685,11 @@ static void __net_exit fib_net_pre_exit(struct net *net)
> >         nl_fib_lookup_exit(net);
> >  }
> >
> > +static void __net_exit fib_net_pre_exit_batch(struct list_head *net_exit_list)
> > +{
> > +       rt_flush_dev(NULL);
> > +}
> > +
> >  static void __net_exit fib_net_exit_rtnl(struct net *net,
> >                                          struct list_head *dev_kill_list)
> >  {
> > @@ -1704,6 +1709,7 @@ static void __net_exit fib_net_exit(struct net *net)
> >  static struct pernet_operations fib_net_ops = {
> >         .init = fib_net_init,
> >         .pre_exit = fib_net_pre_exit,
> > +       .pre_exit_batch = fib_net_pre_exit_batch,
> >         .exit_rtnl = fib_net_exit_rtnl,
> >         .exit = fib_net_exit,
> >  };
> > diff --git a/net/ipv4/route.c b/net/ipv4/route.c
> > index d7da2f1acbb5..cbe328b3f254 100644
> > --- a/net/ipv4/route.c
> > +++ b/net/ipv4/route.c
> > @@ -1554,14 +1554,29 @@ struct uncached_list {
> >
> >  static DEFINE_PER_CPU_ALIGNED(struct uncached_list, rt_uncached_list);
> >
> > +static void rt_replace_uncached_list(struct rtable *rt)
> > +{
> > +       struct net_device *dev = dst_dev(&rt->dst);
> > +
> > +       rcu_assign_pointer(rt->dst.dev_rcu, blackhole_netdev);
> > +       netdev_ref_replace(dev, blackhole_netdev,
> > +                          &rt->dst.dev_tracker, GFP_ATOMIC);
> > +}
> > +
> >  void rt_add_uncached_list(struct rtable *rt)
> >  {
> >         struct uncached_list *ul = raw_cpu_ptr(&rt_uncached_list);
> >
> > +       /* Set once and never cleared: non-NULL marks an uncached route. */
> >         rt->dst.rt_uncached_list = ul;
> >
> >         spin_lock_bh(&ul->lock);
> > -       list_add_tail(&rt->dst.rt_uncached, &ul->head);
> > +
> > +       if (check_net(dst_dev_net_rcu(&rt->dst)))
> > +               list_add_tail(&rt->dst.rt_uncached, &ul->head);
> > +       else
> > +               rt_replace_uncached_list(rt);
> > +
> >         spin_unlock_bh(&ul->lock);
> >  }
>
> Hi Kuniyuki,
>
> Doing this in ->pre_exit_batch() and rt_add_uncached_list() /
> rt6_uncached_list_add() is too early.
>
> 1) check_net(net) becomes false as soon as the last put_net() drops
>    __ns_ref to 0, while devices in the netns are still UP and sockets /
>    packets are still active (until default_device_exit_batch()).
>    Swapping rt->dst.dev_rcu to blackhole_netdev on live routes (while
>    rt->dst.input and rt->dst.output are not dst_discard) changes
>    skb_dst_dev_net_rcu(skb) / dst_dev_net_rcu(&rt->dst) to &init_net.
>    RX paths (__inet_lookup_skb, __inet6_lookup_skb, tcp_v4_rcv,
>    icmp_rcv, ip_error, ...) will then look up sockets in init_net or
>    send RST/ICMP from init_net, and TX paths (ip_route_output_flow,
>    icmp6_dst_alloc, ip6_output) will see blackhole_netdev / init_net.
>
> 2) In patch 3/5, ops_undo_list() is also called from setup_net()
>    (out_undo) when an ->init() callback fails. At that point
>    check_net(net) is still true (__ns_ref == 1), so ->pre_exit_batch()
>    will skip net (or not be called at all if setup_net() failed before
>    fib_net_ops), yet ops_undo_list() sets NET_PRE_EXIT_DONE because
>    ops_list == &pernet_list, causing default_device_exit_batch() to
>    skip rt_flush_dev(dev).
>
> 3) In rt_flush_dev() / rt6_uncached_list_flush_dev(), the lockless
>    if (list_empty(&ul->head)) check can race with a concurrent
>    rt_add_uncached_list() that observed check_net(net) == true just
>    before the last put_net() but has not yet executed list_add_tail().
>    Since NET_PRE_EXIT_DONE disables all subsequent rt_flush_dev(dev)
>    calls (including the retry in netdev_wait_allrefs_any()), that dev
>    reference would never be released.

Ah, 1) & 3) are problematic exactly.


>
> Instead of ->pre_exit_batch(), could we batch this at the
> unregister_netdevice_many_notify() stage (after the NETDEV_UNREGISTER
> loop, when all devices in the batch are already down, unlisted, and
> marked NETREG_UNREGISTERING)? That would avoid ->pre_exit_batch(),
> net->undo_state, and touching rt_add_uncached_list(), while also
> speeding up any batched device unregistration.

I think it's doable, we can skip each dev's call by

if (dev && dev->reg_state == NETREG_UNREGISTERING)
  return;

and do batching later.  I think it can be done in netdev_run_todo()
after unlocking  __rtnl_lock() and if (!list_empty()), then we could
check like

if (!dev && list_empty(&rt->dst.dev.todo_list))
  /* purge */

^ permalink raw reply	[flat|nested] 12+ messages in thread

* Re: [PATCH v1 net-next 1/5] net: Remove net->is_dying.
  2026-09-27 20:23 ` [PATCH v1 net-next 1/5] net: Remove net->is_dying Kuniyuki Iwashima
@ 2026-09-29  6:26   ` netdev-bot+sashiko
  0 siblings, 0 replies; 12+ messages in thread
From: netdev-bot+sashiko @ 2026-09-29  6:26 UTC (permalink / raw)
  To: kuniyu
  Cc: dsahern, idosch, davem, edumazet, kuba, pabeni, horms, carges,
	kuni1840, netdev

Thank you for your contribution! Sashiko AI review found more severe
issue(s) in other patches of this series. This mail is only an FYI, in case
you want to address the nit picks below when you respin.

Low priority / nit picks for your consideration:

- [Low] The commit message for "net: Remove net->is_dying." says
  net->is_dying "is an alias of !check_net(net)".

Pre-existing issues:
- [Medium] This problem wasn't introduced by this patch, but a valid NSID
  of INT_MAX makes the signed cursor in unhash_nsid() overflow, which…
- [Low] This problem wasn't introduced by this patch, but the block
  comment in cleanup_net() right after the loop this patch edits says…

--

net: Remove net->is_dying.

This patch removes the is_dying flag from struct net. cleanup_net() used
to set that flag on the nets in its current kill batch. In unhash_nsid()
the flag check is replaced with check_net(peer), on the grounds that
is_dying is an alias of !check_net(net).

> diff --git a/net/core/net_namespace.c b/net/core/net_namespace.c
> index da5f881fbd3b9..0e13de0cd36b6 100644
> --- a/net/core/net_namespace.c
> +++ b/net/core/net_namespace.c
> @@ -644,7 +644,7 @@ static void unhash_nsid(struct net *last)
>  			int curr_id = id;
>
>  			id++;

[Severity: Medium]
This is a pre-existing issue, not one introduced by this patch. What
happens here when a netns_ids IDR has an entry at INT_MAX?

rtnl_net_newid() accepts any s32 NETNSA_NSID. alloc_netid() then does:

	min = reqid;
	max = reqid + 1;

For reqid == INT_MAX, max wraps to INT_MIN. idr_alloc() treats end <= 0
as unbounded, so ID INT_MAX gets allocated.

When unhash_nsid() later reaches that entry, id++ wraps to INT_MIN,
whether or not the entry is removed. On the next iteration idr_get_next()
turns the cursor into an unsigned long above INT_MAX. idr_get_next_ul()
finds no slot and leaves the value unchanged, so this fires:

lib/idr.c:idr_get_next() {
	unsigned long id = *nextid;
	void *entry = idr_get_next_ul(idr, &id);

	if (WARN_ON_ONCE(id > INT_MAX))
		return NULL;
	...
}

An unprivileged user can reach this with unshare(CLONE_NEWUSER |
CLONE_NEWNET) followed by RTM_NEWNSID with NETNSA_NSID set to INT_MAX.
The next netns teardown anywhere on the system then hits the WARN. With
panic_on_warn=1 that becomes a panic. No entries get skipped, because
INT_MAX is the last possible ID.

Would it make sense to use an unsigned long cursor with idr_get_next_ul(),
or to break out of the loop after processing INT_MAX?

> -			if (!peer->is_dying)
> +			if (check_net(peer))
>  				continue;

[Severity: Low]
The commit message says:

    It is an alias of !check_net(net), and having two ways to represent
    the same state is confusing.

Is that equivalence exact? is_dying was only set on the nets that
cleanup_net() took from cleanup_list with llist_del_all(). check_net()
is false for every net whose refcount has reached zero.

That set also includes nets queued after this cleanup_net() pass took its
snapshot:

put_net()
  ns_ref_put() reaches 0
  __put_net()
    llist_add(&net->cleanup_list, &cleanup_list)

It also includes nets whose refcount has hit zero but which have not yet
reached llist_add(). Those nets are still on net_namespace_list and still
sit in other nets' netns_ids IDRs.

Before this patch, unhash_nsid() skipped them because is_dying was false,
and their own later cleanup pass handled them. With check_net(peer), the
current pass now calls idr_remove() on them and sends RTM_DELNSID
straight away. That means userspace can see the notification one pass
earlier.

The wider condition does look safe. peernet2id_alloc() goes through
maybe_get_net(peer), and rtnl_net_newid() gets its peer from
get_net_ns_by_pid() or get_net_ns_by_fd(). So a zero-refcount peer can
never be added back to an IDR, and the later pass finds nothing left to
remove.

Could the commit message say that the new check covers more nets than the
old flag did, and explain why that is fine?

Later patches in the series ("net: Track state in ops_undo_list()." and
"ipv4: Batch rt_flush_dev() for dying netns.") describe a dying netns
queued for the next cleanup_net() batch as having !check_net() without
being in the current batch. That also contradicts the "alias" wording.

[ ... ]

> @@ -681,7 +681,6 @@ static void cleanup_net(struct work_struct *work)
>  	llist_for_each_entry(net, net_kill_list, cleanup_list) {
>  		ns_tree_remove(net);
>  		list_del_rcu(&net->list);
> -		net->is_dying = true;
>  	}
>  	/* Cache last net. After we unlock rtnl, no one new net
>  	 * added to net_namespace_list can assign nsid pointer

[Severity: Low]
This isn't a bug introduced by this patch, but the comment says "After we
unlock rtnl". This section of cleanup_net() never takes rtnl. It uses
net_rwsem instead:

	down_write(&net_rwsem);
	llist_for_each_entry(net, net_kill_list, cleanup_list) {
	...
	last = list_last_entry(&net_namespace_list, struct net, list);
	up_write(&net_rwsem);

Should the comment say net_rwsem? The guarantee it describes actually
comes from the maybe_get_net() check in peernet2id_alloc().

-- 
Sashiko AI review · https://netdev-ai.bots.linux.dev/sashiko/#/patchset/20260927202429.2452589-1-kuniyu%40google.com

^ permalink raw reply	[flat|nested] 12+ messages in thread

* Re: [PATCH v1 net-next 2/5] net: Add ->pre_exit_batch() to struct pernet_operations.
  2026-09-27 20:23 ` [PATCH v1 net-next 2/5] net: Add ->pre_exit_batch() to struct pernet_operations Kuniyuki Iwashima
@ 2026-09-29  6:26   ` netdev-bot+sashiko
  0 siblings, 0 replies; 12+ messages in thread
From: netdev-bot+sashiko @ 2026-09-29  6:26 UTC (permalink / raw)
  To: kuniyu
  Cc: dsahern, idosch, davem, edumazet, kuba, pabeni, horms, carges,
	kuni1840, netdev

Thank you for your contribution! Sashiko AI review found more severe
issue(s) in other patches of this series. This mail is only an FYI, in case
you want to address the nit picks below when you respin.

Low priority / nit picks for your consideration:

- [Low] The new ->pre_exit_batch() hook is undocumented, and the commit
  message describes its scope inaccurately.

--

net: Add ->pre_exit_batch() to struct pernet_operations.

This commit adds a pre_exit_batch() callback to struct pernet_operations.
ops_pre_exit_list() calls it after the per-net pre_exit() calls. The goal
is to let subsystems batch device-related cleanup, such as flushing
uncached routes, across dying namespaces without holding RTNL and before
ops_exit_rtnl_list().

> diff --git a/include/net/net_namespace.h b/include/net/net_namespace.h
> index b4247e767d135..58b2601bb8691 100644
> --- a/include/net/net_namespace.h
> +++ b/include/net/net_namespace.h
> @@ -488,6 +488,7 @@ struct pernet_operations {
>  	 */
>  	int (*init)(struct net *net);
>  	void (*pre_exit)(struct net *net);
> +	void (*pre_exit_batch)(struct list_head *net_exit_list);
>  	void (*exit)(struct net *net);
>  	void (*exit_batch)(struct list_head *net_exit_list);
>  	/* Following method is called with RTNL held. */
> diff --git a/net/core/net_namespace.c b/net/core/net_namespace.c
> index 0e13de0cd36b6..476fbf913bad8 100644
> --- a/net/core/net_namespace.c
> +++ b/net/core/net_namespace.c
> @@ -160,6 +160,9 @@ static void ops_pre_exit_list(const struct pernet_operations *ops,
>  		list_for_each_entry(net, net_exit_list, exit_list)
>  			ops->pre_exit(net);
>  	}
> +
> +	if (ops->pre_exit_batch)
> +		ops->pre_exit_batch(net_exit_list);

[Severity: Low]
The commit message says the hook is there "To batch such cleanup across
dying netns". Is net_exit_list here always a list of dying namespaces?

ops_pre_exit_list() is also reached through ops_undo_single() from
__unregister_pernet_operations(). That function builds the list from every
live netns, including init_net:

net/core/net_namespace.c:__unregister_pernet_operations() {
    ...
	for_each_net(net)
		list_add_tail(&net->exit_list, &net_exit_list);

	list_del(&ops->list);
	ops_undo_single(ops, &net_exit_list);
}

It is also reached from the error path of __register_pernet_operations().
It is reached from the setup_net() failure path through ops_undo_list() as
well, and there check_net() is still true for the net being torn down.

The two users later in this series seem to cope with this by ignoring the
list and filtering on !check_net(). Could the commit message be updated to
describe the scope the hook actually has?

Separately, the comment block above struct pernet_operations describes
how pre_exit(), exit() and exit_batch() are ordered and where
synchronize_rcu() falls between them. It is not updated for the new
member.

Would it make sense to document the following there?

  - pre_exit_batch() runs after all pre_exit() calls for the same ops.
  - It runs before the synchronize_rcu() that comes ahead of exit_rtnl(),
    exit() and exit_batch().
  - It runs while the netns devices may still be registered and up.
  - net_exit_list may contain live namespaces.

There is a related and more speculative case later in the series. The
change "net: Track state in ops_undo_list()" marks nets as
NET_PRE_EXIT_DONE whenever ops_list == &pernet_list. That includes the
setup_net() out_undo path, where check_net(net) is true.

In that case, rt_flush_dev(NULL) and rt6_uncached_list_flush_dev(NULL)
would skip that net's routes because check_net() is true. The per-device
flush on NETDEV_UNREGISTER would then return early because
net_pre_exit_done() is true.

If an uncached route pointing at a device of the failing netns existed at
that point, would its device reference ever be released? It is not clear
that such a route can exist in a netns that is not yet visible, so this
may not be reachable.

>  }
>  
>  static void ops_exit_rtnl_list(const struct list_head *ops_list,

-- 
Sashiko AI review · https://netdev-ai.bots.linux.dev/sashiko/#/patchset/20260927202429.2452589-1-kuniyu%40google.com

^ permalink raw reply	[flat|nested] 12+ messages in thread

* Re: [PATCH v1 net-next 4/5] ipv4: Batch rt_flush_dev() for dying netns.
  2026-09-27 20:23 ` [PATCH v1 net-next 4/5] ipv4: Batch rt_flush_dev() for dying netns Kuniyuki Iwashima
  2026-09-27 22:51   ` Eric Dumazet
@ 2026-09-29  6:26   ` netdev-bot+sashiko
  1 sibling, 0 replies; 12+ messages in thread
From: netdev-bot+sashiko @ 2026-09-29  6:26 UTC (permalink / raw)
  To: kuniyu
  Cc: dsahern, idosch, davem, edumazet, kuba, pabeni, horms, carges,
	kuni1840, netdev

Thank you for your contribution! Sashiko AI review found 2 potential
issue(s) to consider.

Critical: 0 · High: 0 · Medium: 2 · Low: 0

If you believe none of the issues are legitimate, please reply to each
issue of Medium severity or higher and explain why the AI is wrong.
If any of them are legitimate and you plan to respin, please let patchwork
know by including "pw-bot: cr" as a separate line at the end of your reply
(one such reply per series is enough).

- [Medium] The early return in rt_flush_dev() and the batch purge use two
  different tests for "dying netns", and they disagree when setup_net()…
- [Medium] rt_flush_dev(NULL) now runs from ->pre_exit_batch(), before
  any RCU grace period (ops_undo_list() calls ops_pre_exit_list() before…

--

ipv4: Batch rt_flush_dev() for dying netns.

This patch calls rt_flush_dev(NULL) once from fib_net_ops->pre_exit_batch().
That call purges uncached routes whose device belongs to a netns with a zero
refcount, so the flush no longer runs once per device. rt_add_uncached_list()
now swaps in blackhole_netdev for routes of a dying netns, and
rt_flush_dev(dev) returns early once net_pre_exit_done() is true.

> diff --git a/net/ipv4/route.c b/net/ipv4/route.c
> index d7da2f1acbb59..cbe328b3f2545 100644
> --- a/net/ipv4/route.c
> +++ b/net/ipv4/route.c
> @@ -1554,14 +1554,29 @@ struct uncached_list {

[ ... ]

>  void rt_add_uncached_list(struct rtable *rt)
>  {
>  	struct uncached_list *ul = raw_cpu_ptr(&rt_uncached_list);
>  
> +	/* Set once and never cleared: non-NULL marks an uncached route. */
>  	rt->dst.rt_uncached_list = ul;
>  
>  	spin_lock_bh(&ul->lock);
> -	list_add_tail(&rt->dst.rt_uncached, &ul->head);
> +
> +	if (check_net(dst_dev_net_rcu(&rt->dst)))
> +		list_add_tail(&rt->dst.rt_uncached, &ul->head);
> +	else
> +		rt_replace_uncached_list(rt);
> +
>  	spin_unlock_bh(&ul->lock);
>  }

[Severity: Medium]
Does this check_net() test keep new routes off the list after
rt_flush_dev(NULL) has already run?

The test runs under ul->lock. However, rt_flush_dev() still skips each
per-CPU list with a check that does not take the lock:

net/ipv4/route.c:rt_flush_dev() {
    ...
	for_each_possible_cpu(cpu) {
		struct uncached_list *ul = &per_cpu(rt_uncached_list, cpu);

		if (list_empty(&ul->head))
			continue;

		spin_lock_bh(&ul->lock);
    ...
}

rt_flush_dev(NULL) now runs from ->pre_exit_batch(). In ops_undo_list(),
that is before synchronize_rcu_expedited(). Could this interleaving happen?

CPU1 (softirq RX on a device in netns N)
rt_add_uncached_list()
    spin_lock_bh(&ul->lock);
    check_net(N) returns 1
    <stalls before list_add_tail(), e.g. preempted on PREEMPT_RT
     or a preempted vCPU>

CPU0 (cleanup_net, after the last ns reference is dropped)
ops_undo_list()
    fib_net_pre_exit_batch()
        rt_flush_dev(NULL)
            list_empty(&ul->head) is true for CPU1's list, so continue

CPU1
    list_add_tail(&rt->dst.rt_uncached, &ul->head);
    spin_unlock_bh(&ul->lock);

CPU0
    WRITE_ONCE(net->undo_state, NET_PRE_EXIT_DONE);

After this, every rt_flush_dev(dev) call for a device in N returns at the
new net_pre_exit_done() check. That includes the call from
fib_netdev_event() and the NETDEV_UNREGISTER rebroadcast in
netdev_wait_allrefs_any().

Would the route then keep its dev_tracker reference on a device in N
forever, so that unregister_netdevice waits for that device indefinitely?

The race also needs CPU1's list to be otherwise empty. The window is narrow
on non-RT kernels.

On weakly ordered architectures, the READ_ONCE() in list_empty() is not
ordered against a list_add_tail() that has already completed on another
CPU.

Before this patch, the flush that mattered ran from NETDEV_UNREGISTER after
synchronize_net(), and the rebroadcast ran it again, so a late addition was
still caught. The commit message says:

  Since ->pre_exit_batch() is called before synchronize_rcu(),
  we must prevent adding a new route for dying netns, so now
  rt_add_uncached_list() checks !check_net() and swaps the device
  with blackhole_netdev.

Would taking ul->lock every time in the dev == NULL case close this gap?
Then either the flush sees the route, or rt_add_uncached_list() sees
check_net() return false. This code is unchanged at the end of the series.

> @@ -1587,6 +1602,9 @@ void rt_flush_dev(struct net_device *dev)
>  	struct rtable *rt, *safe;
>  	int cpu;
>  
> +	if (dev && net_pre_exit_done(dev_net(dev)))
> +		return;
> +

[Severity: Medium]
Can this early return skip routes that the rt_flush_dev(NULL) purge never
removed?

The purge filter only picks routes whose netns has check_net() false:

			if (rt->dst.dev != dev &&
			    (dev || check_net(dev_net(rt->dst.dev))))
				continue;

The gate in rt_add_uncached_list() uses the same test. But
net_pre_exit_done() reads undo_state, and ops_undo_list() sets undo_state
for any undo on pernet_list:

net/core/net_namespace.c:ops_undo_list() {
	bool dying = ops_list == &pernet_list;
    ...
	if (dying) {
		list_for_each_entry(net, net_exit_list, exit_list)
			WRITE_ONCE(net->undo_state, NET_PRE_EXIT_DONE);
	}
    ...
}

That includes the setup_net() failure path:

net/core/net_namespace.c:setup_net() {
    ...
	list_add(&net->exit_list, &net_exit_list);
	ops_undo_list(&pernet_list, ops, &net_exit_list, false);
    ...
}

On that path the netns still holds the initial ns reference that
ns_common_init() takes in copy_net_ns(), so check_net() returns true. Is
this sequence possible?

copy_net_ns()
  setup_net()
    <a pernet init after fib_net_ops fails>
    ops_undo_list(&pernet_list, ...)
      fib_net_pre_exit_batch()
        rt_flush_dev(NULL)      <- skips the netns, check_net() is true
      undo_state = NET_PRE_EXIT_DONE
      ops_exit_rtnl_list()
        <unregister lo, fallback tunnels>
          fib_netdev_event(NETDEV_UNREGISTER)
            rt_flush_dev(dev)   <- returns early

Until then, rt_add_uncached_list() also keeps adding this netns's routes to
the list, since check_net() is still true.

If an IPv4 uncached route (an rt_uncached_list entry, an xfrm4 bundle, or a
route replaced in rt_cache_route()) pins one of these devices, does
anything move it to blackhole_netdev?

The NETDEV_UNREGISTER rebroadcast in netdev_wait_allrefs_any() used to
rescue such routes through rt_flush_dev(dev), but it now hits the same
early return. The rollback could then wait forever for the device inside
setup_net() with pernet_ops_rwsem held.

It is not clear that any in-tree path can create an uncached route on a
device of a netns that has not finished setup_net(), because those devices
are down and the netns is not visible yet. Nothing seems to prevent it
either.

The commit message says:

  When rt_flush_dev() is called with NULL from ->pre_exit_batch(),
  it purges every entry in dying netns

and:

  When rt_flush_dev() is later called again from fib_netdev_event()
  via NETDEV_UNREGISTER, it just returns.

Both hold for cleanup_net(), but not for every path that sets
NET_PRE_EXIT_DONE.

Would it make sense to set NET_PRE_EXIT_DONE only from cleanup_net()? The
other option would be for the NULL purge to pick nets by undo_state or by
membership in net_exit_list instead of check_net(). The matching IPv6
change later in the series appears to keep the same mismatch.

[ ... ]

-- 
Sashiko AI review · https://netdev-ai.bots.linux.dev/sashiko/#/patchset/20260927202429.2452589-1-kuniyu%40google.com

^ permalink raw reply	[flat|nested] 12+ messages in thread

* Re: [PATCH v1 net-next 5/5] ipv6: Batch rt6_uncached_list_flush_dev() for dying netns.
  2026-09-27 20:23 ` [PATCH v1 net-next 5/5] ipv6: Batch rt6_uncached_list_flush_dev() " Kuniyuki Iwashima
@ 2026-09-29  6:26   ` netdev-bot+sashiko
  0 siblings, 0 replies; 12+ messages in thread
From: netdev-bot+sashiko @ 2026-09-29  6:26 UTC (permalink / raw)
  To: kuniyu
  Cc: dsahern, idosch, davem, edumazet, kuba, pabeni, horms, carges,
	kuni1840, netdev

Thank you for your contribution! Sashiko AI review found 2 potential
issue(s) to consider.

Critical: 0 · High: 0 · Medium: 2 · Low: 0

If you believe none of the issues are legitimate, please reply to each
issue of Medium severity or higher and explain why the AI is wrong.
If any of them are legitimate and you plan to respin, please let patchwork
know by including "pw-bot: cr" as a separate line at the end of your reply
(one such reply per series is enough).

- [Medium] The new ->pre_exit_batch() callback can crash the IPv6 init
  error path.
- [Medium] Race between the batch flush and rt6_uncached_list_add().

--

ipv6: Batch rt6_uncached_list_flush_dev() for dying netns.

This commit flushes IPv6 uncached routes for dying netns from a new
->pre_exit_batch() callback, which runs about once per cleanup_net().
Once NET_PRE_EXIT_DONE is set, the per-device rt6_uncached_list_flush_dev()
now returns early. rt6_uncached_list_add() now checks check_net() under
ul->lock, and a route for a dying netns is moved straight to
blackhole_netdev instead of being linked.

Closes: https://lore.kernel.org/netdev/20260917-hash-bucket-route-lists-v3-0-30493a37b6eb@cloudflare.com/
Link: https://lore.kernel.org/netdev/aq2B8PSfjn-xau4V@20HS2G4/

> diff --git a/net/ipv6/route.c b/net/ipv6/route.c
> index 475ced827ec57..7747e4f20fee6 100644
> --- a/net/ipv6/route.c
> +++ b/net/ipv6/route.c

[ ... ]

> @@ -143,7 +159,12 @@ void rt6_uncached_list_add(struct rt6_info *rt)
>  	rt->dst.rt_uncached_list = ul;
>  
>  	spin_lock_bh(&ul->lock);
> -	list_add_tail(&rt->dst.rt_uncached, &ul->head);
> +
> +	if (check_net(dst_dev_net_rcu(&rt->dst)))
> +		list_add_tail(&rt->dst.rt_uncached, &ul->head);
> +	else
> +		rt6_uncached_list_replace(rt);
> +
>  	spin_unlock_bh(&ul->lock);
>  }
>  
> @@ -162,6 +183,9 @@ static void rt6_uncached_list_flush_dev(struct net_device *dev)
>  {
>  	int cpu;
>  
> +	if (dev && net_pre_exit_done(dev_net(dev)))
> +		return;
> +
>  	for_each_possible_cpu(cpu) {
>  		struct uncached_list *ul = per_cpu_ptr(&rt6_uncached_list, cpu);
>  		struct rt6_info *rt, *safe;

[Severity: Medium]
Because of this early return, the batch flush from ->pre_exit_batch() is
the only flush for devices in a dying netns. Can it miss a route that is
being added at the same time?

The dev == NULL path still skips a CPU's list with a lockless check,
before it takes ul->lock:

		if (list_empty(&ul->head))
			continue;

Here is the sequence I have in mind:

CPU A, rt6_uncached_list_add() (e.g. from ip6_pol_route() or
icmp6_dst_alloc()):
  spin_lock_bh(&ul->lock);        /* its own per-cpu list, empty */
  check_net() returns true
  <delayed by a hardirq, vCPU preemption or RT preemption>

CPU B, cleanup_net()->ops_undo_list():
  ops_pre_exit_list()
    ip6_route_net_pre_exit_batch()
      rt6_uncached_list_flush_dev(NULL)
        list_empty() on CPU A's list is true, so it is skipped

CPU A:
  list_add_tail(&rt->dst.rt_uncached, &ul->head);
  spin_unlock_bh(&ul->lock);

CPU B:
  synchronize_rcu();
  WRITE_ONCE(net->undo_state, NET_PRE_EXIT_DONE);
  ...
  NETDEV_UNREGISTER: addrconf_ifdown()->rt6_disable_ip()->
    rt6_uncached_list_flush_dev(dev) returns early

->pre_exit_batch() runs before the synchronize_rcu() in ops_undo_list(),
so CPU A's RCU/BH section is not waited for before the scan.

Would the late route then stay on the uncached list? It would still hold
its dst.dev reference (dev_tracker) and its rt6i_idev reference, and it
would never be moved to blackhole_netdev.

Before this patch, the per-device flush always ran at NETDEV_UNREGISTER,
after the synchronize_net() in unregister_netdevice_many().

If something holds that dst for a long time, could netdev_wait_allrefs()
hang and stall netns cleanup?

One option is to drop the list_empty() shortcut when dev is NULL. Another
is to run the batch flush after the grace period and before
NET_PRE_EXIT_DONE is published. Would either of those work?

[ ... ]

> @@ -6801,6 +6819,11 @@ static int __net_init ip6_route_net_init(struct net *net)
>  	goto out;
>  }
>  
> +static void __net_exit ip6_route_net_pre_exit_batch(struct list_head *net_exit_list)
> +{
> +	rt6_uncached_list_flush_dev(NULL);
> +}
> +

[Severity: Medium]
Could this callback crash on the ip6_route_init() error path?

ip6_route_init() registers ip6_route_net_ops before the per-cpu uncached
lists are initialized. They are only set up at the very end of the
function:

	ret = register_pernet_subsys(&ip6_route_net_ops);
	if (ret)
		goto out_register_inetpeer;
	...
	for_each_possible_cpu(cpu) {
		struct uncached_list *ul = per_cpu_ptr(&rt6_uncached_list, cpu);

		INIT_LIST_HEAD(&ul->head);
		spin_lock_init(&ul->lock);
	}

If a later step fails, the unwind calls
unregister_pernet_subsys(&ip6_route_net_ops). Those steps are fib6_init(),
xfrm6_init(), fib6_rules_init(), registering ip6_route_net_late_ops,
rtnl_register_many(), register_netdevice_notifier() and
bpf_iter_register().

With CONFIG_NET_NS=y, the same undo also runs inside
register_pernet_subsys() if ip6_route_net_init() fails for a netns. That
path goes through __register_pernet_operations()->ops_undo_single().

Both paths reach ops_pre_exit_list(). It calls ->pre_exit_batch() even
when net_exit_list is empty:

	if (ops->pre_exit_batch)
		ops->pre_exit_batch(net_exit_list);

which then calls:

ip6_route_net_pre_exit_batch()
  rt6_uncached_list_flush_dev(NULL)

At this point the static per-cpu heads are still zeroed. With head.next
== NULL, list_empty(&ul->head) returns false. spin_lock_bh() then takes a
lock that was never initialized. list_for_each_entry_safe() also derives
pos from the NULL next pointer and dereferences it.

Before this patch, ip6_route_net_ops had no ->pre_exit_batch(), so this
rollback did not touch the lists.

Would it make sense to initialize the per-cpu uncached lists before
registering ip6_route_net_ops?

-- 
Sashiko AI review · https://netdev-ai.bots.linux.dev/sashiko/#/patchset/20260927202429.2452589-1-kuniyu%40google.com

^ permalink raw reply	[flat|nested] 12+ messages in thread

end of thread, other threads:[~2026-09-29  6:26 UTC | newest]

Thread overview: 12+ messages (download: mbox.gz follow: Atom feed
-- links below jump to the message on this page --
2026-09-27 20:23 [PATCH v1 net-next 0/5] ip: Batch flushing uncached routes for dying netns Kuniyuki Iwashima
2026-09-27 20:23 ` [PATCH v1 net-next 1/5] net: Remove net->is_dying Kuniyuki Iwashima
2026-09-29  6:26   ` netdev-bot+sashiko
2026-09-27 20:23 ` [PATCH v1 net-next 2/5] net: Add ->pre_exit_batch() to struct pernet_operations Kuniyuki Iwashima
2026-09-29  6:26   ` netdev-bot+sashiko
2026-09-27 20:23 ` [PATCH v1 net-next 3/5] net: Track state in ops_undo_list() Kuniyuki Iwashima
2026-09-27 20:23 ` [PATCH v1 net-next 4/5] ipv4: Batch rt_flush_dev() for dying netns Kuniyuki Iwashima
2026-09-27 22:51   ` Eric Dumazet
2026-09-28 16:33     ` Kuniyuki Iwashima
2026-09-29  6:26   ` netdev-bot+sashiko
2026-09-27 20:23 ` [PATCH v1 net-next 5/5] ipv6: Batch rt6_uncached_list_flush_dev() " Kuniyuki Iwashima
2026-09-29  6:26   ` netdev-bot+sashiko

This is an external index of several public inboxes,
see mirroring instructions on how to clone and mirror
all data and code used by this external index.