* [PATCH net v2] ipv4: fix use-after-free in fib_nhc_update_mtu()
@ 2026-08-06 13:33 Chengfeng Ye
2026-08-06 14:42 ` Ido Schimmel
0 siblings, 1 reply; 2+ messages in thread
From: Chengfeng Ye @ 2026-08-06 13:33 UTC (permalink / raw)
To: David Ahern, Ido Schimmel, David S. Miller, Eric Dumazet,
Jakub Kicinski, Paolo Abeni, Simon Horman, Stefano Brivio,
Sabrina Dubroca
Cc: netdev, linux-kernel, Chengfeng Ye, stable
fib_nhc_update_mtu() walks the nexthop exception table under RTNL, but
RTNL does not serialize this walk with PMTU exception updates. The walk
uses rcu_dereference_protected() with a constant true condition without
holding fnhe_lock.
The following interleaving can therefore occur:
CPU 0 CPU 1
fib_nhc_update_mtu() update_or_create_fnhe()
load fnhe spin_lock_bh(&fnhe_lock)
fnhe_remove_oldest()
unlink fnhe
kfree_rcu(fnhe, rcu)
<quiescent state>
access fnhe after grace period
KASAN reported:
BUG: KASAN: slab-use-after-free in fib_nhc_update_mtu+0x3df/0x410
Read of size 8 at addr ffff888107d49000 by task poc/90
Call Trace:
fib_nhc_update_mtu+0x3df/0x410
fib_sync_mtu+0x7a/0xd0
fib_netdev_event+0x229/0x3f0
netif_set_mtu_ext+0x33a/0x570
dev_set_mtu+0x88/0x120
Allocated by task 89:
update_or_create_fnhe+0xa80/0x1110
__ip_rt_update_pmtu+0x8f2/0xcd0
ipv4_sk_update_pmtu+0x49e/0x690
udp_err+0xd92/0x1080
Freed by task 0:
__kasan_slab_free+0x43/0x70
kvfree_rcu_cb+0x12f/0x420
rcu_core+0x509/0x18e0
The same walk updates fnhe_pmtu and fnhe_mtu_locked. These fields form a
pair and other writers serialize them with fnhe_lock. RCU alone would
prevent reclamation, but would still allow concurrent writers to leave a
mixed pair.
Expose fnhe_lock to fib_semantics.c and hold it across the exception-table
walk. This prevents entries from being unlinked while they are visited and
serializes the paired PMTU state updates with all other writers.
Fixes: af7d6cce5369 ("net: ipv4: update fnhe_pmtu when first hop's MTU changes")
Cc: stable@vger.kernel.org
Signed-off-by: Chengfeng Ye <nicoyip.dev@gmail.com>
---
Changes in v2:
- Protect the walk with fnhe_lock rather than RCU alone, covering both
exception lifetime and the paired PMTU state updates.
- Keep fib_nhc_update_mtu() in fib_semantics.c and expose fnhe_lock for a
smaller diff.
Link: https://lore.kernel.org/netdev/20260731162938.3388534-1-nicoyip.dev@gmail.com/ [v1]
include/net/ip_fib.h | 4 ++++
net/ipv4/fib_semantics.c | 5 ++++-
net/ipv4/route.c | 2 +-
3 files changed, 9 insertions(+), 2 deletions(-)
diff --git a/include/net/ip_fib.h b/include/net/ip_fib.h
index c63a3c4967ae..634a669ba46b 100644
--- a/include/net/ip_fib.h
+++ b/include/net/ip_fib.h
@@ -12,6 +12,7 @@
#ifndef _NET_IP_FIB_H
#define _NET_IP_FIB_H
+#include <linux/spinlock.h>
#include <net/flow.h>
#include <linux/seq_file.h>
#include <linux/rcupdate.h>
@@ -495,6 +496,9 @@ int fib_sync_up(struct net_device *dev, unsigned char nh_flags);
void fib_sync_mtu(struct net_device *dev, u32 orig_mtu);
void fib_nhc_update_mtu(struct fib_nh_common *nhc, u32 new, u32 orig);
+/* Protects nexthop exception table updates. */
+extern spinlock_t fnhe_lock;
+
/* Fields used for sysctl_fib_multipath_hash_fields.
* Common to IPv4 and IPv6.
*
diff --git a/net/ipv4/fib_semantics.c b/net/ipv4/fib_semantics.c
index 4f3c0740dde9..f091a8ac1974 100644
--- a/net/ipv4/fib_semantics.c
+++ b/net/ipv4/fib_semantics.c
@@ -1879,9 +1879,10 @@ void fib_nhc_update_mtu(struct fib_nh_common *nhc, u32 new, u32 orig)
struct fnhe_hash_bucket *bucket;
int i;
+ spin_lock_bh(&fnhe_lock);
bucket = rcu_dereference_protected(nhc->nhc_exceptions, 1);
if (!bucket)
- return;
+ goto out;
for (i = 0; i < FNHE_HASH_SIZE; i++) {
struct fib_nh_exception *fnhe;
@@ -1900,6 +1901,8 @@ void fib_nhc_update_mtu(struct fib_nh_common *nhc, u32 new, u32 orig)
}
}
}
+out:
+ spin_unlock_bh(&fnhe_lock);
}
void fib_sync_mtu(struct net_device *dev, u32 orig_mtu)
diff --git a/net/ipv4/route.c b/net/ipv4/route.c
index 152d8cb28f65..fc9a51ebd922 100644
--- a/net/ipv4/route.c
+++ b/net/ipv4/route.c
@@ -571,7 +571,7 @@ static void ip_rt_build_flow_key(struct flowi4 *fl4, const struct sock *sk,
build_sk_flow_key(fl4, sk);
}
-static DEFINE_SPINLOCK(fnhe_lock);
+DEFINE_SPINLOCK(fnhe_lock);
static void fnhe_flush_routes(struct fib_nh_exception *fnhe)
{
--
2.43.0
^ permalink raw reply related [flat|nested] 2+ messages in thread
* Re: [PATCH net v2] ipv4: fix use-after-free in fib_nhc_update_mtu()
2026-08-06 13:33 [PATCH net v2] ipv4: fix use-after-free in fib_nhc_update_mtu() Chengfeng Ye
@ 2026-08-06 14:42 ` Ido Schimmel
0 siblings, 0 replies; 2+ messages in thread
From: Ido Schimmel @ 2026-08-06 14:42 UTC (permalink / raw)
To: Chengfeng Ye
Cc: David Ahern, David S. Miller, Eric Dumazet, Jakub Kicinski,
Paolo Abeni, Simon Horman, Stefano Brivio, Sabrina Dubroca,
netdev, linux-kernel, stable
On Thu, Aug 06, 2026 at 09:33:09PM +0800, Chengfeng Ye wrote:
> fib_nhc_update_mtu() walks the nexthop exception table under RTNL, but
> RTNL does not serialize this walk with PMTU exception updates. The walk
> uses rcu_dereference_protected() with a constant true condition without
> holding fnhe_lock.
>
> The following interleaving can therefore occur:
>
> CPU 0 CPU 1
> fib_nhc_update_mtu() update_or_create_fnhe()
> load fnhe spin_lock_bh(&fnhe_lock)
> fnhe_remove_oldest()
> unlink fnhe
> kfree_rcu(fnhe, rcu)
> <quiescent state>
> access fnhe after grace period
[...]
> The same walk updates fnhe_pmtu and fnhe_mtu_locked. These fields form a
> pair and other writers serialize them with fnhe_lock. RCU alone would
> prevent reclamation, but would still allow concurrent writers to leave a
> mixed pair.
>
> Expose fnhe_lock to fib_semantics.c and hold it across the exception-table
> walk. This prevents entries from being unlinked while they are visited and
> serializes the paired PMTU state updates with all other writers.
Looks correct, but can't we use RCU for the traversal and only acquire
the global lock when updating an entry? Otherwise, whenever a device MTU
changes, we acquire the global lock (and disable softIRQs) across a scan
of 2048 buckets and we do that for each nexthop using this device.
How about something like [1] (vibe coded, compile-tested only)?
It also avoids exporting the global lock.
Please wait at least 24h before posting v3.
Thanks
[1]
diff --git a/include/net/route.h b/include/net/route.h
index f90106f383c5..45290177a33c 100644
--- a/include/net/route.h
+++ b/include/net/route.h
@@ -276,6 +276,8 @@ int fib_dump_info_fnhe(struct sk_buff *skb, struct netlink_callback *cb,
u32 table_id, struct fib_info *fi,
int *fa_index, int fa_start, unsigned int flags);
+void fnhe_update_pmtu(struct fib_nh_exception *fnhe, u32 new, u32 orig);
+
static inline void ip_rt_put(struct rtable *rt)
{
/* dst_release() accepts a NULL parameter.
diff --git a/net/ipv4/fib_semantics.c b/net/ipv4/fib_semantics.c
index 4f3c0740dde9..61286948c5d9 100644
--- a/net/ipv4/fib_semantics.c
+++ b/net/ipv4/fib_semantics.c
@@ -1864,42 +1864,30 @@ static int call_fib_nh_notifiers(struct fib_nh *nh,
return NOTIFY_DONE;
}
-/* Update the PMTU of exceptions when:
- * - the new MTU of the first hop becomes smaller than the PMTU
- * - the old MTU was the same as the PMTU, and it limited discovery of
- * larger MTUs on the path. With that limit raised, we can now
- * discover larger MTUs
- * A special case is locked exceptions, for which the PMTU is smaller
- * than the minimal accepted PMTU:
- * - if the new MTU is greater than the PMTU, don't make any change
- * - otherwise, unlock and set PMTU
+/* Walk the exceptions of a nexthop after its first hop MTU changed. The
+ * chain is only RCU protected here, fnhe_update_pmtu() takes fnhe_lock for
+ * the update of each entry.
*/
void fib_nhc_update_mtu(struct fib_nh_common *nhc, u32 new, u32 orig)
{
struct fnhe_hash_bucket *bucket;
int i;
- bucket = rcu_dereference_protected(nhc->nhc_exceptions, 1);
+ rcu_read_lock();
+ bucket = rcu_dereference(nhc->nhc_exceptions);
if (!bucket)
- return;
+ goto out;
for (i = 0; i < FNHE_HASH_SIZE; i++) {
struct fib_nh_exception *fnhe;
- for (fnhe = rcu_dereference_protected(bucket[i].chain, 1);
+ for (fnhe = rcu_dereference(bucket[i].chain);
fnhe;
- fnhe = rcu_dereference_protected(fnhe->fnhe_next, 1)) {
- if (fnhe->fnhe_mtu_locked) {
- if (new <= fnhe->fnhe_pmtu) {
- fnhe->fnhe_pmtu = new;
- fnhe->fnhe_mtu_locked = false;
- }
- } else if (new < fnhe->fnhe_pmtu ||
- orig == fnhe->fnhe_pmtu) {
- fnhe->fnhe_pmtu = new;
- }
- }
+ fnhe = rcu_dereference(fnhe->fnhe_next))
+ fnhe_update_pmtu(fnhe, new, orig);
}
+out:
+ rcu_read_unlock();
}
void fib_sync_mtu(struct net_device *dev, u32 orig_mtu)
diff --git a/net/ipv4/route.c b/net/ipv4/route.c
index fd688e1f879f..604cc51dfd9b 100644
--- a/net/ipv4/route.c
+++ b/net/ipv4/route.c
@@ -741,6 +741,35 @@ static void update_or_create_fnhe(struct fib_nh_common *nhc, __be32 daddr,
spin_unlock_bh(&fnhe_lock);
}
+/* Update the PMTU of an exception when:
+ * - the new MTU of the first hop becomes smaller than the PMTU
+ * - the old MTU was the same as the PMTU, and it limited discovery of
+ * larger MTUs on the path. With that limit raised, we can now
+ * discover larger MTUs
+ * A special case is locked exceptions, for which the PMTU is smaller
+ * than the minimal accepted PMTU:
+ * - if the new MTU is greater than the PMTU, don't make any change
+ * - otherwise, unlock and set PMTU
+ *
+ * fnhe_lock keeps fnhe_pmtu and fnhe_mtu_locked consistent against
+ * update_or_create_fnhe(), which sets both under the same lock.
+ */
+void fnhe_update_pmtu(struct fib_nh_exception *fnhe, u32 new, u32 orig)
+{
+ spin_lock_bh(&fnhe_lock);
+
+ if (fnhe->fnhe_mtu_locked) {
+ if (new <= fnhe->fnhe_pmtu) {
+ fnhe->fnhe_pmtu = new;
+ fnhe->fnhe_mtu_locked = false;
+ }
+ } else if (new < fnhe->fnhe_pmtu || orig == fnhe->fnhe_pmtu) {
+ fnhe->fnhe_pmtu = new;
+ }
+
+ spin_unlock_bh(&fnhe_lock);
+}
+
static void __ip_do_redirect(struct rtable *rt, struct sk_buff *skb, struct flowi4 *fl4,
bool kill_route)
{
^ permalink raw reply related [flat|nested] 2+ messages in thread
end of thread, other threads:[~2026-08-06 14:42 UTC | newest]
Thread overview: 2+ messages (download: mbox.gz follow: Atom feed
-- links below jump to the message on this page --
2026-08-06 13:33 [PATCH net v2] ipv4: fix use-after-free in fib_nhc_update_mtu() Chengfeng Ye
2026-08-06 14:42 ` Ido Schimmel
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox