* [PATCH net 1/2] nexthop: take nh->lock for f6i_list walks in replace check and notify
@ 2026-07-22 0:29 Xiang Mei (Microsoft)
2026-07-22 0:29 ` [PATCH net 2/2] nexthop: avoid unlocked f6i_list walk in nh_rt_cache_flush Xiang Mei (Microsoft)
2026-07-23 19:19 ` [PATCH net 1/2] nexthop: take nh->lock for f6i_list walks in replace check and notify Ido Schimmel
0 siblings, 2 replies; 4+ messages in thread
From: Xiang Mei (Microsoft) @ 2026-07-22 0:29 UTC (permalink / raw)
To: David Ahern, Ido Schimmel, David S . Miller, Eric Dumazet,
Jakub Kicinski, Paolo Abeni, Simon Horman
Cc: Kuniyuki Iwashima, netdev, linux-kernel, AutonomousCodeSecurity,
tgopinath, kys, Xiang Mei (Microsoft)
fib6_check_nh_list() and __nexthop_replace_notify() walk nh->f6i_list
during an RTNL-serialized nexthop replace without holding nh->lock. IPv6
RTM_NEWROUTE/RTM_DELROUTE run without RTNL and mutate that list under
nh->lock (fib6_add_rt2node_nh(), fib6_purge_rt()), so both walks race a
concurrent route delete that unlinks and frees a fib6_info:
BUG: KASAN: slab-use-after-free in rt6_fill_node.isra.0 (net/ipv6/route.c:5799)
Read of size 4 at addr ffff888014607e64 by task exploit/143
rt6_fill_node.isra.0 (net/ipv6/route.c:5799)
fib6_rt_update (net/ipv6/route.c:6412)
__nexthop_replace_notify (net/ipv4/nexthop.c:2542)
rtm_new_nexthop (net/ipv4/nexthop.c:2554)
rtnetlink_rcv_msg (net/core/rtnetlink.c:7076)
BUG: KASAN: slab-use-after-free in fib6_check_nh_list (net/ipv4/nexthop.c:1605)
Read of size 8 at addr ffff888014a7d068 by task exploit/142
fib6_check_nh_list (net/ipv4/nexthop.c:1605)
rtm_new_nexthop (net/ipv4/nexthop.c:2575)
rtnetlink_rcv_msg (net/core/rtnetlink.c:7076)
Both walks only read the entries and take no tb6_lock, so protect them
with nh->lock; fib6_rt_update() uses gfp_any(), which returns GFP_ATOMIC
under the lock.
Fixes: 081efd18326e ("ipv6: Protect nh->f6i_list with spinlock and flag.")
Reported-by: AutonomousCodeSecurity@microsoft.com
Signed-off-by: Xiang Mei (Microsoft) <xmei5@asu.edu>
---
net/ipv4/nexthop.c | 13 +++++++++++--
1 file changed, 11 insertions(+), 2 deletions(-)
diff --git a/net/ipv4/nexthop.c b/net/ipv4/nexthop.c
index 6205bd57aa85..c1d3b7d7100d 100644
--- a/net/ipv4/nexthop.c
+++ b/net/ipv4/nexthop.c
@@ -1597,14 +1597,21 @@ static int fib6_check_nh_list(struct nexthop *old, struct nexthop *new,
struct netlink_ext_ack *extack)
{
struct fib6_info *f6i;
+ int err = 0;
if (list_empty(&old->f6i_list))
return 0;
+ spin_lock_bh(&old->lock);
list_for_each_entry(f6i, &old->f6i_list, nh_list) {
- if (check_src_addr(&f6i->fib6_src.addr, extack) < 0)
- return -EINVAL;
+ err = check_src_addr(&f6i->fib6_src.addr, extack);
+ if (err)
+ break;
}
+ spin_unlock_bh(&old->lock);
+
+ if (err)
+ return err;
return fib6_check_nexthop(new, NULL, extack);
}
@@ -2538,8 +2545,10 @@ static void __nexthop_replace_notify(struct net *net, struct nexthop *nh,
fi->nh_updated = false;
}
+ spin_lock_bh(&nh->lock);
list_for_each_entry(f6i, &nh->f6i_list, nh_list)
fib6_rt_update(net, f6i, info);
+ spin_unlock_bh(&nh->lock);
}
/* send RTM_NEWROUTE with REPLACE flag set for all FIB entries
--
2.43.0
^ permalink raw reply related [flat|nested] 4+ messages in thread
* [PATCH net 2/2] nexthop: avoid unlocked f6i_list walk in nh_rt_cache_flush
2026-07-22 0:29 [PATCH net 1/2] nexthop: take nh->lock for f6i_list walks in replace check and notify Xiang Mei (Microsoft)
@ 2026-07-22 0:29 ` Xiang Mei (Microsoft)
2026-07-23 19:20 ` Ido Schimmel
2026-07-23 19:19 ` [PATCH net 1/2] nexthop: take nh->lock for f6i_list walks in replace check and notify Ido Schimmel
1 sibling, 1 reply; 4+ messages in thread
From: Xiang Mei (Microsoft) @ 2026-07-22 0:29 UTC (permalink / raw)
To: David Ahern, Ido Schimmel, David S . Miller, Eric Dumazet,
Jakub Kicinski, Paolo Abeni, Simon Horman
Cc: Kuniyuki Iwashima, netdev, linux-kernel, AutonomousCodeSecurity,
tgopinath, kys, Xiang Mei (Microsoft)
nh_rt_cache_flush() walks nh->f6i_list during an RTNL-serialized nexthop
replace without holding nh->lock, racing the unlocked IPv6 route
add/delete that mutate the list under nh->lock and free fib6_info
entries (nh_rt_cache_flush() is inlined into rtm_new_nexthop()):
BUG: KASAN: slab-use-after-free in nh_rt_cache_flush (net/ipv4/nexthop.c:2243)
Read of size 8 at addr ffff888012953e18 by task exploit/146
nh_rt_cache_flush (net/ipv4/nexthop.c:2243)
replace_nexthop (net/ipv4/nexthop.c:2610)
rtm_new_nexthop (net/ipv4/nexthop.c:3323)
rtnetlink_rcv_msg (net/core/rtnetlink.c:7076)
Unlike the other f6i_list walks, this one bumps each route's sernum via
fib6_update_sernum_upto_root(), which needs tb6_lock; taking nh->lock
around it would invert the established tb6_lock -> nh->lock order and
deadlock. As the only purpose is to invalidate cached dsts, bump the
IPv6 sernum for the whole netns with rt_genid_bump_ipv6() instead,
mirroring the rt_cache_flush() already done for IPv4 just above.
Fixes: 081efd18326e ("ipv6: Protect nh->f6i_list with spinlock and flag.")
Reported-by: AutonomousCodeSecurity@microsoft.com
Signed-off-by: Xiang Mei (Microsoft) <xmei5@asu.edu>
---
net/ipv4/nexthop.c | 12 ++++++------
1 file changed, 6 insertions(+), 6 deletions(-)
diff --git a/net/ipv4/nexthop.c b/net/ipv4/nexthop.c
index c1d3b7d7100d..1e9ccac6bf0f 100644
--- a/net/ipv4/nexthop.c
+++ b/net/ipv4/nexthop.c
@@ -2240,18 +2240,18 @@ static void remove_one_nexthop(struct net *net, struct nexthop *nh,
static void nh_rt_cache_flush(struct net *net, struct nexthop *nh,
struct nexthop *replaced_nh)
{
- struct fib6_info *f6i;
struct nh_group *nhg;
+ bool have_f6i;
int i;
if (!list_empty(&nh->fi_list))
rt_cache_flush(net);
- list_for_each_entry(f6i, &nh->f6i_list, nh_list) {
- spin_lock_bh(&f6i->fib6_table->tb6_lock);
- fib6_update_sernum_upto_root(net, f6i);
- spin_unlock_bh(&f6i->fib6_table->tb6_lock);
- }
+ spin_lock_bh(&nh->lock);
+ have_f6i = !list_empty(&nh->f6i_list);
+ spin_unlock_bh(&nh->lock);
+ if (have_f6i)
+ rt_genid_bump_ipv6(net);
/* if an IPv6 group was replaced, we have to release all old
* dsts to make sure all refcounts are released
--
2.43.0
^ permalink raw reply related [flat|nested] 4+ messages in thread
* Re: [PATCH net 1/2] nexthop: take nh->lock for f6i_list walks in replace check and notify
2026-07-22 0:29 [PATCH net 1/2] nexthop: take nh->lock for f6i_list walks in replace check and notify Xiang Mei (Microsoft)
2026-07-22 0:29 ` [PATCH net 2/2] nexthop: avoid unlocked f6i_list walk in nh_rt_cache_flush Xiang Mei (Microsoft)
@ 2026-07-23 19:19 ` Ido Schimmel
1 sibling, 0 replies; 4+ messages in thread
From: Ido Schimmel @ 2026-07-23 19:19 UTC (permalink / raw)
To: Xiang Mei (Microsoft)
Cc: David Ahern, David S . Miller, Eric Dumazet, Jakub Kicinski,
Paolo Abeni, Simon Horman, Kuniyuki Iwashima, netdev,
linux-kernel, AutonomousCodeSecurity, tgopinath, kys
On Wed, Jul 22, 2026 at 12:29:50AM +0000, Xiang Mei (Microsoft) wrote:
> fib6_check_nh_list() and __nexthop_replace_notify() walk nh->f6i_list
> during an RTNL-serialized nexthop replace without holding nh->lock. IPv6
> RTM_NEWROUTE/RTM_DELROUTE run without RTNL and mutate that list under
> nh->lock (fib6_add_rt2node_nh(), fib6_purge_rt()), so both walks race a
> concurrent route delete that unlinks and frees a fib6_info:
>
> BUG: KASAN: slab-use-after-free in rt6_fill_node.isra.0 (net/ipv6/route.c:5799)
> Read of size 4 at addr ffff888014607e64 by task exploit/143
> rt6_fill_node.isra.0 (net/ipv6/route.c:5799)
> fib6_rt_update (net/ipv6/route.c:6412)
> __nexthop_replace_notify (net/ipv4/nexthop.c:2542)
> rtm_new_nexthop (net/ipv4/nexthop.c:2554)
> rtnetlink_rcv_msg (net/core/rtnetlink.c:7076)
>
> BUG: KASAN: slab-use-after-free in fib6_check_nh_list (net/ipv4/nexthop.c:1605)
> Read of size 8 at addr ffff888014a7d068 by task exploit/142
> fib6_check_nh_list (net/ipv4/nexthop.c:1605)
> rtm_new_nexthop (net/ipv4/nexthop.c:2575)
> rtnetlink_rcv_msg (net/core/rtnetlink.c:7076)
>
> Both walks only read the entries and take no tb6_lock, so protect them
> with nh->lock; fib6_rt_update() uses gfp_any(), which returns GFP_ATOMIC
> under the lock.
>
> Fixes: 081efd18326e ("ipv6: Protect nh->f6i_list with spinlock and flag.")
> Reported-by: AutonomousCodeSecurity@microsoft.com
> Signed-off-by: Xiang Mei (Microsoft) <xmei5@asu.edu>
> ---
> net/ipv4/nexthop.c | 13 +++++++++++--
> 1 file changed, 11 insertions(+), 2 deletions(-)
>
> diff --git a/net/ipv4/nexthop.c b/net/ipv4/nexthop.c
> index 6205bd57aa85..c1d3b7d7100d 100644
> --- a/net/ipv4/nexthop.c
> +++ b/net/ipv4/nexthop.c
> @@ -1597,14 +1597,21 @@ static int fib6_check_nh_list(struct nexthop *old, struct nexthop *new,
> struct netlink_ext_ack *extack)
> {
> struct fib6_info *f6i;
> + int err = 0;
>
> if (list_empty(&old->f6i_list))
> return 0;
>
> + spin_lock_bh(&old->lock);
> list_for_each_entry(f6i, &old->f6i_list, nh_list) {
> - if (check_src_addr(&f6i->fib6_src.addr, extack) < 0)
> - return -EINVAL;
> + err = check_src_addr(&f6i->fib6_src.addr, extack);
> + if (err)
> + break;
> }
> + spin_unlock_bh(&old->lock);
> +
> + if (err)
> + return err;
>
> return fib6_check_nexthop(new, NULL, extack);
> }
> @@ -2538,8 +2545,10 @@ static void __nexthop_replace_notify(struct net *net, struct nexthop *nh,
> fi->nh_updated = false;
> }
>
> + spin_lock_bh(&nh->lock);
> list_for_each_entry(f6i, &nh->f6i_list, nh_list)
> fib6_rt_update(net, f6i, info);
> + spin_unlock_bh(&nh->lock);
I read the feedback from Sashiko, but even if we convert this list to
RCU we would need GFP_ATOMIC. I don't find __nexthop_replace_notify()
too interesting given that modern systems unset
net.ipv4.nexthop_compat_mode, which means that this function is never
called.
Reviewed-by: Ido Schimmel <idosch@nvidia.com>
> }
>
> /* send RTM_NEWROUTE with REPLACE flag set for all FIB entries
> --
> 2.43.0
>
^ permalink raw reply [flat|nested] 4+ messages in thread
* Re: [PATCH net 2/2] nexthop: avoid unlocked f6i_list walk in nh_rt_cache_flush
2026-07-22 0:29 ` [PATCH net 2/2] nexthop: avoid unlocked f6i_list walk in nh_rt_cache_flush Xiang Mei (Microsoft)
@ 2026-07-23 19:20 ` Ido Schimmel
0 siblings, 0 replies; 4+ messages in thread
From: Ido Schimmel @ 2026-07-23 19:20 UTC (permalink / raw)
To: Xiang Mei (Microsoft)
Cc: David Ahern, David S . Miller, Eric Dumazet, Jakub Kicinski,
Paolo Abeni, Simon Horman, Kuniyuki Iwashima, netdev,
linux-kernel, AutonomousCodeSecurity, tgopinath, kys
On Wed, Jul 22, 2026 at 12:29:51AM +0000, Xiang Mei (Microsoft) wrote:
> nh_rt_cache_flush() walks nh->f6i_list during an RTNL-serialized nexthop
> replace without holding nh->lock, racing the unlocked IPv6 route
> add/delete that mutate the list under nh->lock and free fib6_info
> entries (nh_rt_cache_flush() is inlined into rtm_new_nexthop()):
>
> BUG: KASAN: slab-use-after-free in nh_rt_cache_flush (net/ipv4/nexthop.c:2243)
> Read of size 8 at addr ffff888012953e18 by task exploit/146
> nh_rt_cache_flush (net/ipv4/nexthop.c:2243)
> replace_nexthop (net/ipv4/nexthop.c:2610)
> rtm_new_nexthop (net/ipv4/nexthop.c:3323)
> rtnetlink_rcv_msg (net/core/rtnetlink.c:7076)
>
> Unlike the other f6i_list walks, this one bumps each route's sernum via
> fib6_update_sernum_upto_root(), which needs tb6_lock; taking nh->lock
> around it would invert the established tb6_lock -> nh->lock order and
> deadlock. As the only purpose is to invalidate cached dsts, bump the
> IPv6 sernum for the whole netns with rt_genid_bump_ipv6() instead,
> mirroring the rt_cache_flush() already done for IPv4 just above.
>
> Fixes: 081efd18326e ("ipv6: Protect nh->f6i_list with spinlock and flag.")
> Reported-by: AutonomousCodeSecurity@microsoft.com
> Signed-off-by: Xiang Mei (Microsoft) <xmei5@asu.edu>
> ---
> net/ipv4/nexthop.c | 12 ++++++------
> 1 file changed, 6 insertions(+), 6 deletions(-)
>
> diff --git a/net/ipv4/nexthop.c b/net/ipv4/nexthop.c
> index c1d3b7d7100d..1e9ccac6bf0f 100644
> --- a/net/ipv4/nexthop.c
> +++ b/net/ipv4/nexthop.c
> @@ -2240,18 +2240,18 @@ static void remove_one_nexthop(struct net *net, struct nexthop *nh,
> static void nh_rt_cache_flush(struct net *net, struct nexthop *nh,
> struct nexthop *replaced_nh)
> {
> - struct fib6_info *f6i;
> struct nh_group *nhg;
> + bool have_f6i;
> int i;
>
> if (!list_empty(&nh->fi_list))
> rt_cache_flush(net);
>
> - list_for_each_entry(f6i, &nh->f6i_list, nh_list) {
> - spin_lock_bh(&f6i->fib6_table->tb6_lock);
> - fib6_update_sernum_upto_root(net, f6i);
> - spin_unlock_bh(&f6i->fib6_table->tb6_lock);
> - }
> + spin_lock_bh(&nh->lock);
> + have_f6i = !list_empty(&nh->f6i_list);
> + spin_unlock_bh(&nh->lock);
> + if (have_f6i)
> + rt_genid_bump_ipv6(net);
Can you look into converting this list to RCU in net-next so that we
wouldn't need rt_genid_bump_ipv6() and instead restore
fib6_update_sernum_upto_root()?
Note that rt_genid_bump_ipv6() is much more expensive than
rt_cache_flush().
Reviewed-by: Ido Schimmel <idosch@nvidia.com>
>
> /* if an IPv6 group was replaced, we have to release all old
> * dsts to make sure all refcounts are released
> --
> 2.43.0
>
^ permalink raw reply [flat|nested] 4+ messages in thread
end of thread, other threads:[~2026-07-23 19:20 UTC | newest]
Thread overview: 4+ messages (download: mbox.gz follow: Atom feed
-- links below jump to the message on this page --
2026-07-22 0:29 [PATCH net 1/2] nexthop: take nh->lock for f6i_list walks in replace check and notify Xiang Mei (Microsoft)
2026-07-22 0:29 ` [PATCH net 2/2] nexthop: avoid unlocked f6i_list walk in nh_rt_cache_flush Xiang Mei (Microsoft)
2026-07-23 19:20 ` Ido Schimmel
2026-07-23 19:19 ` [PATCH net 1/2] nexthop: take nh->lock for f6i_list walks in replace check and notify Ido Schimmel
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox