Netdev List
 help / color / mirror / Atom feed
* [PATCH net v2] net/mlx5: Fix slab-out-of-bounds when handling team device events
@ 2026-09-24  6:42 Anirudh Virdi
  2026-09-25  7:33 ` Mark Bloch
  2026-09-29  0:50 ` patchwork-bot+netdevbpf
  0 siblings, 2 replies; 3+ messages in thread
From: Anirudh Virdi @ 2026-09-24  6:42 UTC (permalink / raw)
  To: netdev
  Cc: saeedm, leon, tariqt, mbloch, andrew+netdev, davem, edumazet,
	kuba, pabeni, shayd, ohartoov, maorg, linux-rdma, Anirudh Virdi

The mlx5_handle_changeupper_event() and mlx5_handle_changeinfodata_event()
functions check for LAG masters (which includes both bonding and team
devices) but only call bonding-specific APIs. When processing team device
events, calling bond_slave_get_rcu() and bond_is_slave_inactive() on
team_port structures causes KASAN to detect an out-of-bounds memory access
since team_port is smaller than bond slave.

Fix this by wrapping the bond-specific API calls with a check for bonding
devices. This allows the function to still process LAG events for both
bonding and teams but only calls bond-specific functions when dealing with
actual bonding devices.

For mlx5_handle_changeinfodata_event(), keep the explicit bond check as
mlx5 doesn't handle state information for team ports anyway.

Tested with Mellanox ConnectX-5 on Linux 7.2.0-rc6:
- Bonding: PASS (no regressions)
- Team device: PASS (no KASAN errors)

Fixes: 54493a08e21f ("net/mlx5: Lag, record inactive state of bond device")
Suggested-by: Mark Bloch <mbloch@nvidia.com>
Signed-off-by: Anirudh Virdi <avirdi@redhat.com>
---
Changes in v2:
- Changed mlx5_handle_changeupper_event() to keep netif_is_lag_master() check
  instead of using netif_is_bond_master() at the start
- Wrapped bond-specific API calls (bond_slave_get_rcu, bond_is_slave_inactive)
  with an explicit netif_is_bond_master() check inside the loop
- This allows both bonding and team events to be processed, but only calls
  bond-specific functions for actual bonding devices
- Suggested-by: Mark Bloch <mbloch@nvidia.com>
- Tested on hardware to confirm both bonding and team devices work without KASAN errors

 drivers/net/ethernet/mellanox/mlx5/core/lag/lag.c | 10 ++++++----
 1 file changed, 6 insertions(+), 4 deletions(-)

diff --git a/drivers/net/ethernet/mellanox/mlx5/core/lag/lag.c b/drivers/net/ethernet/mellanox/mlx5/core/lag/lag.c
index 28d16fdc3f06..1c4107d9a408 100644
--- a/drivers/net/ethernet/mellanox/mlx5/core/lag/lag.c
+++ b/drivers/net/ethernet/mellanox/mlx5/core/lag/lag.c
@@ -1918,9 +1918,11 @@ static int mlx5_handle_changeupper_event(struct mlx5_lag *ldev,
 			}
 		}
 		if (i < MLX5_MAX_PORTS) {
-			slave = bond_slave_get_rcu(ndev_tmp);
-			if (slave)
-				has_inactive |= bond_is_slave_inactive(slave);
+			if (netif_is_bond_master(upper)) {
+				slave = bond_slave_get_rcu(ndev_tmp);
+				if (slave)
+					has_inactive |= bond_is_slave_inactive(slave);
+			}
 			bond_status |= (1 << idx);
 		}
 
@@ -2004,7 +2006,7 @@ static int mlx5_handle_changeinfodata_event(struct mlx5_lag *ldev,
 	bool has_inactive = 0;
 	int idx;
 
-	if (!netif_is_lag_master(ndev))
+	if (!netif_is_bond_master(ndev))
 		return 0;
 
 	rcu_read_lock();
-- 
2.55.0


^ permalink raw reply related	[flat|nested] 3+ messages in thread

* Re: [PATCH net v2] net/mlx5: Fix slab-out-of-bounds when handling team device events
  2026-09-24  6:42 [PATCH net v2] net/mlx5: Fix slab-out-of-bounds when handling team device events Anirudh Virdi
@ 2026-09-25  7:33 ` Mark Bloch
  2026-09-29  0:50 ` patchwork-bot+netdevbpf
  1 sibling, 0 replies; 3+ messages in thread
From: Mark Bloch @ 2026-09-25  7:33 UTC (permalink / raw)
  To: Anirudh Virdi, netdev
  Cc: saeedm, leon, tariqt, andrew+netdev, davem, edumazet, kuba,
	pabeni, shayd, ohartoov, maorg, linux-rdma



On 24/09/2026 9:42, Anirudh Virdi wrote:
> The mlx5_handle_changeupper_event() and mlx5_handle_changeinfodata_event()
> functions check for LAG masters (which includes both bonding and team
> devices) but only call bonding-specific APIs. When processing team device
> events, calling bond_slave_get_rcu() and bond_is_slave_inactive() on
> team_port structures causes KASAN to detect an out-of-bounds memory access
> since team_port is smaller than bond slave.
> 
> Fix this by wrapping the bond-specific API calls with a check for bonding
> devices. This allows the function to still process LAG events for both
> bonding and teams but only calls bond-specific functions when dealing with
> actual bonding devices.
> 
> For mlx5_handle_changeinfodata_event(), keep the explicit bond check as
> mlx5 doesn't handle state information for team ports anyway.
> 
> Tested with Mellanox ConnectX-5 on Linux 7.2.0-rc6:
> - Bonding: PASS (no regressions)
> - Team device: PASS (no KASAN errors)
> 
> Fixes: 54493a08e21f ("net/mlx5: Lag, record inactive state of bond device")
> Suggested-by: Mark Bloch <mbloch@nvidia.com>
> Signed-off-by: Anirudh Virdi <avirdi@redhat.com>

Thanks for the patch,

Reviewed-by: Mark Bloch <mbloch@nvidia.com>

Mark

> ---
> Changes in v2:
> - Changed mlx5_handle_changeupper_event() to keep netif_is_lag_master() check
>   instead of using netif_is_bond_master() at the start
> - Wrapped bond-specific API calls (bond_slave_get_rcu, bond_is_slave_inactive)
>   with an explicit netif_is_bond_master() check inside the loop
> - This allows both bonding and team events to be processed, but only calls
>   bond-specific functions for actual bonding devices
> - Suggested-by: Mark Bloch <mbloch@nvidia.com>
> - Tested on hardware to confirm both bonding and team devices work without KASAN errors
> 
>  drivers/net/ethernet/mellanox/mlx5/core/lag/lag.c | 10 ++++++----
>  1 file changed, 6 insertions(+), 4 deletions(-)
> 
> diff --git a/drivers/net/ethernet/mellanox/mlx5/core/lag/lag.c b/drivers/net/ethernet/mellanox/mlx5/core/lag/lag.c
> index 28d16fdc3f06..1c4107d9a408 100644
> --- a/drivers/net/ethernet/mellanox/mlx5/core/lag/lag.c
> +++ b/drivers/net/ethernet/mellanox/mlx5/core/lag/lag.c
> @@ -1918,9 +1918,11 @@ static int mlx5_handle_changeupper_event(struct mlx5_lag *ldev,
>  			}
>  		}
>  		if (i < MLX5_MAX_PORTS) {
> -			slave = bond_slave_get_rcu(ndev_tmp);
> -			if (slave)
> -				has_inactive |= bond_is_slave_inactive(slave);
> +			if (netif_is_bond_master(upper)) {
> +				slave = bond_slave_get_rcu(ndev_tmp);
> +				if (slave)
> +					has_inactive |= bond_is_slave_inactive(slave);
> +			}
>  			bond_status |= (1 << idx);
>  		}
>  
> @@ -2004,7 +2006,7 @@ static int mlx5_handle_changeinfodata_event(struct mlx5_lag *ldev,
>  	bool has_inactive = 0;
>  	int idx;
>  
> -	if (!netif_is_lag_master(ndev))
> +	if (!netif_is_bond_master(ndev))
>  		return 0;
>  
>  	rcu_read_lock();


^ permalink raw reply	[flat|nested] 3+ messages in thread

* Re: [PATCH net v2] net/mlx5: Fix slab-out-of-bounds when handling team device events
  2026-09-24  6:42 [PATCH net v2] net/mlx5: Fix slab-out-of-bounds when handling team device events Anirudh Virdi
  2026-09-25  7:33 ` Mark Bloch
@ 2026-09-29  0:50 ` patchwork-bot+netdevbpf
  1 sibling, 0 replies; 3+ messages in thread
From: patchwork-bot+netdevbpf @ 2026-09-29  0:50 UTC (permalink / raw)
  To: Anirudh Virdi
  Cc: netdev, saeedm, leon, tariqt, mbloch, andrew+netdev, davem,
	edumazet, kuba, pabeni, shayd, ohartoov, maorg, linux-rdma

Hello:

This patch was applied to netdev/net.git (main)
by Jakub Kicinski <kuba@kernel.org>:

On Thu, 24 Sep 2026 12:12:44 +0530 you wrote:
> The mlx5_handle_changeupper_event() and mlx5_handle_changeinfodata_event()
> functions check for LAG masters (which includes both bonding and team
> devices) but only call bonding-specific APIs. When processing team device
> events, calling bond_slave_get_rcu() and bond_is_slave_inactive() on
> team_port structures causes KASAN to detect an out-of-bounds memory access
> since team_port is smaller than bond slave.
> 
> [...]

Here is the summary with links:
  - [net,v2] net/mlx5: Fix slab-out-of-bounds when handling team device events
    https://git.kernel.org/netdev/net/c/04c824940d52

You are awesome, thank you!
-- 
Deet-doot-dot, I am a bot.
https://korg.docs.kernel.org/patchwork/pwbot.html



^ permalink raw reply	[flat|nested] 3+ messages in thread

end of thread, other threads:[~2026-09-29  0:51 UTC | newest]

Thread overview: 3+ messages (download: mbox.gz follow: Atom feed
-- links below jump to the message on this page --
2026-09-24  6:42 [PATCH net v2] net/mlx5: Fix slab-out-of-bounds when handling team device events Anirudh Virdi
2026-09-25  7:33 ` Mark Bloch
2026-09-29  0:50 ` patchwork-bot+netdevbpf

This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox