Linux RDMA and InfiniBand development
 help / color / mirror / Atom feed
* [PATCH net v2] net/mlx5: Fix slab-out-of-bounds when handling team device events
@ 2026-09-24  6:42 Anirudh Virdi
  2026-09-24  6:50 ` sashiko-bot
                   ` (2 more replies)
  0 siblings, 3 replies; 4+ messages in thread
From: Anirudh Virdi @ 2026-09-24  6:42 UTC (permalink / raw)
  To: netdev
  Cc: saeedm, leon, tariqt, mbloch, andrew+netdev, davem, edumazet,
	kuba, pabeni, shayd, ohartoov, maorg, linux-rdma, Anirudh Virdi

The mlx5_handle_changeupper_event() and mlx5_handle_changeinfodata_event()
functions check for LAG masters (which includes both bonding and team
devices) but only call bonding-specific APIs. When processing team device
events, calling bond_slave_get_rcu() and bond_is_slave_inactive() on
team_port structures causes KASAN to detect an out-of-bounds memory access
since team_port is smaller than bond slave.

Fix this by wrapping the bond-specific API calls with a check for bonding
devices. This allows the function to still process LAG events for both
bonding and teams but only calls bond-specific functions when dealing with
actual bonding devices.

For mlx5_handle_changeinfodata_event(), keep the explicit bond check as
mlx5 doesn't handle state information for team ports anyway.

Tested with Mellanox ConnectX-5 on Linux 7.2.0-rc6:
- Bonding: PASS (no regressions)
- Team device: PASS (no KASAN errors)

Fixes: 54493a08e21f ("net/mlx5: Lag, record inactive state of bond device")
Suggested-by: Mark Bloch <mbloch@nvidia.com>
Signed-off-by: Anirudh Virdi <avirdi@redhat.com>
---
Changes in v2:
- Changed mlx5_handle_changeupper_event() to keep netif_is_lag_master() check
  instead of using netif_is_bond_master() at the start
- Wrapped bond-specific API calls (bond_slave_get_rcu, bond_is_slave_inactive)
  with an explicit netif_is_bond_master() check inside the loop
- This allows both bonding and team events to be processed, but only calls
  bond-specific functions for actual bonding devices
- Suggested-by: Mark Bloch <mbloch@nvidia.com>
- Tested on hardware to confirm both bonding and team devices work without KASAN errors

 drivers/net/ethernet/mellanox/mlx5/core/lag/lag.c | 10 ++++++----
 1 file changed, 6 insertions(+), 4 deletions(-)

diff --git a/drivers/net/ethernet/mellanox/mlx5/core/lag/lag.c b/drivers/net/ethernet/mellanox/mlx5/core/lag/lag.c
index 28d16fdc3f06..1c4107d9a408 100644
--- a/drivers/net/ethernet/mellanox/mlx5/core/lag/lag.c
+++ b/drivers/net/ethernet/mellanox/mlx5/core/lag/lag.c
@@ -1918,9 +1918,11 @@ static int mlx5_handle_changeupper_event(struct mlx5_lag *ldev,
 			}
 		}
 		if (i < MLX5_MAX_PORTS) {
-			slave = bond_slave_get_rcu(ndev_tmp);
-			if (slave)
-				has_inactive |= bond_is_slave_inactive(slave);
+			if (netif_is_bond_master(upper)) {
+				slave = bond_slave_get_rcu(ndev_tmp);
+				if (slave)
+					has_inactive |= bond_is_slave_inactive(slave);
+			}
 			bond_status |= (1 << idx);
 		}
 
@@ -2004,7 +2006,7 @@ static int mlx5_handle_changeinfodata_event(struct mlx5_lag *ldev,
 	bool has_inactive = 0;
 	int idx;
 
-	if (!netif_is_lag_master(ndev))
+	if (!netif_is_bond_master(ndev))
 		return 0;
 
 	rcu_read_lock();
-- 
2.55.0


^ permalink raw reply related	[flat|nested] 4+ messages in thread

end of thread, other threads:[~2026-09-29  0:51 UTC | newest]

Thread overview: 4+ messages (download: mbox.gz follow: Atom feed
-- links below jump to the message on this page --
2026-09-24  6:42 [PATCH net v2] net/mlx5: Fix slab-out-of-bounds when handling team device events Anirudh Virdi
2026-09-24  6:50 ` sashiko-bot
2026-09-25  7:33 ` Mark Bloch
2026-09-29  0:50 ` patchwork-bot+netdevbpf

This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox