* [PATCH net V2 0/3] net/mlx5: LAG bug fixes
@ 2026-06-30 11:29 Tariq Toukan
2026-06-30 11:29 ` [PATCH net V2 1/3] net/mlx5: LAG, Fix off-by-one in single-FDB error rollback Tariq Toukan
` (2 more replies)
0 siblings, 3 replies; 4+ messages in thread
From: Tariq Toukan @ 2026-06-30 11:29 UTC (permalink / raw)
To: Andrew Lunn, David S. Miller, Eric Dumazet, Jakub Kicinski,
netdev, Paolo Abeni
Cc: Edward Srouji, Jacob Keller, Kees Cook, Leon Romanovsky,
linux-kernel, linux-rdma, Maher Sanalla, Mark Bloch,
Moshe Shemesh, Or Har-Toov, Rongwei Liu, Saeed Mahameed,
Shay Drori, Simon Horman, Tariq Toukan
Hi,
Three bug fixes by Shay in the mlx5 LAG subsystem.
Patch 1 fixes an off-by-one in the error rollback path of
mlx5_lag_create_single_fdb_filter(): the loop started from the
failed index i, potentially operating on uninitialized state or
double-tearing-down an entry that had already self-rolled-back.
The rollback should start from i - 1.
Patch 2 fixes a hang in mlx5_mpesw_work(): when
mlx5_lag_get_devcom_comp() returns NULL the function returned
early without calling complete(), blocking any caller waiting on
mpesww->comp indefinitely.
Patch 3 fixes a kernel crash during teardown when
mlx5_lag_get_dev_seq() returns an error because no device is
marked as master or the peer is no longer in the LAG. The peer
flow cleanup is now skipped instead of proceeding with a bad
pointer.
This series by Shay fixes three bugs in the mlx5 LAG subsystem.
Regards,
Tariq
V2:
- Rebase.
- Patch 3: simplify to a single 'continue' on seq lookup failure.
V1:
https://lore.kernel.org/all/20260617063204.547427-2-tariqt@nvidia.com/
Find replies to previous Sashiko comments here:
https://lore.kernel.org/all/e18662ac-413e-43f6-ac65-a4e15fd47bb7@nvidia.com/
Shay Drory (3):
net/mlx5: LAG, Fix off-by-one in single-FDB error rollback
net/mlx5: LAG, MPESW, Fix missing complete() on devcom error
net/mlx5e: TC, skip peer flow cleanup when LAG seq is unavailable
drivers/net/ethernet/mellanox/mlx5/core/en_tc.c | 3 +++
drivers/net/ethernet/mellanox/mlx5/core/lag/mpesw.c | 7 +++++--
drivers/net/ethernet/mellanox/mlx5/core/lag/shared_fdb.c | 2 +-
3 files changed, 9 insertions(+), 3 deletions(-)
base-commit: dbf803bc4a8b0522c9a12560c20905a5952d1cb9
--
2.44.0
^ permalink raw reply [flat|nested] 4+ messages in thread
* [PATCH net V2 1/3] net/mlx5: LAG, Fix off-by-one in single-FDB error rollback
2026-06-30 11:29 [PATCH net V2 0/3] net/mlx5: LAG bug fixes Tariq Toukan
@ 2026-06-30 11:29 ` Tariq Toukan
2026-06-30 11:29 ` [PATCH net V2 2/3] net/mlx5: LAG, MPESW, Fix missing complete() on devcom error Tariq Toukan
2026-06-30 11:29 ` [PATCH net V2 3/3] net/mlx5e: TC, skip peer flow cleanup when LAG seq is unavailable Tariq Toukan
2 siblings, 0 replies; 4+ messages in thread
From: Tariq Toukan @ 2026-06-30 11:29 UTC (permalink / raw)
To: Andrew Lunn, David S. Miller, Eric Dumazet, Jakub Kicinski,
netdev, Paolo Abeni
Cc: Edward Srouji, Jacob Keller, Kees Cook, Leon Romanovsky,
linux-kernel, linux-rdma, Maher Sanalla, Mark Bloch,
Moshe Shemesh, Or Har-Toov, Rongwei Liu, Saeed Mahameed,
Shay Drori, Simon Horman, Tariq Toukan
From: Shay Drory <shayd@nvidia.com>
On failure at index i, the reverse cleanup loop in
mlx5_lag_create_single_fdb() starts from i, so the failed index
itself is rolled back. That can operate on uninitialized state or
double-tear-down a rule the add_one path already self-rolled-back.
Start the rollback from i - 1 so only successfully-installed entries
are undone.
Fixes: ddbb5ddc43ad ("net/mlx5: LAG, Refactor lag logic")
Signed-off-by: Shay Drory <shayd@nvidia.com>
Reviewed-by: Mark Bloch <mbloch@nvidia.com>
Signed-off-by: Tariq Toukan <tariqt@nvidia.com>
---
drivers/net/ethernet/mellanox/mlx5/core/lag/shared_fdb.c | 2 +-
1 file changed, 1 insertion(+), 1 deletion(-)
diff --git a/drivers/net/ethernet/mellanox/mlx5/core/lag/shared_fdb.c b/drivers/net/ethernet/mellanox/mlx5/core/lag/shared_fdb.c
index 113866494d16..6b4ad3c53f2f 100644
--- a/drivers/net/ethernet/mellanox/mlx5/core/lag/shared_fdb.c
+++ b/drivers/net/ethernet/mellanox/mlx5/core/lag/shared_fdb.c
@@ -78,7 +78,7 @@ static int mlx5_lag_create_single_fdb_filter(struct mlx5_lag *ldev, u32 filter)
}
return 0;
err:
- mlx5_lag_for_each_reverse(j, i, 0, ldev, filter) {
+ mlx5_lag_for_each_reverse(j, i - 1, 0, ldev, filter) {
struct mlx5_eswitch *slave_esw;
if (j == master_idx)
--
2.44.0
^ permalink raw reply related [flat|nested] 4+ messages in thread
* [PATCH net V2 2/3] net/mlx5: LAG, MPESW, Fix missing complete() on devcom error
2026-06-30 11:29 [PATCH net V2 0/3] net/mlx5: LAG bug fixes Tariq Toukan
2026-06-30 11:29 ` [PATCH net V2 1/3] net/mlx5: LAG, Fix off-by-one in single-FDB error rollback Tariq Toukan
@ 2026-06-30 11:29 ` Tariq Toukan
2026-06-30 11:29 ` [PATCH net V2 3/3] net/mlx5e: TC, skip peer flow cleanup when LAG seq is unavailable Tariq Toukan
2 siblings, 0 replies; 4+ messages in thread
From: Tariq Toukan @ 2026-06-30 11:29 UTC (permalink / raw)
To: Andrew Lunn, David S. Miller, Eric Dumazet, Jakub Kicinski,
netdev, Paolo Abeni
Cc: Edward Srouji, Jacob Keller, Kees Cook, Leon Romanovsky,
linux-kernel, linux-rdma, Maher Sanalla, Mark Bloch,
Moshe Shemesh, Or Har-Toov, Rongwei Liu, Saeed Mahameed,
Shay Drori, Simon Horman, Tariq Toukan
From: Shay Drory <shayd@nvidia.com>
mlx5_mpesw_work() returned without calling complete() when
mlx5_lag_get_devcom_comp() returned NULL. A caller that queued the
work and waited on mpesww->comp would block indefinitely.
Funnel the early-return path through a new "complete" label so the
waiter is always woken.
Fixes: b430c1b4f63b ("net/mlx5: Replace global mlx5_intf_lock with HCA devcom component lock")
Signed-off-by: Shay Drory <shayd@nvidia.com>
Reviewed-by: Mark Bloch <mbloch@nvidia.com>
Signed-off-by: Tariq Toukan <tariqt@nvidia.com>
---
drivers/net/ethernet/mellanox/mlx5/core/lag/mpesw.c | 7 +++++--
1 file changed, 5 insertions(+), 2 deletions(-)
diff --git a/drivers/net/ethernet/mellanox/mlx5/core/lag/mpesw.c b/drivers/net/ethernet/mellanox/mlx5/core/lag/mpesw.c
index 50bfb450c71e..abf72026c751 100644
--- a/drivers/net/ethernet/mellanox/mlx5/core/lag/mpesw.c
+++ b/drivers/net/ethernet/mellanox/mlx5/core/lag/mpesw.c
@@ -194,8 +194,10 @@ static void mlx5_mpesw_work(struct work_struct *work)
struct mlx5_lag *ldev = mpesww->lag;
devcom = mlx5_lag_get_devcom_comp(ldev);
- if (!devcom)
- return;
+ if (!devcom) {
+ mpesww->result = -ENODEV;
+ goto complete;
+ }
mlx5_devcom_comp_lock(devcom);
mlx5_mpesw_sd_devcoms_lock(ldev);
@@ -213,6 +215,7 @@ static void mlx5_mpesw_work(struct work_struct *work)
mutex_unlock(&ldev->lock);
mlx5_mpesw_sd_devcoms_unlock(ldev);
mlx5_devcom_comp_unlock(devcom);
+complete:
complete(&mpesww->comp);
}
--
2.44.0
^ permalink raw reply related [flat|nested] 4+ messages in thread
* [PATCH net V2 3/3] net/mlx5e: TC, skip peer flow cleanup when LAG seq is unavailable
2026-06-30 11:29 [PATCH net V2 0/3] net/mlx5: LAG bug fixes Tariq Toukan
2026-06-30 11:29 ` [PATCH net V2 1/3] net/mlx5: LAG, Fix off-by-one in single-FDB error rollback Tariq Toukan
2026-06-30 11:29 ` [PATCH net V2 2/3] net/mlx5: LAG, MPESW, Fix missing complete() on devcom error Tariq Toukan
@ 2026-06-30 11:29 ` Tariq Toukan
2 siblings, 0 replies; 4+ messages in thread
From: Tariq Toukan @ 2026-06-30 11:29 UTC (permalink / raw)
To: Andrew Lunn, David S. Miller, Eric Dumazet, Jakub Kicinski,
netdev, Paolo Abeni
Cc: Edward Srouji, Jacob Keller, Kees Cook, Leon Romanovsky,
linux-kernel, linux-rdma, Maher Sanalla, Mark Bloch,
Moshe Shemesh, Or Har-Toov, Rongwei Liu, Saeed Mahameed,
Shay Drori, Simon Horman, Tariq Toukan
From: Shay Drory <shayd@nvidia.com>
mlx5_lag_get_dev_seq() will return error when the peer isn't in the LAG
or when no device is marked as master. Result bad memory access and kernel
crash[1].
Hence, skip the peer when lookup fails.
Note: In case there are peer flows, they are cleaned before LAG cleared
the master mark.
[1]
RIP: 0010:mlx5e_tc_del_fdb_peers_flow+0x3d/0x350 [mlx5_core]
Call Trace:
<TASK>
mlx5e_tc_clean_fdb_peer_flows+0xc1/0x130 [mlx5_core]
mlx5_esw_offloads_unpair+0x3a/0x400 [mlx5_core]
mlx5_esw_offloads_devcom_event+0xee/0x360 [mlx5_core]
mlx5_devcom_send_event+0x7a/0x140 [mlx5_core]
mlx5_esw_offloads_devcom_cleanup+0x2f/0x90 [mlx5_core]
mlx5e_tc_esw_cleanup+0x28/0xf0 [mlx5_core]
mlx5e_rep_tc_cleanup+0x19/0x30 [mlx5_core]
mlx5e_cleanup_uplink_rep_tx+0x36/0x40 [mlx5_core]
mlx5e_cleanup_rep_tx+0x55/0x60 [mlx5_core]
mlx5e_detach_netdev+0x96/0xf0 [mlx5_core]
mlx5e_netdev_change_profile+0x5b/0x120 [mlx5_core]
mlx5e_netdev_attach_nic_profile+0x1b/0x30 [mlx5_core]
mlx5e_vport_rep_unload+0xdd/0x110 [mlx5_core]
__esw_offloads_unload_rep+0x81/0xb0 [mlx5_core]
mlx5_eswitch_unregister_vport_reps+0x1d7/0x220 [mlx5_core]
mlx5e_rep_remove+0x22/0x30 [mlx5_core]
device_release_driver_internal+0x194/0x1f0
bus_remove_device+0xe8/0x1b0
device_del+0x159/0x3c0
mlx5_rescan_drivers_locked+0xbc/0x2d0 [mlx5_core]
mlx5_unregister_device+0x54/0x80 [mlx5_core]
mlx5_uninit_one+0x73/0x130 [mlx5_core]
remove_one+0x78/0xe0 [mlx5_core]
pci_device_remove+0x39/0xa0
Fixes: 971b28accc09 ("net/mlx5: LAG, replace mlx5_get_dev_index with LAG sequence number")
Signed-off-by: Shay Drory <shayd@nvidia.com>
Reviewed-by: Mark Bloch <mbloch@nvidia.com>
Signed-off-by: Tariq Toukan <tariqt@nvidia.com>
---
drivers/net/ethernet/mellanox/mlx5/core/en_tc.c | 3 +++
1 file changed, 3 insertions(+)
diff --git a/drivers/net/ethernet/mellanox/mlx5/core/en_tc.c b/drivers/net/ethernet/mellanox/mlx5/core/en_tc.c
index 910492eb51f2..1bc7b9019124 100644
--- a/drivers/net/ethernet/mellanox/mlx5/core/en_tc.c
+++ b/drivers/net/ethernet/mellanox/mlx5/core/en_tc.c
@@ -5547,6 +5547,9 @@ void mlx5e_tc_clean_fdb_peer_flows(struct mlx5_eswitch *esw)
mlx5_devcom_for_each_peer_entry(devcom, peer_esw, pos) {
i = mlx5_lag_get_dev_seq(peer_esw->dev);
+ if (i < 0)
+ continue;
+
list_for_each_entry_safe(flow, tmp, &esw->offloads.peer_flows[i], peer[i])
mlx5e_tc_del_fdb_peers_flow(flow);
}
--
2.44.0
^ permalink raw reply related [flat|nested] 4+ messages in thread
end of thread, other threads:[~2026-06-30 11:30 UTC | newest]
Thread overview: 4+ messages (download: mbox.gz follow: Atom feed
-- links below jump to the message on this page --
2026-06-30 11:29 [PATCH net V2 0/3] net/mlx5: LAG bug fixes Tariq Toukan
2026-06-30 11:29 ` [PATCH net V2 1/3] net/mlx5: LAG, Fix off-by-one in single-FDB error rollback Tariq Toukan
2026-06-30 11:29 ` [PATCH net V2 2/3] net/mlx5: LAG, MPESW, Fix missing complete() on devcom error Tariq Toukan
2026-06-30 11:29 ` [PATCH net V2 3/3] net/mlx5e: TC, skip peer flow cleanup when LAG seq is unavailable Tariq Toukan
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox