From: sashiko-bot@kernel.org
To: "Tariq Toukan" <tariqt@nvidia.com>
Cc: linux-rdma@vger.kernel.org
Subject: Re: [PATCH net-next 04/13] net/mlx5: LAG, allocate port-indexed scratch buffers dynamically
Date: Thu, 24 Sep 2026 10:41:42 +0000 [thread overview]
Message-ID: <20260924104142.8368F1F0089C@smtp.kernel.org> (raw)
In-Reply-To: <20260923103830.1183-5-tariqt@nvidia.com>
Thank you for your contribution! Sashiko AI review found 2 potential issue(s) to consider:
- [Medium] net/mlx5: heap buffer overflow in MLX5 LAG configuration arrays
- [Medium] net/mlx5: partial LAG hardware state update on ENOMEM
--
commit 5ce8c0ac5b16870dc2d8dca3c2afacfca478bc6b
Author: Shay Drory <shayd@nvidia.com>
Subject: net/mlx5: LAG, allocate port-indexed scratch buffers dynamically
This patch changes the MLX5 LAG configuration functions to dynamically
allocate their scratch arrays based on the ldev->ports count instead of
using stack buffers sized to MLX5_MAX_PORTS. It also adds error handling
for these new allocations.
> diff --git a/drivers/net/ethernet/mellanox/mlx5/core/lag/lag.c b/drivers/net/ethernet/mellanox/mlx5/core/lag/lag.c
> --- a/drivers/net/ethernet/mellanox/mlx5/core/lag/lag.c
> +++ b/drivers/net/ethernet/mellanox/mlx5/core/lag/lag.c
> @@ -63,19 +63,25 @@ static int get_port_sel_mode(enum mlx5_lag_mode mode, unsigned long flags)
[ ... ]
> -static u8 lag_active_port_bits(struct mlx5_lag *ldev,
> - struct lag_tracker *tracker)
> +static int lag_active_port_bits(struct mlx5_lag *ldev,
> + struct lag_tracker *tracker)
> {
> - u8 enabled_ports[MLX5_MAX_PORTS] = {};
> u8 active_port = 0;
> + u8 *enabled_ports;
> int num_enabled;
> int idx;
>
> + enabled_ports = kcalloc(ldev->ports, sizeof(*enabled_ports),
> + GFP_KERNEL);
> + if (!enabled_ports)
> + return -ENOMEM;
> +
> mlx5_infer_tx_enabled(tracker, ldev, enabled_ports,
> &num_enabled);
[Severity: Medium]
Can mlx5_infer_tx_enabled() write past the end of enabled_ports?
The array is now strictly sized to ldev->ports. However,
mlx5_infer_tx_enabled() calls mlx5_ldev_for_each(), which iterates over
the pfs xarray. If hardware misbehaves and more PCI functions match the
GUID than the reported capacity, the xarray could contain up to
MLX5_MAX_PORTS items, overflowing this buffer.
[ ... ]
> @@ -240,22 +254,31 @@ static void mlx5_lag_print_mapping(struct mlx5_core_dev *dev,
> struct lag_tracker *tracker,
> unsigned long flags)
> {
> - char buf[MLX5_MAX_PORTS * 10 + 1] = {};
> - u8 enabled_ports[MLX5_MAX_PORTS] = {};
> + u8 *enabled_ports = NULL;
> int written = 0;
> int num_enabled;
> + char *buf;
> int idx;
> int err;
> int i;
> int j;
>
> + buf = kcalloc(ldev->ports * 10 + 1, sizeof(*buf), GFP_KERNEL);
> + if (!buf)
> + return;
> +
> if (test_bit(MLX5_LAG_MODE_FLAG_HASH_BASED, &flags)) {
> + enabled_ports = kcalloc(ldev->ports, sizeof(*enabled_ports),
> + GFP_KERNEL);
[Severity: Medium]
Similar to the above, can the string building loop that uses this buffer
exceed ldev->ports iterations if mlx5_ldev_for_each() encounters more
xarray items than expected?
[ ... ]
> @@ -668,20 +696,28 @@ static bool __mlx5_lag_is_sd_active(struct mlx5_lag *ldev,
[ ... ]
> -static void mlx5_infer_tx_affinity_mapping(struct lag_tracker *tracker,
> - struct mlx5_lag *ldev,
> - u8 buckets,
> - u8 *ports)
> +static int mlx5_infer_tx_affinity_mapping(struct lag_tracker *tracker,
> + struct mlx5_lag *ldev,
> + u8 buckets,
> + u8 *ports)
> {
> - int disabled[MLX5_MAX_PORTS] = {};
> - int enabled[MLX5_MAX_PORTS] = {};
> int disabled_ports_num = 0;
> int enabled_ports_num = 0;
> + int *disabled;
> + int *enabled;
> + int err = 0;
> int idx;
> u32 rand;
> int i;
> int j;
>
> + enabled = kcalloc(ldev->ports, sizeof(*enabled), GFP_KERNEL);
> + disabled = kcalloc(ldev->ports, sizeof(*disabled), GFP_KERNEL);
> + if (!enabled || !disabled) {
> + err = -ENOMEM;
> + goto out;
> + }
> +
> mlx5_ldev_for_each(i, 0, ldev) {
> if (tracker->netdev_state[i].tx_enabled &&
> tracker->netdev_state[i].link_up)
> enabled[enabled_ports_num++] = i;
> else
> disabled[disabled_ports_num++] = i;
[Severity: Medium]
Does this iteration risk writing out of bounds into the enabled and disabled
arrays if mlx5_ldev_for_each() yields more items than ldev->ports?
[ ... ]
> @@ -816,7 +857,10 @@ static int _mlx5_modify_lag(struct mlx5_lag *ldev,
> !MLX5_CAP_PORT_SELECTION(dev0, port_select_flow_table_bypass))
> return ret;
>
> - active_ports = lag_active_port_bits(ldev, tracker);
> + ret = lag_active_port_bits(ldev, tracker);
> + if (ret < 0)
> + return ret;
[Severity: Medium]
Does returning early here leave the hardware and software states out of sync?
Just before this snippet in _mlx5_modify_lag(), mlx5_lag_port_sel_modify()
is called to update the hardware definers. If lag_active_port_bits() now
fails due to ENOMEM, we skip sending the active port command.
This error also propagates up to mlx5_modify_lag(), which then jumps to
the error path and skips updating the software state:
mlx5_modify_lag() {
...
err = _mlx5_modify_lag(ldev, tracker, ports);
if (err) {
mlx5_core_err(dev0, "Failed to modify LAG (%d)\n", err);
goto out;
}
memcpy(ldev->v2p_map, ports, ...);
...
}
Should there be a rollback mechanism for mlx5_lag_port_sel_modify() if
we hit this new error path?
--
Sashiko AI review · https://sashiko.dev/#/patchset/20260923103830.1183-1-tariqt@nvidia.com?part=4
next prev parent reply other threads:[~2026-09-24 10:41 UTC|newest]
Thread overview: 39+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-09-23 10:38 [PATCH net-next 00/13] net/mlx5: Preparations for nested E-switch Tariq Toukan
2026-09-23 10:38 ` [PATCH net-next 01/13] net/mlx5e: Assign a random MAC to any netdev with a zero MAC address Tariq Toukan
2026-09-24 10:41 ` sashiko-bot
2026-09-24 17:45 ` netdev-bot+sashiko
2026-09-23 10:38 ` [PATCH net-next 02/13] net/mlx5: E-switch, do not leave an unpaired devcom registered Tariq Toukan
2026-09-24 10:41 ` sashiko-bot
2026-09-24 17:45 ` netdev-bot+sashiko
2026-09-23 10:38 ` [PATCH net-next 03/13] net/mlx5: LAG, allocate v2p_map dynamically Tariq Toukan
2026-09-24 10:41 ` sashiko-bot
2026-09-24 17:45 ` netdev-bot+sashiko
2026-09-28 14:56 ` Shay Drori
2026-09-23 10:38 ` [PATCH net-next 04/13] net/mlx5: LAG, allocate port-indexed scratch buffers dynamically Tariq Toukan
2026-09-24 10:41 ` sashiko-bot [this message]
2026-09-24 17:45 ` netdev-bot+sashiko
2026-09-28 14:57 ` Shay Drori
2026-09-23 10:38 ` [PATCH net-next 05/13] net/mlx5: LAG, drop per-port scratch array in drop-rule setup Tariq Toukan
2026-09-24 10:41 ` sashiko-bot
2026-09-23 10:38 ` [PATCH net-next 06/13] net/mlx5: LAG, size debugfs buffers by port count Tariq Toukan
2026-09-24 10:41 ` sashiko-bot
2026-09-24 17:45 ` netdev-bot+sashiko
2026-09-23 10:38 ` [PATCH net-next 07/13] net/mlx5e: TC, anchor peer-flow reverse index on the duplicated flow Tariq Toukan
2026-09-24 10:41 ` sashiko-bot
2026-09-24 17:45 ` netdev-bot+sashiko
2026-09-23 10:38 ` [PATCH net-next 08/13] net/mlx5e: TC, track peer flows in a vhca_id xarray Tariq Toukan
2026-09-24 10:41 ` sashiko-bot
2026-09-24 17:46 ` netdev-bot+sashiko
2026-09-23 10:38 ` [PATCH net-next 09/13] net/mlx5: E-switch, derive manager vport from device capability Tariq Toukan
2026-09-24 10:41 ` sashiko-bot
2026-09-24 17:46 ` netdev-bot+sashiko
2026-09-28 15:13 ` Shay Drori
2026-09-23 10:38 ` [PATCH net-next 10/13] net/mlx5: LAG, don't print port mapping to debugfs in MPESW mode Tariq Toukan
2026-09-24 10:41 ` sashiko-bot
2026-09-23 10:38 ` [PATCH net-next 11/13] net/mlx5: LAG, drop stale esw_shared_ingress_acl gate from shared FDB Tariq Toukan
2026-09-24 10:41 ` sashiko-bot
2026-09-23 10:38 ` [PATCH net-next 12/13] net/mlx5: E-switch, correct stale VF/PF wording in esw-allowed comments Tariq Toukan
2026-09-24 10:41 ` sashiko-bot
2026-09-23 10:38 ` [PATCH net-next 13/13] net/mlx5: E-switch, disable host functions for a non PF e-switch manager Tariq Toukan
2026-09-24 10:41 ` sashiko-bot
2026-09-29 0:20 ` [PATCH net-next 00/13] net/mlx5: Preparations for nested E-switch patchwork-bot+netdevbpf
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=20260924104142.8368F1F0089C@smtp.kernel.org \
--to=sashiko-bot@kernel.org \
--cc=linux-rdma@vger.kernel.org \
--cc=sashiko-reviews@lists.linux.dev \
--cc=tariqt@nvidia.com \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox