Linux RDMA and InfiniBand development
 help / color / mirror / Atom feed
From: sashiko-bot@kernel.org
To: "Tariq Toukan" <tariqt@nvidia.com>
Cc: linux-rdma@vger.kernel.org
Subject: Re: [PATCH net-next 04/13] net/mlx5: LAG, allocate port-indexed scratch buffers dynamically
Date: Thu, 24 Sep 2026 10:41:42 +0000	[thread overview]
Message-ID: <20260924104142.8368F1F0089C@smtp.kernel.org> (raw)
In-Reply-To: <20260923103830.1183-5-tariqt@nvidia.com>

Thank you for your contribution! Sashiko AI review found 2 potential issue(s) to consider:
- [Medium] net/mlx5: heap buffer overflow in MLX5 LAG configuration arrays
- [Medium] net/mlx5: partial LAG hardware state update on ENOMEM
--

commit 5ce8c0ac5b16870dc2d8dca3c2afacfca478bc6b
Author: Shay Drory <shayd@nvidia.com>
Subject: net/mlx5: LAG, allocate port-indexed scratch buffers dynamically

This patch changes the MLX5 LAG configuration functions to dynamically
allocate their scratch arrays based on the ldev->ports count instead of
using stack buffers sized to MLX5_MAX_PORTS. It also adds error handling
for these new allocations.

> diff --git a/drivers/net/ethernet/mellanox/mlx5/core/lag/lag.c b/drivers/net/ethernet/mellanox/mlx5/core/lag/lag.c
> --- a/drivers/net/ethernet/mellanox/mlx5/core/lag/lag.c
> +++ b/drivers/net/ethernet/mellanox/mlx5/core/lag/lag.c
> @@ -63,19 +63,25 @@ static int get_port_sel_mode(enum mlx5_lag_mode mode, unsigned long flags)
[ ... ]
> -static u8 lag_active_port_bits(struct mlx5_lag *ldev,
> -			       struct lag_tracker *tracker)
> +static int lag_active_port_bits(struct mlx5_lag *ldev,
> +				struct lag_tracker *tracker)
>  {
> -	u8 enabled_ports[MLX5_MAX_PORTS] = {};
>  	u8 active_port = 0;
> +	u8 *enabled_ports;
>  	int num_enabled;
>  	int idx;
>  
> +	enabled_ports = kcalloc(ldev->ports, sizeof(*enabled_ports),
> +				GFP_KERNEL);
> +	if (!enabled_ports)
> +		return -ENOMEM;
> +
>  	mlx5_infer_tx_enabled(tracker, ldev, enabled_ports,
>  			      &num_enabled);

[Severity: Medium]
Can mlx5_infer_tx_enabled() write past the end of enabled_ports?

The array is now strictly sized to ldev->ports. However,
mlx5_infer_tx_enabled() calls mlx5_ldev_for_each(), which iterates over
the pfs xarray. If hardware misbehaves and more PCI functions match the
GUID than the reported capacity, the xarray could contain up to
MLX5_MAX_PORTS items, overflowing this buffer.

[ ... ]
> @@ -240,22 +254,31 @@ static void mlx5_lag_print_mapping(struct mlx5_core_dev *dev,
>  				   struct lag_tracker *tracker,
>  				   unsigned long flags)
>  {
> -	char buf[MLX5_MAX_PORTS * 10 + 1] = {};
> -	u8 enabled_ports[MLX5_MAX_PORTS] = {};
> +	u8 *enabled_ports = NULL;
>  	int written = 0;
>  	int num_enabled;
> +	char *buf;
>  	int idx;
>  	int err;
>  	int i;
>  	int j;
>  
> +	buf = kcalloc(ldev->ports * 10 + 1, sizeof(*buf), GFP_KERNEL);
> +	if (!buf)
> +		return;
> +
>  	if (test_bit(MLX5_LAG_MODE_FLAG_HASH_BASED, &flags)) {
> +		enabled_ports = kcalloc(ldev->ports, sizeof(*enabled_ports),
> +					GFP_KERNEL);

[Severity: Medium]
Similar to the above, can the string building loop that uses this buffer
exceed ldev->ports iterations if mlx5_ldev_for_each() encounters more
xarray items than expected?

[ ... ]
> @@ -668,20 +696,28 @@ static bool __mlx5_lag_is_sd_active(struct mlx5_lag *ldev,
[ ... ]
> -static void mlx5_infer_tx_affinity_mapping(struct lag_tracker *tracker,
> -					   struct mlx5_lag *ldev,
> -					   u8 buckets,
> -					   u8 *ports)
> +static int mlx5_infer_tx_affinity_mapping(struct lag_tracker *tracker,
> +					  struct mlx5_lag *ldev,
> +					  u8 buckets,
> +					  u8 *ports)
>  {
> -	int disabled[MLX5_MAX_PORTS] = {};
> -	int enabled[MLX5_MAX_PORTS] = {};
>  	int disabled_ports_num = 0;
>  	int enabled_ports_num = 0;
> +	int *disabled;
> +	int *enabled;
> +	int err = 0;
>  	int idx;
>  	u32 rand;
>  	int i;
>  	int j;
>  
> +	enabled = kcalloc(ldev->ports, sizeof(*enabled), GFP_KERNEL);
> +	disabled = kcalloc(ldev->ports, sizeof(*disabled), GFP_KERNEL);
> +	if (!enabled || !disabled) {
> +		err = -ENOMEM;
> +		goto out;
> +	}
> +
>  	mlx5_ldev_for_each(i, 0, ldev) {
>  		if (tracker->netdev_state[i].tx_enabled &&
>  		    tracker->netdev_state[i].link_up)
>  			enabled[enabled_ports_num++] = i;
>  		else
>  			disabled[disabled_ports_num++] = i;

[Severity: Medium]
Does this iteration risk writing out of bounds into the enabled and disabled
arrays if mlx5_ldev_for_each() yields more items than ldev->ports?

[ ... ]
> @@ -816,7 +857,10 @@ static int _mlx5_modify_lag(struct mlx5_lag *ldev,
>  		    !MLX5_CAP_PORT_SELECTION(dev0, port_select_flow_table_bypass))
>  			return ret;
>  
> -		active_ports = lag_active_port_bits(ldev, tracker);
> +		ret = lag_active_port_bits(ldev, tracker);
> +		if (ret < 0)
> +			return ret;

[Severity: Medium]
Does returning early here leave the hardware and software states out of sync?

Just before this snippet in _mlx5_modify_lag(), mlx5_lag_port_sel_modify()
is called to update the hardware definers. If lag_active_port_bits() now
fails due to ENOMEM, we skip sending the active port command.

This error also propagates up to mlx5_modify_lag(), which then jumps to
the error path and skips updating the software state:

mlx5_modify_lag() {
        ...
        err = _mlx5_modify_lag(ldev, tracker, ports);
        if (err) {
                mlx5_core_err(dev0, "Failed to modify LAG (%d)\n", err);
                goto out;
        }
        memcpy(ldev->v2p_map, ports, ...);
        ...
}

Should there be a rollback mechanism for mlx5_lag_port_sel_modify() if
we hit this new error path?

-- 
Sashiko AI review · https://sashiko.dev/#/patchset/20260923103830.1183-1-tariqt@nvidia.com?part=4

  reply	other threads:[~2026-09-24 10:41 UTC|newest]

Thread overview: 39+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-09-23 10:38 [PATCH net-next 00/13] net/mlx5: Preparations for nested E-switch Tariq Toukan
2026-09-23 10:38 ` [PATCH net-next 01/13] net/mlx5e: Assign a random MAC to any netdev with a zero MAC address Tariq Toukan
2026-09-24 10:41   ` sashiko-bot
2026-09-24 17:45   ` netdev-bot+sashiko
2026-09-23 10:38 ` [PATCH net-next 02/13] net/mlx5: E-switch, do not leave an unpaired devcom registered Tariq Toukan
2026-09-24 10:41   ` sashiko-bot
2026-09-24 17:45   ` netdev-bot+sashiko
2026-09-23 10:38 ` [PATCH net-next 03/13] net/mlx5: LAG, allocate v2p_map dynamically Tariq Toukan
2026-09-24 10:41   ` sashiko-bot
2026-09-24 17:45   ` netdev-bot+sashiko
2026-09-28 14:56     ` Shay Drori
2026-09-23 10:38 ` [PATCH net-next 04/13] net/mlx5: LAG, allocate port-indexed scratch buffers dynamically Tariq Toukan
2026-09-24 10:41   ` sashiko-bot [this message]
2026-09-24 17:45   ` netdev-bot+sashiko
2026-09-28 14:57     ` Shay Drori
2026-09-23 10:38 ` [PATCH net-next 05/13] net/mlx5: LAG, drop per-port scratch array in drop-rule setup Tariq Toukan
2026-09-24 10:41   ` sashiko-bot
2026-09-23 10:38 ` [PATCH net-next 06/13] net/mlx5: LAG, size debugfs buffers by port count Tariq Toukan
2026-09-24 10:41   ` sashiko-bot
2026-09-24 17:45   ` netdev-bot+sashiko
2026-09-23 10:38 ` [PATCH net-next 07/13] net/mlx5e: TC, anchor peer-flow reverse index on the duplicated flow Tariq Toukan
2026-09-24 10:41   ` sashiko-bot
2026-09-24 17:45   ` netdev-bot+sashiko
2026-09-23 10:38 ` [PATCH net-next 08/13] net/mlx5e: TC, track peer flows in a vhca_id xarray Tariq Toukan
2026-09-24 10:41   ` sashiko-bot
2026-09-24 17:46   ` netdev-bot+sashiko
2026-09-23 10:38 ` [PATCH net-next 09/13] net/mlx5: E-switch, derive manager vport from device capability Tariq Toukan
2026-09-24 10:41   ` sashiko-bot
2026-09-24 17:46   ` netdev-bot+sashiko
2026-09-28 15:13     ` Shay Drori
2026-09-23 10:38 ` [PATCH net-next 10/13] net/mlx5: LAG, don't print port mapping to debugfs in MPESW mode Tariq Toukan
2026-09-24 10:41   ` sashiko-bot
2026-09-23 10:38 ` [PATCH net-next 11/13] net/mlx5: LAG, drop stale esw_shared_ingress_acl gate from shared FDB Tariq Toukan
2026-09-24 10:41   ` sashiko-bot
2026-09-23 10:38 ` [PATCH net-next 12/13] net/mlx5: E-switch, correct stale VF/PF wording in esw-allowed comments Tariq Toukan
2026-09-24 10:41   ` sashiko-bot
2026-09-23 10:38 ` [PATCH net-next 13/13] net/mlx5: E-switch, disable host functions for a non PF e-switch manager Tariq Toukan
2026-09-24 10:41   ` sashiko-bot
2026-09-29  0:20 ` [PATCH net-next 00/13] net/mlx5: Preparations for nested E-switch patchwork-bot+netdevbpf

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=20260924104142.8368F1F0089C@smtp.kernel.org \
    --to=sashiko-bot@kernel.org \
    --cc=linux-rdma@vger.kernel.org \
    --cc=sashiko-reviews@lists.linux.dev \
    --cc=tariqt@nvidia.com \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox