Linux RDMA and InfiniBand development
 help / color / mirror / Atom feed
From: sashiko-bot@kernel.org
To: "Tariq Toukan" <tariqt@nvidia.com>
Cc: linux-rdma@vger.kernel.org
Subject: Re: [PATCH net] net/mlx5: Lag, only cache max_tx_speed that FW has not accepted
Date: Mon, 05 Oct 2026 07:03:28 +0000	[thread overview]
Message-ID: <20261005070328.AAAFD1F000FF@smtp.kernel.org> (raw)
In-Reply-To: <20261004070246.215239-1-tariqt@nvidia.com>

Thank you for your contribution! Sashiko AI review found 1 potential issue(s) to consider:
- [High] net/mlx5: rate limit dropped due to zeroed max TX speed fallback cache
--

commit e0ff80bc97aeb275a4ecdbc780d58984bb55ac36
Author: Or Har-Toov <ohartoov@nvidia.com>

net/mlx5: Lag, only cache max_tx_speed that FW has not accepted

vport->agg_max_tx_speed caches a max_tx_speed that could not be pushed
to FW, to be applied by mlx5_esw_vport_enable() once the vport comes up.

Link: https://lore.kernel.org/all/20260915015118.875210-1-kuba@kernel.org/

> diff --git a/drivers/net/ethernet/mellanox/mlx5/core/eswitch.c b/drivers/net/ethernet/mellanox/mlx5/core/eswitch.c
> index fc197d1dc9dfe..f8c220bf7b5ca 100644
> --- a/drivers/net/ethernet/mellanox/mlx5/core/eswitch.c
> +++ b/drivers/net/ethernet/mellanox/mlx5/core/eswitch.c
> @@ -951,6 +951,8 @@ static void mlx5_esw_vport_set_max_tx_speed(struct mlx5_eswitch *esw,
>  		mlx5_core_dbg(esw->dev,
>  			      "Failed to set vport %d speed %d, err=%d\n",
>  			      vport->vport, vport->agg_max_tx_speed, ret);
> +	else
> +		vport->agg_max_tx_speed = 0;
>  }

[Severity: High]
This patch sets agg_max_tx_speed to 0 on a successful push, but multiple
fallback paths in the driver rely on this variable containing the actual
configured speed if a query fails.

For example, in mlx5_modify_vport_admin_state() in
drivers/net/ethernet/mellanox/mlx5/core/vport.c, if the query for the max
TX speed fails, the code falls back to the cache:

        err = mlx5_query_vport_max_tx_speed(mdev, opmod, vport,
                                            other_vport, &speed, NULL);
        if (err) {
#ifdef CONFIG_MLX5_ESWITCH
            struct mlx5_vport *esw_vport;

            esw_vport = mlx5_eswitch_get_vport(mdev->priv.eswitch,
                                               vport);
            speed = IS_ERR(esw_vport) ? 0 :
                esw_vport->agg_max_tx_speed;
        ...
        MLX5_SET(modify_vport_state_in, in, max_tx_speed, speed);

If a max TX speed was successfully configured, this patch clears
agg_max_tx_speed to 0. If a later admin state modification triggers
mlx5_modify_vport_admin_state() and the query fails, won't this fallback
logic read 0 and inadvertently un-limit the vport by pushing 0 to the firmware?

A similar fallback path that relies on this cache exists in
mlx5_esw_adj_vport_modify() in
drivers/net/ethernet/mellanox/mlx5/core/esw/adj_vport.c.

> diff --git a/drivers/net/ethernet/mellanox/mlx5/core/lag/lag.c b/drivers/net/ethernet/mellanox/mlx5/core/lag/lag.c
> index 3b34bec559e08..7cee78cb47df9 100644
> --- a/drivers/net/ethernet/mellanox/mlx5/core/lag/lag.c
> +++ b/drivers/net/ethernet/mellanox/mlx5/core/lag/lag.c
> @@ -1502,17 +1502,21 @@ static void mlx5_lag_modify_device_vports_speed(struct mlx5_core_dev *mdev,
>  		if (vport->vport == MLX5_VPORT_UPLINK)
>  			continue;
>  
> -		vport->agg_max_tx_speed = speed;
> -
> -		if (!vport->enabled)
> +		if (!vport->enabled) {
> +			vport->agg_max_tx_speed = speed;
>  			continue;
> +		}
>  
>  		ret = mlx5_modify_vport_max_tx_speed(mdev, op_mod,
>  						     vport->vport, true, speed);
> -		if (ret)
> +		if (ret) {
>  			mlx5_core_dbg(mdev,
>  				      "Failed to set vport %d speed %d, err=%d\n",
>  				      vport->vport, speed, ret);
> +			vport->agg_max_tx_speed = speed;
> +		} else {
> +			vport->agg_max_tx_speed = 0;
> +		}
>  	}

[Severity: High]
Since this successfully clears the cache to 0, wouldn't it break the fallback
mechanisms mentioned above?

Additionally, this logic appears to overload 0 to mean "no cached value".
As seen in mlx5_esw_vport_enable() in
drivers/net/ethernet/mellanox/mlx5/core/eswitch.c:

    if (vport->agg_max_tx_speed)
        mlx5_esw_vport_set_max_tx_speed(esw, vport);

If a user explicitly configures a speed of 0 (unlimited) while the vport is
disabled, wouldn't the cache value be 0, preventing the unlimited setting
from being applied when the vport is later enabled?

-- 
Sashiko AI review · https://sashiko.dev/#/patchset/20261004070246.215239-1-tariqt@nvidia.com?part=1

  parent reply	other threads:[~2026-10-05  7:03 UTC|newest]

Thread overview: 8+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-10-04  7:02 [PATCH net] net/mlx5: Lag, only cache max_tx_speed that FW has not accepted Tariq Toukan
2026-10-04  7:08 ` netdev-bot+sinfo
2026-10-05  7:03 ` sashiko-bot [this message]
2026-10-07  6:38   ` Or Har-Toov
2026-10-05  7:42 ` netdev-bot+sashiko
2026-10-08 16:48   ` Jakub Kicinski
2026-10-08 17:41     ` Tariq Toukan
2026-10-08 19:20       ` Jakub Kicinski

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=20261005070328.AAAFD1F000FF@smtp.kernel.org \
    --to=sashiko-bot@kernel.org \
    --cc=linux-rdma@vger.kernel.org \
    --cc=sashiko-reviews@lists.linux.dev \
    --cc=tariqt@nvidia.com \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox