Linux RDMA and InfiniBand development
 help / color / mirror / Atom feed
From: Or Har-Toov <ohartoov@nvidia.com>
To: sashiko-reviews@lists.linux.dev, Tariq Toukan <tariqt@nvidia.com>
Cc: linux-rdma@vger.kernel.org
Subject: Re: [PATCH net] net/mlx5: Lag, only cache max_tx_speed that FW has not accepted
Date: Wed, 7 Oct 2026 09:38:51 +0300	[thread overview]
Message-ID: <ef3568b9-f5ca-421e-b719-4774e3bbccf7@nvidia.com> (raw)
In-Reply-To: <20261005070328.AAAFD1F000FF@smtp.kernel.org>



On 05/10/2026 10:03, sashiko-bot@kernel.org wrote:
> External email: Use caution opening links or attachments
> 
> 
> Thank you for your contribution! Sashiko AI review found 1 potential issue(s) to consider:
> - [High] net/mlx5: rate limit dropped due to zeroed max TX speed fallback cache
> --
> 
> commit e0ff80bc97aeb275a4ecdbc780d58984bb55ac36
> Author: Or Har-Toov <ohartoov@nvidia.com>
> 
> net/mlx5: Lag, only cache max_tx_speed that FW has not accepted
> 
> vport->agg_max_tx_speed caches a max_tx_speed that could not be pushed
> to FW, to be applied by mlx5_esw_vport_enable() once the vport comes up.
> 
> Link: https://lore.kernel.org/all/20260915015118.875210-1-kuba@kernel.org/
> 
>> diff --git a/drivers/net/ethernet/mellanox/mlx5/core/eswitch.c b/drivers/net/ethernet/mellanox/mlx5/core/eswitch.c
>> index fc197d1dc9dfe..f8c220bf7b5ca 100644
>> --- a/drivers/net/ethernet/mellanox/mlx5/core/eswitch.c
>> +++ b/drivers/net/ethernet/mellanox/mlx5/core/eswitch.c
>> @@ -951,6 +951,8 @@ static void mlx5_esw_vport_set_max_tx_speed(struct mlx5_eswitch *esw,
>>                mlx5_core_dbg(esw->dev,
>>                              "Failed to set vport %d speed %d, err=%d\n",
>>                              vport->vport, vport->agg_max_tx_speed, ret);
>> +     else
>> +             vport->agg_max_tx_speed = 0;
>>   }
> 
> [Severity: High]
> This patch sets agg_max_tx_speed to 0 on a successful push, but multiple
> fallback paths in the driver rely on this variable containing the actual
> configured speed if a query fails.
> 
> For example, in mlx5_modify_vport_admin_state() in
> drivers/net/ethernet/mellanox/mlx5/core/vport.c, if the query for the max
> TX speed fails, the code falls back to the cache:
> 
>          err = mlx5_query_vport_max_tx_speed(mdev, opmod, vport,
>                                              other_vport, &speed, NULL);
>          if (err) {
> #ifdef CONFIG_MLX5_ESWITCH
>              struct mlx5_vport *esw_vport;
> 
>              esw_vport = mlx5_eswitch_get_vport(mdev->priv.eswitch,
>                                                 vport);
>              speed = IS_ERR(esw_vport) ? 0 :
>                  esw_vport->agg_max_tx_speed;
>          ...
>          MLX5_SET(modify_vport_state_in, in, max_tx_speed, speed);
> 
> If a max TX speed was successfully configured, this patch clears
> agg_max_tx_speed to 0. If a later admin state modification triggers
> mlx5_modify_vport_admin_state() and the query fails, won't this fallback
> logic read 0 and inadvertently un-limit the vport by pushing 0 to the firmware?
> 
> A similar fallback path that relies on this cache exists in
> mlx5_esw_adj_vport_modify() in
> drivers/net/ethernet/mellanox/mlx5/core/esw/adj_vport.c.
> 

This code was removed in Revert "net/mlx5: E-Switch, preserve max tx 
speed on vport state modification" -
https://git.kernel.org/pub/scm/linux/kernel/git/netdev/net.git/commit/?id=86b785b18f8f4a6b4a7392d1e2b29d6bb095011a

>> diff --git a/drivers/net/ethernet/mellanox/mlx5/core/lag/lag.c b/drivers/net/ethernet/mellanox/mlx5/core/lag/lag.c
>> index 3b34bec559e08..7cee78cb47df9 100644
>> --- a/drivers/net/ethernet/mellanox/mlx5/core/lag/lag.c
>> +++ b/drivers/net/ethernet/mellanox/mlx5/core/lag/lag.c
>> @@ -1502,17 +1502,21 @@ static void mlx5_lag_modify_device_vports_speed(struct mlx5_core_dev *mdev,
>>                if (vport->vport == MLX5_VPORT_UPLINK)
>>                        continue;
>>
>> -             vport->agg_max_tx_speed = speed;
>> -
>> -             if (!vport->enabled)
>> +             if (!vport->enabled) {
>> +                     vport->agg_max_tx_speed = speed;
>>                        continue;
>> +             }
>>
>>                ret = mlx5_modify_vport_max_tx_speed(mdev, op_mod,
>>                                                     vport->vport, true, speed);
>> -             if (ret)
>> +             if (ret) {
>>                        mlx5_core_dbg(mdev,
>>                                      "Failed to set vport %d speed %d, err=%d\n",
>>                                      vport->vport, speed, ret);
>> +                     vport->agg_max_tx_speed = speed;
>> +             } else {
>> +                     vport->agg_max_tx_speed = 0;
>> +             }
>>        }
> 
> [Severity: High]
> Since this successfully clears the cache to 0, wouldn't it break the fallback
> mechanisms mentioned above?
> 
> Additionally, this logic appears to overload 0 to mean "no cached value".
> As seen in mlx5_esw_vport_enable() in
> drivers/net/ethernet/mellanox/mlx5/core/eswitch.c:
> 
>      if (vport->agg_max_tx_speed)
>          mlx5_esw_vport_set_max_tx_speed(esw, vport);
> 
> If a user explicitly configures a speed of 0 (unlimited) while the vport is
> disabled, wouldn't the cache value be 0, preventing the unlimited setting
> from being applied when the vport is later enabled?
> 

Speed 0 means unsupported and not unlimited by PRM. User can't configure 
max_tx_speed to be 0.

> --
> Sashiko AI review · https://sashiko.dev/#/patchset/20261004070246.215239-1-tariqt@nvidia.com?part=1


  reply	other threads:[~2026-10-07  6:39 UTC|newest]

Thread overview: 8+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-10-04  7:02 [PATCH net] net/mlx5: Lag, only cache max_tx_speed that FW has not accepted Tariq Toukan
2026-10-04  7:08 ` netdev-bot+sinfo
2026-10-05  7:03 ` sashiko-bot
2026-10-07  6:38   ` Or Har-Toov [this message]
2026-10-05  7:42 ` netdev-bot+sashiko
2026-10-08 16:48   ` Jakub Kicinski
2026-10-08 17:41     ` Tariq Toukan
2026-10-08 19:20       ` Jakub Kicinski

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=ef3568b9-f5ca-421e-b719-4774e3bbccf7@nvidia.com \
    --to=ohartoov@nvidia.com \
    --cc=linux-rdma@vger.kernel.org \
    --cc=sashiko-reviews@lists.linux.dev \
    --cc=tariqt@nvidia.com \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox