From: Or Har-Toov <ohartoov@nvidia.com>
To: sashiko-reviews@lists.linux.dev, Tariq Toukan <tariqt@nvidia.com>
Cc: linux-rdma@vger.kernel.org
Subject: Re: [PATCH net] net/mlx5: Lag, only cache max_tx_speed that FW has not accepted
Date: Wed, 7 Oct 2026 09:38:51 +0300 [thread overview]
Message-ID: <ef3568b9-f5ca-421e-b719-4774e3bbccf7@nvidia.com> (raw)
In-Reply-To: <20261005070328.AAAFD1F000FF@smtp.kernel.org>
On 05/10/2026 10:03, sashiko-bot@kernel.org wrote:
> External email: Use caution opening links or attachments
>
>
> Thank you for your contribution! Sashiko AI review found 1 potential issue(s) to consider:
> - [High] net/mlx5: rate limit dropped due to zeroed max TX speed fallback cache
> --
>
> commit e0ff80bc97aeb275a4ecdbc780d58984bb55ac36
> Author: Or Har-Toov <ohartoov@nvidia.com>
>
> net/mlx5: Lag, only cache max_tx_speed that FW has not accepted
>
> vport->agg_max_tx_speed caches a max_tx_speed that could not be pushed
> to FW, to be applied by mlx5_esw_vport_enable() once the vport comes up.
>
> Link: https://lore.kernel.org/all/20260915015118.875210-1-kuba@kernel.org/
>
>> diff --git a/drivers/net/ethernet/mellanox/mlx5/core/eswitch.c b/drivers/net/ethernet/mellanox/mlx5/core/eswitch.c
>> index fc197d1dc9dfe..f8c220bf7b5ca 100644
>> --- a/drivers/net/ethernet/mellanox/mlx5/core/eswitch.c
>> +++ b/drivers/net/ethernet/mellanox/mlx5/core/eswitch.c
>> @@ -951,6 +951,8 @@ static void mlx5_esw_vport_set_max_tx_speed(struct mlx5_eswitch *esw,
>> mlx5_core_dbg(esw->dev,
>> "Failed to set vport %d speed %d, err=%d\n",
>> vport->vport, vport->agg_max_tx_speed, ret);
>> + else
>> + vport->agg_max_tx_speed = 0;
>> }
>
> [Severity: High]
> This patch sets agg_max_tx_speed to 0 on a successful push, but multiple
> fallback paths in the driver rely on this variable containing the actual
> configured speed if a query fails.
>
> For example, in mlx5_modify_vport_admin_state() in
> drivers/net/ethernet/mellanox/mlx5/core/vport.c, if the query for the max
> TX speed fails, the code falls back to the cache:
>
> err = mlx5_query_vport_max_tx_speed(mdev, opmod, vport,
> other_vport, &speed, NULL);
> if (err) {
> #ifdef CONFIG_MLX5_ESWITCH
> struct mlx5_vport *esw_vport;
>
> esw_vport = mlx5_eswitch_get_vport(mdev->priv.eswitch,
> vport);
> speed = IS_ERR(esw_vport) ? 0 :
> esw_vport->agg_max_tx_speed;
> ...
> MLX5_SET(modify_vport_state_in, in, max_tx_speed, speed);
>
> If a max TX speed was successfully configured, this patch clears
> agg_max_tx_speed to 0. If a later admin state modification triggers
> mlx5_modify_vport_admin_state() and the query fails, won't this fallback
> logic read 0 and inadvertently un-limit the vport by pushing 0 to the firmware?
>
> A similar fallback path that relies on this cache exists in
> mlx5_esw_adj_vport_modify() in
> drivers/net/ethernet/mellanox/mlx5/core/esw/adj_vport.c.
>
This code was removed in Revert "net/mlx5: E-Switch, preserve max tx
speed on vport state modification" -
https://git.kernel.org/pub/scm/linux/kernel/git/netdev/net.git/commit/?id=86b785b18f8f4a6b4a7392d1e2b29d6bb095011a
>> diff --git a/drivers/net/ethernet/mellanox/mlx5/core/lag/lag.c b/drivers/net/ethernet/mellanox/mlx5/core/lag/lag.c
>> index 3b34bec559e08..7cee78cb47df9 100644
>> --- a/drivers/net/ethernet/mellanox/mlx5/core/lag/lag.c
>> +++ b/drivers/net/ethernet/mellanox/mlx5/core/lag/lag.c
>> @@ -1502,17 +1502,21 @@ static void mlx5_lag_modify_device_vports_speed(struct mlx5_core_dev *mdev,
>> if (vport->vport == MLX5_VPORT_UPLINK)
>> continue;
>>
>> - vport->agg_max_tx_speed = speed;
>> -
>> - if (!vport->enabled)
>> + if (!vport->enabled) {
>> + vport->agg_max_tx_speed = speed;
>> continue;
>> + }
>>
>> ret = mlx5_modify_vport_max_tx_speed(mdev, op_mod,
>> vport->vport, true, speed);
>> - if (ret)
>> + if (ret) {
>> mlx5_core_dbg(mdev,
>> "Failed to set vport %d speed %d, err=%d\n",
>> vport->vport, speed, ret);
>> + vport->agg_max_tx_speed = speed;
>> + } else {
>> + vport->agg_max_tx_speed = 0;
>> + }
>> }
>
> [Severity: High]
> Since this successfully clears the cache to 0, wouldn't it break the fallback
> mechanisms mentioned above?
>
> Additionally, this logic appears to overload 0 to mean "no cached value".
> As seen in mlx5_esw_vport_enable() in
> drivers/net/ethernet/mellanox/mlx5/core/eswitch.c:
>
> if (vport->agg_max_tx_speed)
> mlx5_esw_vport_set_max_tx_speed(esw, vport);
>
> If a user explicitly configures a speed of 0 (unlimited) while the vport is
> disabled, wouldn't the cache value be 0, preventing the unlimited setting
> from being applied when the vport is later enabled?
>
Speed 0 means unsupported and not unlimited by PRM. User can't configure
max_tx_speed to be 0.
> --
> Sashiko AI review · https://sashiko.dev/#/patchset/20261004070246.215239-1-tariqt@nvidia.com?part=1
next prev parent reply other threads:[~2026-10-07 6:39 UTC|newest]
Thread overview: 8+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-10-04 7:02 [PATCH net] net/mlx5: Lag, only cache max_tx_speed that FW has not accepted Tariq Toukan
2026-10-04 7:08 ` netdev-bot+sinfo
2026-10-05 7:03 ` sashiko-bot
2026-10-07 6:38 ` Or Har-Toov [this message]
2026-10-05 7:42 ` netdev-bot+sashiko
2026-10-08 16:48 ` Jakub Kicinski
2026-10-08 17:41 ` Tariq Toukan
2026-10-08 19:20 ` Jakub Kicinski
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=ef3568b9-f5ca-421e-b719-4774e3bbccf7@nvidia.com \
--to=ohartoov@nvidia.com \
--cc=linux-rdma@vger.kernel.org \
--cc=sashiko-reviews@lists.linux.dev \
--cc=tariqt@nvidia.com \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox