From: sashiko-bot@kernel.org
To: "Tariq Toukan" <tariqt@nvidia.com>
Cc: linux-rdma@vger.kernel.org
Subject: Re: [PATCH net-next 4/6] net/mlx5: HWS, Add timeout mechanism to draining send queues
Date: Wed, 16 Sep 2026 11:20:29 +0000 [thread overview]
Message-ID: <20260916112030.145DD1F00899@smtp.kernel.org> (raw)
In-Reply-To: <20260915111912.3926585-5-tariqt@nvidia.com>
Thank you for your contribution! Sashiko AI review found 2 potential issue(s) to consider:
- [High] mlx5: HW UAF when queue drain timeout leaves pending WQEs
- [Medium] mlx5: false timeout error when send queue drains exactly at timeout boundary
--
commit 61f83bb352e1062985b4a65a41c453e002f769ea
Author: Yevgeny Kliteynik <kliteyn@nvidia.com>
net/mlx5: HWS, Add timeout mechanism to draining send queues
This commit introduces a timeout mechanism for draining send queues in the
hardware steering (HWS) layer, returning an error if polling exceeds the
timeout limit.
[ ... ]
> diff --git a/drivers/net/ethernet/mellanox/mlx5/core/steering/hws/send.c b/drivers/net/ethernet/mellanox/mlx5/core/steering/hws/send.c
> index 49dee44d0bef5..9ff109fc8230a 100644
> --- a/drivers/net/ethernet/mellanox/mlx5/core/steering/hws/send.c
> +++ b/drivers/net/ethernet/mellanox/mlx5/core/steering/hws/send.c
[ ... ]
> @@ -1213,9 +1215,18 @@ int mlx5hws_send_queue_action(struct mlx5hws_context *ctx,
> mlx5hws_send_engine_flush_queue(queue);
>
> /* Poll queue until empty */
> - while (wait_comp && !mlx5hws_send_engine_empty(queue))
> + while (wait_comp && !mlx5hws_send_engine_empty(queue)) {
> hws_send_engine_poll_cq(queue, NULL, &polled, 0);
>
> + if (unlikely(time_after(jiffies, timeout))) {
[Severity: Medium]
If the queue successfully drains on this exact iteration, or if the thread
is preempted after polling and before the timeout check, does this
unconditionally return -ETIMEDOUT?
Since hws_send_engine_poll_cq() processes the completions, checking the
timeout immediately without verifying if the queue is now empty might
cause callers to improperly abort and roll back fully successful operations.
> + mlx5hws_err(ctx,
> + "Error draining send queue %d - TIMEOUT (%d sec)\n",
> + queue_id,
> + MLX5HWS_BWC_POLLING_TIMEOUT);
> + return -ETIMEDOUT;
[Severity: High]
When returning -ETIMEDOUT here, the send queue still contains pending
hardware WQEs that are not cleared or aborted. Can this lead to a
hardware use-after-free?
Callers react to this error by destroying the allocated hardware resources.
For example, if mlx5hws_arg_create() receives the timeout error:
mlx5hws_arg_create()
mlx5hws_arg_write_inline_arg_data()
mlx5hws_send_queue_action() <-- times out
if (err)
mlx5hws_cmd_arg_destroy(ctx->mdev, id); // id is freed
Because the hardware queue is not stopped and the pending WQEs are not
aborted, couldn't the firmware later process the stuck WQE and write to
the freed resource ID, potentially corrupting firmware state or crashing?
> + }
> + }
> +
> break;
> default:
> return -EINVAL;
--
Sashiko AI review · https://sashiko.dev/#/patchset/20260915111912.3926585-1-tariqt@nvidia.com?part=4
next prev parent reply other threads:[~2026-09-16 11:20 UTC|newest]
Thread overview: 19+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-09-15 11:19 [PATCH net-next 0/6] net/mlx5: more HWS misc enhancements Tariq Toukan
2026-09-15 11:19 ` [PATCH net-next 1/6] net/mlx5: HWS, Return meaningful error code in hws_send_wqe_fw Tariq Toukan
2026-09-16 11:20 ` sashiko-bot
2026-09-15 11:19 ` [PATCH net-next 2/6] net/mlx5: HWS, Fix error message in mlx5hws_cmd_generate_wqe Tariq Toukan
2026-09-16 11:20 ` sashiko-bot
2026-09-16 23:42 ` netdev-bot+sashiko
2026-09-15 11:19 ` [PATCH net-next 3/6] net/mlx5: HWS, Replace kzalloc with kzalloc_obj Tariq Toukan
2026-09-16 11:20 ` sashiko-bot
2026-09-16 23:42 ` netdev-bot+sashiko
2026-09-15 11:19 ` [PATCH net-next 4/6] net/mlx5: HWS, Add timeout mechanism to draining send queues Tariq Toukan
2026-09-16 11:20 ` sashiko-bot [this message]
2026-09-16 23:42 ` netdev-bot+sashiko
2026-09-15 11:19 ` [PATCH net-next 5/6] net/mlx5: HWS, Handle timeout draining the send queue for FW STEs Tariq Toukan
2026-09-16 11:20 ` sashiko-bot
2026-09-16 23:42 ` netdev-bot+sashiko
2026-09-15 11:19 ` [PATCH net-next 6/6] net/mlx5: HWS, Remove unneeded WRITE_ONCE in post send Tariq Toukan
2026-09-16 11:20 ` sashiko-bot
2026-09-16 23:42 ` netdev-bot+sashiko
2026-09-19 1:00 ` [PATCH net-next 0/6] net/mlx5: more HWS misc enhancements patchwork-bot+netdevbpf
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=20260916112030.145DD1F00899@smtp.kernel.org \
--to=sashiko-bot@kernel.org \
--cc=linux-rdma@vger.kernel.org \
--cc=sashiko-reviews@lists.linux.dev \
--cc=tariqt@nvidia.com \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox