Linux RDMA and InfiniBand development
 help / color / mirror / Atom feed
From: sashiko-bot@kernel.org
To: "Tariq Toukan" <tariqt@nvidia.com>
Cc: linux-rdma@vger.kernel.org
Subject: Re: [PATCH net-next 06/10] net/mlx5e: Add data direct TX infrastructure
Date: Fri, 09 Oct 2026 13:31:04 +0000	[thread overview]
Message-ID: <sashiko-outbox-165640@kernel.org> (raw)
In-Reply-To: <20261008132815.654147-7-tariqt@nvidia.com>

Thank you for your contribution! Sashiko AI review found 1 potential issue(s) to consider:
- [High] net/mlx5e: GPU unbind causes DMA corruption for active devmem TX
--

commit 2beff8c1365001300351d8ed26d31fd617a37400
Author: Dragos Tatulea <dtatulea@nvidia.com>

net/mlx5e: Add data direct TX infrastructure

This commit introduces the necessary infrastructure for peer-to-peer DMA
between the NIC and another device (e.g., GPU) on the TX path. It creates
a crossing mkey linked to the SQ and adds a new SQ state flag to indicate
when Data Direct is enabled.

> diff --git a/drivers/net/ethernet/mellanox/mlx5/core/en_tx.c b/drivers/net/ethernet/mellanox/mlx5/core/en_tx.c
> index 14479da98f427..62ffb41a64dfe 100644
> --- a/drivers/net/ethernet/mellanox/mlx5/core/en_tx.c
> +++ b/drivers/net/ethernet/mellanox/mlx5/core/en_tx.c
> @@ -182,6 +182,7 @@ mlx5e_txwqe_build_dsegs(struct mlx5e_txqsq *sq, struct sk_buff *skb,
>  			unsigned char *skb_data, u16 headlen,
>  			struct mlx5_wqe_data_seg *dseg)
>  {
> +	bool dd = test_bit(MLX5E_SQ_STATE_DATA_DIRECT, &sq->state);
>  	dma_addr_t dma_addr = 0;
>  	u8 num_dma          = 0;
>  	int i;

[ ... ]

> @@ -211,7 +212,8 @@ mlx5e_txwqe_build_dsegs(struct mlx5e_txqsq *sq, struct sk_buff *skb,
>  			goto dma_unmap_wqe_err;
>  
>  		dseg->addr       = cpu_to_be64(dma_addr);
> -		dseg->lkey       = sq->mkey_be;
> +		dseg->lkey       = dd && skb_frag_is_net_iov(frag) ?
> +				   sq->crossing_mkey_be : sq->mkey_be;

[Severity: High]
If a peer GPU device unbinds while a Data Direct devmem DMABUF is mapped
for TX, the MLX5E_SQ_STATE_DATA_DIRECT flag is forcefully cleared.

Since the asynchronous unbind bypasses the devmem socket binding safety
checks, the application can continue transmitting net_iov fragments that
still contain GPU IOVAs.

When those fragments arrive here, dd is now false, so the code falls
back to using the NIC's local sq->mkey_be for those GPU IOVAs.

Could this cause the NIC to read incorrect local memory or trigger host
IOMMU protection faults? How should active net_iov transmissions be explicitly
fenced or dropped when the queue state reverts?

>  		dseg->byte_count = cpu_to_be32(fsz);
>  
>  		mlx5e_dma_push_netmem(sq, skb_frag_netmem(frag), dma_addr, fsz);

-- 
Sashiko AI review · https://sashiko.dev/#/patchset/20261008132815.654147-1-tariqt@nvidia.com?part=6

  reply	other threads:[~2026-10-09 13:31 UTC|newest]

Thread overview: 22+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-10-08 13:28 [PATCH net-next 00/10] net/mlx5e: Add netdev support for data direct Tariq Toukan
2026-10-08 13:28 ` [PATCH net-next 01/10] net/mlx5: Log the data direct to PF device mapping Tariq Toukan
2026-10-09 13:31   ` sashiko-bot
2026-10-08 13:28 ` [PATCH net-next 02/10] net/mlx5e: Register supported netdevs as data direct users Tariq Toukan
2026-10-09 13:31   ` sashiko-bot
2026-10-08 13:28 ` [PATCH net-next 03/10] net/mlx5e: Pre-calculate UMR padding and entry size Tariq Toukan
2026-10-09 13:31   ` sashiko-bot
2026-10-08 13:28 ` [PATCH net-next 04/10] net/mlx5e: Add data direct ethtool private flag Tariq Toukan
2026-10-09 13:31   ` sashiko-bot
2026-10-08 13:28 ` [PATCH net-next 05/10] net/mlx5e: Add data direct RX infrastructure Tariq Toukan
2026-10-09 13:31   ` sashiko-bot
2026-10-08 13:28 ` [PATCH net-next 06/10] net/mlx5e: Add data direct TX infrastructure Tariq Toukan
2026-10-09 13:31   ` sashiko-bot [this message]
2026-10-08 13:28 ` [PATCH net-next 07/10] net/mlx5e: Use the correct DMA dev when data_direct pdev enabled Tariq Toukan
2026-10-09 13:31   ` sashiko-bot
2026-10-08 13:28 ` [PATCH net-next 08/10] net/mlx5e: Recreate netdev channels on data direct device unbind Tariq Toukan
2026-10-09 13:31   ` sashiko-bot
2026-10-08 13:28 ` [PATCH net-next 09/10] net: devmem: add netdev_has_dmabuf_binding() helper Tariq Toukan
2026-10-09 13:31   ` sashiko-bot
2026-10-08 13:28 ` [PATCH net-next 10/10] net/mlx5e: Enable the data direct netdev feature Tariq Toukan
2026-10-09 13:31   ` sashiko-bot
2026-10-09 16:42 ` [PATCH net-next 00/10] net/mlx5e: Add netdev support for data direct Mina Almasry

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=sashiko-outbox-165640@kernel.org \
    --to=sashiko-bot@kernel.org \
    --cc=linux-rdma@vger.kernel.org \
    --cc=sashiko-reviews@lists.linux.dev \
    --cc=tariqt@nvidia.com \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox