From: sashiko-bot@kernel.org
To: "Tariq Toukan" <tariqt@nvidia.com>
Cc: linux-rdma@vger.kernel.org
Subject: Re: [PATCH net-next 05/10] net/mlx5e: Add data direct RX infrastructure
Date: Fri, 09 Oct 2026 13:31:04 +0000 [thread overview]
Message-ID: <sashiko-outbox-165639@kernel.org> (raw)
In-Reply-To: <20261008132815.654147-6-tariqt@nvidia.com>
Thank you for your contribution! Sashiko AI review found 2 potential issue(s) to consider:
- [High] net/mlx5e: data direct with SHAMPO disabled causes peer-to-peer DMA to use host memory addresses
- [High] net/mlx5e: MKey PD mismatch for wqe_overflow causes HW protection fault
--
commit 6520b18194c553450bad712fcea9bc323c0c23ce
Author: Dragos Tatulea <dtatulea@nvidia.com>
net/mlx5e: Add data direct RX infrastructure
This commit adds RX infrastructure for data direct, which enables
peer-to-peer DMA between the NIC and another device like a GPU. It
introduces a new RQ state flag, a crossing mkey, and a KSM mkey
pointing to the crossed mkey.
> diff --git a/drivers/net/ethernet/mellanox/mlx5/core/en_main.c b/drivers/net/ethernet/mellanox/mlx5/core/en_main.c
> index b86c7b52774a0..e0d46e067c720 100644
> --- a/drivers/net/ethernet/mellanox/mlx5/core/en_main.c
> +++ b/drivers/net/ethernet/mellanox/mlx5/core/en_main.c
[ ... ]
> @@ -546,13 +548,80 @@ static int mlx5e_create_rq_umr_mkey(struct mlx5_core_dev *mdev, struct mlx5e_rq
> __func__, wq_size, rq->mpwqe.mtts_per_wqe,
> max_num_entries);
>
> + pdn = dd ? mdev->data_direct->pdn : mdev->mlx5e_res.hw_objs.pdn;
> err = mlx5e_create_umr_mkey(mdev, num_entries, rq->mpwqe.page_shift,
> &umr_mkey, rq->wqe_overflow.addr,
> - rq->mpwqe.umr_mode, xsk_chunk_size);
> + rq->mpwqe.umr_mode, pdn,
> + xsk_chunk_size);
> rq->mpwqe.umr_mkey_be = cpu_to_be32(umr_mkey);
[Severity: High]
Could this lead to a hardware protection fault due to a protection domain
mismatch?
When data direct is enabled, the indirect umr_mkey is explicitly created in
the data direct pdn. However, the gap (padding) entries for oversized
packets are initialized in mlx5e_create_umr_mkey() using an mkey that belongs
to the hardware objects protection domain:
ksm[i] = (struct mlx5_ksm) {
.key = cpu_to_be32(mdev->mlx5e_res.hw_objs.mkey),
.va = cpu_to_be64(filler_addr),
};
Because mlx5e_alloc_rx_mpwqe() issues a UMR WQE that only overwrites the valid
pages_per_wqe entries, these gap entries between WQEs remain untouched. If an
oversized packet overflows into the gap, the hardware will access the KSM and
encounter a mismatch between the indirect umr_mkey and the target
hw_objs.mkey.
Would this cause the hardware to drop the queue into an error state,
resulting in a potential denial of service from remotely sent oversized
packets?
> diff --git a/drivers/net/ethernet/mellanox/mlx5/core/en_rx.c b/drivers/net/ethernet/mellanox/mlx5/core/en_rx.c
> index 9e1224930e067..c992da08caaee 100644
> --- a/drivers/net/ethernet/mellanox/mlx5/core/en_rx.c
> +++ b/drivers/net/ethernet/mellanox/mlx5/core/en_rx.c
[ ... ]
> @@ -669,9 +669,15 @@ static int mlx5e_alloc_rx_mpwqe(struct mlx5e_rq *rq, u16 ix)
> goto err_unmap;
>
> addr = page_pool_get_dma_addr_netmem(frag_page->netmem);
> - umr_wqe->inline_mtts[i] = (struct mlx5_mtt) {
> - .ptag = cpu_to_be64(addr | MLX5_EN_WR),
> - };
> + if (test_bit(MLX5E_RQ_STATE_DATA_DIRECT, &rq->state))
> + umr_wqe->inline_ksms[i] = (struct mlx5_ksm) {
> + .key = rq->mpwqe.ksm_child_mkey_be,
> + .va = cpu_to_be64(addr),
> + };
> + else
> + umr_wqe->inline_mtts[i] = (struct mlx5_mtt) {
> + .ptag = cpu_to_be64(addr | MLX5_EN_WR),
> + };
> }
[Severity: High]
Can this write out of bounds and corrupt peer device memory?
If the administrator enables data direct but later disables hardware GRO
(SHAMPO) via ethtool, or if it is otherwise disabled, mlx5e_alloc_rq() does
not set PP_FLAG_ALLOW_UNREADABLE_NETMEM during page pool creation.
Without this flag, page_pool_create() ignores the dmabuf memory provider
and allocates standard host memory pages instead. Here in
mlx5e_alloc_rx_mpwqe(), the function fetches a host page, retrieves its host
DMA address, and configures the hardware using the data direct crossing mkey
along with the host DMA address as the virtual address (va).
Since the crossing mkey points to the peer device memory (such as a GPU),
the hardware will treat the large host DMA address as an offset into the
peer memory. Will this lead to out-of-bounds peer-to-peer DMA writes that
corrupt the GPU memory or trigger an IOMMU fault?
--
Sashiko AI review · https://sashiko.dev/#/patchset/20261008132815.654147-1-tariqt@nvidia.com?part=5
next prev parent reply other threads:[~2026-10-09 13:31 UTC|newest]
Thread overview: 22+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-10-08 13:28 [PATCH net-next 00/10] net/mlx5e: Add netdev support for data direct Tariq Toukan
2026-10-08 13:28 ` [PATCH net-next 01/10] net/mlx5: Log the data direct to PF device mapping Tariq Toukan
2026-10-09 13:31 ` sashiko-bot
2026-10-08 13:28 ` [PATCH net-next 02/10] net/mlx5e: Register supported netdevs as data direct users Tariq Toukan
2026-10-09 13:31 ` sashiko-bot
2026-10-08 13:28 ` [PATCH net-next 03/10] net/mlx5e: Pre-calculate UMR padding and entry size Tariq Toukan
2026-10-09 13:31 ` sashiko-bot
2026-10-08 13:28 ` [PATCH net-next 04/10] net/mlx5e: Add data direct ethtool private flag Tariq Toukan
2026-10-09 13:31 ` sashiko-bot
2026-10-08 13:28 ` [PATCH net-next 05/10] net/mlx5e: Add data direct RX infrastructure Tariq Toukan
2026-10-09 13:31 ` sashiko-bot [this message]
2026-10-08 13:28 ` [PATCH net-next 06/10] net/mlx5e: Add data direct TX infrastructure Tariq Toukan
2026-10-09 13:31 ` sashiko-bot
2026-10-08 13:28 ` [PATCH net-next 07/10] net/mlx5e: Use the correct DMA dev when data_direct pdev enabled Tariq Toukan
2026-10-09 13:31 ` sashiko-bot
2026-10-08 13:28 ` [PATCH net-next 08/10] net/mlx5e: Recreate netdev channels on data direct device unbind Tariq Toukan
2026-10-09 13:31 ` sashiko-bot
2026-10-08 13:28 ` [PATCH net-next 09/10] net: devmem: add netdev_has_dmabuf_binding() helper Tariq Toukan
2026-10-09 13:31 ` sashiko-bot
2026-10-08 13:28 ` [PATCH net-next 10/10] net/mlx5e: Enable the data direct netdev feature Tariq Toukan
2026-10-09 13:31 ` sashiko-bot
2026-10-09 16:42 ` [PATCH net-next 00/10] net/mlx5e: Add netdev support for data direct Mina Almasry
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=sashiko-outbox-165639@kernel.org \
--to=sashiko-bot@kernel.org \
--cc=linux-rdma@vger.kernel.org \
--cc=sashiko-reviews@lists.linux.dev \
--cc=tariqt@nvidia.com \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox