linux-rdma.vger.kernel.org archive mirror
 help / color / mirror / Atom feed
From: Yishai Hadas <yishaih@nvidia.com>
To: <jgg@ziepe.ca>, <leon@kernel.org>
Cc: <linux-rdma@vger.kernel.org>, <selvin.xavier@broadcom.com>,
	<kalesh-anakkur.purayil@broadcom.com>,
	<chengyou@linux.alibaba.com>, <kaishen@linux.alibaba.com>,
	<tangchengchang@huawei.com>, <huangjunxian6@hisilicon.com>,
	<abhijit.gangurde@amd.com>, <allen.hubbe@amd.com>,
	<longli@microsoft.com>, <kotaranov@microsoft.com>,
	<mkalderon@marvell.com>, <bryan-bt.tan@broadcom.com>,
	<vishnu.dasa@broadcom.com>, <yishaih@nvidia.com>,
	<maorg@nvidia.com>
Subject: [PATCH rdma-next 03/15] RDMA/vmw_pvrdma: Pin QP and SRQ rings writable to match device DMA write access
Date: Tue, 8 Sep 2026 18:28:39 +0300	[thread overview]
Message-ID: <20260908152851.1307294-4-yishaih@nvidia.com> (raw)
In-Reply-To: <20260908152851.1307294-1-yishaih@nvidia.com>

pvrdma_create_qp() and pvrdma_create_srq() pin their ring buffers with
access=0. PVRDMA embeds a pvrdma_ring_state header directly inside these
buffers:

  qp->sq.ring = qp->pdir.pages[0];
  qp->rq.ring = is_srq ? NULL : &qp->sq.ring[1];

and the hypervisor writes cons_head into that header -- a field the
driver never writes itself outside of QP reset. With access=0,
ib_access_writable() is false, so ib_umem_get_va() pins the pages
without FOLL_WRITE and does not mark them dirty on unpin.

Without FOLL_WRITE, pin_user_pages_fast() can hand back a page this
process does not exclusively own (e.g. the shared zero page for an
untouched anonymous mapping) instead of forcing a private copy. The
hypervisor is then free to write cons_head into that shared physical
page, corrupting memory visible to every other mapper of it. Missing the
dirty mark on unpin also risks the hypervisor's writes being silently
discarded on reclaim.

This is not gated by any IOMMU permission check: PVRDMA is a fully
emulated device, not a real PCIe device sitting behind a guest-facing
IOMMU, so the DMA mapping direction has no enforcement effect here.  The
hypervisor already owns the guest's entire physical address space and
writes to it directly, without going through this process's page tables
or any virtual-address permission check -- the CPU's page protection
bits only gate CPU-issued load/store instructions via
virtual-to-physical translation, not a hypervisor writing to guest RAM
it already controls. FOLL_WRITE is a one-time decision made at pin time
about which physical page GUP hands back for this VA; once that page is
chosen (or the wrong one is chosen, as happens today), nothing stops a
later write to it.

PVRDMA's SRQ ring does not currently wire up its ring_state field (no
post_srq_recv is implemented), so nothing depends on it being written by
the hypervisor today. Pin it writable anyway, both for symmetry with the
QP rings this same driver embeds ring state in, and because the umem
pinning here has no way to know if that will still be true tomorrow.

Pass IB_ACCESS_LOCAL_WRITE for all three rings so they are pinned
consistently with how the hypervisor actually uses them.

Fixes: 29c8d9eba550 ("IB: Add vmw_pvrdma driver")
Signed-off-by: Yishai Hadas <yishaih@nvidia.com>
---
 drivers/infiniband/hw/vmw_pvrdma/pvrdma_qp.c  | 6 ++++--
 drivers/infiniband/hw/vmw_pvrdma/pvrdma_srq.c | 3 ++-
 2 files changed, 6 insertions(+), 3 deletions(-)

diff --git a/drivers/infiniband/hw/vmw_pvrdma/pvrdma_qp.c b/drivers/infiniband/hw/vmw_pvrdma/pvrdma_qp.c
index e939cd5ce40b..53cc49f2b8c9 100644
--- a/drivers/infiniband/hw/vmw_pvrdma/pvrdma_qp.c
+++ b/drivers/infiniband/hw/vmw_pvrdma/pvrdma_qp.c
@@ -270,7 +270,8 @@ int pvrdma_create_qp(struct ib_qp *ibqp, struct ib_qp_init_attr *init_attr,
 				/* set qp->sq.wqe_cnt, shift, buf_size.. */
 				qp->rumem = ib_umem_get_va(ibqp->device,
 							   ucmd.rbuf_addr,
-							   ucmd.rbuf_size, 0);
+							   ucmd.rbuf_size,
+							   IB_ACCESS_LOCAL_WRITE);
 				if (IS_ERR(qp->rumem)) {
 					ret = PTR_ERR(qp->rumem);
 					goto err_qp;
@@ -282,7 +283,8 @@ int pvrdma_create_qp(struct ib_qp *ibqp, struct ib_qp_init_attr *init_attr,
 			}
 
 			qp->sumem = ib_umem_get_va(ibqp->device, ucmd.sbuf_addr,
-						   ucmd.sbuf_size, 0);
+						   ucmd.sbuf_size,
+						   IB_ACCESS_LOCAL_WRITE);
 			if (IS_ERR(qp->sumem)) {
 				if (!is_srq)
 					ib_umem_release(qp->rumem);
diff --git a/drivers/infiniband/hw/vmw_pvrdma/pvrdma_srq.c b/drivers/infiniband/hw/vmw_pvrdma/pvrdma_srq.c
index 345ec486a223..3252c2eb405a 100644
--- a/drivers/infiniband/hw/vmw_pvrdma/pvrdma_srq.c
+++ b/drivers/infiniband/hw/vmw_pvrdma/pvrdma_srq.c
@@ -146,7 +146,8 @@ int pvrdma_create_srq(struct ib_srq *ibsrq, struct ib_srq_init_attr *init_attr,
 	if (ret)
 		goto err_srq;
 
-	srq->umem = ib_umem_get_va(ibsrq->device, ucmd.buf_addr, ucmd.buf_size, 0);
+	srq->umem = ib_umem_get_va(ibsrq->device, ucmd.buf_addr, ucmd.buf_size,
+				   IB_ACCESS_LOCAL_WRITE);
 	if (IS_ERR(srq->umem)) {
 		ret = PTR_ERR(srq->umem);
 		goto err_srq;
-- 
2.18.1


  parent reply	other threads:[~2026-09-08 15:30 UTC|newest]

Thread overview: 16+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-09-08 15:28 [PATCH rdma-next 00/15] DMA direction and mlx5 driver correctness fixes Yishai Hadas
2026-09-08 15:28 ` [PATCH rdma-next 01/15] RDMA/erdma: Pin CQ buffer writable to match device DMA write access Yishai Hadas
2026-09-08 15:28 ` [PATCH rdma-next 02/15] RDMA/hns: " Yishai Hadas
2026-09-08 15:28 ` Yishai Hadas [this message]
2026-09-08 15:28 ` [PATCH rdma-next 04/15] RDMA/umem: Reuse ib_umem_get_cq_buf_or_va() for VA-only CQ pinning Yishai Hadas
2026-09-08 15:28 ` [PATCH rdma-next 05/15] RDMA/umem: Support an explicit DMA direction other than DMA_BIDIRECTIONAL Yishai Hadas
2026-09-08 15:28 ` [PATCH rdma-next 06/15] RDMA/umem: Map CQ buffers DMA_FROM_DEVICE Yishai Hadas
2026-09-08 15:28 ` [PATCH rdma-next 07/15] RDMA/umem: Derive DMA direction from IB access flags Yishai Hadas
2026-09-08 15:28 ` [PATCH rdma-next 08/15] RDMA/mlx5: Fix mlx5_ib_dev_res_init() failure when XRC cap is absent Yishai Hadas
2026-09-08 15:28 ` [PATCH rdma-next 09/15] RDMA/mlx5: Put resource reference in mlx5_ib_wq_event() Yishai Hadas
2026-09-08 15:28 ` [PATCH rdma-next 10/15] RDMA/mlx5: Set WQ event handler before firmware RQ insertion Yishai Hadas
2026-09-08 15:28 ` [PATCH rdma-next 11/15] RDMA/mlx5: Set SRQ event handler before xarray insertion Yishai Hadas
2026-09-08 15:28 ` [PATCH rdma-next 12/15] RDMA/mlx5: Set RQ event handler for raw-packet QP Yishai Hadas
2026-09-08 15:28 ` [PATCH rdma-next 13/15] RDMA/mlx5: Set QP event handler before firmware QPC insertion Yishai Hadas
2026-09-08 15:28 ` [PATCH rdma-next 14/15] RDMA/mlx5: Initialize QP/RQ/SQ resource refcount before publishing it Yishai Hadas
2026-09-08 15:28 ` [PATCH rdma-next 15/15] RDMA/mlx5: Fix signed integer overflow in EQE qp_srq type shift Yishai Hadas

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=20260908152851.1307294-4-yishaih@nvidia.com \
    --to=yishaih@nvidia.com \
    --cc=abhijit.gangurde@amd.com \
    --cc=allen.hubbe@amd.com \
    --cc=bryan-bt.tan@broadcom.com \
    --cc=chengyou@linux.alibaba.com \
    --cc=huangjunxian6@hisilicon.com \
    --cc=jgg@ziepe.ca \
    --cc=kaishen@linux.alibaba.com \
    --cc=kalesh-anakkur.purayil@broadcom.com \
    --cc=kotaranov@microsoft.com \
    --cc=leon@kernel.org \
    --cc=linux-rdma@vger.kernel.org \
    --cc=longli@microsoft.com \
    --cc=maorg@nvidia.com \
    --cc=mkalderon@marvell.com \
    --cc=selvin.xavier@broadcom.com \
    --cc=tangchengchang@huawei.com \
    --cc=vishnu.dasa@broadcom.com \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox;
as well as URLs for NNTP newsgroup(s).