From: Pranjal Shrivastava <praan@google.com>
To: Matt Evans <matt@ozlabs.org>
Cc: "Alex Williamson" <alex@shazbot.org>,
"Leon Romanovsky" <leon@kernel.org>,
"Jason Gunthorpe" <jgg@nvidia.com>,
"Alex Mastro" <amastro@fb.com>,
"Christian König" <christian.koenig@amd.com>,
"Bjorn Helgaas" <bhelgaas@google.com>,
"Logan Gunthorpe" <logang@deltatee.com>,
"Kevin Tian" <kevin.tian@intel.com>,
"Longfang Liu" <liulongfang@huawei.com>,
"Mahmoud Adam" <mngyadam@amazon.de>,
"David Matlack" <dmatlack@google.com>,
"Björn Töpel" <bjorn@kernel.org>,
"Sumit Semwal" <sumit.semwal@linaro.org>,
"Ankit Agrawal" <ankita@nvidia.com>,
"Alistair Popple" <apopple@nvidia.com>,
"Vivek Kasireddy" <vivek.kasireddy@intel.com>,
linux-kernel@vger.kernel.org, linux-media@vger.kernel.org,
dri-devel@lists.freedesktop.org, linaro-mm-sig@lists.linaro.org,
kvm@vger.kernel.org, linux-pci@vger.kernel.org
Subject: Re: [PATCH v5 3/9] vfio/pci: Add a helper to look up PFNs for DMABUFs
Date: Thu, 30 Jul 2026 22:55:24 +0000 [thread overview]
Message-ID: <amvWXNRtviKEuUoV@google.com> (raw)
In-Reply-To: <20260715174737.15287-4-matt@ozlabs.org>
On Wed, Jul 15, 2026 at 06:47:26PM +0100, Matt Evans wrote:
> Add vfio_pci_dma_buf_find_pfn(), which a VMA fault handler can use to
> find a PFN.
>
> This supports multi-range DMABUFs, which typically would be used to
> represent scattered spans but might even represent overlapping or
> aliasing spans of PFNs.
>
> Because this is intended to be used in vfio_pci_core.c, we also need
> to expose the struct vfio_pci_dma_buf in the vfio_pci_priv.h header.
>
> Signed-off-by: Matt Evans <matt@ozlabs.org>
> ---
> drivers/vfio/pci/vfio_pci_dmabuf.c | 153 ++++++++++++++++++++++++++---
> drivers/vfio/pci/vfio_pci_priv.h | 20 ++++
> 2 files changed, 160 insertions(+), 13 deletions(-)
>
> diff --git a/drivers/vfio/pci/vfio_pci_dmabuf.c b/drivers/vfio/pci/vfio_pci_dmabuf.c
> index c16f460c01d6..7c047400dfd1 100644
> --- a/drivers/vfio/pci/vfio_pci_dmabuf.c
> +++ b/drivers/vfio/pci/vfio_pci_dmabuf.c
> @@ -9,19 +9,6 @@
>
> MODULE_IMPORT_NS("DMA_BUF");
>
> -struct vfio_pci_dma_buf {
> - struct dma_buf *dmabuf;
> - struct vfio_pci_core_device *vdev;
> - struct list_head dmabufs_elm;
> - size_t size;
> - struct phys_vec *phys_vec;
> - struct p2pdma_provider *provider;
> - u32 nr_ranges;
> - struct kref kref;
> - struct completion comp;
> - u8 revoked : 1;
> -};
> -
> static int vfio_pci_dma_buf_attach(struct dma_buf *dmabuf,
> struct dma_buf_attachment *attachment)
> {
> @@ -106,6 +93,146 @@ static const struct dma_buf_ops vfio_pci_dmabuf_ops = {
> .release = vfio_pci_dma_buf_release,
> };
>
> +int vfio_pci_dma_buf_find_pfn(struct vfio_pci_dma_buf *priv,
> + struct vm_area_struct *vma,
> + unsigned long fault_addr,
> + unsigned int order,
> + unsigned long *out_pfn)
> +{
> + /*
> + * Given a VMA (start, end, pgoffs) and a fault address,
> + * search the corresponding DMABUF's phys_vec[] to find the
> + * range representing the address's offset into the VMA, and
> + * its PFN.
> + *
> + * The phys_vec[] ranges represent contiguous spans of VAs
> + * upwards from the buffer offset 0; the actual PFNs might be
> + * in any order, overlap/alias, etc. Calculate an offset of
> + * the desired page given VMA start/pgoff and address, then
> + * search upwards from 0 to find which span contains it.
> + *
> + * On success, a valid PFN for a page sized by 'order' is
> + * returned into out_pfn.
> + *
> + * Failure occurs if:
> + * - A hugepage would cross the edge of the VMA,
> + * - A hugepage isn't entirely contained within a range
> + * (including where it straddles the boundary between
> + * ranges),
> + * - We find a range, but the final PFN isn't aligned to the
> + * requested order.
> + *
> + * Upon failure, -EAGAIN is returned and the caller is
> + * expected to try again with a smaller order, which will
> + * eventually succeed (order=0 will always work).
> + *
> + * It's suboptimal if DMABUFs are created with neighbouring
> + * ranges that are physically contiguous, since hugepages
> + * can't straddle range boundaries. (The construction of the
> + * ranges should merge them in this case.)
> + *
> + * Finally, vma_pgoff_adjust is used with a DMABUF created for
> + * a VFIO BAR mmap: a BAR mapped with vm_pgoff > 0 creates a
> + * DMABUF such that byte 0 of the VMA corresponds to byte 0 of
> + * the DMABUF and byte 'vm_pgoff << PAGE_SHIFT' into the BAR.
> + * To avoid double-offsetting in this scenario, subtracting
> + * vma_pgoff_adjust from this (non-zero) vm_pgoff generates
> + * the effective offset.
> + */
> +
> + const unsigned long pagesize = PAGE_SIZE << order;
> + unsigned long vma_off = ((vma->vm_pgoff - priv->vma_pgoff_adjust) <<
> + PAGE_SHIFT) & VFIO_PCI_OFFSET_MASK;
Maybe I'm getting ahead of myself here.. but it seems like this
restricts us to only mapping DMABUFs at offsets < 1TB due to the
VFIO_PCI_OFFSET_MASK (since we have HBMs on PCI devices now, hitting 1TB
may not be a very distant future).
While I understand this mask is needed to drop the BAR encoding in the
high bits.
My worry is, if in the future a user were to export a massive
contiguous DMABUF (e.g., >1TB of aggregated HBM) and tried to mmap deep
into it (passing an offset >= 1TB), this bitwise AND would silently drop
the high bits, leading to silent data corruption.
I think we should explicitly reject such an mmap with -EINVAL like:
+const unsigned long pagesize = PAGE_SIZE << order;
+unsigned long vma_off = (vma->vm_pgoff - priv->vma_pgoff_adjust) << PAGE_SHIFT;
+/*
+ * Prevent silent wrap-around if the user mmaps a DMABUF at an
+ * offset greater than the VFIO index mask allows.
+ */
+if (unlikely(vma_off > VFIO_PCI_OFFSET_MASK))
+ return -EINVAL;
+vma_off &= VFIO_PCI_OFFSET_MASK;
> + unsigned long rounded_page_addr = ALIGN_DOWN(fault_addr, pagesize);
> + unsigned long rounded_page_end = rounded_page_addr + pagesize;
> + unsigned long fault_offset;
> + unsigned long fault_offset_end;
> + unsigned long range_start_offset = 0;
> + unsigned int i;
> + int ret;
> +
Thanks,
Praan
next prev parent reply other threads:[~2026-07-30 22:55 UTC|newest]
Thread overview: 50+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-07-15 17:47 [PATCH v5 0/9] vfio/pci: Add mmap() for DMABUFs Matt Evans
2026-07-15 17:47 ` [PATCH v5 1/9] PCI/P2PDMA: Split pool-related cleanup out of pci_p2pdma_release() Matt Evans
2026-07-15 18:05 ` sashiko-bot
2026-07-17 8:02 ` Tian, Kevin
2026-07-28 22:33 ` Alex Williamson
2026-07-29 10:08 ` Leon Romanovsky
2026-07-30 16:45 ` Pranjal Shrivastava
2026-07-15 17:47 ` [PATCH v5 2/9] PCI/P2PDMA: Add CONFIG_PCI_P2PDMA_CORE Matt Evans
2026-07-15 18:03 ` sashiko-bot
2026-07-17 8:02 ` Tian, Kevin
2026-07-30 19:37 ` Pranjal Shrivastava
2026-07-15 17:47 ` [PATCH v5 3/9] vfio/pci: Add a helper to look up PFNs for DMABUFs Matt Evans
2026-07-15 18:08 ` sashiko-bot
2026-07-17 8:02 ` Tian, Kevin
2026-07-29 17:52 ` Alex Williamson
2026-07-30 17:34 ` Matt Evans
2026-07-30 22:55 ` Pranjal Shrivastava [this message]
2026-07-15 17:47 ` [PATCH v5 4/9] vfio/pci: Add a helper to create a DMABUF for a BAR-map VMA Matt Evans
2026-07-15 18:12 ` sashiko-bot
2026-07-27 13:45 ` Matt Evans
2026-07-15 17:47 ` [PATCH v5 5/9] vfio/pci: Convert BAR mmap() to use a DMABUF Matt Evans
2026-07-15 18:13 ` sashiko-bot
2026-07-27 13:45 ` Matt Evans
2026-07-15 17:47 ` [PATCH v5 6/9] vfio/pci: Provide a user-facing name for BAR mappings Matt Evans
2026-07-15 18:01 ` sashiko-bot
2026-07-17 8:03 ` Tian, Kevin
2026-07-29 17:52 ` Alex Williamson
2026-07-30 14:37 ` Matt Evans
2026-07-15 17:47 ` [PATCH v5 7/9] vfio/pci: Clean up BAR zap and revocation Matt Evans
2026-07-15 18:00 ` sashiko-bot
2026-07-17 8:03 ` Tian, Kevin
2026-07-29 17:52 ` Alex Williamson
2026-07-30 14:47 ` Matt Evans
2026-07-30 23:20 ` Pranjal Shrivastava
2026-07-15 17:47 ` [PATCH v5 8/9] vfio/pci: Support mmap() of a VFIO DMABUF Matt Evans
2026-07-15 18:17 ` sashiko-bot
2026-07-17 8:03 ` Tian, Kevin
2026-07-30 23:33 ` Pranjal Shrivastava
2026-07-15 17:47 ` [PATCH v5 9/9] vfio/pci: Permanently revoke a DMABUF on request Matt Evans
2026-07-15 18:11 ` sashiko-bot
2026-07-27 13:45 ` Matt Evans
2026-07-17 8:03 ` Tian, Kevin
2026-07-30 23:43 ` Pranjal Shrivastava
2026-07-15 18:12 ` [PATCH v5 0/9] vfio/pci: Add mmap() for DMABUFs David Matlack
2026-07-16 14:51 ` Matt Evans
2026-07-16 21:23 ` David Matlack
2026-07-17 8:42 ` David Laight
2026-07-17 16:30 ` David Matlack
2026-07-17 17:12 ` Matt Evans
2026-07-20 21:50 ` David Matlack
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=amvWXNRtviKEuUoV@google.com \
--to=praan@google.com \
--cc=alex@shazbot.org \
--cc=amastro@fb.com \
--cc=ankita@nvidia.com \
--cc=apopple@nvidia.com \
--cc=bhelgaas@google.com \
--cc=bjorn@kernel.org \
--cc=christian.koenig@amd.com \
--cc=dmatlack@google.com \
--cc=dri-devel@lists.freedesktop.org \
--cc=jgg@nvidia.com \
--cc=kevin.tian@intel.com \
--cc=kvm@vger.kernel.org \
--cc=leon@kernel.org \
--cc=linaro-mm-sig@lists.linaro.org \
--cc=linux-kernel@vger.kernel.org \
--cc=linux-media@vger.kernel.org \
--cc=linux-pci@vger.kernel.org \
--cc=liulongfang@huawei.com \
--cc=logang@deltatee.com \
--cc=matt@ozlabs.org \
--cc=mngyadam@amazon.de \
--cc=sumit.semwal@linaro.org \
--cc=vivek.kasireddy@intel.com \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is an external index of several public inboxes,
see mirroring instructions on how to clone and mirror
all data and code used by this external index.