Linux Documentation
 help / color / mirror / Atom feed
From: Richard Cheng <icheng@nvidia.com>
To: mhonap@nvidia.com
Cc: alex@shazbot.org, jgg@ziepe.ca, ankita@nvidia.com,
	jic23@kernel.org,  dave.jiang@intel.com,
	alejandro.lucero-palau@amd.com, smadhavan@nvidia.com,
	 corbet@lwn.net, skhan@linuxfoundation.org, dave@stgolabs.net,
	 alison.schofield@intel.com, vishal.l.verma@intel.com,
	iweiny@kernel.org,  ming.li@zohomail.com, yishaih@nvidia.com,
	skolothumtho@nvidia.com,  kevin.tian@intel.com,
	bhelgaas@google.com, dmatlack@google.com, kees@kernel.org,
	 gustavoars@kernel.org, cjia@nvidia.com, kjaju@nvidia.com,
	vsethi@nvidia.com,  zhiw@nvidia.com, linux-doc@vger.kernel.org,
	linux-kernel@vger.kernel.org,  kvm@vger.kernel.org,
	linux-cxl@vger.kernel.org, linux-pci@vger.kernel.org,
	 linux-kselftest@vger.kernel.org,
	linux-hardening@vger.kernel.org
Subject: Re: [PATCH v5 24/27] vfio/cxl: Export the HDM memory region as a dma-buf
Date: Thu, 17 Sep 2026 15:55:22 +0800	[thread overview]
Message-ID: <aquaxVHh2O8SUCvj@MWDK4CY14F> (raw)
In-Reply-To: <20260916183540.3813685-25-mhonap@nvidia.com>

On Thu, Sep 17, 2026 at 12:05:37AM +0800, mhonap@nvidia.com wrote:
> From: Manish Honap <mhonap@nvidia.com>
> 
> A Type-2 accelerator issues ATS-translated DMA to addresses inside its
> own HDM window, so that coherent host range must be present in the
> guest's IOAS (the iommufd IOAS backing the nested SMMU stage-2). iommufd
> maps a struct-page-less range only by fd, via IOMMU_IOAS_MAP_FILE over a
> dma-buf; a userspace-VA IOMMU_IOAS_MAP of the HDM mmap is rejected
> because the VMA is VM_IO | VM_PFNMAP. Without a dma-buf the range could
> only be mapped through an out-of-tree PFNMAP work-around.
> 
> vfio-pci already exports BAR memory as a P2P dma-buf, but the exporter
> is BAR-only: vfio_pci_core_feature_dma_buf() rejects any region index at
> or above the ROM index, and vfio_pci_core_get_dmabuf_phys() resolves the
> physical range from a PCI BAR. The HDM memory region is a dynamic
> device-specific region, not a BAR.
> 
> Let a device-specific region reach the device's get_dmabuf_phys(): a
> region index at or above VFIO_PCI_NUM_REGIONS skips the BAR-resource
> check and is validated by the driver instead, bounded to the regions
> that exist. Install a CXL-aware get_dmabuf_phys() in the vfio-cxl
> provider that returns cxl->hpa_range for the HDM memory region and
> delegates real BARs to the core, keeping the BAR path unchanged and the
> core free of CXL knowledge.
> 
> The HDM window is coherent host memory with no p2pdma provider of its
> own, so borrow BAR 0's, matching nvgrace-gpu's handling of its non-BAR
> device memory. The iommufd importer does not consume the provider; the
> scatterlist map path (real peer DMA) is left to a follow-up once
> upstream grows a negotiated interconnect for coherent CXL memory.
> 
> Assisted-by: LLM
> Signed-off-by: Manish Honap <mhonap@nvidia.com>
> ---
>  drivers/vfio/pci/cxl/vfio_cxl_core.c | 60 ++++++++++++++++++++++++++++
>  drivers/vfio/pci/vfio_pci_dmabuf.c   | 27 +++++++++++--
>  2 files changed, 83 insertions(+), 4 deletions(-)
> 
> diff --git a/drivers/vfio/pci/cxl/vfio_cxl_core.c b/drivers/vfio/pci/cxl/vfio_cxl_core.c
> index 5fe8e35c63c4..55fa1f86850d 100644
> --- a/drivers/vfio/pci/cxl/vfio_cxl_core.c
> +++ b/drivers/vfio/pci/cxl/vfio_cxl_core.c
> @@ -11,6 +11,7 @@
>  #include <linux/mm.h>
>  #include <linux/module.h>
>  #include <linux/pci.h>
> +#include <linux/pci-p2pdma.h>
>  #include <linux/range.h>
>  #include <linux/slab.h>
>  #include <linux/uaccess.h>
> @@ -27,6 +28,7 @@
>   * @hdm_regs: mapped HDM decoder registers, read live by the decoder region
>   * @hdm_len: length of the HDM decoder register block
>   * @hdm_valid: true when host CPU access to the HDM range is safe; under memory_lock
> + * @mem_region_index: vfio region index of the mmap-able HDM memory region
>   */
>  struct vfio_cxl_state {
>  	struct cxl_dev_state cxlds;
> @@ -36,6 +38,7 @@ struct vfio_cxl_state {
>  	void __iomem *hdm_regs;
>  	u32 hdm_len;
>  	bool hdm_valid;
> +	unsigned int mem_region_index;
>  };
>  
>  static unsigned long vfio_cxl_mem_pgoff(struct vm_area_struct *vma,
> @@ -289,6 +292,50 @@ static void vfio_cxl_release_hpa(void *data)
>  	release_mem_region(cxl->hpa_range.start, range_len(&cxl->hpa_range));
>  }
>  
> +/*
> + * Resolve the physical range that backs a dma-buf export. The core exporter
> + * only knows BARs; teach it the HDM memory region so a guest IOAS can map the
> + * coherent window by fd (IOMMU_IOAS_MAP_FILE) instead of the removed PFNMAP
> + * work-around. Real BARs stay on the byte-identical core path.
> + */
> +static int vfio_cxl_get_dmabuf_phys(struct vfio_pci_core_device *vdev,
> +				    struct p2pdma_provider **provider,
> +				    unsigned int region_index,
> +				    struct phys_vec *phys_vec,
> +				    struct vfio_region_dma_range *dma_ranges,
> +				    size_t nr_ranges)
> +{
> +	struct vfio_cxl_state *cxl = vdev->cxl;
> +
> +	/* Real BARs go through the core P2P exporter unchanged. */
> +	if (region_index < VFIO_PCI_NUM_REGIONS)
> +		return vfio_pci_core_get_dmabuf_phys(vdev, provider,
> +						     region_index, phys_vec,
> +						     dma_ranges, nr_ranges);
> +
> +	/* Of the device regions, only the HDM memory window is exportable. */
> +	if (region_index != cxl->mem_region_index)
> +		return -EINVAL;
> +
> +	/*
> +	 * The HDM window is coherent host memory, not BAR MMIO, so it has no
> +	 * p2pdma provider of its own. Borrow BAR 0's: the P2P properties match
> +	 * and the iommufd importer does not consume the provider. The sgt map
> +	 * path (real peer DMA) is not supported for the HDM window.
> +	 */
> +	*provider = pcim_p2pdma_provider(vdev->pdev, 0);

I think of a weird scenario which might make peer DMA work, not sure if that's possible.

userspace can give this dma-buf fd to an NIC driver or so to do RDMA ? use HDM memory as
network buffer ?

If that's the case and direct P2P addressing is selected, it applies BAR 0's bus offset to
HDM physical address.

Maybe explicitly reject peer-DMA mapping would be safer ?

Best regards,
Richard Cheng.


> +	if (!*provider)
> +		return -EINVAL;
> +
> +	return vfio_pci_core_fill_phys_vec(phys_vec, dma_ranges, nr_ranges,
> +					  cxl->hpa_range.start,
> +					  range_len(&cxl->hpa_range));
> +}
> +
> +static const struct vfio_pci_device_ops vfio_cxl_pci_dev_ops = {
> +	.get_dmabuf_phys = vfio_cxl_get_dmabuf_phys,
> +};
> +
>  static int vfio_cxl_init_device(struct vfio_pci_core_device *vdev)
>  {
>  	struct pci_dev *pdev = vdev->pdev;
> @@ -483,6 +530,19 @@ static int vfio_cxl_open_device(struct vfio_pci_core_device *vdev)
>  	if (ret)
>  		return ret;
>  
> +	/* Record where the HDM memory region landed for the dma-buf export. */
> +	cxl->mem_region_index = VFIO_PCI_NUM_REGIONS + vdev->num_regions - 1;
> +
> +	/*
> +	 * Override the device ops so a dma-buf export of the HDM memory region
> +	 * resolves to the coherent host range. This is done at open, not init:
> +	 * vfio_pci_probe() resets pci_ops after vfio_alloc_device() returns, so
> +	 * an override installed during init would be clobbered. Only a CXL device
> +	 * reaches this hook (cxl_ops is set on init success), so a fallback to
> +	 * plain vfio-pci keeps the core ops.
> +	 */
> +	vdev->pci_ops = &vfio_cxl_pci_dev_ops;
> +
>  	ret = vfio_cxl_add_region(vdev, VFIO_REGION_SUBTYPE_CXL_COMP_REGS,
>  				  &vfio_cxl_comp_regops, cxl->hdm_len,
>  				  VFIO_REGION_INFO_FLAG_READ |
> diff --git a/drivers/vfio/pci/vfio_pci_dmabuf.c b/drivers/vfio/pci/vfio_pci_dmabuf.c
> index c16f460c01d6..436c616d5b66 100644
> --- a/drivers/vfio/pci/vfio_pci_dmabuf.c
> +++ b/drivers/vfio/pci/vfio_pci_dmabuf.c
> @@ -178,6 +178,15 @@ int vfio_pci_core_get_dmabuf_phys(struct vfio_pci_core_device *vdev,
>  {
>  	struct pci_dev *pdev = vdev->pdev;
>  
> +	/*
> +	 * This resolver only handles PCI BARs. A device-specific region index
> +	 * (>= PCI_STD_NUM_BARS) would index pdev->resource[] out of bounds via
> +	 * pcim_p2pdma_provider(), so reject it; a driver that exports such a
> +	 * region installs its own get_dmabuf_phys.
> +	 */
> +	if (region_index >= PCI_STD_NUM_BARS)
> +		return -EINVAL;
> +
>  	*provider = pcim_p2pdma_provider(pdev, region_index);
>  	if (!*provider)
>  		return -EINVAL;
> @@ -227,6 +236,7 @@ int vfio_pci_core_feature_dma_buf(struct vfio_pci_core_device *vdev, u32 flags,
>  	DEFINE_DMA_BUF_EXPORT_INFO(exp_info);
>  	struct vfio_pci_dma_buf *priv;
>  	size_t length;
> +	u32 index;
>  	int ret;
>  
>  	if (!vdev->pci_ops || !vdev->pci_ops->get_dmabuf_phys)
> @@ -243,13 +253,22 @@ int vfio_pci_core_feature_dma_buf(struct vfio_pci_core_device *vdev, u32 flags,
>  	if (!get_dma_buf.nr_ranges || get_dma_buf.flags)
>  		return -EINVAL;
>  
> +	index = get_dma_buf.region_index;
> +
>  	/*
> -	 * For PCI the region_index is the BAR number like everything
> -	 * else.  Check that PCI resources have been claimed for it.
> +	 * A fixed region index is the BAR number; only a BAR can be exported
> +	 * and its PCI resource must be claimed. A device-specific region (index
> +	 * >= VFIO_PCI_NUM_REGIONS) has no BAR resource and is validated by the
> +	 * device's get_dmabuf_phys instead, but the index must name a region
> +	 * that exists.
>  	 */
> -	if (get_dma_buf.region_index >= VFIO_PCI_ROM_REGION_INDEX ||
> -	    IS_ERR(vfio_pci_core_get_iomap(vdev, get_dma_buf.region_index)))
> +	if (index < VFIO_PCI_NUM_REGIONS) {
> +		if (index >= VFIO_PCI_ROM_REGION_INDEX ||
> +		    IS_ERR(vfio_pci_core_get_iomap(vdev, index)))
> +			return -ENODEV;
> +	} else if (index - VFIO_PCI_NUM_REGIONS >= vdev->num_regions) {
>  		return -ENODEV;
> +	}
>  
>  	dma_ranges = memdup_array_user(&arg->dma_ranges, get_dma_buf.nr_ranges,
>  				       sizeof(*dma_ranges));
> -- 
> 2.25.1
> 
> 

  reply	other threads:[~2026-09-17  7:55 UTC|newest]

Thread overview: 57+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-09-16 18:35 [PATCH v5 00/27] vfio/pci: Add CXL Type-2 device passthrough support mhonap
2026-09-16 18:35 ` [PATCH v5 01/27] cxl/regs: Split the BAR block request and ioremap helpers mhonap
2026-09-22  1:28   ` Jonathan Cameron
2026-09-16 18:35 ` [PATCH v5 02/27] cxl/regs: Let a BAR-owning driver own the component register block mhonap
2026-09-22  1:36   ` Jonathan Cameron
2026-09-16 18:35 ` [PATCH v5 03/27] cxl: Move component register defines to uapi/cxl/cxl_regs.h mhonap
2026-09-22  1:43   ` Jonathan Cameron
2026-09-16 18:35 ` [PATCH v5 04/27] cxl: Add cxl_reset_dvsec_sequence() for vfio-pci mhonap
2026-09-25 20:41   ` Jonathan Cameron
2026-09-16 18:35 ` [PATCH v5 05/27] vfio/pci: Add the CXL provider ops registration interface mhonap
2026-09-16 18:35 ` [PATCH v5 06/27] vfio/pci: Detect CXL devices and load the CXL provider on demand mhonap
2026-09-17  8:48   ` Richard Cheng
2026-09-21 10:05     ` Manish Honap
2026-09-16 18:35 ` [PATCH v5 07/27] vfio/pci: Honor -EPROBE_DEFER from CXL provider probe mhonap
2026-09-16 18:35 ` [PATCH v5 08/27] vfio/pci: Fall back to plain vfio-pci when CXL init fails mhonap
2026-09-22  2:15   ` Alex Williamson
2026-09-16 18:35 ` [PATCH v5 09/27] vfio/pci: Add a generic excluded-range list mhonap
2026-09-22  2:13   ` Alex Williamson
2026-09-25 22:05   ` Jonathan Cameron
2026-09-16 18:35 ` [PATCH v5 10/27] vfio/pci: Migrate MSI-X exclusion onto the " mhonap
2026-09-16 18:35 ` [PATCH v5 11/27] vfio/pci: Virtualize the CXL DVSEC in vfio_pci_config.c mhonap
2026-09-22  2:14   ` Alex Williamson
2026-09-25 22:15   ` Jonathan Cameron
2026-09-16 18:35 ` [PATCH v5 12/27] vfio/pci: Call the CXL open and close hooks around device use mhonap
2026-09-25 22:17   ` Jonathan Cameron
2026-09-16 18:35 ` [PATCH v5 13/27] vfio/pci: Bracket PCI resets with the CXL reset hooks mhonap
2026-09-22  2:14   ` Alex Williamson
2026-09-16 18:35 ` [PATCH v5 14/27] vfio/pci: Provide an opt-out for the CXL Type-2 extensions mhonap
2026-09-16 18:35 ` [PATCH v5 15/27] vfio/cxl: Add the vfio-cxl provider module skeleton mhonap
2026-09-16 18:35 ` [PATCH v5 16/27] vfio/cxl: Create the CXL memdev and set media ready at bind mhonap
2026-09-22  2:14   ` Alex Williamson
2026-09-25 22:22   ` Jonathan Cameron
2026-09-16 18:35 ` [PATCH v5 17/27] vfio/cxl: Own the whole component register BAR mhonap
2026-09-22  2:14   ` Alex Williamson
2026-09-16 18:35 ` [PATCH v5 18/27] vfio/cxl: Expose the HDM memory region to the guest mhonap
2026-09-22  2:14   ` Alex Williamson
2026-09-16 18:35 ` [PATCH v5 19/27] vfio/cxl: Contain HDM memory errors with memory_failure() mhonap
2026-09-16 18:35 ` [PATCH v5 20/27] vfio/cxl: Expose the HDM decoder registers read-only to the guest mhonap
2026-09-22  2:13   ` Alex Williamson
2026-09-16 18:35 ` [PATCH v5 21/27] vfio/cxl: Exclude the HDM decoder registers from direct BAR access mhonap
2026-09-17  7:28   ` Richard Cheng
2026-09-21  9:52     ` Manish Honap
2026-09-16 18:35 ` [PATCH v5 22/27] vfio/cxl: Clear the HDM access gate after a hot reset mhonap
2026-09-22  2:13   ` Alex Williamson
2026-09-16 18:35 ` [PATCH v5 23/27] vfio/cxl: Describe the CXL device and decoder geometry to userspace mhonap
2026-09-16 18:35 ` [PATCH v5 24/27] vfio/cxl: Export the HDM memory region as a dma-buf mhonap
2026-09-17  7:55   ` Richard Cheng [this message]
2026-09-21  9:58     ` Manish Honap
2026-09-16 18:35 ` [PATCH v5 25/27] vfio/cxl: Run the CXL reset at the vfio reset points mhonap
2026-09-17  8:11   ` Richard Cheng
2026-09-21 10:01     ` Manish Honap
2026-09-22  2:13   ` Alex Williamson
2026-09-16 18:35 ` [PATCH v5 26/27] Documentation: vfio-pci: Document CXL Type-2 device passthrough mhonap
2026-09-16 19:33   ` Gregory Price
2026-09-21  9:43     ` Manish Honap
2026-09-24  0:54       ` Jonathan Cameron
2026-09-16 18:35 ` [PATCH v5 27/27] selftests/vfio: Add CXL Type-2 passthrough tests mhonap

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=aquaxVHh2O8SUCvj@MWDK4CY14F \
    --to=icheng@nvidia.com \
    --cc=alejandro.lucero-palau@amd.com \
    --cc=alex@shazbot.org \
    --cc=alison.schofield@intel.com \
    --cc=ankita@nvidia.com \
    --cc=bhelgaas@google.com \
    --cc=cjia@nvidia.com \
    --cc=corbet@lwn.net \
    --cc=dave.jiang@intel.com \
    --cc=dave@stgolabs.net \
    --cc=dmatlack@google.com \
    --cc=gustavoars@kernel.org \
    --cc=iweiny@kernel.org \
    --cc=jgg@ziepe.ca \
    --cc=jic23@kernel.org \
    --cc=kees@kernel.org \
    --cc=kevin.tian@intel.com \
    --cc=kjaju@nvidia.com \
    --cc=kvm@vger.kernel.org \
    --cc=linux-cxl@vger.kernel.org \
    --cc=linux-doc@vger.kernel.org \
    --cc=linux-hardening@vger.kernel.org \
    --cc=linux-kernel@vger.kernel.org \
    --cc=linux-kselftest@vger.kernel.org \
    --cc=linux-pci@vger.kernel.org \
    --cc=mhonap@nvidia.com \
    --cc=ming.li@zohomail.com \
    --cc=skhan@linuxfoundation.org \
    --cc=skolothumtho@nvidia.com \
    --cc=smadhavan@nvidia.com \
    --cc=vishal.l.verma@intel.com \
    --cc=vsethi@nvidia.com \
    --cc=yishaih@nvidia.com \
    --cc=zhiw@nvidia.com \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox