Kernel KVM virtualization development
 help / color / mirror / Atom feed
From: Jonathan Cameron <jic23@kernel.org>
To: <mhonap@nvidia.com>
Cc: <alex@shazbot.org>, <jgg@ziepe.ca>, <ankita@nvidia.com>,
	<dave.jiang@intel.com>, <alejandro.lucero-palau@amd.com>,
	<smadhavan@nvidia.com>, <corbet@lwn.net>,
	<skhan@linuxfoundation.org>, <dave@stgolabs.net>,
	<alison.schofield@intel.com>, <vishal.l.verma@intel.com>,
	<iweiny@kernel.org>, <ming.li@zohomail.com>, <yishaih@nvidia.com>,
	<skolothumtho@nvidia.com>, <kevin.tian@intel.com>,
	<bhelgaas@google.com>, <dmatlack@google.com>, <kees@kernel.org>,
	<gustavoars@kernel.org>, <cjia@nvidia.com>, <kjaju@nvidia.com>,
	<vsethi@nvidia.com>, <zhiw@nvidia.com>,
	<linux-doc@vger.kernel.org>, <linux-kernel@vger.kernel.org>,
	<kvm@vger.kernel.org>, <linux-cxl@vger.kernel.org>,
	<linux-pci@vger.kernel.org>, <linux-kselftest@vger.kernel.org>,
	<linux-hardening@vger.kernel.org>
Subject: Re: [PATCH v5 16/27] vfio/cxl: Create the CXL memdev and set media ready at bind
Date: Fri, 25 Sep 2026 23:22:50 +0100	[thread overview]
Message-ID: <20260925232250.7d408311@jic23-hlaptop> (raw)
In-Reply-To: <20260916183540.3813685-17-mhonap@nvidia.com>

On Thu, 17 Sep 2026 00:05:29 +0530
<mhonap@nvidia.com> wrote:

> From: Manish Honap <mhonap@nvidia.com>
> 
> At bind, build the CXL memory device for the passed-through Type-2
> accelerator so it joins the CXL topology and its HDM region resolves to a
> host physical range. A Type-2 device has no mailbox, so there is no
> media-ready register to poll: set media ready directly once the component
> registers validate (mirroring drivers/net/ethernet/sfc/efx_cxl.c)
> 
> As per current vfio-cxl support, reject a device with:
> - more than one HDM decoder
> - interleaving enabled
> - whose reset the host cannot service
> 
> The CXL-core allocations are grouped with devres so a failed bind unwinds
> them: init failure falls back to plain vfio-pci with the device still
> bound, so devm would otherwise hold them until unbind.
> 
> A low-power transition would reset the CXL Type-2 function and lose
> its CXL.mem contents, so keep it in D0 while it is assigned.
> 
> Assisted-by: LLM
> Signed-off-by: Manish Honap <mhonap@nvidia.com>
A small thing inline.

> ---
>  drivers/vfio/pci/cxl/vfio_cxl_core.c | 118 +++++++++++++++++++++++++++
>  drivers/vfio/pci/vfio_pci_core.c     |  16 ++++
>  2 files changed, 134 insertions(+)
> 
> diff --git a/drivers/vfio/pci/cxl/vfio_cxl_core.c b/drivers/vfio/pci/cxl/vfio_cxl_core.c
> index cd5d41856404..0d92e6a409c1 100644
> --- a/drivers/vfio/pci/cxl/vfio_cxl_core.c
> +++ b/drivers/vfio/pci/cxl/vfio_cxl_core.c
> @@ -6,15 +6,132 @@
>   */
>  
>  #include <linux/module.h>
> +#include <linux/pci.h>
> +#include <linux/range.h>
>  #include <linux/vfio_pci_core.h>
> +#include <cxl/cxl.h>
> +#include <cxl/pci.h>
> +
> +/**
> + * struct vfio_cxl_state - per-device state for a vfio-cxl device
> + * @cxlds: CXL device state; kept first for devm_cxl_dev_state_create()
> + * @cxlmd: memory device joined to the CXL topology at bind
> + * @hpa_range: host physical range of the HDM region
> + */
> +struct vfio_cxl_state {
> +	struct cxl_dev_state cxlds;
> +	struct cxl_memdev *cxlmd;
> +	struct range hpa_range;
> +};
>  
>  static int vfio_cxl_init_device(struct vfio_pci_core_device *vdev)
>  {
> +	struct pci_dev *pdev = vdev->pdev;
> +	struct vfio_cxl_state *cxl;
> +	struct cxl_memdev *cxlmd;
> +	u64 hdm_size, serial;
> +	u16 dvsec;
> +	int ret;
> +
> +	/*
> +	 * pdev->hdm is cached at PCI enumeration, before any driver binds, so a
> +	 * device without it has no usable HDM decoder. Fall back to plain
> +	 * vfio-pci rather than deferring the bind forever.
> +	 */
> +	if (!pdev->hdm)
> +		return -ENODEV;
> +
> +	/* The guest drives one virtual decoder; multiple are unsupported. */
> +	if (pdev->hdm->decoder_count != 1)
> +		return -EOPNOTSUPP;
> +
> +	/* An interleaved decoder cannot be mapped 1:1 to the guest. */
> +	if (pdev->hdm->settings[0].interleave_ways != 1)
> +		return -EOPNOTSUPP;
> +
> +	/*
> +	 * The guest drives resets through the CXL Device DVSEC and polls the
> +	 * shadow for completion. If the host cannot service a function-scoped
> +	 * CXL reset, that request could never complete, so refuse the device
> +	 * rather than advertise a reset the guest would poll on forever.
> +	 */
> +	if (!cxl_reset_capable(pdev))
> +		return -EOPNOTSUPP;
> +
> +	hdm_size = range_len(&pdev->hdm->settings[0].hpa_range);
> +	if (!hdm_size)
> +		return -ENXIO;
> +
> +	dvsec = pci_find_dvsec_capability(pdev, PCI_VENDOR_ID_CXL,
> +					  PCI_DVSEC_CXL_DEVICE);
> +	serial = pci_get_dsn(pdev);
> +
> +	/*
> +	 * Group the CXL-core allocations so a later failure unwinds them here.
> +	 * A failed init falls back to plain vfio-pci with the device still
> +	 * bound, so devm would otherwise hold them until unbind.
> +	 */
> +	if (!devres_open_group(&pdev->dev, NULL, GFP_KERNEL))
> +		return -ENOMEM;
> +

I'd be tempted to factor out the stuff that is unwound via the
group. If this were simple a call to a helper function, that helper could
return directly on error giving us simpler code flow.  Then
you'd just check the helper return to find out if you should close
or release the group.

> +	cxl = devm_cxl_dev_state_create(&pdev->dev, CXL_DEVTYPE_DEVMEM, serial,
> +					dvsec, struct vfio_cxl_state, cxlds,
> +					false);
> +	if (!cxl) {
> +		ret = -ENOMEM;
> +		goto err;
> +	}
> +
> +	ret = cxl_pci_setup_regs(pdev, CXL_REGLOC_RBI_COMPONENT,
> +				 &cxl->cxlds.reg_map);
> +	if (ret) {
> +		pci_err(pdev, "vfio-cxl: no component registers\n");
> +		goto err;
> +	}
> +
> +	if (!cxl->cxlds.reg_map.component_map.hdm_decoder.valid) {
> +		pci_err(pdev, "vfio-cxl: HDM decoder registers not found\n");
> +		ret = -ENODEV;
> +		goto err;
> +	}
> +
> +	/*
> +	 * A Type-2 accelerator has no mailbox and no media-ready register, so
> +	 * set media ready directly.
> +	 */
> +	cxl->cxlds.media_ready = true;
> +
> +	ret = cxl_set_capacity(&cxl->cxlds, hdm_size);
> +	if (ret)
> +		goto err;
> +
> +	cxlmd = devm_cxl_probe_mem(&cxl->cxlds, &cxl->hpa_range);
> +	if (IS_ERR(cxlmd)) {
> +		ret = PTR_ERR(cxlmd);
> +		goto err;
> +	}
> +
> +	cxl->cxlmd = cxlmd;
> +	devres_close_group(&pdev->dev, NULL);
> +
> +	/*
> +	 * Powering a CXL Type-2 function down and back up reinitializes its
> +	 * device state and discards the contents of its coherent memory. Pin
> +	 * it in D0 for as long as it is assigned so CXL.mem stays intact.
> +	 */
> +	vdev->disable_idle_d3 = true;
> +	vdev->cxl = cxl;
> +
>  	return 0;
> +
> +err:
> +	devres_release_group(&pdev->dev, NULL);
> +	return ret;
>  }
>  
>  static void vfio_cxl_release_device(struct vfio_pci_core_device *vdev)
>  {
> +	vdev->cxl = NULL;
>  }
>  
>  static int vfio_cxl_open_device(struct vfio_pci_core_device *vdev)
> @@ -59,3 +176,4 @@ module_exit(vfio_cxl_exit);
>  
>  MODULE_LICENSE("GPL");
>  MODULE_DESCRIPTION("VFIO support for CXL Type-2 devices");
> +MODULE_IMPORT_NS("CXL");

  parent reply	other threads:[~2026-09-25 22:22 UTC|newest]

Thread overview: 57+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-09-16 18:35 [PATCH v5 00/27] vfio/pci: Add CXL Type-2 device passthrough support mhonap
2026-09-16 18:35 ` [PATCH v5 01/27] cxl/regs: Split the BAR block request and ioremap helpers mhonap
2026-09-22  1:28   ` Jonathan Cameron
2026-09-16 18:35 ` [PATCH v5 02/27] cxl/regs: Let a BAR-owning driver own the component register block mhonap
2026-09-22  1:36   ` Jonathan Cameron
2026-09-16 18:35 ` [PATCH v5 03/27] cxl: Move component register defines to uapi/cxl/cxl_regs.h mhonap
2026-09-22  1:43   ` Jonathan Cameron
2026-09-16 18:35 ` [PATCH v5 04/27] cxl: Add cxl_reset_dvsec_sequence() for vfio-pci mhonap
2026-09-25 20:41   ` Jonathan Cameron
2026-09-16 18:35 ` [PATCH v5 05/27] vfio/pci: Add the CXL provider ops registration interface mhonap
2026-09-16 18:35 ` [PATCH v5 06/27] vfio/pci: Detect CXL devices and load the CXL provider on demand mhonap
2026-09-17  8:48   ` Richard Cheng
2026-09-21 10:05     ` Manish Honap
2026-09-16 18:35 ` [PATCH v5 07/27] vfio/pci: Honor -EPROBE_DEFER from CXL provider probe mhonap
2026-09-16 18:35 ` [PATCH v5 08/27] vfio/pci: Fall back to plain vfio-pci when CXL init fails mhonap
2026-09-22  2:15   ` Alex Williamson
2026-09-16 18:35 ` [PATCH v5 09/27] vfio/pci: Add a generic excluded-range list mhonap
2026-09-22  2:13   ` Alex Williamson
2026-09-25 22:05   ` Jonathan Cameron
2026-09-16 18:35 ` [PATCH v5 10/27] vfio/pci: Migrate MSI-X exclusion onto the " mhonap
2026-09-16 18:35 ` [PATCH v5 11/27] vfio/pci: Virtualize the CXL DVSEC in vfio_pci_config.c mhonap
2026-09-22  2:14   ` Alex Williamson
2026-09-25 22:15   ` Jonathan Cameron
2026-09-16 18:35 ` [PATCH v5 12/27] vfio/pci: Call the CXL open and close hooks around device use mhonap
2026-09-25 22:17   ` Jonathan Cameron
2026-09-16 18:35 ` [PATCH v5 13/27] vfio/pci: Bracket PCI resets with the CXL reset hooks mhonap
2026-09-22  2:14   ` Alex Williamson
2026-09-16 18:35 ` [PATCH v5 14/27] vfio/pci: Provide an opt-out for the CXL Type-2 extensions mhonap
2026-09-16 18:35 ` [PATCH v5 15/27] vfio/cxl: Add the vfio-cxl provider module skeleton mhonap
2026-09-16 18:35 ` [PATCH v5 16/27] vfio/cxl: Create the CXL memdev and set media ready at bind mhonap
2026-09-22  2:14   ` Alex Williamson
2026-09-25 22:22   ` Jonathan Cameron [this message]
2026-09-16 18:35 ` [PATCH v5 17/27] vfio/cxl: Own the whole component register BAR mhonap
2026-09-22  2:14   ` Alex Williamson
2026-09-16 18:35 ` [PATCH v5 18/27] vfio/cxl: Expose the HDM memory region to the guest mhonap
2026-09-22  2:14   ` Alex Williamson
2026-09-16 18:35 ` [PATCH v5 19/27] vfio/cxl: Contain HDM memory errors with memory_failure() mhonap
2026-09-16 18:35 ` [PATCH v5 20/27] vfio/cxl: Expose the HDM decoder registers read-only to the guest mhonap
2026-09-22  2:13   ` Alex Williamson
2026-09-16 18:35 ` [PATCH v5 21/27] vfio/cxl: Exclude the HDM decoder registers from direct BAR access mhonap
2026-09-17  7:28   ` Richard Cheng
2026-09-21  9:52     ` Manish Honap
2026-09-16 18:35 ` [PATCH v5 22/27] vfio/cxl: Clear the HDM access gate after a hot reset mhonap
2026-09-22  2:13   ` Alex Williamson
2026-09-16 18:35 ` [PATCH v5 23/27] vfio/cxl: Describe the CXL device and decoder geometry to userspace mhonap
2026-09-16 18:35 ` [PATCH v5 24/27] vfio/cxl: Export the HDM memory region as a dma-buf mhonap
2026-09-17  7:55   ` Richard Cheng
2026-09-21  9:58     ` Manish Honap
2026-09-16 18:35 ` [PATCH v5 25/27] vfio/cxl: Run the CXL reset at the vfio reset points mhonap
2026-09-17  8:11   ` Richard Cheng
2026-09-21 10:01     ` Manish Honap
2026-09-22  2:13   ` Alex Williamson
2026-09-16 18:35 ` [PATCH v5 26/27] Documentation: vfio-pci: Document CXL Type-2 device passthrough mhonap
2026-09-16 19:33   ` Gregory Price
2026-09-21  9:43     ` Manish Honap
2026-09-24  0:54       ` Jonathan Cameron
2026-09-16 18:35 ` [PATCH v5 27/27] selftests/vfio: Add CXL Type-2 passthrough tests mhonap

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=20260925232250.7d408311@jic23-hlaptop \
    --to=jic23@kernel.org \
    --cc=alejandro.lucero-palau@amd.com \
    --cc=alex@shazbot.org \
    --cc=alison.schofield@intel.com \
    --cc=ankita@nvidia.com \
    --cc=bhelgaas@google.com \
    --cc=cjia@nvidia.com \
    --cc=corbet@lwn.net \
    --cc=dave.jiang@intel.com \
    --cc=dave@stgolabs.net \
    --cc=dmatlack@google.com \
    --cc=gustavoars@kernel.org \
    --cc=iweiny@kernel.org \
    --cc=jgg@ziepe.ca \
    --cc=kees@kernel.org \
    --cc=kevin.tian@intel.com \
    --cc=kjaju@nvidia.com \
    --cc=kvm@vger.kernel.org \
    --cc=linux-cxl@vger.kernel.org \
    --cc=linux-doc@vger.kernel.org \
    --cc=linux-hardening@vger.kernel.org \
    --cc=linux-kernel@vger.kernel.org \
    --cc=linux-kselftest@vger.kernel.org \
    --cc=linux-pci@vger.kernel.org \
    --cc=mhonap@nvidia.com \
    --cc=ming.li@zohomail.com \
    --cc=skhan@linuxfoundation.org \
    --cc=skolothumtho@nvidia.com \
    --cc=smadhavan@nvidia.com \
    --cc=vishal.l.verma@intel.com \
    --cc=vsethi@nvidia.com \
    --cc=yishaih@nvidia.com \
    --cc=zhiw@nvidia.com \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox