From: Alex Williamson <alex@shazbot.org>
To: <mhonap@nvidia.com>
Cc: <jgg@ziepe.ca>, <ankita@nvidia.com>, <jic23@kernel.org>,
<dave.jiang@intel.com>, <alejandro.lucero-palau@amd.com>,
<smadhavan@nvidia.com>, <corbet@lwn.net>,
<skhan@linuxfoundation.org>, <dave@stgolabs.net>,
<alison.schofield@intel.com>, <vishal.l.verma@intel.com>,
<iweiny@kernel.org>, <ming.li@zohomail.com>, <yishaih@nvidia.com>,
<skolothumtho@nvidia.com>, <kevin.tian@intel.com>,
<bhelgaas@google.com>, <dmatlack@google.com>, <kees@kernel.org>,
<gustavoars@kernel.org>, <cjia@nvidia.com>, <kjaju@nvidia.com>,
<vsethi@nvidia.com>, <zhiw@nvidia.com>,
<linux-doc@vger.kernel.org>, <linux-kernel@vger.kernel.org>,
<kvm@vger.kernel.org>, <linux-cxl@vger.kernel.org>,
<linux-pci@vger.kernel.org>, <linux-kselftest@vger.kernel.org>,
<linux-hardening@vger.kernel.org>,
alex@shazbot.org
Subject: Re: [PATCH v4 07/27] vfio/pci: Detect CXL devices and load vfio-cxl on demand
Date: Wed, 26 Aug 2026 16:17:25 -0600 [thread overview]
Message-ID: <20260826161725.7af1915b@shazbot.org> (raw)
In-Reply-To: <20260813093631.2288172-8-mhonap@nvidia.com>
On Thu, 13 Aug 2026 15:06:11 +0530
<mhonap@nvidia.com> wrote:
> From: Manish Honap <mhonap@nvidia.com>
>
> A CXL device needs the vfio-cxl callbacks, but pulling vfio-cxl and the
> CXL core in unconditionally would bloat every vfio-pci setup. At bind,
> detect a CXL device with pcie_is_cxl() and request_module("vfio-cxl")
> only then, and hand the device to the registered ops.
>
> Each bound CXL device pins vfio-cxl through try_module_get() and drops
> the reference at release, so vfio-cxl can unload once no CXL device is
> bound. If vfio-cxl is absent the device is driven as plain vfio-pci.
>
> Signed-off-by: Manish Honap <mhonap@nvidia.com>
> ---
> drivers/vfio/pci/vfio_pci_core.c | 81 ++++++++++++++++++++++++++++++--
> include/linux/vfio_pci_core.h | 3 ++
> 2 files changed, 81 insertions(+), 3 deletions(-)
>
> diff --git a/drivers/vfio/pci/vfio_pci_core.c b/drivers/vfio/pci/vfio_pci_core.c
> index 88e68d43af9a..0f9b5dfeea66 100644
> --- a/drivers/vfio/pci/vfio_pci_core.c
> +++ b/drivers/vfio/pci/vfio_pci_core.c
> @@ -2176,6 +2176,44 @@ static void vfio_pci_vga_uninit(struct vfio_pci_core_device *vdev)
> VGA_RSRC_LEGACY_MEM);
> }
>
> +static const struct vfio_cxl_ops *vfio_pci_cxl_ops;
> +static DEFINE_MUTEX(vfio_pci_cxl_ops_lock);
> +
> +static const struct vfio_cxl_ops *vfio_pci_get_cxl_ops(void)
> +{
> + const struct vfio_cxl_ops *ops;
> +
> + mutex_lock(&vfio_pci_cxl_ops_lock);
> + ops = vfio_pci_cxl_ops;
> + if (ops && !try_module_get(ops->owner))
> + ops = NULL;
> + mutex_unlock(&vfio_pci_cxl_ops_lock);
> +
> + return ops;
> +}
Awkward flow, resolved with guards:
guard(rwsem_read)(&vfio_pci_cxl_ops_lock);
ops = vfio_pci_cxl_ops;
if (!ops || !try_module_get(ops->owner))
return NULL;
return ops;
> +
> +/*
> + * A CXL Type-2 device advertises both CXL.cache and CXL.mem in its CXL DVSEC.
> + * pcie_is_cxl() is also true for Type-1 (cache only) and Type-3 (mem only)
> + * devices, which the vfio-cxl provider does not handle, so confirm the Type-2
> + * identity before engaging it.
> + */
> +static bool vfio_pci_is_cxl_type2(struct pci_dev *pdev)
> +{
> + u16 dvsec, cap;
> +
> + dvsec = pci_find_dvsec_capability(pdev, PCI_VENDOR_ID_CXL,
> + PCI_DVSEC_CXL_DEVICE);
> + if (!dvsec)
> + return false;
> +
> + if (pci_read_config_word(pdev, dvsec + PCI_DVSEC_CXL_CAP, &cap))
> + return false;
> +
> + return (cap & PCI_DVSEC_CXL_CACHE_CAPABLE) &&
> + (cap & PCI_DVSEC_CXL_MEM_CAPABLE);
> +}
> +
> int vfio_pci_core_init_dev(struct vfio_device *core_vdev)
> {
> struct vfio_pci_core_device *vdev =
> @@ -2197,6 +2235,41 @@ int vfio_pci_core_init_dev(struct vfio_device *core_vdev)
> init_rwsem(&vdev->memory_lock);
> xa_init(&vdev->ctx);
>
> + /*
> + * Load vfio-cxl on demand for a CXL device. If it is absent, drive the
> + * device as plain vfio-pci rather than failing the bind.
> + */
> + if (pcie_is_cxl(vdev->pdev) && vfio_pci_is_cxl_type2(vdev->pdev)) {
Nit, embed the pcie_is_cxl() test in vfio_pci_is_cxl_type2().
> + const struct vfio_cxl_ops *ops;
> +
> + request_module("vfio-cxl");
> + ops = vfio_pci_get_cxl_ops();
> + if (ops) {
> + ret = ops->init_device(vdev);
> + if (ret) {
> + module_put(ops->owner);
Create a trivial vfio_pci_put_cxl_ops() for consistency.
> + return ret;
This looks like a regression, a device that previously worked with
vfio-pci now fails if vfio-cxl .init returns an error. It should
continue with a log message.
The disable_cxl option that comes later is a global opt-out, not an
opt-in. Users can opt-in to a feature that might fail previous
behavior but they should not be required to opt-out to retain existing
functionality, especially with only global granularity.
> + }
> + vdev->cxl_ops = ops;
> + /*
> + * Pin the device in D0 while bound rather than let
> + * the host power it down between opens.
> + */
> + vdev->disable_idle_d3 = true;
Why? Letting the host power down the device between opens is exactly
what we want for non-cxl devices. If we're trying to do something
around preserving the coherent memory configuration, it needs to be
justified as such, and should happen at the point where it's relevant,
ie. in the vfio-cxl .init path.
However, this alone doesn't prevent the user from using low power
states, so it also seems insufficient by itself.
> + } else if (IS_BUILTIN(CONFIG_VFIO_CXL)) {
> + /*
> + * Only DEFER for a built-in provider so the bind
> + * retries once vfio-cxl registers its ops.
> + * A modular provider was already loaded synchronously
> + * by request_module() above, so if it is still absent
> + * it is missing, blocked, or failed to init; drive the
> + * device as plain vfio-pci then rather than defer the
> + * bind forever.
> + */
> + return -EPROBE_DEFER;
LLM asks if the registration function should call
driver_deferred_probe_trigger() to make the retry explicit?
> + }
> + }
> +
> return 0;
> }
> EXPORT_SYMBOL_GPL(vfio_pci_core_init_dev);
> @@ -2206,6 +2279,11 @@ void vfio_pci_core_release_dev(struct vfio_device *core_vdev)
> struct vfio_pci_core_device *vdev =
> container_of(core_vdev, struct vfio_pci_core_device, vdev);
>
> + if (vdev->cxl_ops) {
> + vdev->cxl_ops->release_device(vdev);
> + module_put(vdev->cxl_ops->owner);
> + }
Turn both of these into helpers:
static int vfio_pci_core_cxl_init(struct vfio_device *core_vdev);
static void vfio_pci_core_cxl_release(struct vfio_device *core_vdev);
Include the tests is-cxl/cxl_ops tests in the helpers to
compartmentalize cxl init/release. Thanks,
Alex
> +
> mutex_destroy(&vdev->igate);
> mutex_destroy(&vdev->ioeventfds_lock);
> kfree(vdev->region);
> @@ -2670,9 +2748,6 @@ static void vfio_pci_dev_set_try_reset(struct vfio_device_set *dev_set)
> }
> }
>
> -static const struct vfio_cxl_ops *vfio_pci_cxl_ops;
> -static DEFINE_MUTEX(vfio_pci_cxl_ops_lock);
> -
> int vfio_pci_core_register_cxl_ops(const struct vfio_cxl_ops *ops)
> {
> int ret = 0;
> diff --git a/include/linux/vfio_pci_core.h b/include/linux/vfio_pci_core.h
> index 14753972e714..117cd67995d8 100644
> --- a/include/linux/vfio_pci_core.h
> +++ b/include/linux/vfio_pci_core.h
> @@ -29,6 +29,7 @@ struct vfio_pci_core_device;
> struct vfio_pci_region;
> struct p2pdma_provider;
> struct dma_buf_attachment;
> +struct vfio_cxl_state;
>
> struct vfio_pci_eventfd {
> struct eventfd_ctx *ctx;
> @@ -109,6 +110,8 @@ struct vfio_pci_core_device {
> struct vfio_device vdev;
> struct pci_dev *pdev;
> const struct vfio_pci_device_ops *pci_ops;
> + const struct vfio_cxl_ops *cxl_ops;
> + struct vfio_cxl_state *cxl;
> void __iomem *barmap[PCI_STD_NUM_BARS];
> bool bar_mmap_supported[PCI_STD_NUM_BARS];
> /* Flags modified at runtime - dedicated storage unit */
next prev parent reply other threads:[~2026-08-26 22:17 UTC|newest]
Thread overview: 38+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-08-13 9:36 [PATCH v4 00/27] vfio/pci: Add CXL Type-2 device passthrough support mhonap
2026-08-13 9:36 ` [PATCH v4 01/27] cxl: Fix resource.c include path and export cxl_restore_hdm_after_pci_reset mhonap
2026-08-21 22:52 ` Jonathan Cameron
2026-08-22 1:22 ` Manish Honap
2026-08-13 9:36 ` [PATCH v4 02/27] cxl/regs: Skip sub-block region request for BAR-owning drivers mhonap
2026-08-25 21:26 ` Alex Williamson
2026-08-13 9:36 ` [PATCH v4 03/27] cxl: Move component register defines to uapi/cxl/cxl_regs.h mhonap
2026-08-13 9:36 ` [PATCH v4 04/27] cxl: Establish media readiness in cxl_mem_probe() mhonap
2026-08-25 22:18 ` Alex Williamson
2026-08-13 9:36 ` [PATCH v4 05/27] cxl: Add a function-scoped reset entry for vfio-pci mhonap
2026-08-25 23:11 ` Alex Williamson
2026-08-13 9:36 ` [PATCH v4 06/27] vfio/pci: Add CXL ops registration interface mhonap
2026-08-26 21:11 ` Alex Williamson
2026-08-13 9:36 ` [PATCH v4 07/27] vfio/pci: Detect CXL devices and load vfio-cxl on demand mhonap
2026-08-26 22:17 ` Alex Williamson [this message]
2026-08-13 9:36 ` [PATCH v4 08/27] vfio/cxl: Add the vfio-cxl module skeleton mhonap
2026-08-26 22:50 ` Alex Williamson
2026-08-13 9:36 ` [PATCH v4 09/27] vfio/cxl: Create the CXL memory device at bind mhonap
2026-08-13 9:36 ` [PATCH v4 10/27] vfio/cxl: Reject unsupported decoder topologies " mhonap
2026-08-13 9:36 ` [PATCH v4 11/27] vfio/cxl: Own the whole component register BAR mhonap
2026-08-13 9:36 ` [PATCH v4 12/27] vfio/pci: Let a provider exclude a BAR sub-range from mmap mhonap
2026-08-13 9:36 ` [PATCH v4 13/27] vfio/pci: Refuse read/write to an excluded BAR sub-range mhonap
2026-08-13 9:36 ` [PATCH v4 14/27] vfio: Add CXL region type for the HDM region mhonap
2026-08-13 9:36 ` [PATCH v4 15/27] vfio/pci: Call CXL open and close hooks around device use mhonap
2026-08-13 9:36 ` [PATCH v4 16/27] vfio/cxl: Shadow the CXL DVSEC body at open mhonap
2026-08-13 9:36 ` [PATCH v4 17/27] vfio/cxl: Virtualize the CXL DVSEC mhonap
2026-08-13 9:36 ` [PATCH v4 18/27] vfio/cxl: Expose the HDM memory and trap the decoder registers mhonap
2026-08-13 9:36 ` [PATCH v4 19/27] vfio/cxl: Keep the HDM decoder block off the direct BAR mapping mhonap
2026-08-13 9:36 ` [PATCH v4 20/27] vfio/cxl: Emulate the HDM decoder commit handshake mhonap
2026-08-13 9:36 ` [PATCH v4 21/27] vfio/cxl: Describe the CXL device and decoder geometry to userspace mhonap
2026-08-13 9:36 ` [PATCH v4 22/27] vfio/cxl: Revoke the HDM mapping on reset and power transitions mhonap
2026-08-13 9:36 ` [PATCH v4 23/27] vfio/cxl: Refresh the decoder snapshot after a device reset mhonap
2026-08-13 9:36 ` [PATCH v4 24/27] vfio/cxl: Service a guest-triggered CXL reset mhonap
2026-08-13 9:36 ` [PATCH v4 25/27] vfio/pci: Provide an opt-out for the CXL Type-2 extensions mhonap
2026-08-13 9:36 ` [PATCH v4 26/27] Documentation: vfio-pci: Document CXL Type-2 device passthrough mhonap
2026-08-13 9:36 ` [PATCH v4 27/27] selftests/vfio: Add CXL Type-2 passthrough corner-case tests mhonap
2026-08-26 7:28 ` Shuai Xue
2026-08-26 16:17 ` Manish Honap
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=20260826161725.7af1915b@shazbot.org \
--to=alex@shazbot.org \
--cc=alejandro.lucero-palau@amd.com \
--cc=alison.schofield@intel.com \
--cc=ankita@nvidia.com \
--cc=bhelgaas@google.com \
--cc=cjia@nvidia.com \
--cc=corbet@lwn.net \
--cc=dave.jiang@intel.com \
--cc=dave@stgolabs.net \
--cc=dmatlack@google.com \
--cc=gustavoars@kernel.org \
--cc=iweiny@kernel.org \
--cc=jgg@ziepe.ca \
--cc=jic23@kernel.org \
--cc=kees@kernel.org \
--cc=kevin.tian@intel.com \
--cc=kjaju@nvidia.com \
--cc=kvm@vger.kernel.org \
--cc=linux-cxl@vger.kernel.org \
--cc=linux-doc@vger.kernel.org \
--cc=linux-hardening@vger.kernel.org \
--cc=linux-kernel@vger.kernel.org \
--cc=linux-kselftest@vger.kernel.org \
--cc=linux-pci@vger.kernel.org \
--cc=mhonap@nvidia.com \
--cc=ming.li@zohomail.com \
--cc=skhan@linuxfoundation.org \
--cc=skolothumtho@nvidia.com \
--cc=smadhavan@nvidia.com \
--cc=vishal.l.verma@intel.com \
--cc=vsethi@nvidia.com \
--cc=yishaih@nvidia.com \
--cc=zhiw@nvidia.com \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is an external index of several public inboxes,
see mirroring instructions on how to clone and mirror
all data and code used by this external index.