Kernel KVM virtualization development
 help / color / mirror / Atom feed
From: Jonathan Cameron <jic23@kernel.org>
To: <mhonap@nvidia.com>
Cc: <alex@shazbot.org>, <jgg@ziepe.ca>, <ankita@nvidia.com>,
	<dave.jiang@intel.com>, <alejandro.lucero-palau@amd.com>,
	<smadhavan@nvidia.com>, <corbet@lwn.net>,
	<skhan@linuxfoundation.org>, <dave@stgolabs.net>,
	<alison.schofield@intel.com>, <vishal.l.verma@intel.com>,
	<iweiny@kernel.org>, <ming.li@zohomail.com>, <yishaih@nvidia.com>,
	<skolothumtho@nvidia.com>, <kevin.tian@intel.com>,
	<bhelgaas@google.com>, <dmatlack@google.com>, <kees@kernel.org>,
	<gustavoars@kernel.org>, <cjia@nvidia.com>, <kjaju@nvidia.com>,
	<vsethi@nvidia.com>, <zhiw@nvidia.com>,
	<linux-doc@vger.kernel.org>, <linux-kernel@vger.kernel.org>,
	<kvm@vger.kernel.org>, <linux-cxl@vger.kernel.org>,
	<linux-pci@vger.kernel.org>, <linux-kselftest@vger.kernel.org>,
	<linux-hardening@vger.kernel.org>
Subject: Re: [PATCH v5 11/27] vfio/pci: Virtualize the CXL DVSEC in vfio_pci_config.c
Date: Fri, 25 Sep 2026 23:15:41 +0100	[thread overview]
Message-ID: <20260925231541.492d5407@jic23-hlaptop> (raw)
In-Reply-To: <20260916183540.3813685-12-mhonap@nvidia.com>

On Thu, 17 Sep 2026 00:05:24 +0530
<mhonap@nvidia.com> wrote:

> From: Manish Honap <mhonap@nvidia.com>
> 
> A CXL Type-2 device is reprogrammable through its CXL DVSEC: a guest could
> set Config Lock, toggle CXL.cache and CXL.mem enable, or rewrite the HDM
> range registers that govern host memory decode. Virtualize the DVSEC so
> the guest sees a shadow it cannot use to reprogram the hardware.
> 
> Build a per-device cxl_perm permission map, modeled on msi_perm, when the
> device is bound through the CXL provider (e.g. vfio-cxl). The whole CXL
> DVSEC is served from the vconfig shadow; only Control and Control2 are
> guest programmable, while Capability, Status, Lock and the Range registers
> keep their firmware snapshot. A vendor DVSEC on the same device is
> unaffected: the map is selected only for the CXL DVSEC offset.
> 
> Control2 carries the CXL reset and cache write-back-invalidate initiate
> bits as self-clearing doorbells. vfio never forwards them to hardware, so
> a custom writefn synthesizes their completion in the shadow: the initiate
> bit self-clears and the matching Status2 bit (Cache Invalid or Reset Done)
> is set, so a guest following the spec reset sequence
> (INIT_CACHE_WBI, poll Cache Invalid, INIT_CXL_RST, poll Reset Done)
> progresses instead of timing out. The host performs the real cache
> write-back and reset at the vfio reset points.
> 
> Assisted-by: LLM
> Signed-off-by: Manish Honap <mhonap@nvidia.com>
A few minor comments inline.

J
> ---
>  drivers/vfio/pci/vfio_pci_config.c | 132 ++++++++++++++++++++++++++++-
>  include/linux/vfio_pci_core.h      |   3 +
>  2 files changed, 133 insertions(+), 2 deletions(-)
> 
> diff --git a/drivers/vfio/pci/vfio_pci_config.c b/drivers/vfio/pci/vfio_pci_config.c
> index 9914f3ac69ae..9a020a768055 100644
> --- a/drivers/vfio/pci/vfio_pci_config.c
> +++ b/drivers/vfio/pci/vfio_pci_config.c

> +
> +static int init_cxl_dvsec_perm(struct perm_bits *perm, int len)
> +{
> +	int i;
> +
> +	if (alloc_perm_bits(perm, len))
> +		return -ENOMEM;
> +
> +	perm->writefn = vfio_cxl_dvsec_write;
> +
> +	/* Serve the whole CXL DVSEC from the shadow. */
> +	for (i = 0; i < len; i++)
	for (int i = 0;

Is mostly acceptable in the kernel these days and keeps
the scope tightly defined.

> +		p_setb(perm, i, (u8)ALL_VIRT, NO_WRITE);
> +
> +	/*
> +	 * Control and Control2 are guest programmable; Capability, Status,
> +	 * Lock and the Range registers keep their firmware snapshot, so the
> +	 * guest cannot set Config Lock or rewrite the capability and ranges.
> +	 */
> +	p_setw(perm, PCI_DVSEC_CXL_CTRL, (u16)ALL_VIRT, (u16)ALL_WRITE);
> +	p_setw(perm, PCI_DVSEC_CXL_CTRL2, (u16)ALL_VIRT, (u16)ALL_WRITE);
> +
> +	return 0;
> +}
> +
> +/* Virtualize the CXL DVSEC so a guest cannot reprogram the device through it. */
> +static int vfio_cxl_dvsec_init(struct vfio_pci_core_device *vdev)
> +{
> +	struct pci_dev *pdev = vdev->pdev;
> +	u32 dword;
> +	u16 dvsec;
> +	int len, ret;
> +
> +	dvsec = pci_find_dvsec_capability(pdev, PCI_VENDOR_ID_CXL,
> +					  PCI_DVSEC_CXL_DEVICE);
> +	if (!dvsec)
> +		return 0;
> +
> +	ret = pci_read_config_dword(pdev, dvsec + PCI_DVSEC_HEADER1, &dword);
> +	if (ret)
> +		return pcibios_err_to_errno(ret);
> +	len = PCI_DVSEC_HEADER1_LEN(dword);
> +
> +	/*
> +	 * The virtualization writes fixed DVSEC offsets up to Status2 (the reset
> +	 * doorbell stamps it). A device that reports a shorter DVSEC is not a
> +	 * usable Type-2 function; leave it as plain vfio-pci rather than index the
> +	 * device-length-sized perm allocation past its end.
> +	 */
> +	if (len < PCI_DVSEC_CXL_STATUS2 + 2)
> +		return 0;
> +
> +	vdev->cxl_perm = kmalloc_obj(struct perm_bits, GFP_KERNEL_ACCOUNT);

I'd use *vdev->cxl_perm instead of struct perm_bits just because
that saves anyone checking types.

> +	if (!vdev->cxl_perm)
> +		return -ENOMEM;
> +
> +	ret = init_cxl_dvsec_perm(vdev->cxl_perm, len);
> +	if (ret) {
> +		kfree(vdev->cxl_perm);
> +		vdev->cxl_perm = NULL;
> +		return ret;
> +	}
> +
> +	vdev->cxl_dvsec = dvsec;
> +	vdev->cxl_dvsec_len = len;
> +
> +	return 0;
> +}
> +
>  int vfio_config_init(struct vfio_pci_core_device *vdev)
>  {
>  	struct pci_dev *pdev = vdev->pdev;
> @@ -1842,6 +1955,12 @@ int vfio_config_init(struct vfio_pci_core_device *vdev)
>  	if (ret)
>  		goto out;
>  
> +	if (vdev->cxl_ops) {
> +		ret = vfio_cxl_dvsec_init(vdev);
> +		if (ret)
> +			goto out;
> +	}
> +
>  	return 0;
>  
>  out:
> @@ -1863,6 +1982,12 @@ void vfio_config_free(struct vfio_pci_core_device *vdev)
>  		kfree(vdev->msi_perm);
>  		vdev->msi_perm = NULL;
>  	}
> +	if (vdev->cxl_perm) {
> +		free_perm_bits(vdev->cxl_perm);
> +		kfree(vdev->cxl_perm);
> +		vdev->cxl_perm = NULL;
> +		vdev->cxl_dvsec = 0;

I'd define a vfio_cxl_dvsec_exit() or _unint() for this so
it is clear it pairs with vfio_cxl_dvsec_init() above.


> +	}
>  }
>  
>  /*
> @@ -1926,12 +2051,15 @@ ssize_t vfio_pci_config_rw_single(struct vfio_pci_core_device *vdev,
>  			 * of the extended capability list.  Use default, ro
>  			 * access, which will virtualize the id and next values.
>  			 */
> +			cap_start = vfio_find_cap_start(vdev, *ppos);
> +
>  			if (cap_id > PCI_EXT_CAP_ID_MAX)
>  				perm = &direct_ro_perms;
> +			else if (cap_id == PCI_EXT_CAP_ID_DVSEC && vdev->cxl_perm &&
> +				 cap_start == vdev->cxl_dvsec)

If it is only used here (I haven't read on in series)
				 vfio_find_cap_start(vdev, *ppos) == vdev->cxl_dvsec)
seems fine to me.  It's only just over 80 chars and hopeful this bit of the kernel
is flexible on that!

> +				perm = vdev->cxl_perm;
>  			else
>  				perm = &ecap_perms[cap_id];
> -
> -			cap_start = vfio_find_cap_start(vdev, *ppos);
>  		} else {
>  			WARN_ON(cap_id > PCI_CAP_ID_MAX);
>  


  parent reply	other threads:[~2026-09-25 22:15 UTC|newest]

Thread overview: 57+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-09-16 18:35 [PATCH v5 00/27] vfio/pci: Add CXL Type-2 device passthrough support mhonap
2026-09-16 18:35 ` [PATCH v5 01/27] cxl/regs: Split the BAR block request and ioremap helpers mhonap
2026-09-22  1:28   ` Jonathan Cameron
2026-09-16 18:35 ` [PATCH v5 02/27] cxl/regs: Let a BAR-owning driver own the component register block mhonap
2026-09-22  1:36   ` Jonathan Cameron
2026-09-16 18:35 ` [PATCH v5 03/27] cxl: Move component register defines to uapi/cxl/cxl_regs.h mhonap
2026-09-22  1:43   ` Jonathan Cameron
2026-09-16 18:35 ` [PATCH v5 04/27] cxl: Add cxl_reset_dvsec_sequence() for vfio-pci mhonap
2026-09-25 20:41   ` Jonathan Cameron
2026-09-16 18:35 ` [PATCH v5 05/27] vfio/pci: Add the CXL provider ops registration interface mhonap
2026-09-16 18:35 ` [PATCH v5 06/27] vfio/pci: Detect CXL devices and load the CXL provider on demand mhonap
2026-09-17  8:48   ` Richard Cheng
2026-09-21 10:05     ` Manish Honap
2026-09-16 18:35 ` [PATCH v5 07/27] vfio/pci: Honor -EPROBE_DEFER from CXL provider probe mhonap
2026-09-16 18:35 ` [PATCH v5 08/27] vfio/pci: Fall back to plain vfio-pci when CXL init fails mhonap
2026-09-22  2:15   ` Alex Williamson
2026-09-16 18:35 ` [PATCH v5 09/27] vfio/pci: Add a generic excluded-range list mhonap
2026-09-22  2:13   ` Alex Williamson
2026-09-25 22:05   ` Jonathan Cameron
2026-09-16 18:35 ` [PATCH v5 10/27] vfio/pci: Migrate MSI-X exclusion onto the " mhonap
2026-09-16 18:35 ` [PATCH v5 11/27] vfio/pci: Virtualize the CXL DVSEC in vfio_pci_config.c mhonap
2026-09-22  2:14   ` Alex Williamson
2026-09-25 22:15   ` Jonathan Cameron [this message]
2026-09-16 18:35 ` [PATCH v5 12/27] vfio/pci: Call the CXL open and close hooks around device use mhonap
2026-09-25 22:17   ` Jonathan Cameron
2026-09-16 18:35 ` [PATCH v5 13/27] vfio/pci: Bracket PCI resets with the CXL reset hooks mhonap
2026-09-22  2:14   ` Alex Williamson
2026-09-16 18:35 ` [PATCH v5 14/27] vfio/pci: Provide an opt-out for the CXL Type-2 extensions mhonap
2026-09-16 18:35 ` [PATCH v5 15/27] vfio/cxl: Add the vfio-cxl provider module skeleton mhonap
2026-09-16 18:35 ` [PATCH v5 16/27] vfio/cxl: Create the CXL memdev and set media ready at bind mhonap
2026-09-22  2:14   ` Alex Williamson
2026-09-25 22:22   ` Jonathan Cameron
2026-09-16 18:35 ` [PATCH v5 17/27] vfio/cxl: Own the whole component register BAR mhonap
2026-09-22  2:14   ` Alex Williamson
2026-09-16 18:35 ` [PATCH v5 18/27] vfio/cxl: Expose the HDM memory region to the guest mhonap
2026-09-22  2:14   ` Alex Williamson
2026-09-16 18:35 ` [PATCH v5 19/27] vfio/cxl: Contain HDM memory errors with memory_failure() mhonap
2026-09-16 18:35 ` [PATCH v5 20/27] vfio/cxl: Expose the HDM decoder registers read-only to the guest mhonap
2026-09-22  2:13   ` Alex Williamson
2026-09-16 18:35 ` [PATCH v5 21/27] vfio/cxl: Exclude the HDM decoder registers from direct BAR access mhonap
2026-09-17  7:28   ` Richard Cheng
2026-09-21  9:52     ` Manish Honap
2026-09-16 18:35 ` [PATCH v5 22/27] vfio/cxl: Clear the HDM access gate after a hot reset mhonap
2026-09-22  2:13   ` Alex Williamson
2026-09-16 18:35 ` [PATCH v5 23/27] vfio/cxl: Describe the CXL device and decoder geometry to userspace mhonap
2026-09-16 18:35 ` [PATCH v5 24/27] vfio/cxl: Export the HDM memory region as a dma-buf mhonap
2026-09-17  7:55   ` Richard Cheng
2026-09-21  9:58     ` Manish Honap
2026-09-16 18:35 ` [PATCH v5 25/27] vfio/cxl: Run the CXL reset at the vfio reset points mhonap
2026-09-17  8:11   ` Richard Cheng
2026-09-21 10:01     ` Manish Honap
2026-09-22  2:13   ` Alex Williamson
2026-09-16 18:35 ` [PATCH v5 26/27] Documentation: vfio-pci: Document CXL Type-2 device passthrough mhonap
2026-09-16 19:33   ` Gregory Price
2026-09-21  9:43     ` Manish Honap
2026-09-24  0:54       ` Jonathan Cameron
2026-09-16 18:35 ` [PATCH v5 27/27] selftests/vfio: Add CXL Type-2 passthrough tests mhonap

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=20260925231541.492d5407@jic23-hlaptop \
    --to=jic23@kernel.org \
    --cc=alejandro.lucero-palau@amd.com \
    --cc=alex@shazbot.org \
    --cc=alison.schofield@intel.com \
    --cc=ankita@nvidia.com \
    --cc=bhelgaas@google.com \
    --cc=cjia@nvidia.com \
    --cc=corbet@lwn.net \
    --cc=dave.jiang@intel.com \
    --cc=dave@stgolabs.net \
    --cc=dmatlack@google.com \
    --cc=gustavoars@kernel.org \
    --cc=iweiny@kernel.org \
    --cc=jgg@ziepe.ca \
    --cc=kees@kernel.org \
    --cc=kevin.tian@intel.com \
    --cc=kjaju@nvidia.com \
    --cc=kvm@vger.kernel.org \
    --cc=linux-cxl@vger.kernel.org \
    --cc=linux-doc@vger.kernel.org \
    --cc=linux-hardening@vger.kernel.org \
    --cc=linux-kernel@vger.kernel.org \
    --cc=linux-kselftest@vger.kernel.org \
    --cc=linux-pci@vger.kernel.org \
    --cc=mhonap@nvidia.com \
    --cc=ming.li@zohomail.com \
    --cc=skhan@linuxfoundation.org \
    --cc=skolothumtho@nvidia.com \
    --cc=smadhavan@nvidia.com \
    --cc=vishal.l.verma@intel.com \
    --cc=vsethi@nvidia.com \
    --cc=yishaih@nvidia.com \
    --cc=zhiw@nvidia.com \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox