Linux Hardening
 help / color / mirror / Atom feed
From: Alex Williamson <alex@shazbot.org>
To: <mhonap@nvidia.com>
Cc: <jgg@ziepe.ca>, <ankita@nvidia.com>, <jic23@kernel.org>,
	<dave.jiang@intel.com>, <alejandro.lucero-palau@amd.com>,
	<smadhavan@nvidia.com>, <corbet@lwn.net>,
	<skhan@linuxfoundation.org>, <dave@stgolabs.net>,
	<alison.schofield@intel.com>, <vishal.l.verma@intel.com>,
	<iweiny@kernel.org>, <ming.li@zohomail.com>, <yishaih@nvidia.com>,
	<skolothumtho@nvidia.com>, <kevin.tian@intel.com>,
	<bhelgaas@google.com>, <dmatlack@google.com>, <kees@kernel.org>,
	<gustavoars@kernel.org>, <cjia@nvidia.com>, <kjaju@nvidia.com>,
	<vsethi@nvidia.com>, <zhiw@nvidia.com>,
	<linux-doc@vger.kernel.org>, <linux-kernel@vger.kernel.org>,
	<kvm@vger.kernel.org>, <linux-cxl@vger.kernel.org>,
	<linux-pci@vger.kernel.org>, <linux-kselftest@vger.kernel.org>,
	<linux-hardening@vger.kernel.org>,
	alex@shazbot.org
Subject: Re: [PATCH v4 23/27] vfio/cxl: Refresh the decoder snapshot after a device reset
Date: Fri, 28 Aug 2026 16:33:24 -0600	[thread overview]
Message-ID: <20260828163324.72a685d1@shazbot.org> (raw)
In-Reply-To: <20260813093631.2288172-24-mhonap@nvidia.com>

On Thu, 13 Aug 2026 15:06:27 +0530
<mhonap@nvidia.com> wrote:

> From: Manish Honap <mhonap@nvidia.com>
> 
> A reset clears the HDM decoder registers, so the guest snapshot has to be
> resampled once the reset settles, on every path that can reset the
> function: the reset ioctl, an FLR driven through config space, and a bus
> hot reset. Re-enable Memory Space first, since a config restore can leave
> it off and the component-BAR read would then take an Unsupported Request.
> 
> The bus hot reset zaps BARs directly rather than through
> vfio_pci_zap_and_down_write_memory_lock(), so it also needs the HDM
> window zapped by hand; route both zap sites through a common helper.
> 
> Signed-off-by: Manish Honap <mhonap@nvidia.com>
> ---
>  drivers/vfio/pci/cxl/vfio_cxl_core.c | 82 ++++++++++++++++++++++++++++
>  drivers/vfio/pci/vfio_pci_config.c   |  2 +
>  drivers/vfio/pci/vfio_pci_core.c     | 25 ++++++++-
>  drivers/vfio/pci/vfio_pci_priv.h     | 20 +++++++
>  include/linux/vfio_pci_core.h        |  4 ++
>  5 files changed, 131 insertions(+), 2 deletions(-)
> 
> diff --git a/drivers/vfio/pci/cxl/vfio_cxl_core.c b/drivers/vfio/pci/cxl/vfio_cxl_core.c
> index f1c6bf06c408..f45eaa60bad2 100644
> --- a/drivers/vfio/pci/cxl/vfio_cxl_core.c
> +++ b/drivers/vfio/pci/cxl/vfio_cxl_core.c
> @@ -555,6 +555,86 @@ static void vfio_cxl_zap(struct vfio_pci_core_device *vdev)
>  			    range_len(&cxl->hpa_range), true);
>  }
>  
> +static void vfio_cxl_post_reset(struct vfio_pci_core_device *vdev)
> +{
> +	struct vfio_cxl_state *cxl = vdev->cxl;
> +	struct pci_dev *pdev = vdev->pdev;
> +	bool re_enabled = false;
> +	int i, dwords;
> +	u16 cmd;
> +
> +	lockdep_assert_held_write(&vdev->memory_lock);
> +
> +	if (!cxl || !cxl->hdm_shadow)
> +		return;
> +
> +	/*
> +	 * The decoder registers are read through the component BAR. A config
> +	 * restore can leave Memory Space disabled, and the read would then
> +	 * return an Unsupported Request, so re-enable it before sampling.
> +	 */
> +	pci_read_config_word(pdev, PCI_COMMAND, &cmd);
> +	if (!(cmd & PCI_COMMAND_MEMORY)) {
> +		pci_write_config_word(pdev, PCI_COMMAND,
> +				      cmd | PCI_COMMAND_MEMORY);
> +		re_enabled = true;
> +	}
> +
> +	dwords = cxl->hdm_len / sizeof(u32);
> +	for (i = 0; i < dwords; i++)
> +		cxl->hdm_shadow[i] = cpu_to_le32(readl(cxl->hdm_regs +
> +						       i * sizeof(u32)));
> +	/*
> +	 * Leave Memory Space as it was found. The guest owns Memory Space
> +	 * through vconfig, so a physical enable done only to sample must not
> +	 * outlive the sampling or the function would decode while vconfig
> +	 * reports it off.
> +	 */
> +	if (re_enabled)
> +		pci_write_config_word(pdev, PCI_COMMAND, cmd);

This is usually done by just saving the original and re-writing it, but
this whole patch is really just ammunition for why it should be read
live rather than shadowed.

See also the reset_done callback in pci_error_handlers, we shouldn't be
open coding calls to this everywhere, but of course this goes away if
we expose it live rather than shadow it.  Thanks,

Alex

> +}
> +
> +static int vfio_cxl_pm_restore(struct vfio_pci_core_device *vdev)
> +{
> +	struct vfio_cxl_state *cxl = vdev->cxl;
> +	struct pci_dev *pdev = vdev->pdev;
> +	int rc;
> +
> +	lockdep_assert_held_write(&vdev->memory_lock);
> +
> +	if (!cxl || !cxl->hdm_shadow) {
> +		pci_dbg(pdev, "vfio-cxl: pm_restore: no shadow (device not open), skipping\n");
> +		return 0;
> +	}
> +
> +	/*
> +	 * A D3hot->D0 transition can soft-reset the function and clear the HDM
> +	 * decoder. Restore the physical decoder before the fault gate re-inserts
> +	 * the mapping. The restore needs the device lock, taken here after
> +	 * memory_lock to match the reset path ordering. On failure the decoder is
> +	 * left unrestored, so close the access gate (zap no longer clears it) and
> +	 * return the error so the caller keeps the HDM range inaccessible.
> +	 */
> +	if (!pci_dev_trylock(pdev)) {
> +		pci_warn(pdev, "vfio-cxl: pm_restore: could not lock device, HDM not restored\n");
> +		cxl->hdm_valid = false;
> +		return -EBUSY;
> +	}
> +
> +	rc = cxl_restore_hdm_after_pci_reset(pdev);
> +	pci_dev_unlock(pdev);
> +	if (rc) {
> +		pci_err(pdev, "vfio-cxl: pm_restore: HDM restore failed: %d\n", rc);
> +		cxl->hdm_valid = false;
> +		return rc;
> +	}
> +
> +	vfio_cxl_post_reset(vdev);
> +	/* The decoder is restored and re-sampled, so reopen the access gate. */
> +	cxl->hdm_valid = true;
> +	return 0;
> +}
> +
>  static int vfio_cxl_open_device(struct vfio_pci_core_device *vdev)
>  {
>  	struct vfio_cxl_state *cxl = vdev->cxl;
> @@ -767,6 +847,8 @@ static const struct vfio_cxl_ops vfio_cxl_ops = {
>  	.config_read	= vfio_cxl_config_read,
>  	.config_write	= vfio_cxl_config_write,
>  	.zap		= vfio_cxl_zap,
> +	.post_reset	= vfio_cxl_post_reset,
> +	.pm_restore	= vfio_cxl_pm_restore,
>  	.owner		= THIS_MODULE,
>  };
>  
> diff --git a/drivers/vfio/pci/vfio_pci_config.c b/drivers/vfio/pci/vfio_pci_config.c
> index f088e4ce5e07..01d808546a4c 100644
> --- a/drivers/vfio/pci/vfio_pci_config.c
> +++ b/drivers/vfio/pci/vfio_pci_config.c
> @@ -911,6 +911,7 @@ static int vfio_exp_config_write(struct vfio_pci_core_device *vdev, int pos,
>  			vfio_pci_zap_and_down_write_memory_lock(vdev);
>  			vfio_pci_dma_buf_move(vdev, true);
>  			pci_try_reset_function(vdev->pdev);
> +			vfio_pci_cxl_post_reset(vdev);
>  			if (__vfio_pci_memory_enabled(vdev))
>  				vfio_pci_dma_buf_move(vdev, false);
>  			up_write(&vdev->memory_lock);
> @@ -996,6 +997,7 @@ static int vfio_af_config_write(struct vfio_pci_core_device *vdev, int pos,
>  			vfio_pci_zap_and_down_write_memory_lock(vdev);
>  			vfio_pci_dma_buf_move(vdev, true);
>  			pci_try_reset_function(vdev->pdev);
> +			vfio_pci_cxl_post_reset(vdev);
>  			if (__vfio_pci_memory_enabled(vdev))
>  				vfio_pci_dma_buf_move(vdev, false);
>  			up_write(&vdev->memory_lock);
> diff --git a/drivers/vfio/pci/vfio_pci_core.c b/drivers/vfio/pci/vfio_pci_core.c
> index 1a54f15d1c2c..fc8235c8b4fc 100644
> --- a/drivers/vfio/pci/vfio_pci_core.c
> +++ b/drivers/vfio/pci/vfio_pci_core.c
> @@ -362,6 +362,14 @@ int vfio_pci_set_power_state(struct vfio_pci_core_device *vdev, pci_power_t stat
>  		} else if (needs_restore) {
>  			pci_load_and_free_saved_state(pdev, &vdev->pm_save);
>  			pci_restore_state(pdev);
> +			/*
> +			 * A NoSoftRst- device soft-resets on D3hot->D0, which can
> +			 * clear a CXL HDM decoder. Restore it before the fault
> +			 * gate re-inserts the HDM mapping. memory_lock is held on
> +			 * this path (the PM config write and runtime PM entry both
> +			 * take it before the D0 transition).
> +			 */
> +			vfio_pci_cxl_pm_restore(vdev);
>  		}
>  	}
>  
> @@ -529,6 +537,13 @@ static int vfio_pci_core_runtime_resume(struct device *dev)
>  		eventfd_signal(vdev->pm_wake_eventfd_ctx);
>  		__vfio_pci_runtime_pm_exit(vdev);
>  	}
> +	/*
> +	 * A NoSoftRst- function can soft-reset on the runtime D3hot->D0
> +	 * transition and clear a CXL HDM decoder. Restore it while memory_lock
> +	 * is held, before the fault gate can re-insert the HDM mapping. PCI
> +	 * config restore alone does not restore the component decoder registers.
> +	 */
> +	vfio_pci_cxl_pm_restore(vdev);
>  	up_write(&vdev->memory_lock);
>  
>  	if (vdev->pm_intx_masked)
> @@ -1449,6 +1464,7 @@ static int vfio_pci_ioctl_reset(struct vfio_pci_core_device *vdev,
>  
>  	vfio_pci_dma_buf_move(vdev, true);
>  	ret = pci_try_reset_function(vdev->pdev);
> +	vfio_pci_cxl_post_reset(vdev);
>  	if (__vfio_pci_memory_enabled(vdev))
>  		vfio_pci_dma_buf_move(vdev, false);
>  	up_write(&vdev->memory_lock);
> @@ -1838,8 +1854,7 @@ void vfio_pci_zap_and_down_write_memory_lock(struct vfio_pci_core_device *vdev)
>  	 * a runtime-PM entry, D3 transition, or reset would leave the guest
>  	 * with live mappings into a quiesced device.
>  	 */
> -	if (vdev->cxl_ops && vdev->cxl_ops->zap)
> -		vdev->cxl_ops->zap(vdev);
> +	vfio_pci_cxl_zap(vdev);
>  }
>  
>  u16 vfio_pci_memory_lock_and_enable(struct vfio_pci_core_device *vdev)
> @@ -2780,6 +2795,8 @@ static int vfio_pci_dev_set_hot_reset(struct vfio_device_set *dev_set,
>  
>  		vfio_pci_dma_buf_move(vdev, true);
>  		vfio_pci_zap_bars(vdev);
> +		/* zap_bars misses the HDM window; bus reset needs it too */
> +		vfio_pci_cxl_zap(vdev);
>  	}
>  
>  	if (!list_entry_is_head(vdev,
> @@ -2802,6 +2819,10 @@ static int vfio_pci_dev_set_hot_reset(struct vfio_device_set *dev_set,
>  
>  	ret = pci_reset_bus(pdev);
>  
> +	/* Re-sample decoder state for any CXL device the bus reset touched. */
> +	list_for_each_entry(vdev, &dev_set->device_list, vdev.dev_set_list)
> +		vfio_pci_cxl_post_reset(vdev);
> +
>  	vdev = list_last_entry(&dev_set->device_list,
>  			       struct vfio_pci_core_device, vdev.dev_set_list);
>  
> diff --git a/drivers/vfio/pci/vfio_pci_priv.h b/drivers/vfio/pci/vfio_pci_priv.h
> index 902d17815ab6..46e67573d264 100644
> --- a/drivers/vfio/pci/vfio_pci_priv.h
> +++ b/drivers/vfio/pci/vfio_pci_priv.h
> @@ -82,6 +82,26 @@ int vfio_pci_set_power_state(struct vfio_pci_core_device *vdev,
>  			     pci_power_t state);
>  
>  void vfio_pci_zap_and_down_write_memory_lock(struct vfio_pci_core_device *vdev);
> +
> +static inline void vfio_pci_cxl_zap(struct vfio_pci_core_device *vdev)
> +{
> +	if (vdev->cxl_ops && vdev->cxl_ops->zap)
> +		vdev->cxl_ops->zap(vdev);
> +}
> +
> +static inline void vfio_pci_cxl_post_reset(struct vfio_pci_core_device *vdev)
> +{
> +	if (vdev->cxl_ops && vdev->cxl_ops->post_reset)
> +		vdev->cxl_ops->post_reset(vdev);
> +}
> +
> +static inline int vfio_pci_cxl_pm_restore(struct vfio_pci_core_device *vdev)
> +{
> +	if (vdev->cxl_ops && vdev->cxl_ops->pm_restore)
> +		return vdev->cxl_ops->pm_restore(vdev);
> +	return 0;
> +}
> +
>  u16 vfio_pci_memory_lock_and_enable(struct vfio_pci_core_device *vdev);
>  void vfio_pci_memory_unlock_and_restore(struct vfio_pci_core_device *vdev,
>  					u16 cmd);
> diff --git a/include/linux/vfio_pci_core.h b/include/linux/vfio_pci_core.h
> index 8b93949d4484..c438d968dc59 100644
> --- a/include/linux/vfio_pci_core.h
> +++ b/include/linux/vfio_pci_core.h
> @@ -78,6 +78,10 @@ struct vfio_cxl_ops {
>  				int count, __le32 val);
>  	/* Revoke the HDM mapping; paired with the BAR zap */
>  	void    (*zap)(struct vfio_pci_core_device *vdev);
> +	/* Re-sample the decoder state once a reset has settled */
> +	void    (*post_reset)(struct vfio_pci_core_device *vdev);
> +	/* Restore the HDM decoder after a D3hot->D0 soft reset */
> +	int     (*pm_restore)(struct vfio_pci_core_device *vdev);
>  
>  	/* Pinned per bound CXL device so vfio-cxl cannot unload under usage */
>  	struct module *owner;


  reply	other threads:[~2026-08-28 22:33 UTC|newest]

Thread overview: 53+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-08-13  9:36 [PATCH v4 00/27] vfio/pci: Add CXL Type-2 device passthrough support mhonap
2026-08-13  9:36 ` [PATCH v4 01/27] cxl: Fix resource.c include path and export cxl_restore_hdm_after_pci_reset mhonap
2026-08-21 22:52   ` Jonathan Cameron
2026-08-22  1:22     ` Manish Honap
2026-08-13  9:36 ` [PATCH v4 02/27] cxl/regs: Skip sub-block region request for BAR-owning drivers mhonap
2026-08-25 21:26   ` Alex Williamson
2026-08-13  9:36 ` [PATCH v4 03/27] cxl: Move component register defines to uapi/cxl/cxl_regs.h mhonap
2026-08-28 15:27   ` Dave Jiang
2026-08-13  9:36 ` [PATCH v4 04/27] cxl: Establish media readiness in cxl_mem_probe() mhonap
2026-08-25 22:18   ` Alex Williamson
2026-08-28 15:59   ` Dave Jiang
2026-08-13  9:36 ` [PATCH v4 05/27] cxl: Add a function-scoped reset entry for vfio-pci mhonap
2026-08-25 23:11   ` Alex Williamson
2026-08-13  9:36 ` [PATCH v4 06/27] vfio/pci: Add CXL ops registration interface mhonap
2026-08-26 21:11   ` Alex Williamson
2026-08-13  9:36 ` [PATCH v4 07/27] vfio/pci: Detect CXL devices and load vfio-cxl on demand mhonap
2026-08-26 22:17   ` Alex Williamson
2026-08-13  9:36 ` [PATCH v4 08/27] vfio/cxl: Add the vfio-cxl module skeleton mhonap
2026-08-26 22:50   ` Alex Williamson
2026-08-13  9:36 ` [PATCH v4 09/27] vfio/cxl: Create the CXL memory device at bind mhonap
2026-08-27 20:43   ` Alex Williamson
2026-08-13  9:36 ` [PATCH v4 10/27] vfio/cxl: Reject unsupported decoder topologies " mhonap
2026-08-27 20:58   ` Alex Williamson
2026-08-13  9:36 ` [PATCH v4 11/27] vfio/cxl: Own the whole component register BAR mhonap
2026-08-27 21:13   ` Alex Williamson
2026-08-13  9:36 ` [PATCH v4 12/27] vfio/pci: Let a provider exclude a BAR sub-range from mmap mhonap
2026-08-27 22:37   ` Alex Williamson
2026-08-13  9:36 ` [PATCH v4 13/27] vfio/pci: Refuse read/write to an excluded BAR sub-range mhonap
2026-08-13  9:36 ` [PATCH v4 14/27] vfio: Add CXL region type for the HDM region mhonap
2026-08-27 22:43   ` Alex Williamson
2026-08-13  9:36 ` [PATCH v4 15/27] vfio/pci: Call CXL open and close hooks around device use mhonap
2026-08-27 23:03   ` Alex Williamson
2026-08-13  9:36 ` [PATCH v4 16/27] vfio/cxl: Shadow the CXL DVSEC body at open mhonap
2026-08-28 15:09   ` Alex Williamson
2026-08-13  9:36 ` [PATCH v4 17/27] vfio/cxl: Virtualize the CXL DVSEC mhonap
2026-08-28 16:39   ` Alex Williamson
2026-08-13  9:36 ` [PATCH v4 18/27] vfio/cxl: Expose the HDM memory and trap the decoder registers mhonap
2026-08-28 20:53   ` Alex Williamson
2026-08-13  9:36 ` [PATCH v4 19/27] vfio/cxl: Keep the HDM decoder block off the direct BAR mapping mhonap
2026-08-13  9:36 ` [PATCH v4 20/27] vfio/cxl: Emulate the HDM decoder commit handshake mhonap
2026-08-13  9:36 ` [PATCH v4 21/27] vfio/cxl: Describe the CXL device and decoder geometry to userspace mhonap
2026-08-13  9:36 ` [PATCH v4 22/27] vfio/cxl: Revoke the HDM mapping on reset and power transitions mhonap
2026-08-28 21:54   ` Alex Williamson
2026-08-13  9:36 ` [PATCH v4 23/27] vfio/cxl: Refresh the decoder snapshot after a device reset mhonap
2026-08-28 22:33   ` Alex Williamson [this message]
2026-08-13  9:36 ` [PATCH v4 24/27] vfio/cxl: Service a guest-triggered CXL reset mhonap
2026-08-28 23:09   ` Alex Williamson
2026-08-13  9:36 ` [PATCH v4 25/27] vfio/pci: Provide an opt-out for the CXL Type-2 extensions mhonap
2026-08-28 22:56   ` Alex Williamson
2026-08-13  9:36 ` [PATCH v4 26/27] Documentation: vfio-pci: Document CXL Type-2 device passthrough mhonap
2026-08-13  9:36 ` [PATCH v4 27/27] selftests/vfio: Add CXL Type-2 passthrough corner-case tests mhonap
2026-08-26  7:28   ` Shuai Xue
2026-08-26 16:17     ` Manish Honap

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=20260828163324.72a685d1@shazbot.org \
    --to=alex@shazbot.org \
    --cc=alejandro.lucero-palau@amd.com \
    --cc=alison.schofield@intel.com \
    --cc=ankita@nvidia.com \
    --cc=bhelgaas@google.com \
    --cc=cjia@nvidia.com \
    --cc=corbet@lwn.net \
    --cc=dave.jiang@intel.com \
    --cc=dave@stgolabs.net \
    --cc=dmatlack@google.com \
    --cc=gustavoars@kernel.org \
    --cc=iweiny@kernel.org \
    --cc=jgg@ziepe.ca \
    --cc=jic23@kernel.org \
    --cc=kees@kernel.org \
    --cc=kevin.tian@intel.com \
    --cc=kjaju@nvidia.com \
    --cc=kvm@vger.kernel.org \
    --cc=linux-cxl@vger.kernel.org \
    --cc=linux-doc@vger.kernel.org \
    --cc=linux-hardening@vger.kernel.org \
    --cc=linux-kernel@vger.kernel.org \
    --cc=linux-kselftest@vger.kernel.org \
    --cc=linux-pci@vger.kernel.org \
    --cc=mhonap@nvidia.com \
    --cc=ming.li@zohomail.com \
    --cc=skhan@linuxfoundation.org \
    --cc=skolothumtho@nvidia.com \
    --cc=smadhavan@nvidia.com \
    --cc=vishal.l.verma@intel.com \
    --cc=vsethi@nvidia.com \
    --cc=yishaih@nvidia.com \
    --cc=zhiw@nvidia.com \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox