From: Alex Williamson <alex@shazbot.org>
To: <mhonap@nvidia.com>
Cc: <jgg@ziepe.ca>, <ankita@nvidia.com>, <jic23@kernel.org>,
<dave.jiang@intel.com>, <alejandro.lucero-palau@amd.com>,
<smadhavan@nvidia.com>, <corbet@lwn.net>,
<skhan@linuxfoundation.org>, <dave@stgolabs.net>,
<alison.schofield@intel.com>, <vishal.l.verma@intel.com>,
<iweiny@kernel.org>, <ming.li@zohomail.com>, <yishaih@nvidia.com>,
<skolothumtho@nvidia.com>, <kevin.tian@intel.com>,
<bhelgaas@google.com>, <dmatlack@google.com>, <kees@kernel.org>,
<gustavoars@kernel.org>, <cjia@nvidia.com>, <kjaju@nvidia.com>,
<vsethi@nvidia.com>, <zhiw@nvidia.com>,
<linux-doc@vger.kernel.org>, <linux-kernel@vger.kernel.org>,
<kvm@vger.kernel.org>, <linux-cxl@vger.kernel.org>,
<linux-pci@vger.kernel.org>, <linux-kselftest@vger.kernel.org>,
<linux-hardening@vger.kernel.org>,
alex@shazbot.org
Subject: Re: [PATCH v4 04/27] cxl: Establish media readiness in cxl_mem_probe()
Date: Tue, 25 Aug 2026 16:18:39 -0600 [thread overview]
Message-ID: <20260825161839.2cb3a6c7@shazbot.org> (raw)
In-Reply-To: <20260813093631.2288172-5-mhonap@nvidia.com>
On Thu, 13 Aug 2026 15:06:08 +0530
<mhonap@nvidia.com> wrote:
> From: Manish Honap <mhonap@nvidia.com>
>
> media_ready was only ever set by cxl_pci, in advance of registering a
> memdev and with CXL Memory Device register assumptions. A consumer that
> creates a memdev without cxl_pci, such as a mailbox-less Type-2
> accelerator brought up through devm_cxl_probe_mem(), therefore handed
> cxl_mem a device with media_ready still false, and cxl_mem_probe()
> rejected it with -EBUSY. __devm_cxl_add_memdev() turns that into -ENXIO
> back to the caller and the bind fails.
>
> Move the readiness wait into cxl_mem_probe() so every memdev consumer
> gets a ready resource regardless of how the memdev was created. When
> media_ready is not already set, wait on the device's DVSEC
> Mem_Info_Valid and Mem_Active bits and mark it ready. cxl_pci keeps
> setting media_ready before it registers its memdev, so that path skips
> the wait.
>
> The CXL Memory Device register group is optional and many Type-2 devices
> do not implement it, so cxl_await_media_ready() must not read the Memdev
> Status register unless the group is mapped. Reading regs.memdev on a
> device that lacks it would fault. Gate that read on regs.memdev; the
> DVSEC bits already prove readiness for such devices.
>
> Signed-off-by: Manish Honap <mhonap@nvidia.com>
> ---
> drivers/cxl/core/pci.c | 15 +++++++++++----
> drivers/cxl/mem.c | 9 +++++++--
> 2 files changed, 18 insertions(+), 6 deletions(-)
>
> diff --git a/drivers/cxl/core/pci.c b/drivers/cxl/core/pci.c
> index 08d4c955137d..9b372d5a1aa4 100644
> --- a/drivers/cxl/core/pci.c
> +++ b/drivers/cxl/core/pci.c
> @@ -151,7 +151,6 @@ int cxl_await_media_ready(struct cxl_dev_state *cxlds)
> struct pci_dev *pdev = to_pci_dev(cxlds->dev);
> int d = cxlds->cxl_dvsec;
> int rc, i, hdm_count;
> - u64 md_status;
> u16 cap;
>
> rc = pci_read_config_word(pdev,
> @@ -172,9 +171,17 @@ int cxl_await_media_ready(struct cxl_dev_state *cxlds)
> return rc;
> }
>
> - md_status = readq(cxlds->regs.memdev + CXLMDEV_STATUS_OFFSET);
> - if (!CXLMDEV_READY(md_status))
> - return -EIO;
> + /*
> + * It is possible some Type-2 devices (CXL_DEVTYPE_DEVMEM) do not
> + * implement regs.memdev; only consult the Memdev Status register when
> + * the group is actually present.
> + */
> + if (cxlds->regs.memdev) {
> + u64 md_status = readq(cxlds->regs.memdev + CXLMDEV_STATUS_OFFSET);
> +
> + if (!CXLMDEV_READY(md_status))
> + return -EIO;
> + }
>
> return 0;
> }
This looks like it should be two separate patches. The change below
depends on the above, but the above change stands on its own.
> diff --git a/drivers/cxl/mem.c b/drivers/cxl/mem.c
> index 798e5c369cfc..9c6e99b9124c 100644
> --- a/drivers/cxl/mem.c
> +++ b/drivers/cxl/mem.c
> @@ -105,8 +105,13 @@ static int cxl_mem_probe(struct device *dev)
> struct dentry *dentry;
> int rc;
>
> - if (!cxlds->media_ready)
> - return -EBUSY;
> + if (!cxlds->media_ready) {
> + rc = cxl_await_media_ready(cxlds);
> + if (rc)
> + return rc;
> + cxlds->media_ready = true;
> + dev_dbg(dev, "CXL media ready\n");
> + }
>
> /*
> * Someone is trying to reattach this device after it lost its port
LLM review is noting a plausible behavioral change here; if there is an
actual media issue, it seems it's flagged in cxl_pci_probe() by leaving
cxlds.media_ready false. That path can then go on to find memory
devices, add a CXL_DEVICE_MEMORY_EXPANDER, and via the .probe callback
reach cxl_mem_probe() and wait another timeout delay here for the same
media issue.
Should the mailbox-less path test and set media_ready before adding the
CXL_DEVICE_MEMORY_EXPANDER and getting to this .probe callback, making
it consistently tested upstream of this function? Thanks,
Alex
next prev parent reply other threads:[~2026-08-25 22:18 UTC|newest]
Thread overview: 38+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-08-13 9:36 [PATCH v4 00/27] vfio/pci: Add CXL Type-2 device passthrough support mhonap
2026-08-13 9:36 ` [PATCH v4 01/27] cxl: Fix resource.c include path and export cxl_restore_hdm_after_pci_reset mhonap
2026-08-21 22:52 ` Jonathan Cameron
2026-08-22 1:22 ` Manish Honap
2026-08-13 9:36 ` [PATCH v4 02/27] cxl/regs: Skip sub-block region request for BAR-owning drivers mhonap
2026-08-25 21:26 ` Alex Williamson
2026-08-13 9:36 ` [PATCH v4 03/27] cxl: Move component register defines to uapi/cxl/cxl_regs.h mhonap
2026-08-13 9:36 ` [PATCH v4 04/27] cxl: Establish media readiness in cxl_mem_probe() mhonap
2026-08-25 22:18 ` Alex Williamson [this message]
2026-08-13 9:36 ` [PATCH v4 05/27] cxl: Add a function-scoped reset entry for vfio-pci mhonap
2026-08-25 23:11 ` Alex Williamson
2026-08-13 9:36 ` [PATCH v4 06/27] vfio/pci: Add CXL ops registration interface mhonap
2026-08-26 21:11 ` Alex Williamson
2026-08-13 9:36 ` [PATCH v4 07/27] vfio/pci: Detect CXL devices and load vfio-cxl on demand mhonap
2026-08-26 22:17 ` Alex Williamson
2026-08-13 9:36 ` [PATCH v4 08/27] vfio/cxl: Add the vfio-cxl module skeleton mhonap
2026-08-26 22:50 ` Alex Williamson
2026-08-13 9:36 ` [PATCH v4 09/27] vfio/cxl: Create the CXL memory device at bind mhonap
2026-08-13 9:36 ` [PATCH v4 10/27] vfio/cxl: Reject unsupported decoder topologies " mhonap
2026-08-13 9:36 ` [PATCH v4 11/27] vfio/cxl: Own the whole component register BAR mhonap
2026-08-13 9:36 ` [PATCH v4 12/27] vfio/pci: Let a provider exclude a BAR sub-range from mmap mhonap
2026-08-13 9:36 ` [PATCH v4 13/27] vfio/pci: Refuse read/write to an excluded BAR sub-range mhonap
2026-08-13 9:36 ` [PATCH v4 14/27] vfio: Add CXL region type for the HDM region mhonap
2026-08-13 9:36 ` [PATCH v4 15/27] vfio/pci: Call CXL open and close hooks around device use mhonap
2026-08-13 9:36 ` [PATCH v4 16/27] vfio/cxl: Shadow the CXL DVSEC body at open mhonap
2026-08-13 9:36 ` [PATCH v4 17/27] vfio/cxl: Virtualize the CXL DVSEC mhonap
2026-08-13 9:36 ` [PATCH v4 18/27] vfio/cxl: Expose the HDM memory and trap the decoder registers mhonap
2026-08-13 9:36 ` [PATCH v4 19/27] vfio/cxl: Keep the HDM decoder block off the direct BAR mapping mhonap
2026-08-13 9:36 ` [PATCH v4 20/27] vfio/cxl: Emulate the HDM decoder commit handshake mhonap
2026-08-13 9:36 ` [PATCH v4 21/27] vfio/cxl: Describe the CXL device and decoder geometry to userspace mhonap
2026-08-13 9:36 ` [PATCH v4 22/27] vfio/cxl: Revoke the HDM mapping on reset and power transitions mhonap
2026-08-13 9:36 ` [PATCH v4 23/27] vfio/cxl: Refresh the decoder snapshot after a device reset mhonap
2026-08-13 9:36 ` [PATCH v4 24/27] vfio/cxl: Service a guest-triggered CXL reset mhonap
2026-08-13 9:36 ` [PATCH v4 25/27] vfio/pci: Provide an opt-out for the CXL Type-2 extensions mhonap
2026-08-13 9:36 ` [PATCH v4 26/27] Documentation: vfio-pci: Document CXL Type-2 device passthrough mhonap
2026-08-13 9:36 ` [PATCH v4 27/27] selftests/vfio: Add CXL Type-2 passthrough corner-case tests mhonap
2026-08-26 7:28 ` Shuai Xue
2026-08-26 16:17 ` Manish Honap
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=20260825161839.2cb3a6c7@shazbot.org \
--to=alex@shazbot.org \
--cc=alejandro.lucero-palau@amd.com \
--cc=alison.schofield@intel.com \
--cc=ankita@nvidia.com \
--cc=bhelgaas@google.com \
--cc=cjia@nvidia.com \
--cc=corbet@lwn.net \
--cc=dave.jiang@intel.com \
--cc=dave@stgolabs.net \
--cc=dmatlack@google.com \
--cc=gustavoars@kernel.org \
--cc=iweiny@kernel.org \
--cc=jgg@ziepe.ca \
--cc=jic23@kernel.org \
--cc=kees@kernel.org \
--cc=kevin.tian@intel.com \
--cc=kjaju@nvidia.com \
--cc=kvm@vger.kernel.org \
--cc=linux-cxl@vger.kernel.org \
--cc=linux-doc@vger.kernel.org \
--cc=linux-hardening@vger.kernel.org \
--cc=linux-kernel@vger.kernel.org \
--cc=linux-kselftest@vger.kernel.org \
--cc=linux-pci@vger.kernel.org \
--cc=mhonap@nvidia.com \
--cc=ming.li@zohomail.com \
--cc=skhan@linuxfoundation.org \
--cc=skolothumtho@nvidia.com \
--cc=smadhavan@nvidia.com \
--cc=vishal.l.verma@intel.com \
--cc=vsethi@nvidia.com \
--cc=yishaih@nvidia.com \
--cc=zhiw@nvidia.com \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox