All of lore.kernel.org
 help / color / mirror / Atom feed
From: Dave Jiang <dave.jiang@intel.com>
To: Guixin Liu <kanie@linux.alibaba.com>,
	Davidlohr Bueso <dave@stgolabs.net>,
	Jonathan Cameron <jic23@kernel.org>,
	Alison Schofield <alison.schofield@intel.com>,
	Vishal Verma <vishal.l.verma@intel.com>,
	Dan Williams <djbw@kernel.org>, Ira Weiny <iweiny@kernel.org>,
	Li Ming <ming.li@zohomail.com>
Cc: linux-cxl@vger.kernel.org
Subject: Re: [PATCH v5] cxl/pci: Skip reset detection for DVSEC emulated decoders
Date: Wed, 9 Sep 2026 09:21:04 -0700	[thread overview]
Message-ID: <ec6c61c3-5c28-40c4-be86-4e6ee90bb2eb@intel.com> (raw)
In-Reply-To: <20260831110449.719086-1-kanie@linux.alibaba.com>



On 8/31/26 4:04 AM, Guixin Liu wrote:
> HDM decoders are emulated from the DVSEC range registers in two cases:
> (a) the component registers expose no HDM decoder capability, or (b) the
> capability is present but the DVSEC ranges were the ones in use at driver
> load.
> 
> After an FLR or SBR, __cxl_endpoint_decoder_reset_detected() reads the HDM
> decoder Committed bit for every decoder marked enabled, emulated ones
> included. In case (a) regs.hdm_decoder is NULL and the read oopses. In case
> (b) the Committed bit was never set, so a reset gets reported that never
> happened.
> 
> Use the absence of cxld->commit to elide the check for emulated decoders.
> 
> Case (b) tested under QEMU: a reset on an endpoint driven down the DVSEC
> emulation path no longer reports a reset or strips the decoder flags.
> 
> Fixes: 934edcd436dc ("cxl: Add post-reset warning if reset results in loss of previously committed HDM decoders")
> Signed-off-by: Guixin Liu <kanie@linux.alibaba.com>
> Reviewed-by: Jonathan Cameron <jonathan.cameron@oss.qualcomm.com>
> ---
> This was patch 4/8 of the "cxl: Assorted fixes" series [1], resent
> individually per review feedback.
> 
> The two synchronisation concerns the review bot raised on v2 - the missing
> exclusion against a cxl_port unbind freeing the devm allocated cxl_hdm, and
> the plain read-modify-write of cxld->flags in
> cxl_endpoint_decoder_clear_reset_flags() while the region paths update the
> same word under cxl_rwsem.region - are still untouched here. Both are
> about the reset handler's locking rather than about which registers it
> reads, and neither fix is local; happy to follow up with separate patches.
> 
> Testing
> 
> Reaching case (b) needs an endpoint whose DVSEC ranges are live while the
> global HDM decoder enable is clear. Unbinding cxl_pci runs the
> disable_hdm() and clear_mem_enable() devm actions, so both bits start
> clean. setpci then programs DVSEC Range 1 Base to 0x890000000, which sits
> inside a RAM capable CFMWS, and sets Mem_Enable. Rebinding takes the
> emulation path:
> 
>   cxl_port endpoint3: Fallback map 1 range register
>   cxl_pci 0000:35:00.0: DVSEC Range0 allowed by platform
> 
> The endpoint ends up with a single decoder, enabled and locked, with no
> Committed bit behind it:
> 
>   decoder3.0: locked=1 start=0x890000000 size=0x100000000 mode=ram
> 
> The device's only available reset method is cxl_bus, which unmasks SBR
> through the port DVSEC. Before this change:
> 
>   cxl_pci 0000:35:00.0: resetting
>   cxl_pci 0000:35:00.0: SBR happened without memory regions removal.
>   cxl_pci 0000:35:00.0: System may be unstable if regions hosted system memory.
> 
> /proc/sys/kernel/tainted picked up TAINT_USER (8192 -> 8256) and locked
> dropped from 1 to 0. After this change the same reset logs only
> "resetting", locked stays 1, and no taint is added. The decoder is
> decoder6.0 in the second run because reloading cxl_core renumbered the
> memdevs.
> 
> QEMU does not clear the DVSEC range registers on SBR, so the emulated
> decode is still fully programmed by the time the reset completes. The old
> report was spurious here too.
> 
> Not covered: case (a) needs an endpoint with no component register block,
> which QEMU's type3 always provides, and the FLR variant needs FLR support
> that QEMU's type3 does not advertise (FLReset-). I also did not run a
> positive control confirming that a normally committed HDM decoder is still
> detected after this change.
> 
> v1->v2:
> - rebase onto cxl/next
> - rewrite the commit message to describe the behaviour rather than narrate
>   the code change (Alison Schofield)
> 
> v2->v3:
> - test cxld->commit instead of cxlhdm->regs.hdm_decoder, so that decoders
>   emulated from the DVSEC ranges are also skipped when the HDM decoder
>   registers exist but are globally disabled (Richard Cheng)
> - update the subject and the commit message for the widened scope
> 
> v3->v4:
> - shorten the commit message (Dave Jiang)
> 
> v4->v5:
> - condense the commit message further, wording taken from Jonathan
>   Cameron's suggestion on v4
> - note the case (b) test in the commit message
> 
> [1] https://lore.kernel.org/linux-cxl/20260811113608.2815625-1-kanie@linux.alibaba.com/
> 
>  drivers/cxl/core/pci.c | 7 +++++++
>  1 file changed, 7 insertions(+)
> 
> diff --git a/drivers/cxl/core/pci.c b/drivers/cxl/core/pci.c
> index 9d807c1a002c..d8b07f86bab0 100644
> --- a/drivers/cxl/core/pci.c
> +++ b/drivers/cxl/core/pci.c
> @@ -683,6 +683,13 @@ static int __cxl_endpoint_decoder_reset_detected(struct device *dev, void *data)
>  	if ((cxld->flags & CXL_DECODER_F_ENABLE) == 0)
>  		return 0;
>  
> +	/*
> +	 * Decoders emulated from the DVSEC range registers have no commit
> +	 * callback and no HDM decoder registers to consult.
> +	 */
> +	if (!cxld->commit)
> +		return 0;

Do we want a warning here given that if SBR is issued and we end up skipping? I think this is the issue sashiko raised.

DJ

> +
>  	cxlhdm = dev_get_drvdata(&port->dev);
>  	hdm = cxlhdm->regs.hdm_decoder;
>  	ctrl = readl(hdm + CXL_HDM_DECODER0_CTRL_OFFSET(cxld->id));
> 
> base-commit: 7098e9cd98a05c0c5de2fae0c2465f9d966fdd07


  parent reply	other threads:[~2026-09-09 16:21 UTC|newest]

Thread overview: 5+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-08-31 11:04 [PATCH v5] cxl/pci: Skip reset detection for DVSEC emulated decoders Guixin Liu
2026-08-31 11:17 ` sashiko-bot
2026-09-09  9:22 ` Guixin Liu
2026-09-09 16:21 ` Dave Jiang [this message]
2026-09-10  8:32   ` Guixin Liu

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=ec6c61c3-5c28-40c4-be86-4e6ee90bb2eb@intel.com \
    --to=dave.jiang@intel.com \
    --cc=alison.schofield@intel.com \
    --cc=dave@stgolabs.net \
    --cc=djbw@kernel.org \
    --cc=iweiny@kernel.org \
    --cc=jic23@kernel.org \
    --cc=kanie@linux.alibaba.com \
    --cc=linux-cxl@vger.kernel.org \
    --cc=ming.li@zohomail.com \
    --cc=vishal.l.verma@intel.com \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is an external index of several public inboxes,
see mirroring instructions on how to clone and mirror
all data and code used by this external index.