From: "Luck, Tony" <tony.luck@intel.com>
To: Terry Bowman <terry.bowman@amd.com>
Cc: Jonathan Cameron <jic23@kernel.org>,
Dave Jiang <dave.jiang@intel.com>,
Alison Schofield <alison.schofield@intel.com>,
Vishal Verma <vishal.l.verma@intel.com>,
Davidlohr Bueso <dave@stgolabs.net>,
"Bjorn Helgaas" <bhelgaas@google.com>,
"Rafael J . Wysocki" <rafael@kernel.org>,
Jonathan Corbet <corbet@lwn.net>, <linux-cxl@vger.kernel.org>,
"Borislav Petkov" <bp@alien8.de>,
Hanjun Guo <guohanjun@huawei.com>,
"Mauro Carvalho Chehab" <mchehab@kernel.org>,
Shuai Xue <xueshuai@linux.alibaba.com>,
"Len Brown" <lenb@kernel.org>, Ira Weiny <iweiny@kernel.org>,
Li Ming <ming.li@zohomail.com>,
Shuah Khan <skhan@linuxfoundation.org>,
Ben Cheatham <Benjamin.Cheatham@amd.com>,
Richard Cheng <icheng@nvidia.com>,
"Robert Richter" <rrichter@amd.com>, <linux-pci@vger.kernel.org>,
<linux-acpi@vger.kernel.org>, <linux-doc@vger.kernel.org>,
<linux-kernel@vger.kernel.org>
Subject: Re: [PATCH v19 03/14] acpi/apei/ghes: Use raw_spinlock_t for CXL CPER work locks
Date: Wed, 5 Aug 2026 11:41:25 -0700 [thread overview]
Message-ID: <anOD1casyoxFOgG7@agluck-desk3> (raw)
In-Reply-To: <20260803221810.3685703-4-terry.bowman@amd.com>
On Mon, Aug 03, 2026 at 05:17:59PM -0500, Terry Bowman wrote:
> The CXL CPER work registration and unregistration helpers acquire
> cxl_cper_work_lock and cxl_cper_prot_err_work_lock with a spinlock
> guard(), which leaves local interrupts enabled. The corresponding post
> paths (cxl_cper_post_event(), cxl_cper_post_prot_err()) execute in hard
> IRQ context (they are called from the GHES error notification path) and
> acquire the same locks with an irqsave guard().
>
> If a CPU is holding one of these locks via a spinlock guard() when a GHES
> interrupt arrives on the same CPU, the IRQ handler spins on the held lock
> waiting for it to release, while the lock holder is preempted by the IRQ.
> The result is a deadlock.
>
> Convert both locks from spinlock_t to raw_spinlock_t and use guard() at
> all call sites. On PREEMPT_RT kernels spinlock_t is backed by rt_mutex and
> sleeping from hard IRQ context is not permitted; raw_spinlock_t is safe in
> both contexts.
>
> Add WARN_ONCE to both register functions to surface double-registration
> bugs at runtime.
>
> Restructure both unregister functions to clear the global work pointer
> under the lock before calling cancel_work_sync(), closing the window
> where a CPER interrupt could schedule work on a pointer about to be
> freed. Add kfifo_reset() after cancel_work_sync() so stale entries
> are not replayed on next module load.
>
> Both kfifos are single-consumer: only one work_struct is registered at
> a time, enforced by the WARN_ONCE guard in the register functions.
> kfifo_reset() is safe outside the lock because cancel_work_sync() has
> already quiesced the consumer, and no new consumer can register until
> the current module exit completes and a fresh module init runs.
>
> Remove the redundant cancel_work_sync() call from cxl_ras_exit() and
> cxl_pci_driver_exit(). The CPER unregister functions now quiesce
> the work internally.
>
> Reported-by: Sashiko <sashiko@linuxfoundation.org>
> Signed-off-by: Terry Bowman <terry.bowman@amd.com>
> Fixes: 5e4a264bf8b5 ("acpi/ghes: Process CXL Component Events")
> Fixes: 36f257e3b0ba ("acpi/ghes, cxl/pci: Process CXL CPER Protocol Errors")
> Cc: stable@vger.kernel.org
> Reviewed-by: Dave Jiang <dave.jiang@intel.com>
> Reviewed-by: Jonathan Cameron <jonathan.cameron@oss.qualcomm.com>
Reviewed-by: Tony Luck <tony.luck@intel.com>
[But could this be split ... the commit message feels like a list
of three changes. No strong feelings about this]
-Tony
next prev parent reply other threads:[~2026-08-05 18:41 UTC|newest]
Thread overview: 36+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-08-03 22:17 [PATCH v19 00/14] Enable CXL PCIe Port Protocol Error handling and logging Terry Bowman
2026-08-03 22:17 ` [PATCH v19 01/14] cxl/ras: Fix cxl_rch_get_aer_info() out-of-bounds AER register read Terry Bowman
2026-08-04 2:10 ` Alison Schofield
2026-08-03 22:17 ` [PATCH v19 02/14] cxl/ras: Fix cxl_rch_get_aer_severity() wrong severity register Terry Bowman
2026-08-04 2:11 ` Alison Schofield
2026-08-09 15:57 ` Lukas Wunner
2026-09-03 8:35 ` Lukas Wunner
2026-09-07 14:52 ` Terry Bowman
2026-09-09 16:13 ` Bowman, Terry
2026-08-03 22:17 ` [PATCH v19 03/14] acpi/apei/ghes: Use raw_spinlock_t for CXL CPER work locks Terry Bowman
2026-08-05 18:41 ` Luck, Tony [this message]
2026-08-03 22:18 ` [PATCH v19 04/14] cxl: Tighten CPER kfifo registration API and symbol visibility Terry Bowman
2026-08-04 2:13 ` Alison Schofield
2026-08-03 22:18 ` [PATCH v19 05/14] cxl: Rename find_cxl_port() to find_cxl_port_by_dport() Terry Bowman
2026-08-04 2:14 ` Alison Schofield
2026-08-03 22:18 ` [PATCH v19 06/14] PCI/AER: Introduce AER-CXL protocol error kfifo Terry Bowman
2026-08-04 8:15 ` Richard Cheng
2026-08-04 14:10 ` Bowman, Terry
2026-08-03 22:18 ` [PATCH v19 07/14] PCI: Establish common CXL Port protocol error flow Terry Bowman
2026-08-03 22:18 ` [PATCH v19 08/14] cxl/ras: Handle RCH correctable and uncorrectable errors in one pass Terry Bowman
2026-08-04 2:16 ` Alison Schofield
2026-08-03 22:18 ` [PATCH v19 09/14] cxl/pci: Thread port and dport through RAS handling helpers Terry Bowman
2026-08-04 2:16 ` Alison Schofield
2026-08-03 22:18 ` [PATCH v19 10/14] cxl: Update CXL Endpoint AER handler Terry Bowman
2026-08-04 2:17 ` Alison Schofield
2026-08-03 22:18 ` [PATCH v19 11/14] PCI: Cache PCI DSN into pci_dev->dsn during probe Terry Bowman
2026-08-04 2:26 ` Alison Schofield
2026-08-03 22:18 ` [PATCH v19 12/14] cxl: Add port and dport identifiers to CXL AER trace events Terry Bowman
2026-08-04 2:27 ` Alison Schofield
2026-08-04 7:56 ` Richard Cheng
2026-08-04 13:46 ` Bowman, Terry
2026-08-03 22:18 ` [PATCH v19 13/14] PCI/CXL: Mask/Unmask CXL protocol errors Terry Bowman
2026-08-04 2:29 ` Alison Schofield
2026-08-03 22:18 ` [PATCH v19 14/14] Documentation: cxl: Document CXL protocol error handling Terry Bowman
2026-08-04 2:30 ` Alison Schofield
2026-08-05 21:19 ` [PATCH v19 00/14] Enable CXL PCIe Port Protocol Error handling and logging Dave Jiang
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=anOD1casyoxFOgG7@agluck-desk3 \
--to=tony.luck@intel.com \
--cc=Benjamin.Cheatham@amd.com \
--cc=alison.schofield@intel.com \
--cc=bhelgaas@google.com \
--cc=bp@alien8.de \
--cc=corbet@lwn.net \
--cc=dave.jiang@intel.com \
--cc=dave@stgolabs.net \
--cc=guohanjun@huawei.com \
--cc=icheng@nvidia.com \
--cc=iweiny@kernel.org \
--cc=jic23@kernel.org \
--cc=lenb@kernel.org \
--cc=linux-acpi@vger.kernel.org \
--cc=linux-cxl@vger.kernel.org \
--cc=linux-doc@vger.kernel.org \
--cc=linux-kernel@vger.kernel.org \
--cc=linux-pci@vger.kernel.org \
--cc=mchehab@kernel.org \
--cc=ming.li@zohomail.com \
--cc=rafael@kernel.org \
--cc=rrichter@amd.com \
--cc=skhan@linuxfoundation.org \
--cc=terry.bowman@amd.com \
--cc=vishal.l.verma@intel.com \
--cc=xueshuai@linux.alibaba.com \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox