Linux CXL
 help / color / mirror / Atom feed
* [PATCH v2] cxl/events: Return IRQ_NONE when no event is pending
@ 2026-09-12  9:38 Shaikh Kamaluddin
  2026-09-12  9:48 ` sashiko-bot
                   ` (5 more replies)
  0 siblings, 6 replies; 9+ messages in thread
From: Shaikh Kamaluddin @ 2026-09-12  9:38 UTC (permalink / raw)
  To: Davidlohr Bueso, Jonathan Cameron, Dave Jiang, Alison Schofield,
	Vishal Verma, Dan Williams, Ira Weiny, Li Ming
  Cc: anisa.su887, benjamin.cheatham, linux-cxl, linux-kernel

CXL event interrupts may share an MSI/MSI-X vector with other event
logs or device features. Consequently, cxl_event_thread() is registered
with IRQF_SHARED and must determine whether the event-log facility
claimed the interrupt.

The handler masks the Device Event Status register to the event logs
supported by the driver. However, when no supported status bit is set,
it exits the processing loop and still returns IRQ_HANDLED. This
incorrectly reports that the interrupt was claimed and prevents the
generic spurious-interrupt detector from accounting it as unhandled.

Track whether the CXL event handler claimed the interrupt and return
IRQ_NONE when no supported event is detected. If no handler sharing the
vector claims a sustained interrupt storm, the generic spurious-
interrupt detector can eventually identify and disable the affected
IRQ.

Signed-off-by: Shaikh Kamaluddin <shaikhkamal2012@gmail.com>
---
Changes in V2:
- Emphasize that returning IRQ_NONE enables generic spurious-interrupt
  detection and drop the Fixes tag. (Jonathan Cameron)
- Check whether the event interrupt is already disabled by Anisa Su's
  patch; no code changes were needed. (Ben Cheatham)
- Rebase onto v7.3-rc2; no code changes.

 drivers/cxl/pci.c | 5 ++++-
 1 file changed, 4 insertions(+), 1 deletion(-)

diff --git a/drivers/cxl/pci.c b/drivers/cxl/pci.c
index c7c91e8dc51d..8b560cae91f2 100644
--- a/drivers/cxl/pci.c
+++ b/drivers/cxl/pci.c
@@ -515,6 +515,7 @@ static irqreturn_t cxl_event_thread(int irq, void *id)
 	struct cxl_dev_id *dev_id = id;
 	struct cxl_dev_state *cxlds = dev_id->cxlds;
 	struct cxl_memdev_state *mds = to_cxl_memdev_state(cxlds);
+	bool handled = false;
 	u32 status;
 
 	do {
@@ -527,11 +528,13 @@ static irqreturn_t cxl_event_thread(int irq, void *id)
 		status &= CXLDEV_EVENT_STATUS_ALL;
 		if (!status)
 			break;
+
+		handled = true;
 		cxl_mem_get_event_records(mds, status);
 		cond_resched();
 	} while (status);
 
-	return IRQ_HANDLED;
+	return handled ? IRQ_HANDLED : IRQ_NONE;
 }
 
 static int cxl_event_req_irq(struct cxl_dev_state *cxlds, u8 setting)

base-commit: df2908090cda368b01ff43709f51890076c56157
-- 
2.43.0


^ permalink raw reply related	[flat|nested] 9+ messages in thread

* Re: [PATCH v2] cxl/events: Return IRQ_NONE when no event is pending
  2026-09-12  9:38 [PATCH v2] cxl/events: Return IRQ_NONE when no event is pending Shaikh Kamaluddin
@ 2026-09-12  9:48 ` sashiko-bot
  2026-09-16 14:12   ` Shaikh Kamaluddin
  2026-09-15 23:28 ` Jonathan Cameron
                   ` (4 subsequent siblings)
  5 siblings, 1 reply; 9+ messages in thread
From: sashiko-bot @ 2026-09-12  9:48 UTC (permalink / raw)
  To: Shaikh Kamaluddin; +Cc: linux-cxl

Thank you for your contribution! Sashiko AI review found 2 potential issue(s) to consider:

Pre-existing issues:
- [High] The threaded IRQ handler `cxl_event_thread()` can enter an infinite loop if mailbox commands fail to clear event records on the device, or if a PCIe surprise removal causes the MMIO read to persistently return all 1s (~0U).
- [High] The driver blindly trusts the `record_count` and `count` fields returned by the device in mailbox payloads, leading to out-of-bounds kernel memory reads and potential information leaks.
--

commit 659b756fc25aa4e783ce4a882044a537d7b5d643
Author: Shaikh Kamaluddin <shaikhkamal2012@gmail.com>

cxl/events: Return IRQ_NONE when no event is pending

This patch tracks whether the CXL event handler claimed the interrupt and
returns IRQ_NONE when no supported event is detected. This allows the generic
spurious-interrupt detector to identify and disable the affected IRQ during
interrupt storms where no handler claims the event.

> diff --git a/drivers/cxl/pci.c b/drivers/cxl/pci.c
> index c7c91e8dc51dc..8b560cae91f2e 100644
> --- a/drivers/cxl/pci.c
> +++ b/drivers/cxl/pci.c
[ ... ]
> @@ -515,6 +515,7 @@ static irqreturn_t cxl_event_thread(int irq, void *id)
>  	struct cxl_dev_id *dev_id = id;
>  	struct cxl_dev_state *cxlds = dev_id->cxlds;
>  	struct cxl_memdev_state *mds = to_cxl_memdev_state(cxlds);
> +	bool handled = false;
>  	u32 status;
>  
>  	do {
>  		/*
>  		 * CXL 3.0 8.2.8.3.1: The lower 32 bits are the status;
>  		 * ignore the reserved upper 32 bits
>  		 */
>  		status = readl(cxlds->regs.status + CXLDEV_DEV_EVENT_STATUS_OFFSET);
>  		/* Ignore logs unknown to the driver */
>  		status &= CXLDEV_EVENT_STATUS_ALL;
>  		if (!status)
>  			break;

[Severity: High]
This is a pre-existing issue, but can this loop become infinite if a PCIe
surprise removal or fatal device error occurs?

If readl() returns all 1s (~0U), status & CXLDEV_EVENT_STATUS_ALL will
still evaluate to a non-zero value, passing the if (!status) check. The
handler will then call into cxl_mem_get_event_records() where mailbox
commands will fail on the removed device, exiting without clearing the
hardware event status.

Since status is never cleared and the error handling doesn't break the
loop, this evaluates to true at the while (status) condition and repeats
indefinitely. Does this IRQF_ONESHOT threaded handler need to check for ~0U
to avoid monopolizing a CPU and preventing device teardown during hotplug
removal?

> +
> +		handled = true;
>  		cxl_mem_get_event_records(mds, status);

[Severity: High]
This isn't a bug introduced by this patch, but does this code blindly trust
the record_count returned by the device hardware?

When cxl_mem_get_event_records() calls into cxl_mem_get_records_log(), the
driver reads the count directly from the untrusted payload:

drivers/cxl/core/mbox.c:cxl_mem_get_records_log() {
    ...
    nr_rec = le16_to_cpu(payload->record_count);
    ...
    for (i = 0; i < nr_rec; i++)
        __cxl_event_trace_record(cxlmd, type, &payload->records[i]);
    ...
}

If a compromised or malicious CXL device returns an artificially large
record_count, could this read out-of-bounds kernel memory and leak it to
userspace via tracing? 

The same pattern appears in cxl_clear_event_record() and
cxl_mem_get_poison(), where the untrusted count is used as a loop bound.
Should the payload size be validated against mbox_cmd.size_out or the
maximum allocated buffer size before iterating?

>  		cond_resched();
>  	} while (status);
>  
> -	return IRQ_HANDLED;
> +	return handled ? IRQ_HANDLED : IRQ_NONE;
>  }

-- 
Sashiko AI review · https://sashiko.dev/#/patchset/20260912093815.14455-1-shaikhkamal2012@gmail.com?part=1

^ permalink raw reply	[flat|nested] 9+ messages in thread

* Re: [PATCH v2] cxl/events: Return IRQ_NONE when no event is pending
  2026-09-12  9:38 [PATCH v2] cxl/events: Return IRQ_NONE when no event is pending Shaikh Kamaluddin
  2026-09-12  9:48 ` sashiko-bot
@ 2026-09-15 23:28 ` Jonathan Cameron
  2026-09-16  0:38 ` Anisa Su
                   ` (3 subsequent siblings)
  5 siblings, 0 replies; 9+ messages in thread
From: Jonathan Cameron @ 2026-09-15 23:28 UTC (permalink / raw)
  To: Shaikh Kamaluddin
  Cc: Davidlohr Bueso, Dave Jiang, Alison Schofield, Vishal Verma,
	Dan Williams, Ira Weiny, Li Ming, anisa.su887, benjamin.cheatham,
	linux-cxl, linux-kernel

On Sat, 12 Sep 2026 15:08:15 +0530
Shaikh Kamaluddin <shaikhkamal2012@gmail.com> wrote:

> CXL event interrupts may share an MSI/MSI-X vector with other event
> logs or device features. Consequently, cxl_event_thread() is registered
> with IRQF_SHARED and must determine whether the event-log facility
> claimed the interrupt.
> 
> The handler masks the Device Event Status register to the event logs
> supported by the driver. However, when no supported status bit is set,
> it exits the processing loop and still returns IRQ_HANDLED. This
> incorrectly reports that the interrupt was claimed and prevents the
> generic spurious-interrupt detector from accounting it as unhandled.
> 
> Track whether the CXL event handler claimed the interrupt and return
> IRQ_NONE when no supported event is detected. If no handler sharing the
> vector claims a sustained interrupt storm, the generic spurious-
> interrupt detector can eventually identify and disable the affected
> IRQ.
> 
> Signed-off-by: Shaikh Kamaluddin <shaikhkamal2012@gmail.com>
Reviewed-by: Jonathan Cameron <jonathan.cameron@oss.qualcomm.com>

^ permalink raw reply	[flat|nested] 9+ messages in thread

* Re: [PATCH v2] cxl/events: Return IRQ_NONE when no event is pending
  2026-09-12  9:38 [PATCH v2] cxl/events: Return IRQ_NONE when no event is pending Shaikh Kamaluddin
  2026-09-12  9:48 ` sashiko-bot
  2026-09-15 23:28 ` Jonathan Cameron
@ 2026-09-16  0:38 ` Anisa Su
  2026-09-16 14:30   ` Shaikh Kamaluddin
  2026-09-16 12:34 ` Li Ming
                   ` (2 subsequent siblings)
  5 siblings, 1 reply; 9+ messages in thread
From: Anisa Su @ 2026-09-16  0:38 UTC (permalink / raw)
  To: Shaikh Kamaluddin
  Cc: Davidlohr Bueso, Jonathan Cameron, Dave Jiang, Alison Schofield,
	Vishal Verma, Dan Williams, Ira Weiny, Li Ming, anisa.su887,
	benjamin.cheatham, linux-cxl, linux-kernel

On Sat, Sep 12, 2026 at 03:08:15PM +0530, Shaikh Kamaluddin wrote:
> CXL event interrupts may share an MSI/MSI-X vector with other event
> logs or device features. Consequently, cxl_event_thread() is registered
> with IRQF_SHARED and must determine whether the event-log facility
> claimed the interrupt.
> 
> The handler masks the Device Event Status register to the event logs
> supported by the driver. However, when no supported status bit is set,
> it exits the processing loop and still returns IRQ_HANDLED. This
> incorrectly reports that the interrupt was claimed and prevents the
> generic spurious-interrupt detector from accounting it as unhandled.
> 
> Track whether the CXL event handler claimed the interrupt and return
> IRQ_NONE when no supported event is detected. If no handler sharing the
> vector claims a sustained interrupt storm, the generic spurious-
> interrupt detector can eventually identify and disable the affected
> IRQ.
> 
> Signed-off-by: Shaikh Kamaluddin <shaikhkamal2012@gmail.com>
Reviewed-by: Anisa Su <anisa.su@samsung.com>
> ---
> Changes in V2:
> - Emphasize that returning IRQ_NONE enables generic spurious-interrupt
>   detection and drop the Fixes tag. (Jonathan Cameron)
> - Check whether the event interrupt is already disabled by Anisa Su's
>   patch; no code changes were needed. (Ben Cheatham)

One thing I want to correct: The change requested was to say a
spurious IRQ will eventually be disabled in the commit message, not
check whether the event interrupts is already disabled...
See https://lore.kernel.org/linux-cxl/aqLZP3ORYzCLfSgR@acer-nitro-anv15-41/T/#m32c74ebf05eee9a3ee49c57dd24848db29919031
I do see that it's been added to the commit message though, so no need
to do anything.

Thanks,
Anisa
> - Rebase onto v7.3-rc2; no code changes.
> 
>  drivers/cxl/pci.c | 5 ++++-
>  1 file changed, 4 insertions(+), 1 deletion(-)
> 
> diff --git a/drivers/cxl/pci.c b/drivers/cxl/pci.c
> index c7c91e8dc51d..8b560cae91f2 100644
> --- a/drivers/cxl/pci.c
> +++ b/drivers/cxl/pci.c
> @@ -515,6 +515,7 @@ static irqreturn_t cxl_event_thread(int irq, void *id)
>  	struct cxl_dev_id *dev_id = id;
>  	struct cxl_dev_state *cxlds = dev_id->cxlds;
>  	struct cxl_memdev_state *mds = to_cxl_memdev_state(cxlds);
> +	bool handled = false;
>  	u32 status;
>  
>  	do {
> @@ -527,11 +528,13 @@ static irqreturn_t cxl_event_thread(int irq, void *id)
>  		status &= CXLDEV_EVENT_STATUS_ALL;
>  		if (!status)
>  			break;
> +
> +		handled = true;
>  		cxl_mem_get_event_records(mds, status);
>  		cond_resched();
>  	} while (status);
>  
> -	return IRQ_HANDLED;
> +	return handled ? IRQ_HANDLED : IRQ_NONE;
>  }
>  
>  static int cxl_event_req_irq(struct cxl_dev_state *cxlds, u8 setting)
> 
> base-commit: df2908090cda368b01ff43709f51890076c56157
> -- 
> 2.43.0
> 

^ permalink raw reply	[flat|nested] 9+ messages in thread

* Re: [PATCH v2] cxl/events: Return IRQ_NONE when no event is pending
  2026-09-12  9:38 [PATCH v2] cxl/events: Return IRQ_NONE when no event is pending Shaikh Kamaluddin
                   ` (2 preceding siblings ...)
  2026-09-16  0:38 ` Anisa Su
@ 2026-09-16 12:34 ` Li Ming
  2026-09-16 20:49 ` Alison Schofield
  2026-09-18 15:40 ` Dave Jiang
  5 siblings, 0 replies; 9+ messages in thread
From: Li Ming @ 2026-09-16 12:34 UTC (permalink / raw)
  To: Shaikh Kamaluddin, Davidlohr Bueso, Jonathan Cameron, Dave Jiang,
	Alison Schofield, Vishal Verma, Dan Williams, Ira Weiny
  Cc: anisa.su887, benjamin.cheatham, linux-cxl, linux-kernel

On 9/12/2026 5:38 PM, Shaikh Kamaluddin wrote:
> CXL event interrupts may share an MSI/MSI-X vector with other event
> logs or device features. Consequently, cxl_event_thread() is registered
> with IRQF_SHARED and must determine whether the event-log facility
> claimed the interrupt.
>
> The handler masks the Device Event Status register to the event logs
> supported by the driver. However, when no supported status bit is set,
> it exits the processing loop and still returns IRQ_HANDLED. This
> incorrectly reports that the interrupt was claimed and prevents the
> generic spurious-interrupt detector from accounting it as unhandled.
>
> Track whether the CXL event handler claimed the interrupt and return
> IRQ_NONE when no supported event is detected. If no handler sharing the
> vector claims a sustained interrupt storm, the generic spurious-
> interrupt detector can eventually identify and disable the affected
> IRQ.
>
> Signed-off-by: Shaikh Kamaluddin <shaikhkamal2012@gmail.com>
Reviewed-by: Li Ming <ming.li@zohomail.com>
> ---
> Changes in V2:
> - Emphasize that returning IRQ_NONE enables generic spurious-interrupt
>   detection and drop the Fixes tag. (Jonathan Cameron)
> - Check whether the event interrupt is already disabled by Anisa Su's
>   patch; no code changes were needed. (Ben Cheatham)
> - Rebase onto v7.3-rc2; no code changes.
>
>  drivers/cxl/pci.c | 5 ++++-
>  1 file changed, 4 insertions(+), 1 deletion(-)
>
> diff --git a/drivers/cxl/pci.c b/drivers/cxl/pci.c
> index c7c91e8dc51d..8b560cae91f2 100644
> --- a/drivers/cxl/pci.c
> +++ b/drivers/cxl/pci.c
> @@ -515,6 +515,7 @@ static irqreturn_t cxl_event_thread(int irq, void *id)
>  	struct cxl_dev_id *dev_id = id;
>  	struct cxl_dev_state *cxlds = dev_id->cxlds;
>  	struct cxl_memdev_state *mds = to_cxl_memdev_state(cxlds);
> +	bool handled = false;
>  	u32 status;
>  
>  	do {
> @@ -527,11 +528,13 @@ static irqreturn_t cxl_event_thread(int irq, void *id)
>  		status &= CXLDEV_EVENT_STATUS_ALL;
>  		if (!status)
>  			break;
> +
> +		handled = true;
>  		cxl_mem_get_event_records(mds, status);
>  		cond_resched();
>  	} while (status);
>  
> -	return IRQ_HANDLED;
> +	return handled ? IRQ_HANDLED : IRQ_NONE;
>  }
>  
>  static int cxl_event_req_irq(struct cxl_dev_state *cxlds, u8 setting)
>
> base-commit: df2908090cda368b01ff43709f51890076c56157



^ permalink raw reply	[flat|nested] 9+ messages in thread

* Re: [PATCH v2] cxl/events: Return IRQ_NONE when no event is pending
  2026-09-12  9:48 ` sashiko-bot
@ 2026-09-16 14:12   ` Shaikh Kamaluddin
  0 siblings, 0 replies; 9+ messages in thread
From: Shaikh Kamaluddin @ 2026-09-16 14:12 UTC (permalink / raw)
  To: sashiko-reviews; +Cc: linux-cxl

On Sat, Sep 12, 2026 at 09:48:06AM +0000, sashiko-bot@kernel.org wrote:
> Thank you for your contribution! Sashiko AI review found 2 potential issue(s) to consider:
> 
> Pre-existing issues:
> - [High] The threaded IRQ handler `cxl_event_thread()` can enter an infinite loop if mailbox commands fail to clear event records on the device, or if a PCIe surprise removal causes the MMIO read to persistently return all 1s (~0U).
> - [High] The driver blindly trusts the `record_count` and `count` fields returned by the device in mailbox payloads, leading to out-of-bounds kernel memory reads and potential information leaks.
> --
> 
> commit 659b756fc25aa4e783ce4a882044a537d7b5d643
> Author: Shaikh Kamaluddin <shaikhkamal2012@gmail.com>
> 
> cxl/events: Return IRQ_NONE when no event is pending
> 
> This patch tracks whether the CXL event handler claimed the interrupt and
> returns IRQ_NONE when no supported event is detected. This allows the generic
> spurious-interrupt detector to identify and disable the affected IRQ during
> interrupt storms where no handler claims the event.
> 
> > diff --git a/drivers/cxl/pci.c b/drivers/cxl/pci.c
> > index c7c91e8dc51dc..8b560cae91f2e 100644
> > --- a/drivers/cxl/pci.c
> > +++ b/drivers/cxl/pci.c
> [ ... ]
> > @@ -515,6 +515,7 @@ static irqreturn_t cxl_event_thread(int irq, void *id)
> >  	struct cxl_dev_id *dev_id = id;
> >  	struct cxl_dev_state *cxlds = dev_id->cxlds;
> >  	struct cxl_memdev_state *mds = to_cxl_memdev_state(cxlds);
> > +	bool handled = false;
> >  	u32 status;
> >  
> >  	do {
> >  		/*
> >  		 * CXL 3.0 8.2.8.3.1: The lower 32 bits are the status;
> >  		 * ignore the reserved upper 32 bits
> >  		 */
> >  		status = readl(cxlds->regs.status + CXLDEV_DEV_EVENT_STATUS_OFFSET);
> >  		/* Ignore logs unknown to the driver */
> >  		status &= CXLDEV_EVENT_STATUS_ALL;
> >  		if (!status)
> >  			break;
> 
> [Severity: High]
> This is a pre-existing issue, but can this loop become infinite if a PCIe
> surprise removal or fatal device error occurs?
> 
> If readl() returns all 1s (~0U), status & CXLDEV_EVENT_STATUS_ALL will
> still evaluate to a non-zero value, passing the if (!status) check. The
> handler will then call into cxl_mem_get_event_records() where mailbox
> commands will fail on the removed device, exiting without clearing the
> hardware event status.
> 
> Since status is never cleared and the error handling doesn't break the
> loop, this evaluates to true at the while (status) condition and repeats
> indefinitely. Does this IRQF_ONESHOT threaded handler need to check for ~0U
> to avoid monopolizing a CPU and preventing device teardown during hotplug
> removal?
> 
> > +
> > +		handled = true;
> >  		cxl_mem_get_event_records(mds, status);
>

On the drain loop: Jonathan already responded to the same report on v1 — surprise removal isn't supported today, so there's no evidence this is reachable in practice. 
link : https://lore.kernel.org/linux-cxl/20260907194400.17bb5f5c@jic23-huawei/

Li Ming and Davidlohr made the same point on Anisa Su's series: the Event Status register is device-owned and the spec requires the device to clear it, so a device that doesn't is a firmware bug, not something the driver should work around.


> [Severity: High]
> This isn't a bug introduced by this patch, but does this code blindly trust
> the record_count returned by the device hardware?
> 
> When cxl_mem_get_event_records() calls into cxl_mem_get_records_log(), the
> driver reads the count directly from the untrusted payload:
> 
> drivers/cxl/core/mbox.c:cxl_mem_get_records_log() {
>     ...
>     nr_rec = le16_to_cpu(payload->record_count);
>     ...
>     for (i = 0; i < nr_rec; i++)
>         __cxl_event_trace_record(cxlmd, type, &payload->records[i]);
>     ...
> }
> 
> If a compromised or malicious CXL device returns an artificially large
> record_count, could this read out-of-bounds kernel memory and leak it to
> userspace via tracing? 
> 

Valid, and not introduced here -- different file, different Fixes: tag, so
I'm not folding it into this patch.

record_count comes from the device and records[] lives in mds->event.buf,
which is sized at probe. Nothing checks one against the other, so a bad
count walks off the end of the buffer.

> The same pattern appears in cxl_clear_event_record() and
> cxl_mem_get_poison(), where the untrusted count is used as a loop bound.
> Should the payload size be validated against mbox_cmd.size_out or the
> maximum allocated buffer size before iterating?
> 
> >  		cond_resched();
> >  	} while (status);
> >  
> > -	return IRQ_HANDLED;
> > +	return handled ? IRQ_HANDLED : IRQ_NONE;
> >  }
>

These are pre-existing issues and independent of the IRQ_NONE versus
IRQ_HANDLED change in this patch.

Anisa is already taking care of these checks in her upcoming patch series, 
so this can be handled there.

Anisa's patch link: https://lore.kernel.org/linux-cxl/20260901002912.958-3-anisa.su@samsung.com/


> -- 
> Sashiko AI review · https://sashiko.dev/#/patchset/20260912093815.14455-1-shaikhkamal2012@gmail.com?part=1

^ permalink raw reply	[flat|nested] 9+ messages in thread

* Re: [PATCH v2] cxl/events: Return IRQ_NONE when no event is pending
  2026-09-16  0:38 ` Anisa Su
@ 2026-09-16 14:30   ` Shaikh Kamaluddin
  0 siblings, 0 replies; 9+ messages in thread
From: Shaikh Kamaluddin @ 2026-09-16 14:30 UTC (permalink / raw)
  To: Anisa Su
  Cc: Davidlohr Bueso, Jonathan Cameron, Dave Jiang, Alison Schofield,
	Vishal Verma, Dan Williams, Ira Weiny, Li Ming, benjamin.cheatham,
	linux-cxl, linux-kernel

On Wed, Sep 16, 2026 at 09:38:37AM +0900, Anisa Su wrote:
> On Sat, Sep 12, 2026 at 03:08:15PM +0530, Shaikh Kamaluddin wrote:
> > CXL event interrupts may share an MSI/MSI-X vector with other event
> > logs or device features. Consequently, cxl_event_thread() is registered
> > with IRQF_SHARED and must determine whether the event-log facility
> > claimed the interrupt.
> > 
> > The handler masks the Device Event Status register to the event logs
> > supported by the driver. However, when no supported status bit is set,
> > it exits the processing loop and still returns IRQ_HANDLED. This
> > incorrectly reports that the interrupt was claimed and prevents the
> > generic spurious-interrupt detector from accounting it as unhandled.
> > 
> > Track whether the CXL event handler claimed the interrupt and return
> > IRQ_NONE when no supported event is detected. If no handler sharing the
> > vector claims a sustained interrupt storm, the generic spurious-
> > interrupt detector can eventually identify and disable the affected
> > IRQ.
> > 
> > Signed-off-by: Shaikh Kamaluddin <shaikhkamal2012@gmail.com>
> Reviewed-by: Anisa Su <anisa.su@samsung.com>
> > ---
> > Changes in V2:
> > - Emphasize that returning IRQ_NONE enables generic spurious-interrupt
> >   detection and drop the Fixes tag. (Jonathan Cameron)
> > - Check whether the event interrupt is already disabled by Anisa Su's
> >   patch; no code changes were needed. (Ben Cheatham)
> 
> One thing I want to correct: The change requested was to say a
> spurious IRQ will eventually be disabled in the commit message, not
> check whether the event interrupts is already disabled...
> See https://lore.kernel.org/linux-cxl/aqLZP3ORYzCLfSgR@acer-nitro-anv15-41/T/#m32c74ebf05eee9a3ee49c57dd24848db29919031
> I do see that it's been added to the commit message though, so no need
> to do anything.
>

Hi Anisa,

I got your point, yes commit message already covered this.

Thanks,
Shaikh 

> Thanks,
> Anisa
> > - Rebase onto v7.3-rc2; no code changes.
> > 
> >  drivers/cxl/pci.c | 5 ++++-
> >  1 file changed, 4 insertions(+), 1 deletion(-)
> > 
> > diff --git a/drivers/cxl/pci.c b/drivers/cxl/pci.c
> > index c7c91e8dc51d..8b560cae91f2 100644
> > --- a/drivers/cxl/pci.c
> > +++ b/drivers/cxl/pci.c
> > @@ -515,6 +515,7 @@ static irqreturn_t cxl_event_thread(int irq, void *id)
> >  	struct cxl_dev_id *dev_id = id;
> >  	struct cxl_dev_state *cxlds = dev_id->cxlds;
> >  	struct cxl_memdev_state *mds = to_cxl_memdev_state(cxlds);
> > +	bool handled = false;
> >  	u32 status;
> >  
> >  	do {
> > @@ -527,11 +528,13 @@ static irqreturn_t cxl_event_thread(int irq, void *id)
> >  		status &= CXLDEV_EVENT_STATUS_ALL;
> >  		if (!status)
> >  			break;
> > +
> > +		handled = true;
> >  		cxl_mem_get_event_records(mds, status);
> >  		cond_resched();
> >  	} while (status);
> >  
> > -	return IRQ_HANDLED;
> > +	return handled ? IRQ_HANDLED : IRQ_NONE;
> >  }
> >  
> >  static int cxl_event_req_irq(struct cxl_dev_state *cxlds, u8 setting)
> > 
> > base-commit: df2908090cda368b01ff43709f51890076c56157
> > -- 
> > 2.43.0
> > 

^ permalink raw reply	[flat|nested] 9+ messages in thread

* Re: [PATCH v2] cxl/events: Return IRQ_NONE when no event is pending
  2026-09-12  9:38 [PATCH v2] cxl/events: Return IRQ_NONE when no event is pending Shaikh Kamaluddin
                   ` (3 preceding siblings ...)
  2026-09-16 12:34 ` Li Ming
@ 2026-09-16 20:49 ` Alison Schofield
  2026-09-18 15:40 ` Dave Jiang
  5 siblings, 0 replies; 9+ messages in thread
From: Alison Schofield @ 2026-09-16 20:49 UTC (permalink / raw)
  To: Shaikh Kamaluddin
  Cc: Davidlohr Bueso, Jonathan Cameron, Dave Jiang, Vishal Verma,
	Dan Williams, Ira Weiny, Li Ming, anisa.su887, benjamin.cheatham,
	linux-cxl, linux-kernel

On Sat, Sep 12, 2026 at 03:08:15PM +0530, Shaikh Kamaluddin wrote:
> CXL event interrupts may share an MSI/MSI-X vector with other event
> logs or device features. Consequently, cxl_event_thread() is registered
> with IRQF_SHARED and must determine whether the event-log facility
> claimed the interrupt.
> 
> The handler masks the Device Event Status register to the event logs
> supported by the driver. However, when no supported status bit is set,
> it exits the processing loop and still returns IRQ_HANDLED. This
> incorrectly reports that the interrupt was claimed and prevents the
> generic spurious-interrupt detector from accounting it as unhandled.
> 
> Track whether the CXL event handler claimed the interrupt and return
> IRQ_NONE when no supported event is detected. If no handler sharing the
> vector claims a sustained interrupt storm, the generic spurious-
> interrupt detector can eventually identify and disable the affected
> IRQ.

Reviewed-by: Alison Schofield <alison.schofield@intel.com>


^ permalink raw reply	[flat|nested] 9+ messages in thread

* Re: [PATCH v2] cxl/events: Return IRQ_NONE when no event is pending
  2026-09-12  9:38 [PATCH v2] cxl/events: Return IRQ_NONE when no event is pending Shaikh Kamaluddin
                   ` (4 preceding siblings ...)
  2026-09-16 20:49 ` Alison Schofield
@ 2026-09-18 15:40 ` Dave Jiang
  5 siblings, 0 replies; 9+ messages in thread
From: Dave Jiang @ 2026-09-18 15:40 UTC (permalink / raw)
  To: Shaikh Kamaluddin, Davidlohr Bueso, Jonathan Cameron,
	Alison Schofield, Vishal Verma, Dan Williams, Ira Weiny, Li Ming
  Cc: anisa.su887, benjamin.cheatham, linux-cxl, linux-kernel



On 9/12/26 2:38 AM, Shaikh Kamaluddin wrote:
> CXL event interrupts may share an MSI/MSI-X vector with other event
> logs or device features. Consequently, cxl_event_thread() is registered
> with IRQF_SHARED and must determine whether the event-log facility
> claimed the interrupt.
> 
> The handler masks the Device Event Status register to the event logs
> supported by the driver. However, when no supported status bit is set,
> it exits the processing loop and still returns IRQ_HANDLED. This
> incorrectly reports that the interrupt was claimed and prevents the
> generic spurious-interrupt detector from accounting it as unhandled.
> 
> Track whether the CXL event handler claimed the interrupt and return
> IRQ_NONE when no supported event is detected. If no handler sharing the
> vector claims a sustained interrupt storm, the generic spurious-
> interrupt detector can eventually identify and disable the affected
> IRQ.
> 
> Signed-off-by: Shaikh Kamaluddin <shaikhkamal2012@gmail.com>

Applied to cxl/next:
a8c458611a7d

> ---
> Changes in V2:
> - Emphasize that returning IRQ_NONE enables generic spurious-interrupt
>   detection and drop the Fixes tag. (Jonathan Cameron)
> - Check whether the event interrupt is already disabled by Anisa Su's
>   patch; no code changes were needed. (Ben Cheatham)
> - Rebase onto v7.3-rc2; no code changes.
> 
>  drivers/cxl/pci.c | 5 ++++-
>  1 file changed, 4 insertions(+), 1 deletion(-)
> 
> diff --git a/drivers/cxl/pci.c b/drivers/cxl/pci.c
> index c7c91e8dc51d..8b560cae91f2 100644
> --- a/drivers/cxl/pci.c
> +++ b/drivers/cxl/pci.c
> @@ -515,6 +515,7 @@ static irqreturn_t cxl_event_thread(int irq, void *id)
>  	struct cxl_dev_id *dev_id = id;
>  	struct cxl_dev_state *cxlds = dev_id->cxlds;
>  	struct cxl_memdev_state *mds = to_cxl_memdev_state(cxlds);
> +	bool handled = false;
>  	u32 status;
>  
>  	do {
> @@ -527,11 +528,13 @@ static irqreturn_t cxl_event_thread(int irq, void *id)
>  		status &= CXLDEV_EVENT_STATUS_ALL;
>  		if (!status)
>  			break;
> +
> +		handled = true;
>  		cxl_mem_get_event_records(mds, status);
>  		cond_resched();
>  	} while (status);
>  
> -	return IRQ_HANDLED;
> +	return handled ? IRQ_HANDLED : IRQ_NONE;
>  }
>  
>  static int cxl_event_req_irq(struct cxl_dev_state *cxlds, u8 setting)
> 
> base-commit: df2908090cda368b01ff43709f51890076c56157


^ permalink raw reply	[flat|nested] 9+ messages in thread

end of thread, other threads:[~2026-09-18 15:40 UTC | newest]

Thread overview: 9+ messages (download: mbox.gz follow: Atom feed
-- links below jump to the message on this page --
2026-09-12  9:38 [PATCH v2] cxl/events: Return IRQ_NONE when no event is pending Shaikh Kamaluddin
2026-09-12  9:48 ` sashiko-bot
2026-09-16 14:12   ` Shaikh Kamaluddin
2026-09-15 23:28 ` Jonathan Cameron
2026-09-16  0:38 ` Anisa Su
2026-09-16 14:30   ` Shaikh Kamaluddin
2026-09-16 12:34 ` Li Ming
2026-09-16 20:49 ` Alison Schofield
2026-09-18 15:40 ` Dave Jiang

This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox