From: sashiko-bot@kernel.org
To: "Dave Jiang" <dave.jiang@intel.com>
Cc: linux-perf-users@vger.kernel.org, linux-cxl@vger.kernel.org
Subject: Re: [PATCH v3 6/9] perf/cxl: Don't share the overflow interrupt, and keep it pinned
Date: Fri, 31 Jul 2026 23:50:27 +0000 [thread overview]
Message-ID: <20260731235027.B8DCA1F00ACA@smtp.kernel.org> (raw)
In-Reply-To: <20260731232827.401447-7-dave.jiang@intel.com>
Thank you for your contribution! Sashiko AI review found 1 potential issue(s) to consider:
Pre-existing issues:
- [High] CPU hotplug migration race causes cross-CPU data corruption on in-flight interrupts.
--
--- Patch [6]: [PATCH v3 6/9] perf/cxl: Don't share the overflow interrupt, and keep it pinned ---
commit 0c52a628eb5f3fcd46d4c3f0c11a1448687ffd14
Author: Dave Jiang <dave.jiang@intel.com>
perf/cxl: Don't share the overflow interrupt, and keep it pinned
The PMU pins its overflow interrupt to info->on_cpu in the hotplug
callbacks, but requests it with only IRQF_SHARED | IRQF_NO_THREAD. Without
IRQF_NOBALANCING, irqbalance or a userspace smp_affinity write can move the
interrupt to another CPU. cxl_pmu_irq() then runs local64_cmpxchg() and
local64_add() on hwc->prev_count and event->count there, at the same time
as the managing CPU. local64_t is only atomic against same-CPU access, so
the counts get corrupted.
Link: https://sashiko.dev/#/patchset/20260715191454.459673-1-dave@stgolabs.net?part=1
> diff --git a/drivers/perf/cxl_pmu.c b/drivers/perf/cxl_pmu.c
> index c9e30cb149dfa..6fdc66a01fb67 100644
> --- a/drivers/perf/cxl_pmu.c
> +++ b/drivers/perf/cxl_pmu.c
[ ... ]
> @@ -784,7 +784,7 @@ static irqreturn_t cxl_pmu_irq(int irq, void *data)
>
> overflowed = readq(base + CXL_PMU_OVERFLOW_REG);
>
> - /* Interrupt may be shared, so maybe it isn't ours */
> + /* Nothing overflowed, so the device did not raise this */
> if (!overflowed)
> return IRQ_NONE;
[Severity: High]
This is a pre-existing issue, but can this IRQ handler race with CPU hotplug
migration and cause cross-CPU data corruption?
Looking at cxl_pmu_offline_cpu(), the context is migrated to a new CPU before
the IRQ affinity is updated:
cxl_pmu_offline_cpu() {
...
perf_pmu_migrate_context(&info->pmu, cpu, target);
info->on_cpu = target;
/*
* CPU HP lock is held so we should be guaranteed that this CPU hasn't
* yet gone away.
*/
WARN_ON(irq_set_affinity(info->irq, cpumask_of(target)));
...
}
During installation on the new CPU, the hardware counters and interrupts are
re-enabled. If the hardware generates a PMU overflow interrupt in this
window before irq_set_affinity() completes, the interrupt is delivered to
the old CPU.
The old CPU's IRQ handler (cxl_pmu_irq()) will then see the newly installed
event in the globally shared info->hw_events[i] array:
cxl_pmu_irq() {
...
for_each_set_bit(i, overflowedbm, info->num_counters) {
struct perf_event *event = info->hw_events[i];
...
__cxl_pmu_read(event, true);
}
...
}
This causes the old CPU to execute non-atomic local64_cmpxchg() and
local64_add() operations concurrently with the new managing CPU. Since
local64_t operations are only safe for same-CPU access, won't this corrupt
the hardware event counts?
--
Sashiko AI review · https://sashiko.dev/#/patchset/20260731232827.401447-1-dave.jiang@intel.com?part=6
next prev parent reply other threads:[~2026-07-31 23:50 UTC|newest]
Thread overview: 17+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-07-31 23:28 [PATCH v3 0/9] perf/cxlpmu: Misc sashiko raised issues fixes Dave Jiang
2026-07-31 23:28 ` [PATCH v3 1/9] perf/cxl: Program the requested event group on configurable counters Dave Jiang
2026-07-31 23:38 ` sashiko-bot
2026-07-31 23:28 ` [PATCH v3 2/9] perf/cxl: Clear stale event fields before reprogramming a counter Dave Jiang
2026-07-31 23:40 ` sashiko-bot
2026-07-31 23:28 ` [PATCH v3 3/9] perf/cxl: Fix the counter overflow delta fixup Dave Jiang
2026-07-31 23:37 ` sashiko-bot
2026-07-31 23:28 ` [PATCH v3 4/9] perf/cxl: Split the MSI vector out of info->irq Dave Jiang
2026-07-31 23:45 ` sashiko-bot
2026-07-31 23:28 ` [PATCH v3 5/9] perf/cxl: Accept an overflow interrupt on MSI message number 0 Dave Jiang
2026-07-31 23:28 ` [PATCH v3 6/9] perf/cxl: Don't share the overflow interrupt, and keep it pinned Dave Jiang
2026-07-31 23:50 ` sashiko-bot [this message]
2026-07-31 23:28 ` [PATCH v3 7/9] perf/cxl: Unfreeze counters after handling an overflow interrupt Dave Jiang
2026-07-31 23:40 ` sashiko-bot
2026-07-31 23:28 ` [PATCH v3 8/9] perf/cxl: Validate the hardware-reported counter width Dave Jiang
2026-07-31 23:28 ` [PATCH v3 9/9] perf/cxl: Don't log through pmu.dev in the overflow interrupt handler Dave Jiang
2026-07-31 23:46 ` sashiko-bot
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=20260731235027.B8DCA1F00ACA@smtp.kernel.org \
--to=sashiko-bot@kernel.org \
--cc=dave.jiang@intel.com \
--cc=linux-cxl@vger.kernel.org \
--cc=linux-perf-users@vger.kernel.org \
--cc=sashiko-reviews@lists.linux.dev \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is an external index of several public inboxes,
see mirroring instructions on how to clone and mirror
all data and code used by this external index.