Linux Perf Users
 help / color / mirror / Atom feed
* [PATCH v15 0/2] RISC-V IOMMU HPM support
@ 2026-10-02  6:33 Zong Li
  2026-10-02  6:33 ` [PATCH v15 1/2] drivers/perf: riscv-iommu: add risc-v iommu pmu driver Zong Li
  2026-10-02  6:33 ` [PATCH v15 2/2] iommu/riscv: create a auxiliary device for HPM Zong Li
  0 siblings, 2 replies; 5+ messages in thread
From: Zong Li @ 2026-10-02  6:33 UTC (permalink / raw)
  To: tomasz.jeznach, joro, will, robin.murphy, pjw, palmer, aou, alex,
	mark.rutland, andrew.jones, guoren, david.laight.linux,
	zhangzhanpeng.jasper, yang.yicong, nutty.liu, iommu, linux-riscv,
	linux-kernel, linux-perf-users
  Cc: Zong Li

This series implements support for the RISC-V IOMMU hardware performance
monitor.

The RISC-V IOMMU PMU driver is implemented as an auxiliary device driver
created by the parent RISC-V IOMMU driver. Therefore, the child driver
can obtain resources and information from the parent device, such as
the MMIO base address and IRQ number.

It seems sashiko-bot is giving conflicting advice about whether to add
a lock to ->read(). No matter if we add it or not, the sashiko-bot always
suggests the opposite. Looking at no lock case again, sashiko-bot
concerns about a BPF program calling bpf_perf_event_read() from NMI
context and self-deadlocking against the IRQ handler, but that scenario
cannot occur here:

  - bpf_perf_event_read() goes through perf_event_read_local(), which
    refuses to call ->read() unless event_cpu == smp_processor_id()
    (kernel/events/core.c). All events on this PMU are CPU-bound and
    non-task-bound (->event_init() forces event->cpu = pmu->on_cpu and
    rejects sampling events), so PERF_ATTACH_TASK is never set.
  - More fundamentally, perf overflow interrupts are not NMIs on RISC-V,
    and every pmu->lock critical section in this driver is held under
    raw_spin_lock_irqsave(). There is no context that can interrupt them
    and re-enter ->read().

No other driver under drivers/perf/ uses a trylock or in_nmi() check in
its ->read() callback, so dropping it also brings this driver in line
with the rest.

sashiko-bot also reported that migration calling trigger a NULL pointer
dereference. I don't think this can happen on current mainline.
perf_pmu_migrate_context() does not touch pmu->cpu_pmu_context unless it
actually finds events to migrate. This used to be a real hazard: before
commit bd2756811766 ("perf: Rewrite core context handling", v6.2-rc1)
the first two lines were per_cpu_ptr(pmu->pmu_cpu_context, ...), which
did dereference a pmu-owned pointer unconditionally. That is also why
several in-tree drivers pre-seed thier ->cpu field with a real CPU number
rather than a sentinel, and so do call perf_pmu_migrate_context() in
this window, without any reported oops.

Thank Joreg for reporting some differences with the spec.
One issue is that the spec does not say if the counters must be
continuous, or if the mask must be the same for all counters.
Since non-continuous counters and different widths are very rare, the
original driver assumed that all counters were continuous and had the
same width. However, this does not follow the spec. I also fixed this
problem in this version

Changed in v14:
- Rebased onto v7.3-rc5
- Support sparse counters and the different width of each counter
- Accecpt custom event ranges
- Verify IDT supported in event
- Add warning message if riscv_iommu_hpm_enable is failure
- Remove raw_spin_lock for ->read() flow

Changed in v13:
- Reorder the registration of cpuhp action

Changed in v12:
- Rebased onto v7.3-rc4
- Add raw_spin_trylock_irqsave for ->read() flow

Changed in v11:
- Rebased onto v7.3-rc3
- Add riscv_iommu_hpm_disaable to destroy aux dev before MSI is freed
- Fix CPU hotplug race risks reported by sashiko-bot as follows
- Re-validate pre_count after reading hw counter
- Move hwc-state into atomic critical section (pmu->lock)

Changed in v10:
- Optimize hi-lo-hi by do while for hypervisor case
- Add raw spinlock for cpu hotplug race and IRQCHIP_MOVE_DEFERRED
- Remove irq work mechanism for IRQCHIP_MOVE_DEFERRED

Changed in v9:
- Clear PMIP in irq handler on wrong CPU for re-triggering IRQ
- Add a lock in offline_cpu to avoid cpu hotplug race condition

Changed in v8:
- Rebased onto v7.3-rc2
- Add irq work mechanism for IRQCHIP_MOVE_DEFERRED case
- Filter multiple cycle event case

Changed in v7:
- Rebased onto the v7.3-rc1
- Remove raw spinlock
- Check CPU matching at the beginning of irq handler
- Add PERF_HES_STOPPED check before overflow handling

Changed in v6:
- Rebased onto the latest v7.3-rc
- Use sysfs_emit instead of cpumap_print_to_pagebuf
- Set up on_cpu and irq affinity by cpuhp callbacks
- Change type of on_cpu from unsigned int to int
- Reject filter operands of cycle event in event_init
- Check return value of counter number and  masks in probe
- Add raw spinlock for race condition (third commit)

Changed in v5:
- Pick up suggestions from sashiko-bot as follows
- Fix event group validation for sw event
- Bind IRQ to aux PMU dev instead of parent IOMMU dev
- Clear OF bit when event is NULL
- Improve hi-lo-hi patten
- Add back IRQF_SHARED flag due to mismatch
- Manage cpuhp and pmu register by devre

Changed in v4:
- Rebased onto v7.3-rc
- Use is_sampling_event() instead of accessing vairable directly
- Rename the matching name from "iommu.pmu" to "riscv-iommu.pmu"
- Change the naming of PMU device for avoid ":" in PCIe case
- Add suppress_bind_attrs attribute
- Remove IRQF_SHARED flag
- Set irq affinity to local CPU of IOMMU
- Allocate ID by IDA for auxiliary device
- Pick up suggestions from sashiko-bot

Changed in v3:
- Rebased onto v7.2-rc3
- Use hi_lo_writeq/readq to access register
- Pick comments from sashiko-bot as follows
- Set IRQ CPU affinity
- Remove IRQF_ONESHOT flag when request irq
- Adjust cycle event check by checking event_id field only
- Fix bug for group events verificaiton
- Fix KASAN issue about casting 32-bit variable to unsigned long pointer
- Clear IPSR pending bit before starting counter
- Clear OF bit in event selector register in irq handler
- Release irq by devm instead of explicit free_irq

Changed in v2:
- Rebased onto v7.2-rc1
- Use hi-lo-hi mechanism to read counter.
  Suggested by Guo Ren and David Laight

Changed in v1:
- Rebased onto v6.19-rc8
- Pick all suggestions and feedbacks from v1 series
- Add cpu hotplug implementation to avoid race enablement
- Move PMU-related definition from header to c file
- Change PMU driver to auxiliary device driver

Changed in RFC:
- Rebase onto v6.13-rc7
- Clear interrupt pending before handling interrupt
- Fix the counter value issue caused by OF bit in the cycle counter.
- Invoke riscv_iommu_hpm_disable() instead of riscv_iommu_pmu_uninit()
  in riscv_iommu_remove()

Zong Li (2):
  drivers/perf: riscv-iommu: add risc-v iommu pmu driver
  iommu/riscv: create a auxiliary device for HPM

 drivers/iommu/riscv/Kconfig      |    1 +
 drivers/iommu/riscv/iommu-bits.h |   61 --
 drivers/iommu/riscv/iommu.c      |   58 ++
 drivers/iommu/riscv/iommu.h      |    4 +
 drivers/perf/Kconfig             |   12 +
 drivers/perf/Makefile            |    1 +
 drivers/perf/riscv_iommu_pmu.c   | 1081 ++++++++++++++++++++++++++++++
 7 files changed, 1157 insertions(+), 61 deletions(-)
 create mode 100644 drivers/perf/riscv_iommu_pmu.c

-- 
2.43.7


^ permalink raw reply	[flat|nested] 5+ messages in thread

* [PATCH v15 1/2] drivers/perf: riscv-iommu: add risc-v iommu pmu driver
  2026-10-02  6:33 [PATCH v15 0/2] RISC-V IOMMU HPM support Zong Li
@ 2026-10-02  6:33 ` Zong Li
  2026-10-02  9:14   ` sashiko-bot
  2026-10-02  6:33 ` [PATCH v15 2/2] iommu/riscv: create a auxiliary device for HPM Zong Li
  1 sibling, 1 reply; 5+ messages in thread
From: Zong Li @ 2026-10-02  6:33 UTC (permalink / raw)
  To: tomasz.jeznach, joro, will, robin.murphy, pjw, palmer, aou, alex,
	mark.rutland, andrew.jones, guoren, david.laight.linux,
	zhangzhanpeng.jasper, yang.yicong, nutty.liu, iommu, linux-riscv,
	linux-kernel, linux-perf-users
  Cc: Zong Li, Chen Pei, Fangyu Yu

Add a new driver to support the RISC-V IOMMU PMU. This is an auxiliary
device driver created by the parent RISC-V IOMMU driver.

The performance monitor provides counters with filtering support to
collect events for specific device ID/process ID, or GSCID/PSCID.

The RISC-V IOMMU PMU separates the cycle counter from the event counters.
The cycle counter is not associated with iohpmevt0, so a software-defined
cycle event is required for the perf subsystem.

The number and width of the counters are hardware-implemented and must
be detected at runtime.

Leave out all the dead cleanup code (i.e. .remove() operation) if the
PMU driver is tied to the IOMMU driver and can never realistically be
removed.

PMU-related definitions are moved into the perf driver, where they are
used exclusively.

According to RISC-V IOMMU specification Chapter 6:
Whether an 8 byte access to an IOMMU register is single-copy atomic is
UNSPECIFIED. Use two separate 4 byte accesses for hardware
compatibility.

Use raw spin lock to avoid CPU hotplug race and IRQCHIP_MOVE_DEFERRED

 - CPU hotplug race: serialises used_counters/events[]/IOCOUNTINH against
   riscv_iommu_pmu_offline_cpu()'s perf_pmu_migrate_context()
 - PCI MSI/MSI-X on IMSIC sets IRQCHIP_MOVE_DEFERRED, so
   irq_set_affinity() reports success while only recording the request
   and the move is applied later, in interrupt context. Until then the
   interrupt is still routed to the CPU the irqchip picked initially,
   which is not the CPU the events are bound to.

Tested-by: Chen Pei <cp0613@linux.alibaba.com>
Tested-by: Fangyu Yu <fangyu.yu@linux.alibaba.com>
Reviewed-by: Nutty Liu <nutty.liu@hotmail.com>
Reviewed-by: Guo Ren (Alibaba DAMO Academy) <guoren@kernel.org>
Reviewed-by: Yicong Yang <yang.yicong@picoheart.com>
Suggested-by: David Laight <david.laight.linux@gmail.com>
Suggested-by: Guo Ren <guoren@kernel.org>
Link: https://lore.kernel.org/linux-riscv/20260618143634.7f3dd6c5@pumpkin/
Signed-off-by: Zong Li <zong.li@sifive.com>
---
 drivers/iommu/riscv/iommu-bits.h |   61 --
 drivers/perf/Kconfig             |   12 +
 drivers/perf/Makefile            |    1 +
 drivers/perf/riscv_iommu_pmu.c   | 1081 ++++++++++++++++++++++++++++++
 4 files changed, 1094 insertions(+), 61 deletions(-)
 create mode 100644 drivers/perf/riscv_iommu_pmu.c

diff --git a/drivers/iommu/riscv/iommu-bits.h b/drivers/iommu/riscv/iommu-bits.h
index f2ef9bd3cde9..6b5de913a032 100644
--- a/drivers/iommu/riscv/iommu-bits.h
+++ b/drivers/iommu/riscv/iommu-bits.h
@@ -192,67 +192,6 @@ enum riscv_iommu_ddtp_modes {
 #define RISCV_IOMMU_IPSR_PMIP		BIT(RISCV_IOMMU_INTR_PM)
 #define RISCV_IOMMU_IPSR_PIP		BIT(RISCV_IOMMU_INTR_PQ)
 
-/* 5.19 Performance monitoring counter overflow status (32bits) */
-#define RISCV_IOMMU_REG_IOCOUNTOVF	0x0058
-#define RISCV_IOMMU_IOCOUNTOVF_CY	BIT(0)
-#define RISCV_IOMMU_IOCOUNTOVF_HPM	GENMASK_ULL(31, 1)
-
-/* 5.20 Performance monitoring counter inhibits (32bits) */
-#define RISCV_IOMMU_REG_IOCOUNTINH	0x005C
-#define RISCV_IOMMU_IOCOUNTINH_CY	BIT(0)
-#define RISCV_IOMMU_IOCOUNTINH_HPM	GENMASK(31, 1)
-
-/* 5.21 Performance monitoring cycles counter (64bits) */
-#define RISCV_IOMMU_REG_IOHPMCYCLES     0x0060
-#define RISCV_IOMMU_IOHPMCYCLES_COUNTER	GENMASK_ULL(62, 0)
-#define RISCV_IOMMU_IOHPMCYCLES_OF	BIT_ULL(63)
-
-/* 5.22 Performance monitoring event counters (31 * 64bits) */
-#define RISCV_IOMMU_REG_IOHPMCTR_BASE	0x0068
-#define RISCV_IOMMU_REG_IOHPMCTR(_n)	(RISCV_IOMMU_REG_IOHPMCTR_BASE + ((_n) * 0x8))
-
-/* 5.23 Performance monitoring event selectors (31 * 64bits) */
-#define RISCV_IOMMU_REG_IOHPMEVT_BASE	0x0160
-#define RISCV_IOMMU_REG_IOHPMEVT(_n)	(RISCV_IOMMU_REG_IOHPMEVT_BASE + ((_n) * 0x8))
-#define RISCV_IOMMU_IOHPMEVT_EVENTID	GENMASK_ULL(14, 0)
-#define RISCV_IOMMU_IOHPMEVT_DMASK	BIT_ULL(15)
-#define RISCV_IOMMU_IOHPMEVT_PID_PSCID	GENMASK_ULL(35, 16)
-#define RISCV_IOMMU_IOHPMEVT_DID_GSCID	GENMASK_ULL(59, 36)
-#define RISCV_IOMMU_IOHPMEVT_PV_PSCV	BIT_ULL(60)
-#define RISCV_IOMMU_IOHPMEVT_DV_GSCV	BIT_ULL(61)
-#define RISCV_IOMMU_IOHPMEVT_IDT	BIT_ULL(62)
-#define RISCV_IOMMU_IOHPMEVT_OF		BIT_ULL(63)
-
-/* Number of defined performance-monitoring event selectors */
-#define RISCV_IOMMU_IOHPMEVT_CNT	31
-
-/**
- * enum riscv_iommu_hpmevent_id - Performance-monitoring event identifier
- *
- * @RISCV_IOMMU_HPMEVENT_INVALID: Invalid event, do not count
- * @RISCV_IOMMU_HPMEVENT_URQ: Untranslated requests
- * @RISCV_IOMMU_HPMEVENT_TRQ: Translated requests
- * @RISCV_IOMMU_HPMEVENT_ATS_RQ: ATS translation requests
- * @RISCV_IOMMU_HPMEVENT_TLB_MISS: TLB misses
- * @RISCV_IOMMU_HPMEVENT_DD_WALK: Device directory walks
- * @RISCV_IOMMU_HPMEVENT_PD_WALK: Process directory walks
- * @RISCV_IOMMU_HPMEVENT_S_VS_WALKS: First-stage page table walks
- * @RISCV_IOMMU_HPMEVENT_G_WALKS: Second-stage page table walks
- * @RISCV_IOMMU_HPMEVENT_MAX: Value to denote maximum Event IDs
- */
-enum riscv_iommu_hpmevent_id {
-	RISCV_IOMMU_HPMEVENT_INVALID    = 0,
-	RISCV_IOMMU_HPMEVENT_URQ        = 1,
-	RISCV_IOMMU_HPMEVENT_TRQ        = 2,
-	RISCV_IOMMU_HPMEVENT_ATS_RQ     = 3,
-	RISCV_IOMMU_HPMEVENT_TLB_MISS   = 4,
-	RISCV_IOMMU_HPMEVENT_DD_WALK    = 5,
-	RISCV_IOMMU_HPMEVENT_PD_WALK    = 6,
-	RISCV_IOMMU_HPMEVENT_S_VS_WALKS = 7,
-	RISCV_IOMMU_HPMEVENT_G_WALKS    = 8,
-	RISCV_IOMMU_HPMEVENT_MAX        = 9
-};
-
 /* 5.24 Translation request IOVA (64bits) */
 #define RISCV_IOMMU_REG_TR_REQ_IOVA     0x0258
 #define RISCV_IOMMU_TR_REQ_IOVA_VPN	GENMASK_ULL(63, 12)
diff --git a/drivers/perf/Kconfig b/drivers/perf/Kconfig
index 245e7bb763b9..8cce6c2ea626 100644
--- a/drivers/perf/Kconfig
+++ b/drivers/perf/Kconfig
@@ -105,6 +105,18 @@ config RISCV_PMU_SBI
 	  full perf feature support i.e. counter overflow, privilege mode
 	  filtering, counter configuration.
 
+config RISCV_IOMMU_PMU
+	depends on RISCV || COMPILE_TEST
+	depends on RISCV_IOMMU
+	bool "RISC-V IOMMU Hardware Performance Monitor"
+	default y
+	help
+	  Say Y if you want to use the RISC-V IOMMU performance monitor
+	  implementation. The performance monitor is an optional hardware
+	  feature, and whether it is actually enabled depends on IOMMU
+	  hardware support. If the underlying hardware does not implement
+	  the PMU, this option will have no effect.
+
 config STARFIVE_STARLINK_PMU
 	depends on ARCH_STARFIVE || COMPILE_TEST
 	depends on 64BIT
diff --git a/drivers/perf/Makefile b/drivers/perf/Makefile
index eb8a022dad9a..90c75f3c0ac1 100644
--- a/drivers/perf/Makefile
+++ b/drivers/perf/Makefile
@@ -20,6 +20,7 @@ obj-$(CONFIG_QCOM_L3_PMU) += qcom_l3_pmu.o
 obj-$(CONFIG_RISCV_PMU) += riscv_pmu.o
 obj-$(CONFIG_RISCV_PMU_LEGACY) += riscv_pmu_legacy.o
 obj-$(CONFIG_RISCV_PMU_SBI) += riscv_pmu_sbi.o
+obj-$(CONFIG_RISCV_IOMMU_PMU) += riscv_iommu_pmu.o
 obj-$(CONFIG_STARFIVE_STARLINK_PMU) += starfive_starlink_pmu.o
 obj-$(CONFIG_THUNDERX2_PMU) += thunderx2_pmu.o
 obj-$(CONFIG_XGENE_PMU) += xgene_pmu.o
diff --git a/drivers/perf/riscv_iommu_pmu.c b/drivers/perf/riscv_iommu_pmu.c
new file mode 100644
index 000000000000..25f217ac97e7
--- /dev/null
+++ b/drivers/perf/riscv_iommu_pmu.c
@@ -0,0 +1,1081 @@
+// SPDX-License-Identifier: GPL-2.0-only
+/*
+ * Copyright (C) 2026 SiFive
+ *
+ * Authors
+ *	Zong Li <zong.li@sifive.com>
+ */
+
+#include <linux/auxiliary_bus.h>
+#include <linux/cpu.h>
+#include <linux/cpumask.h>
+#include <linux/io-64-nonatomic-hi-lo.h>
+#include <linux/perf_event.h>
+
+#include "../iommu/riscv/iommu.h"
+
+/* 5.19 Performance monitoring counter overflow status (32bits) */
+#define RISCV_IOMMU_REG_IOCOUNTOVF	0x0058
+#define RISCV_IOMMU_IOCOUNTOVF_CY	BIT(0)
+#define RISCV_IOMMU_IOCOUNTOVF_HPM	GENMASK_ULL(31, 1)
+
+/* 5.20 Performance monitoring counter inhibits (32bits) */
+#define RISCV_IOMMU_REG_IOCOUNTINH	0x005C
+#define RISCV_IOMMU_IOCOUNTINH_CY	BIT(0)
+#define RISCV_IOMMU_IOCOUNTINH_HPM	GENMASK(31, 0)
+
+/* 5.21 Performance monitoring cycles counter (64bits) */
+#define RISCV_IOMMU_REG_IOHPMCYCLES	0x0060
+#define RISCV_IOMMU_IOHPMCYCLES_COUNTER	GENMASK_ULL(62, 0)
+#define RISCV_IOMMU_IOHPMCYCLES_OF	BIT_ULL(63)
+#define RISCV_IOMMU_REG_IOHPMCTR(_n)	(RISCV_IOMMU_REG_IOHPMCYCLES + ((_n) * 0x8))
+
+/* 5.22 Performance monitoring event counters (31 * 64bits) */
+#define RISCV_IOMMU_REG_IOHPMCTR_BASE	0x0068
+#define RISCV_IOMMU_IOHPMCTR_COUNTER	GENMASK_ULL(63, 0)
+
+/* 5.23 Performance monitoring event selectors (31 * 64bits) */
+#define RISCV_IOMMU_REG_IOHPMEVT_BASE	0x0160
+#define RISCV_IOMMU_REG_IOHPMEVT(_n)	(RISCV_IOMMU_REG_IOHPMEVT_BASE + ((_n) * 0x8))
+#define RISCV_IOMMU_IOHPMEVT_EVENTID	GENMASK_ULL(14, 0)
+#define RISCV_IOMMU_IOHPMEVT_DMASK	BIT_ULL(15)
+#define RISCV_IOMMU_IOHPMEVT_PID_PSCID	GENMASK_ULL(35, 16)
+#define RISCV_IOMMU_IOHPMEVT_DID_GSCID	GENMASK_ULL(59, 36)
+#define RISCV_IOMMU_IOHPMEVT_PV_PSCV	BIT_ULL(60)
+#define RISCV_IOMMU_IOHPMEVT_DV_GSCV	BIT_ULL(61)
+#define RISCV_IOMMU_IOHPMEVT_IDT	BIT_ULL(62)
+#define RISCV_IOMMU_IOHPMEVT_OF		BIT_ULL(63)
+#define RISCV_IOMMU_IOHPMEVT_EVENT	GENMASK_ULL(62, 0)
+
+/* The total number of counters is 31 event counters plus 1 cycle counter */
+#define RISCV_IOMMU_HPM_COUNTER_NUM	32
+
+/* Counter index 0 is the cycle counter, the event counters start at index 1 */
+#define RISCV_IOMMU_HPM_CYCLE_IDX	0
+
+static int cpuhp_state;
+
+/**
+ * enum riscv_iommu_hpmevent_id - Performance-monitoring event identifier
+ *
+ * @RISCV_IOMMU_HPMEVENT_CYCLE: Clock cycle counter
+ * @RISCV_IOMMU_HPMEVENT_URQ: Untranslated requests
+ * @RISCV_IOMMU_HPMEVENT_TRQ: Translated requests
+ * @RISCV_IOMMU_HPMEVENT_ATS_RQ: ATS translation requests
+ * @RISCV_IOMMU_HPMEVENT_TLB_MISS: TLB misses
+ * @RISCV_IOMMU_HPMEVENT_DD_WALK: Device directory walks
+ * @RISCV_IOMMU_HPMEVENT_PD_WALK: Process directory walks
+ * @RISCV_IOMMU_HPMEVENT_S_VS_WALKS: First-stage page table walks
+ * @RISCV_IOMMU_HPMEVENT_G_WALKS: Second-stage page table walks
+ * @RISCV_IOMMU_HPMEVENT_MAX: Value to denote maximum Event IDs
+ *
+ * The specification does not define an event ID for counting the
+ * number of clock cycles, meaning there is no associated 'iohpmevt0'.
+ * Event ID 0 is an invalid event and does not overlap with any valid
+ * event ID. Let's repurpose ID 0 as the cycle for perf, the cycle
+ * event is not actually written into any register, it serves solely
+ * as an identifier.
+ */
+enum riscv_iommu_hpmevent_id {
+	RISCV_IOMMU_HPMEVENT_CYCLE	= 0,
+	RISCV_IOMMU_HPMEVENT_URQ        = 1,
+	RISCV_IOMMU_HPMEVENT_TRQ        = 2,
+	RISCV_IOMMU_HPMEVENT_ATS_RQ     = 3,
+	RISCV_IOMMU_HPMEVENT_TLB_MISS   = 4,
+	RISCV_IOMMU_HPMEVENT_DD_WALK    = 5,
+	RISCV_IOMMU_HPMEVENT_PD_WALK    = 6,
+	RISCV_IOMMU_HPMEVENT_S_VS_WALKS = 7,
+	RISCV_IOMMU_HPMEVENT_G_WALKS    = 8,
+	RISCV_IOMMU_HPMEVENT_MAX        = 9
+};
+
+/*
+ * Section 6.23 reserves event IDs 16384-32767 for custom events. IDs between
+ * RISCV_IOMMU_HPMEVENT_MAX and this are reserved by the specification.
+ */
+#define RISCV_IOMMU_HPMEVENT_CUSTOM_FIRST	16384
+
+struct riscv_iommu_pmu {
+	struct pmu pmu;
+	struct hlist_node node;
+	void __iomem *reg;
+	int on_cpu;
+	unsigned int irq;
+	int numa_node;
+	unsigned int num_counters;
+	/* Per-counter writable width; index 0 is IOHPMCYCLES */
+	u64 cntr_mask[RISCV_IOMMU_HPM_COUNTER_NUM];
+	/* IOCOUNTINH value which inhibits every implemented counter */
+	u32 inhibit_all;
+	struct perf_event *events[RISCV_IOMMU_HPM_COUNTER_NUM];
+	/* Counter indices the hardware actually implements */
+	DECLARE_BITMAP(implemented_counters, RISCV_IOMMU_HPM_COUNTER_NUM);
+	DECLARE_BITMAP(used_counters, RISCV_IOMMU_HPM_COUNTER_NUM);
+	/*
+	 * Serialises used_counters/events[]/IOCOUNTINH for
+	 * CPU hotplug race and IRQCHIP_MOVE_DEFERRED (PCI MSI/MSI-X on IMSIC)
+	 */
+	raw_spinlock_t lock;
+};
+
+#define to_riscv_iommu_pmu(p) (container_of(p, struct riscv_iommu_pmu, pmu))
+
+#define RISCV_IOMMU_PMU_ATTR_EXTRACTOR(_name, _mask)			\
+	static inline u32 get_##_name(struct perf_event *event)		\
+	{								\
+		return FIELD_GET(_mask, event->attr.config);		\
+	}								\
+
+RISCV_IOMMU_PMU_ATTR_EXTRACTOR(event, RISCV_IOMMU_IOHPMEVT_EVENTID);
+RISCV_IOMMU_PMU_ATTR_EXTRACTOR(partial_matching, RISCV_IOMMU_IOHPMEVT_DMASK);
+RISCV_IOMMU_PMU_ATTR_EXTRACTOR(pid_pscid, RISCV_IOMMU_IOHPMEVT_PID_PSCID);
+RISCV_IOMMU_PMU_ATTR_EXTRACTOR(did_gscid, RISCV_IOMMU_IOHPMEVT_DID_GSCID);
+RISCV_IOMMU_PMU_ATTR_EXTRACTOR(filter_pid_pscid, RISCV_IOMMU_IOHPMEVT_PV_PSCV);
+RISCV_IOMMU_PMU_ATTR_EXTRACTOR(filter_did_gscid, RISCV_IOMMU_IOHPMEVT_DV_GSCV);
+RISCV_IOMMU_PMU_ATTR_EXTRACTOR(filter_id_type, RISCV_IOMMU_IOHPMEVT_IDT);
+
+/* Formats */
+PMU_FORMAT_ATTR(event,            "config:0-14");
+PMU_FORMAT_ATTR(partial_matching, "config:15");
+PMU_FORMAT_ATTR(pid_pscid,        "config:16-35");
+PMU_FORMAT_ATTR(did_gscid,        "config:36-59");
+PMU_FORMAT_ATTR(filter_pid_pscid, "config:60");
+PMU_FORMAT_ATTR(filter_did_gscid, "config:61");
+PMU_FORMAT_ATTR(filter_id_type,   "config:62");
+
+static struct attribute *riscv_iommu_pmu_formats[] = {
+	&format_attr_event.attr,
+	&format_attr_partial_matching.attr,
+	&format_attr_pid_pscid.attr,
+	&format_attr_did_gscid.attr,
+	&format_attr_filter_pid_pscid.attr,
+	&format_attr_filter_did_gscid.attr,
+	&format_attr_filter_id_type.attr,
+	NULL,
+};
+
+static const struct attribute_group riscv_iommu_pmu_format_group = {
+	.name = "format",
+	.attrs = riscv_iommu_pmu_formats,
+};
+
+/* Events */
+static ssize_t riscv_iommu_pmu_event_show(struct device *dev,
+					  struct device_attribute *attr,
+					  char *page)
+{
+	struct perf_pmu_events_attr *pmu_attr;
+
+	pmu_attr = container_of(attr, struct perf_pmu_events_attr, attr);
+
+	return sysfs_emit(page, "event=0x%02llx\n", pmu_attr->id);
+}
+
+#define RISCV_IOMMU_PMU_EVENT_ATTR(name, id)			\
+	PMU_EVENT_ATTR_ID(name, riscv_iommu_pmu_event_show, id)
+
+static struct attribute *riscv_iommu_pmu_events[] = {
+	RISCV_IOMMU_PMU_EVENT_ATTR(cycle, RISCV_IOMMU_HPMEVENT_CYCLE),
+	RISCV_IOMMU_PMU_EVENT_ATTR(untranslated_req, RISCV_IOMMU_HPMEVENT_URQ),
+	RISCV_IOMMU_PMU_EVENT_ATTR(translated_req, RISCV_IOMMU_HPMEVENT_TRQ),
+	RISCV_IOMMU_PMU_EVENT_ATTR(ats_trans_req, RISCV_IOMMU_HPMEVENT_ATS_RQ),
+	RISCV_IOMMU_PMU_EVENT_ATTR(tlb_miss, RISCV_IOMMU_HPMEVENT_TLB_MISS),
+	RISCV_IOMMU_PMU_EVENT_ATTR(ddt_walks, RISCV_IOMMU_HPMEVENT_DD_WALK),
+	RISCV_IOMMU_PMU_EVENT_ATTR(pdt_walks, RISCV_IOMMU_HPMEVENT_PD_WALK),
+	RISCV_IOMMU_PMU_EVENT_ATTR(s_vs_pt_walks, RISCV_IOMMU_HPMEVENT_S_VS_WALKS),
+	RISCV_IOMMU_PMU_EVENT_ATTR(g_pt_walks, RISCV_IOMMU_HPMEVENT_G_WALKS),
+	NULL,
+};
+
+static const struct attribute_group riscv_iommu_pmu_events_group = {
+	.name = "events",
+	.attrs = riscv_iommu_pmu_events,
+};
+
+/* cpumask */
+static ssize_t riscv_iommu_cpumask_show(struct device *dev,
+					struct device_attribute *attr,
+					char *buf)
+{
+	struct riscv_iommu_pmu *pmu = to_riscv_iommu_pmu(dev_get_drvdata(dev));
+	int on_cpu = pmu->on_cpu;
+
+	/*
+	 * riscv_iommu_pmu_offline_cpu() leaves on_cpu at -1 when it cannot
+	 * find another online CPU to migrate to. Report an empty mask rather
+	 * than feeding -1 to cpumask_of(), which indexes out of bounds.
+	 */
+	if (on_cpu < 0)
+		return sysfs_emit(buf, "%*pbl\n", cpumask_pr_args(cpu_none_mask));
+
+	return sysfs_emit(buf, "%*pbl\n", cpumask_pr_args(cpumask_of(on_cpu)));
+}
+
+static struct device_attribute riscv_iommu_cpumask_attr =
+	__ATTR(cpumask, 0444, riscv_iommu_cpumask_show, NULL);
+
+static struct attribute *riscv_iommu_cpumask_attrs[] = {
+	&riscv_iommu_cpumask_attr.attr,
+	NULL
+};
+
+static const struct attribute_group riscv_iommu_pmu_cpumask_group = {
+	.attrs = riscv_iommu_cpumask_attrs,
+};
+
+static const struct attribute_group *riscv_iommu_pmu_attr_grps[] = {
+	&riscv_iommu_pmu_cpumask_group,
+	&riscv_iommu_pmu_format_group,
+	&riscv_iommu_pmu_events_group,
+	NULL,
+};
+
+/*
+ * Register access wrapper
+ *
+ * According to RISC-V IOMMU specification Chapter 6:
+ * A 4 byte access to an IOMMU register must be single-copy atomic.
+ * Whether an 8 byte access to an IOMMU register is single-copy atomic is UNSPECIFIED
+ *
+ * Use two separate 4 byte accesses for hardware compatibility
+ */
+static u64 riscv_iommu_pmu_readq(void __iomem *addr)
+{
+	return hi_lo_readq(addr);
+}
+
+static void riscv_iommu_pmu_writeq(u64 value, void __iomem *addr)
+{
+	hi_lo_writeq(value, addr);
+}
+
+/* PMU Operations */
+static void riscv_iommu_pmu_set_counter(struct riscv_iommu_pmu *pmu, u32 idx,
+					u64 value)
+{
+	u64 counter_mask = pmu->cntr_mask[idx];
+
+	riscv_iommu_pmu_writeq(value & counter_mask, pmu->reg + RISCV_IOMMU_REG_IOHPMCTR(idx));
+}
+
+/*
+ * As stated in the RISC-V IOMMU Specification, Chapter 6:
+ * Whether an 8 byte access to an IOMMU register is single-copy atomic
+ * is UNSPECIFIED, and such an access may appear, internally to the
+ * IOMMU, as if two separate 4 byte accesses - first to the high half
+ * and second to the low half - were performed
+ *
+ * To make sure the driver works correctly on different hardware,
+ * the software will always use two 4-byte access for the counter.
+ *
+ * This function implements the hi-lo-hi pattern to detect and handle
+ * wraparound during the read operation:
+ *   1. Read high half (hi)
+ *   2. Read low half (lo)
+ *   3. Read high half again (hi_again)
+ *
+ * If both reads of the high half agree, then the low half did not carry
+ * into the high half in between, so the two halves belong together. If
+ * they differ, the low half wrapped during the read and the whole
+ * sequence is retried. A single re-read of the low half is not enough:
+ * if the caller (or a hypervisor running it) is preempted for long enough,
+ * the counter may wrap again before that re-read completes, pairing a
+ * stale high half with a low half from yet another wrap. Retrying the
+ * full hi/lo/hi sequence until two consecutive high-half reads agree
+ * converges on a consistent pair regardless of how long the preemption
+ * lasts.
+ */
+static u64 riscv_iommu_pmu_get_counter(struct riscv_iommu_pmu *pmu, u32 idx)
+{
+	void __iomem *addr = pmu->reg + RISCV_IOMMU_REG_IOHPMCTR(idx);
+	u64 value, counter_mask = pmu->cntr_mask[idx];
+	u32 hi, lo, hi_again;
+
+	do {
+		hi = readl(addr + 4);
+		lo = readl(addr);
+		hi_again = readl(addr + 4);
+	} while (hi_again != hi);
+
+	value = (((u64)hi << 32) | lo) & counter_mask;
+
+	/* The bit 63 of cycle counter (i.e., idx == 0) is OF bit */
+	return idx ? value : (value & ~RISCV_IOMMU_IOHPMCYCLES_OF);
+}
+
+static bool is_cycle_event(u64 event)
+{
+	return FIELD_GET(RISCV_IOMMU_IOHPMEVT_EVENTID, event) ==
+	       RISCV_IOMMU_HPMEVENT_CYCLE;
+}
+
+static void riscv_iommu_pmu_set_event(struct riscv_iommu_pmu *pmu, u32 idx,
+				      u64 value)
+{
+	/* There is no associated IOHPMEVT0 for IOHPMCYCLES */
+	if (is_cycle_event(value))
+		return;
+
+	/* Event counter start from idx 1 */
+	riscv_iommu_pmu_writeq(FIELD_GET(RISCV_IOMMU_IOHPMEVT_EVENT, value),
+			       pmu->reg + RISCV_IOMMU_REG_IOHPMEVT(idx - 1));
+}
+
+static void riscv_iommu_pmu_enable_counter(struct riscv_iommu_pmu *pmu, u32 idx)
+{
+	void __iomem *addr = pmu->reg + RISCV_IOMMU_REG_IOCOUNTINH;
+	u32 value = readl(addr);
+
+	writel(value & ~BIT(idx), addr);
+}
+
+static void riscv_iommu_pmu_disable_counter(struct riscv_iommu_pmu *pmu, u32 idx)
+{
+	void __iomem *addr = pmu->reg + RISCV_IOMMU_REG_IOCOUNTINH;
+	u32 value = readl(addr);
+
+	writel(value | BIT(idx), addr);
+}
+
+static void riscv_iommu_pmu_clear_ovf(struct riscv_iommu_pmu *pmu, u32 idx)
+{
+	u64 value;
+
+	/* Counter is disabled here, making it safe to read and write registers */
+	if (idx == RISCV_IOMMU_HPM_CYCLE_IDX) {
+		value = riscv_iommu_pmu_readq(pmu->reg + RISCV_IOMMU_REG_IOHPMCYCLES) &
+					      ~RISCV_IOMMU_IOHPMCYCLES_OF;
+		riscv_iommu_pmu_writeq(value, pmu->reg + RISCV_IOMMU_REG_IOHPMCYCLES);
+	} else {
+		/* Event counter start from idx 1 */
+		value = riscv_iommu_pmu_readq(pmu->reg + RISCV_IOMMU_REG_IOHPMEVT(idx - 1)) &
+					      ~RISCV_IOMMU_IOHPMEVT_OF;
+		riscv_iommu_pmu_writeq(value, pmu->reg + RISCV_IOMMU_REG_IOHPMEVT(idx - 1));
+	}
+}
+
+static void riscv_iommu_pmu_start_all(struct riscv_iommu_pmu *pmu, u32 inhibit)
+{
+	writel(inhibit, pmu->reg + RISCV_IOMMU_REG_IOCOUNTINH);
+}
+
+/* Returns the inhibit state prior to stopping, so callers can restore it later */
+static u32 riscv_iommu_pmu_stop_all(struct riscv_iommu_pmu *pmu)
+{
+	void __iomem *addr = pmu->reg + RISCV_IOMMU_REG_IOCOUNTINH;
+	u32 inhibit = readl(addr);
+
+	/*
+	 * Inhibit exactly the implemented counters. A truncated mask would
+	 * clear the inhibit bit of an implemented but out-of-range counter,
+	 * letting it run unprogrammed until it overflows.
+	 */
+	writel(pmu->inhibit_all, addr);
+
+	return inhibit;
+}
+
+/* PMU APIs */
+static void riscv_iommu_pmu_set_period(struct perf_event *event)
+{
+	struct riscv_iommu_pmu *pmu = to_riscv_iommu_pmu(event->pmu);
+	struct hw_perf_event *hwc = &event->hw;
+	u64 counter_mask = pmu->cntr_mask[hwc->idx];
+	u64 period;
+
+	/*
+	 * Limit the maximum period to prevent the counter value
+	 * from overtaking the one we are about to program.
+	 * In effect we are reducing max_period to account for
+	 * interrupt latency (and we are being very conservative).
+	 */
+	period = counter_mask >> 1;
+	riscv_iommu_pmu_set_counter(pmu, hwc->idx, period);
+	local64_set(&hwc->prev_count, period);
+}
+
+/*
+ * Tally @config against what the hardware implements: one cycle counter plus
+ * pmu->num_counters - 1 event counters. Returns false once the group would need
+ * more of either than exist, so that groups which could never be scheduled are
+ * rejected in ->event_init() instead of failing with -EAGAIN in ->add() forever.
+ */
+static bool riscv_iommu_pmu_claim_counter(struct riscv_iommu_pmu *pmu, u64 config,
+					  unsigned int *nr_cycles,
+					  unsigned int *nr_events)
+{
+	if (is_cycle_event(config))
+		return ++(*nr_cycles) <= 1;
+
+	return ++(*nr_events) <= pmu->num_counters - 1;
+}
+
+/*
+ * Section 6.23 Table 19 only defines IDT=1 for these events. For the others an
+ * unsupported IDT setting means the counter does not increment at all.
+ */
+static bool riscv_iommu_pmu_idt_supported(u32 event_id)
+{
+	switch (event_id) {
+	case RISCV_IOMMU_HPMEVENT_TLB_MISS:
+	case RISCV_IOMMU_HPMEVENT_S_VS_WALKS:
+	case RISCV_IOMMU_HPMEVENT_G_WALKS:
+		return true;
+	default:
+		return false;
+	}
+}
+
+static int riscv_iommu_pmu_event_init(struct perf_event *event)
+{
+	struct riscv_iommu_pmu *pmu = to_riscv_iommu_pmu(event->pmu);
+	struct hw_perf_event *hwc = &event->hw;
+	struct perf_event *sibling;
+	unsigned int nr_cycles = 0;
+	unsigned int nr_events = 0;
+	int on_cpu;
+
+	if (event->attr.type != event->pmu->type)
+		return -ENOENT;
+
+	if (is_sampling_event(event))
+		return -EOPNOTSUPP;
+
+	if (event->cpu < 0)
+		return -EOPNOTSUPP;
+
+	/*
+	 * Reject the range reserved between the standard events and the custom
+	 * event space.
+	 * Custom IDs are passed through: their meaning is implementation-defined,
+	 * so this driver cannot validate them any further.
+	 */
+	if (get_event(event) >= RISCV_IOMMU_HPMEVENT_MAX &&
+	    get_event(event) < RISCV_IOMMU_HPMEVENT_CUSTOM_FIRST)
+		return -EINVAL;
+
+	/*
+	 * An IDT setting the event does not support makes the hardware simply
+	 * not increment the counter, which is indistinguishable from an idle
+	 * one. Reject it rather than silently reporting zero. Custom events are
+	 * exempt: their IDT semantics are implementation-defined.
+	 */
+	if (get_filter_id_type(event) &&
+	    get_event(event) < RISCV_IOMMU_HPMEVENT_MAX &&
+	    !riscv_iommu_pmu_idt_supported(get_event(event)))
+		return -EINVAL;
+
+	/*
+	 * There is no IOHPMEVT register associated with IOHPMCYCLES, so none
+	 * of the filtering fields can be programmed for the cycle event.
+	 * Reject them here instead of counting unfiltered cycles behind the
+	 * user's back.
+	 */
+	if (is_cycle_event(event->attr.config) &&
+	    (event->attr.config & ~RISCV_IOMMU_IOHPMEVT_EVENTID))
+		return -EINVAL;
+
+	/*
+	 * All events are bound to the CPU the interrupt is affine to. That
+	 * CPU is unset while no online CPU could be found for this PMU, and
+	 * assigning -1 here would turn this into a task bound event, which is
+	 * not something this PMU can serve.
+	 */
+	on_cpu = pmu->on_cpu;
+	if (on_cpu < 0)
+		return -ENODEV;
+
+	event->cpu = on_cpu;
+
+	hwc->idx = -1;
+	hwc->config = event->attr.config;
+
+	/*
+	 * Account for this event itself first. It has to be done before the
+	 * check below, otherwise an event which is on its own would never be
+	 * matched against the number of counters the hardware implements.
+	 */
+	if (!riscv_iommu_pmu_claim_counter(pmu, event->attr.config,
+					   &nr_cycles, &nr_events))
+		return -EINVAL;
+
+	/* On its own, so there is no group to validate */
+	if (event->group_leader == event)
+		return 0;
+
+	/*
+	 * Software events never occupy a hardware counter, so they do not have
+	 * to sit on this pmu and do not consume any of its budget. Anything
+	 * else in the group does, starting with the leader.
+	 */
+	if (!is_software_event(event->group_leader)) {
+		/* A hardware leader has to share this pmu's counters */
+		if (event->group_leader->pmu != event->pmu)
+			return -EINVAL;
+
+		if (!riscv_iommu_pmu_claim_counter(pmu,
+						   event->group_leader->attr.config,
+						   &nr_cycles, &nr_events))
+			return -EINVAL;
+	}
+
+	/*
+	 * Then the rest of the group. This walks group_leader->sibling_list,
+	 * which does not contain the event being initialised yet - hence
+	 * accounting for it separately above.
+	 */
+	for_each_sibling_event(sibling, event->group_leader) {
+		if (is_software_event(sibling))
+			continue;
+
+		if (sibling->pmu != event->pmu)
+			return -EINVAL;
+
+		if (!riscv_iommu_pmu_claim_counter(pmu, sibling->attr.config,
+						   &nr_cycles, &nr_events))
+			return -EINVAL;
+	}
+
+	return 0;
+}
+
+static void riscv_iommu_pmu_update(struct perf_event *event)
+{
+	struct hw_perf_event *hwc = &event->hw;
+	struct riscv_iommu_pmu *pmu = to_riscv_iommu_pmu(event->pmu);
+	u64 delta, prev, now;
+	u32 idx = hwc->idx;
+	u64 counter_mask = pmu->cntr_mask[idx];
+
+	/*
+	 * riscv_iommu_pmu_set_period() resets the hardware counter and
+	 * prev_count as two separate writes.
+	 * Pairing a "prev" read from before that reset with a "now" read from
+	 * after it would produce a nonsensical, possibly huge, delta.
+	 * Re-checking prev_count after the hardware read detects that torn
+	 * pairing and retries.
+	 */
+	do {
+		prev = local64_read(&hwc->prev_count);
+		now = riscv_iommu_pmu_get_counter(pmu, idx);
+	} while (prev != local64_read(&hwc->prev_count) ||
+		 local64_cmpxchg(&hwc->prev_count, prev, now) != prev);
+
+	delta = (now - prev) & counter_mask;
+	local64_add(delta, &event->count);
+}
+
+static void riscv_iommu_pmu_start(struct perf_event *event, int flags)
+{
+	struct riscv_iommu_pmu *pmu = to_riscv_iommu_pmu(event->pmu);
+	struct hw_perf_event *hwc = &event->hw;
+	unsigned long irqflags;
+
+	if (WARN_ON_ONCE(!(event->hw.state & PERF_HES_STOPPED)))
+		return;
+
+	if (flags & PERF_EF_RELOAD)
+		WARN_ON_ONCE(!(event->hw.state & PERF_HES_UPTODATE));
+
+	raw_spin_lock_irqsave(&pmu->lock, irqflags);
+	hwc->state = 0;
+	riscv_iommu_pmu_set_period(event);
+	riscv_iommu_pmu_set_event(pmu, hwc->idx, hwc->config);
+	riscv_iommu_pmu_enable_counter(pmu, hwc->idx);
+	raw_spin_unlock_irqrestore(&pmu->lock, irqflags);
+
+	perf_event_update_userpage(event);
+}
+
+static void riscv_iommu_pmu_stop(struct perf_event *event, int flags)
+{
+	struct riscv_iommu_pmu *pmu = to_riscv_iommu_pmu(event->pmu);
+	struct hw_perf_event *hwc = &event->hw;
+	unsigned long irqflags;
+	int idx = hwc->idx;
+
+	if (hwc->state & PERF_HES_STOPPED)
+		return;
+
+	raw_spin_lock_irqsave(&pmu->lock, irqflags);
+	riscv_iommu_pmu_disable_counter(pmu, idx);
+	if ((flags & PERF_EF_UPDATE) && !(hwc->state & PERF_HES_UPTODATE))
+		riscv_iommu_pmu_update(event);
+	hwc->state |= PERF_HES_STOPPED | PERF_HES_UPTODATE;
+	raw_spin_unlock_irqrestore(&pmu->lock, irqflags);
+}
+
+static int riscv_iommu_pmu_add(struct perf_event *event, int flags)
+{
+	struct riscv_iommu_pmu *pmu = to_riscv_iommu_pmu(event->pmu);
+	struct hw_perf_event *hwc = &event->hw;
+	unsigned long irqflags;
+	unsigned int idx;
+
+	raw_spin_lock_irqsave(&pmu->lock, irqflags);
+
+	/* Reserve index zero for iohpmcycles */
+	if (is_cycle_event(event->attr.config))
+		idx = RISCV_IOMMU_HPM_CYCLE_IDX;
+	else
+		/*
+		 * Pick a counter which is implemented and not already taken.
+		 * The implemented set may be sparse, so the free index cannot be
+		 * searched for within a consecutive range.
+		 */
+		idx = find_next_andnot_bit(pmu->implemented_counters,
+					   pmu->used_counters,
+					   RISCV_IOMMU_HPM_COUNTER_NUM,
+					   RISCV_IOMMU_HPM_CYCLE_IDX + 1);
+
+	/* All event counters or cycle counter are in use */
+	if (idx >= RISCV_IOMMU_HPM_COUNTER_NUM || pmu->events[idx]) {
+		raw_spin_unlock_irqrestore(&pmu->lock, irqflags);
+		return -EAGAIN;
+	}
+
+	set_bit(idx, pmu->used_counters);
+
+	pmu->events[idx] = event;
+	hwc->idx = idx;
+	hwc->state = PERF_HES_STOPPED | PERF_HES_UPTODATE;
+	local64_set(&hwc->prev_count, 0);
+
+	raw_spin_unlock_irqrestore(&pmu->lock, irqflags);
+
+	if (flags & PERF_EF_START)
+		riscv_iommu_pmu_start(event, flags);
+
+	/* Propagate changes to the userspace mapping. */
+	perf_event_update_userpage(event);
+
+	return 0;
+}
+
+static void riscv_iommu_pmu_read(struct perf_event *event)
+{
+	riscv_iommu_pmu_update(event);
+}
+
+static void riscv_iommu_pmu_del(struct perf_event *event, int flags)
+{
+	struct riscv_iommu_pmu *pmu = to_riscv_iommu_pmu(event->pmu);
+	struct hw_perf_event *hwc = &event->hw;
+	unsigned long irqflags;
+	int idx = hwc->idx;
+
+	riscv_iommu_pmu_stop(event, PERF_EF_UPDATE);
+
+	raw_spin_lock_irqsave(&pmu->lock, irqflags);
+	pmu->events[idx] = NULL;
+	clear_bit(idx, pmu->used_counters);
+	raw_spin_unlock_irqrestore(&pmu->lock, irqflags);
+
+	perf_event_update_userpage(event);
+}
+
+/*
+ * Pick the CPU perf assigns these events to: riscv_iommu_pmu_event_init()
+ * sets event->cpu to pmu->on_cpu, and perf_pmu_migrate_context() moves
+ * already-installed events to a new on_cpu when this changes. This has no
+ * bearing on which CPU riscv_iommu_pmu_irq_handler() itself runs on -
+ * pmu->lock serialises the two regardless of that, rather than requiring
+ * them to run on the same CPU. Only consider CPUs which the irqchip
+ * actually accepts - committing on_cpu to a CPU which irq_set_affinity()
+ * then rejects would leave events bound to a CPU whose interrupt the
+ * irqchip refuses to deliver.
+ *
+ * cpumask_local_spread() walks the online CPUs in order of NUMA distance from
+ * the iommu, so the closest usable one wins.
+ *
+ * Returns the chosen CPU, or nr_cpu_ids if none could be used.
+ */
+static unsigned int riscv_iommu_pmu_bind_cpu(struct riscv_iommu_pmu *pmu,
+					     unsigned int skip_cpu)
+{
+	unsigned int cpu, i;
+
+	for (i = 0; i < num_online_cpus(); i++) {
+		cpu = cpumask_local_spread(i, pmu->numa_node);
+		if (cpu == skip_cpu)
+			continue;
+		if (!irq_set_affinity(pmu->irq, cpumask_of(cpu)))
+			return cpu;
+	}
+
+	return nr_cpu_ids;
+}
+
+static int riscv_iommu_pmu_online_cpu(unsigned int cpu, struct hlist_node *node)
+{
+	struct riscv_iommu_pmu *iommu_pmu;
+	unsigned int target_cpu;
+
+	iommu_pmu = hlist_entry_safe(node, struct riscv_iommu_pmu, node);
+
+	if (READ_ONCE(iommu_pmu->on_cpu) != -1)
+		return 0;
+
+	target_cpu = riscv_iommu_pmu_bind_cpu(iommu_pmu, nr_cpu_ids);
+	if (target_cpu >= nr_cpu_ids) {
+		/* on_cpu stays unset, so a later callback tries again */
+		pr_debug("failed to point irq %u at any online cpu\n",
+			 iommu_pmu->irq);
+		return 0;
+	}
+
+	WRITE_ONCE(iommu_pmu->on_cpu, target_cpu);
+
+	return 0;
+}
+
+static int riscv_iommu_pmu_offline_cpu(unsigned int cpu, struct hlist_node *node)
+{
+	struct riscv_iommu_pmu *iommu_pmu;
+	unsigned int target_cpu;
+
+	iommu_pmu = hlist_entry_safe(node, struct riscv_iommu_pmu, node);
+
+	if (READ_ONCE(iommu_pmu->on_cpu) != (int)cpu)
+		return 0;
+
+	/*
+	 * The last online CPU cannot be taken offline, and this callback runs
+	 * before __cpu_disable() clears the outgoing CPU from cpu_online_mask,
+	 * so there is always another online CPU to move to.
+	 *
+	 * Should that ever fail, unset on_cpu rather than leaving it pointing at
+	 * the CPU which is going away. That keeps events from being bound to a
+	 * dead CPU and lets riscv_iommu_pmu_online_cpu() pick again once a CPU
+	 * comes back.
+	 */
+	target_cpu = riscv_iommu_pmu_bind_cpu(iommu_pmu, cpu);
+	if (WARN_ON_ONCE(target_cpu >= nr_cpu_ids)) {
+		WRITE_ONCE(iommu_pmu->on_cpu, -1);
+	} else {
+		WRITE_ONCE(iommu_pmu->on_cpu, target_cpu);
+		/*
+		 * perf_pmu_migrate_context() runs ->del() on cpu and ->add()
+		 * on target_cpu with a synchronize_rcu() gap in between.
+		 * riscv_iommu_pmu_irq_handler() can run concurrently with
+		 * either step, on whichever CPU the interrupt physically
+		 * lands on - pmu->lock serialises it against them instead of
+		 * racing.
+		 */
+		perf_pmu_migrate_context(&iommu_pmu->pmu, cpu, target_cpu);
+	}
+
+	return 0;
+}
+
+/*
+ * pmu->lock serialises this against ->add()/->del()/->start()/->stop(),
+ * which can run on a different CPU than this while
+ * riscv_iommu_pmu_offline_cpu() is migrating events across a
+ * perf_pmu_migrate_context() call. ->read() does not take the lock: it
+ * only calls riscv_iommu_pmu_update(), whose retry loop already resolves
+ * concurrency against the set_period() done here.
+ */
+static irqreturn_t riscv_iommu_pmu_irq_handler(int irq, void *dev_id)
+{
+	struct riscv_iommu_pmu *pmu = (struct riscv_iommu_pmu *)dev_id;
+	DECLARE_BITMAP(ovf_bitmap, BITS_PER_TYPE(u64));
+	unsigned long irqflags;
+	u32 ovf, idx, inhibit;
+
+	/* Check whether this interrupt is for PMU */
+	if (!(readl_relaxed(pmu->reg + RISCV_IOMMU_REG_IPSR) & RISCV_IOMMU_IPSR_PMIP))
+		return IRQ_NONE;
+
+	raw_spin_lock_irqsave(&pmu->lock, irqflags);
+
+	inhibit = riscv_iommu_pmu_stop_all(pmu);
+
+	ovf = readl(pmu->reg + RISCV_IOMMU_REG_IOCOUNTOVF);
+	if (ovf) {
+		bitmap_from_u64(ovf_bitmap, ovf);
+
+		for_each_set_bit(idx, ovf_bitmap, RISCV_IOMMU_HPM_COUNTER_NUM) {
+			struct perf_event *event = pmu->events[idx];
+
+			/*
+			 * pmu->events[idx] only means the counter is allocated,
+			 * not that it is counting. A counter which has not been
+			 * started has no valid prev_count to compute a delta
+			 * against, and one which has been stopped must not be
+			 * reprogrammed here. The overflow bit still has to be
+			 * cleared below in either case, including when the event
+			 * was already removed by riscv_iommu_pmu_del(),
+			 * otherwise the interrupt would stay pending forever.
+			 */
+			if (event && !(event->hw.state & PERF_HES_STOPPED)) {
+				riscv_iommu_pmu_update(event);
+				riscv_iommu_pmu_set_period(event);
+			}
+
+			riscv_iommu_pmu_clear_ovf(pmu, idx);
+		}
+	}
+
+	/* Clear performance monitoring interrupt pending bit */
+	writel_relaxed(RISCV_IOMMU_IPSR_PMIP, pmu->reg + RISCV_IOMMU_REG_IPSR);
+
+	riscv_iommu_pmu_start_all(pmu, inhibit);
+
+	raw_spin_unlock_irqrestore(&pmu->lock, irqflags);
+
+	return IRQ_HANDLED;
+}
+
+static unsigned int riscv_iommu_pmu_get_irq_num(struct riscv_iommu_device *iommu)
+{
+	/* Reuse ICVEC.CIV mask for all interrupt vectors mapping */
+	int vec = (iommu->icvec >> (RISCV_IOMMU_INTR_PM * 4)) & RISCV_IOMMU_ICVEC_CIV;
+
+	return iommu->irqs[vec];
+}
+
+static int riscv_iommu_pmu_request_irq(struct auxiliary_device *auxdev,
+				       struct riscv_iommu_device *iommu,
+				       struct riscv_iommu_pmu *pmu)
+{
+	/*
+	 * Bind the handler to the auxiliary device, which is the same devres
+	 * scope that frees @pmu. Requesting it on the parent iommu device
+	 * would keep the handler registered with a dangling dev_id once @pmu
+	 * is freed, either on a later probe failure or on device removal.
+	 *
+	 * IRQF_SHARED is required because ICVEC maps the performance
+	 * monitoring source onto one of the vectors the iommu driver already
+	 * requested for its queues whenever fewer than RISCV_IOMMU_INTR_COUNT
+	 * vectors are available. Both requesters have to agree on sharing, or
+	 * this one is rejected with -EBUSY. IRQF_ONESHOT does not have to be
+	 * matched by hand: devm_request_irq() adds IRQF_COND_ONESHOT, so this
+	 * handler adopts whatever the first requester picked.
+	 */
+	return devm_request_irq(&auxdev->dev, pmu->irq, riscv_iommu_pmu_irq_handler,
+				IRQF_SHARED | IRQF_NOBALANCING,
+				dev_name(iommu->dev), pmu);
+}
+
+static void riscv_iommu_pmu_remove_cpuhp_instance(void *data)
+{
+	struct riscv_iommu_pmu *pmu = data;
+
+	cpuhp_state_remove_instance_nocalls(cpuhp_state, &pmu->node);
+}
+
+static void riscv_iommu_pmu_do_unregister(void *data)
+{
+	struct riscv_iommu_pmu *pmu = data;
+
+	perf_pmu_unregister(&pmu->pmu);
+}
+
+static int riscv_iommu_pmu_probe(struct auxiliary_device *auxdev,
+				 const struct auxiliary_device_id *id)
+{
+	struct riscv_iommu_device *iommu_dev = dev_get_platdata(&auxdev->dev);
+	struct riscv_iommu_pmu *iommu_pmu;
+	void __iomem *addr;
+	unsigned int idx;
+	char *name;
+	u32 inhibit;
+	int ret;
+
+	iommu_pmu = devm_kzalloc(&auxdev->dev, sizeof(*iommu_pmu), GFP_KERNEL);
+	if (!iommu_pmu)
+		return -ENOMEM;
+
+	iommu_pmu->reg = iommu_dev->reg;
+
+	raw_spin_lock_init(&iommu_pmu->lock);
+
+	/*
+	 * Which counters exist and how wide each one is are both
+	 * hardware-implemented, detected by writing 1s and reading back which
+	 * bits stuck.
+	 *
+	 * Section 6.22 defines each IOHPMCTR as an independent 64-bit WARL
+	 * register, so neither the set of implemented counters nor their widths
+	 * can be assumed uniform. Keep an implemented-counter bitmap and a
+	 * per-counter mask rather than a count and one shared width: a
+	 * consecutive index bound derived from a popcount would both hand out
+	 * counters which do not exist and leave overflows of implemented
+	 * counters above the bound unacknowledged, which makes pmip reassert
+	 * forever.
+	 *
+	 * The per-counter masks are still assumed to be a contiguous run of
+	 * bits starting at bit 0, which is what riscv_iommu_pmu_update() relies
+	 * on when it masks the difference of two samples to handle wraparound.
+	 *
+	 * The specification requires a minimum of one programmable event
+	 * counter besides the cycles counter when capabilities.HPM is 1, which
+	 * is the condition under which this device is created. So a compliant
+	 * implementation always provides at least two counters, and both
+	 * IOHPMCYCLES and the first IOHPMCTR are always present. A readback
+	 * which says otherwise is non-compliant hardware and is rejected rather
+	 * than worked around.
+	 */
+	addr = iommu_pmu->reg + RISCV_IOMMU_REG_IOCOUNTINH;
+	writel(RISCV_IOMMU_IOCOUNTINH_HPM, addr);
+	inhibit = readl(addr);
+	bitmap_from_arr32(iommu_pmu->implemented_counters, &inhibit,
+			  RISCV_IOMMU_HPM_COUNTER_NUM);
+
+	/* Bit 63 of IOHPMCYCLES is the OF bit, not part of the counter */
+	addr = iommu_pmu->reg + RISCV_IOMMU_REG_IOHPMCYCLES;
+	riscv_iommu_pmu_writeq(RISCV_IOMMU_IOHPMCYCLES_COUNTER, addr);
+	iommu_pmu->cntr_mask[RISCV_IOMMU_HPM_CYCLE_IDX] =
+		riscv_iommu_pmu_readq(addr) & RISCV_IOMMU_IOHPMCYCLES_COUNTER;
+	if (!iommu_pmu->cntr_mask[RISCV_IOMMU_HPM_CYCLE_IDX]) {
+		dev_err(&auxdev->dev, "cycles counter is not implemented\n");
+		return -ENODEV;
+	}
+
+	/* Detect the width of each implemented event counter individually */
+	idx = RISCV_IOMMU_HPM_CYCLE_IDX + 1;
+	for_each_set_bit_from(idx, iommu_pmu->implemented_counters,
+			      RISCV_IOMMU_HPM_COUNTER_NUM) {
+		addr = iommu_pmu->reg + RISCV_IOMMU_REG_IOHPMCTR(idx);
+		riscv_iommu_pmu_writeq(RISCV_IOMMU_IOHPMCTR_COUNTER, addr);
+		iommu_pmu->cntr_mask[idx] = riscv_iommu_pmu_readq(addr);
+
+		/* A counter with no writable bits is not really there */
+		if (!iommu_pmu->cntr_mask[idx])
+			clear_bit(idx, iommu_pmu->implemented_counters);
+	}
+
+	iommu_pmu->num_counters = bitmap_weight(iommu_pmu->implemented_counters,
+						RISCV_IOMMU_HPM_COUNTER_NUM);
+	if (iommu_pmu->num_counters < 2) {
+		dev_err(&auxdev->dev, "hardware reports %u counter(s)\n",
+			iommu_pmu->num_counters);
+		return -ENODEV;
+	}
+
+	bitmap_to_arr32(&iommu_pmu->inhibit_all, iommu_pmu->implemented_counters,
+			RISCV_IOMMU_HPM_COUNTER_NUM);
+
+	iommu_pmu->pmu = (struct pmu) {
+		.module		= THIS_MODULE,
+		.parent		= &auxdev->dev,
+		.task_ctx_nr	= perf_invalid_context,
+		.event_init	= riscv_iommu_pmu_event_init,
+		.add		= riscv_iommu_pmu_add,
+		.del		= riscv_iommu_pmu_del,
+		.start		= riscv_iommu_pmu_start,
+		.stop		= riscv_iommu_pmu_stop,
+		.read		= riscv_iommu_pmu_read,
+		.attr_groups	= riscv_iommu_pmu_attr_grps,
+		.capabilities	= PERF_PMU_CAP_NO_EXCLUDE,
+	};
+
+	auxiliary_set_drvdata(auxdev, iommu_pmu);
+
+	name = devm_kasprintf(&auxdev->dev, GFP_KERNEL,
+			      "riscv_iommu_pmu_%u", auxdev->id);
+	if (!name) {
+		dev_err(&auxdev->dev, "Failed to create name riscv_iommu_pmu_%u\n",
+			auxdev->id);
+		return -ENOMEM;
+	}
+
+	iommu_pmu->numa_node = dev_to_node(iommu_dev->dev);
+	iommu_pmu->irq = riscv_iommu_pmu_get_irq_num(iommu_dev);
+
+	ret = riscv_iommu_pmu_request_irq(auxdev, iommu_dev, iommu_pmu);
+	if (ret) {
+		dev_err(&auxdev->dev, "Failed to request irq %s: %d\n", name, ret);
+		return ret;
+	}
+
+	/*
+	 * Bind all events to the same cpu context to avoid race enabling.
+	 * riscv_iommu_pmu_online_cpu() picks the CPU and sets the irq
+	 * affinity for us once the instance is registered below.
+	 */
+	iommu_pmu->on_cpu = -1;
+
+	ret = cpuhp_state_add_instance(cpuhp_state, &iommu_pmu->node);
+	if (ret) {
+		dev_err(&auxdev->dev, "Failed to register hotplug %s: %d\n", name, ret);
+		return ret;
+	}
+
+	ret = perf_pmu_register(&iommu_pmu->pmu, name, -1);
+	if (ret) {
+		dev_err(&auxdev->dev, "Failed to register %s: %d\n", name, ret);
+		cpuhp_state_remove_instance_nocalls(cpuhp_state, &iommu_pmu->node);
+		return ret;
+	}
+
+	ret = devm_add_action_or_reset(&auxdev->dev,
+				       riscv_iommu_pmu_do_unregister,
+				       iommu_pmu);
+	if (ret) {
+		/*
+		 * The action above has already run, so the pmu is unregistered
+		 * and pmu->cpu_pmu_context is freed. Drop the cpuhp instance
+		 * with _nocalls() so the teardown callback cannot reach it, and
+		 * so it does not outlive iommu_pmu on the global cpuhp list.
+		 */
+		cpuhp_state_remove_instance_nocalls(cpuhp_state, &iommu_pmu->node);
+		return ret;
+	}
+
+	/*
+	 * Registered after do_unregister so it runs first (LIFO) on unbind:
+	 * the cpuhp instance must be gone before perf_pmu_unregister() runs.
+	 */
+	ret = devm_add_action_or_reset(&auxdev->dev,
+				       riscv_iommu_pmu_remove_cpuhp_instance,
+				       iommu_pmu);
+	if (ret)
+		return ret;
+
+	/*
+	 * The PMU name only carries the aux dev id, not the iommu dev name, so
+	 * find the iommu dev name here to map this PMU back to its iommu dev.
+	 */
+	dev_info(&auxdev->dev, "%s: Registered with %u counters (iommu %s)\n",
+		 name, iommu_pmu->num_counters, dev_name(iommu_dev->dev));
+
+	return 0;
+}
+
+static const struct auxiliary_device_id riscv_iommu_pmu_id_table[] = {
+	{ .name = "riscv-iommu.pmu" },
+	{}
+};
+MODULE_DEVICE_TABLE(auxiliary, riscv_iommu_pmu_id_table);
+
+static struct auxiliary_driver iommu_pmu_driver = {
+	.driver = {
+		.suppress_bind_attrs = true,
+	},
+	.probe		= riscv_iommu_pmu_probe,
+	.id_table	= riscv_iommu_pmu_id_table,
+};
+
+static int __init riscv_iommu_pmu_init(void)
+{
+	int ret;
+
+	cpuhp_state = cpuhp_setup_state_multi(CPUHP_AP_ONLINE_DYN,
+					      "perf/riscv/iommu:online",
+					      riscv_iommu_pmu_online_cpu,
+					      riscv_iommu_pmu_offline_cpu);
+	if (cpuhp_state < 0)
+		return cpuhp_state;
+
+	ret = auxiliary_driver_register(&iommu_pmu_driver);
+	if (ret)
+		cpuhp_remove_multi_state(cpuhp_state);
+
+	return ret;
+}
+module_init(riscv_iommu_pmu_init);
+
+MODULE_DESCRIPTION("RISC-V IOMMU PMU");
+MODULE_LICENSE("GPL");
-- 
2.43.7


^ permalink raw reply related	[flat|nested] 5+ messages in thread

* [PATCH v15 2/2] iommu/riscv: create a auxiliary device for HPM
  2026-10-02  6:33 [PATCH v15 0/2] RISC-V IOMMU HPM support Zong Li
  2026-10-02  6:33 ` [PATCH v15 1/2] drivers/perf: riscv-iommu: add risc-v iommu pmu driver Zong Li
@ 2026-10-02  6:33 ` Zong Li
  2026-10-02  9:14   ` sashiko-bot
  1 sibling, 1 reply; 5+ messages in thread
From: Zong Li @ 2026-10-02  6:33 UTC (permalink / raw)
  To: tomasz.jeznach, joro, will, robin.murphy, pjw, palmer, aou, alex,
	mark.rutland, andrew.jones, guoren, david.laight.linux,
	zhangzhanpeng.jasper, yang.yicong, nutty.liu, iommu, linux-riscv,
	linux-kernel, linux-perf-users
  Cc: Zong Li, Chen Pei, Fangyu Yu, Samuel Holland

Create an auxiliary device for HPM when the IOMMU supports a
hardware performance monitor.

Tested-by: Chen Pei <cp0613@linux.alibaba.com>
Tested-by: Fangyu Yu <fangyu.yu@linux.alibaba.com>
Reviewed-by: Nutty Liu <nutty.liu@hotmail.com>
Reviewed-by: Guo Ren <guoren@kernel.org>
Reviewed-by: Yicong Yang <yang.yicong@picoheart.com>
Suggested-by: Samuel Holland <samuel.holland@sifive.com>
Signed-off-by: Zong Li <zong.li@sifive.com>
---
 drivers/iommu/riscv/Kconfig |  1 +
 drivers/iommu/riscv/iommu.c | 58 +++++++++++++++++++++++++++++++++++++
 drivers/iommu/riscv/iommu.h |  4 +++
 3 files changed, 63 insertions(+)

diff --git a/drivers/iommu/riscv/Kconfig b/drivers/iommu/riscv/Kconfig
index b86e5ab94183..8025bf0fb67f 100644
--- a/drivers/iommu/riscv/Kconfig
+++ b/drivers/iommu/riscv/Kconfig
@@ -10,6 +10,7 @@ config RISCV_IOMMU
 	select GENERIC_PT
 	select IOMMU_PT
 	select IOMMU_PT_RISCV64
+	select AUXILIARY_BUS
 	help
 	  Support for implementations of the RISC-V IOMMU architecture that
 	  complements the RISC-V MMU capabilities, providing similar address
diff --git a/drivers/iommu/riscv/iommu.c b/drivers/iommu/riscv/iommu.c
index fe8e6d0f8a23..1cb7369bd962 100644
--- a/drivers/iommu/riscv/iommu.c
+++ b/drivers/iommu/riscv/iommu.c
@@ -14,6 +14,7 @@
 
 #include <linux/acpi.h>
 #include <linux/acpi_rimt.h>
+#include <linux/auxiliary_bus.h>
 #include <linux/compiler.h>
 #include <linux/crash_dump.h>
 #include <linux/init.h>
@@ -48,6 +49,9 @@
 static DEFINE_IDA(riscv_iommu_pscids);
 #define RISCV_IOMMU_MAX_PSCID		(BIT(20) - 1)
 
+/* IOMMU PMU auxiliary device id allocation namespace. */
+static DEFINE_IDA(riscv_iommu_pmu_ida);
+
 /* Device resource-managed allocations */
 struct riscv_iommu_devres {
 	void *addr;
@@ -587,6 +591,46 @@ static irqreturn_t riscv_iommu_fltq_process(int irq, void *data)
 	return IRQ_HANDLED;
 }
 
+/*
+ * IOMMU Hardware performance monitor
+ */
+static void riscv_iommu_pmu_id_free(void *data)
+{
+	ida_free(&riscv_iommu_pmu_ida, (unsigned long)data);
+}
+
+static int riscv_iommu_hpm_enable(struct riscv_iommu_device *iommu)
+{
+	struct auxiliary_device *auxdev;
+	int id, ret;
+
+	id = ida_alloc(&riscv_iommu_pmu_ida, GFP_KERNEL);
+	if (id < 0)
+		return id;
+
+	ret = devm_add_action_or_reset(iommu->dev, riscv_iommu_pmu_id_free,
+				       (void *)(unsigned long)id);
+	if (ret)
+		return ret;
+
+	auxdev = auxiliary_device_create(iommu->dev, "riscv-iommu",
+					 "pmu", iommu, id);
+	if (!auxdev)
+		return -ENODEV;
+
+	iommu->pmu_dev = auxdev;
+	return 0;
+}
+
+static void riscv_iommu_hpm_disable(struct riscv_iommu_device *iommu)
+{
+	if (!iommu->pmu_dev)
+		return;
+
+	auxiliary_device_destroy(iommu->pmu_dev);
+	iommu->pmu_dev = NULL;
+}
+
 /* Lookup and initialize device context info structure. */
 static struct riscv_iommu_dc *riscv_iommu_get_dc(struct riscv_iommu_device *iommu,
 						 unsigned int devid)
@@ -1570,6 +1614,7 @@ static int riscv_iommu_init_check(struct riscv_iommu_device *iommu)
 
 void riscv_iommu_remove(struct riscv_iommu_device *iommu)
 {
+	riscv_iommu_hpm_disable(iommu);
 	iommu_device_unregister(&iommu->iommu);
 	iommu_device_sysfs_remove(&iommu->iommu);
 	riscv_iommu_iodir_set_mode(iommu, RISCV_IOMMU_DDTP_IOMMU_MODE_OFF);
@@ -1636,6 +1681,19 @@ int riscv_iommu_init(struct riscv_iommu_device *iommu)
 		goto err_remove_sysfs;
 	}
 
+	if (iommu->caps & RISCV_IOMMU_CAPABILITIES_HPM) {
+		/*
+		 * The IOMMU is fully functional without the PMU, so keep it
+		 * operational if the HPM device cannot be created - but do not
+		 * leave the failure unreported either.
+		 */
+		rc = riscv_iommu_hpm_enable(iommu);
+		if (rc)
+			dev_warn(iommu->dev,
+				 "cannot create HPM device, PMU unavailable: %d\n",
+				 rc);
+	}
+
 	return 0;
 
 err_remove_sysfs:
diff --git a/drivers/iommu/riscv/iommu.h b/drivers/iommu/riscv/iommu.h
index 5676001548cc..7fcd79fa47c8 100644
--- a/drivers/iommu/riscv/iommu.h
+++ b/drivers/iommu/riscv/iommu.h
@@ -19,6 +19,7 @@
 #include "iommu-bits.h"
 
 struct riscv_iommu_device;
+struct auxiliary_device;
 
 struct riscv_iommu_queue {
 	atomic_t prod;				/* unbounded producer allocation index */
@@ -62,6 +63,9 @@ struct riscv_iommu_device {
 	unsigned int ddt_mode;
 	dma_addr_t ddt_phys;
 	u64 *ddt_root;
+
+	/* PMU auxiliary device; torn down explicitly before MSI teardown */
+	struct auxiliary_device *pmu_dev;
 };
 
 int riscv_iommu_init(struct riscv_iommu_device *iommu);
-- 
2.43.7


^ permalink raw reply related	[flat|nested] 5+ messages in thread

* Re: [PATCH v15 2/2] iommu/riscv: create a auxiliary device for HPM
  2026-10-02  6:33 ` [PATCH v15 2/2] iommu/riscv: create a auxiliary device for HPM Zong Li
@ 2026-10-02  9:14   ` sashiko-bot
  0 siblings, 0 replies; 5+ messages in thread
From: sashiko-bot @ 2026-10-02  9:14 UTC (permalink / raw)
  To: Zong Li; +Cc: linux-perf-users

> Create an auxiliary device for HPM when the IOMMU supports a
> hardware performance monitor.
> 
> Tested-by: Chen Pei <cp0613@linux.alibaba.com>
> Tested-by: Fangyu Yu <fangyu.yu@linux.alibaba.com>
> Reviewed-by: Nutty Liu <nutty.liu@hotmail.com>
> Reviewed-by: Guo Ren <guoren@kernel.org>
> Reviewed-by: Yicong Yang <yang.yicong@picoheart.com>
> Suggested-by: Samuel Holland <samuel.holland@sifive.com>
> Signed-off-by: Zong Li <zong.li@sifive.com>

Sashiko has reviewed this patch and found no issues. It looks great!

-- 
Sashiko AI review · https://sashiko.dev/#/patchset/20261002063315.1535817-1-zong.li@sifive.com?part=2


^ permalink raw reply	[flat|nested] 5+ messages in thread

* Re: [PATCH v15 1/2] drivers/perf: riscv-iommu: add risc-v iommu pmu driver
  2026-10-02  6:33 ` [PATCH v15 1/2] drivers/perf: riscv-iommu: add risc-v iommu pmu driver Zong Li
@ 2026-10-02  9:14   ` sashiko-bot
  0 siblings, 0 replies; 5+ messages in thread
From: sashiko-bot @ 2026-10-02  9:14 UTC (permalink / raw)
  To: Zong Li; +Cc: linux-perf-users

Thank you for your contribution! Sashiko AI review found 2 potential issue(s) to consider:
- [High] perf/riscv: Torn read race between PMU read and IRQ handler corrupts event count
- [Medium] perf/riscv: UAF in CPU hotplug callback during PMU probe error path
--

commit 472769839e43688676e3929143858e4d52f609c7
Author: Zong Li <zong.li@sifive.com>

    drivers/perf: riscv-iommu: add risc-v iommu pmu driver

This commit introduces a new driver for the RISC-V IOMMU PMU, created as an
auxiliary device by the parent RISC-V IOMMU driver. The PMU provides event
and cycle counters with filtering support to collect performance metrics.

Link: https://lore.kernel.org/linux-riscv/20260618143634.7f3dd6c5@pumpkin/

> diff --git a/drivers/perf/riscv_iommu_pmu.c b/drivers/perf/riscv_iommu_pmu.c
> index 0000000000000..25f217ac97e7f
> --- /dev/null
> +++ b/drivers/perf/riscv_iommu_pmu.c

[ ... ]

> +static void riscv_iommu_pmu_set_period(struct perf_event *event)
> +{
> +	struct riscv_iommu_pmu *pmu = to_riscv_iommu_pmu(event->pmu);
> +	struct hw_perf_event *hwc = &event->hw;
> +	u64 counter_mask = pmu->cntr_mask[hwc->idx];
> +	u64 period;
> +
> +	/*
> +	 * Limit the maximum period to prevent the counter value
> +	 * from overtaking the one we are about to program.
> +	 * In effect we are reducing max_period to account for
> +	 * interrupt latency (and we are being very conservative).
> +	 */
> +	period = counter_mask >> 1;
> +	riscv_iommu_pmu_set_counter(pmu, hwc->idx, period);
> +	local64_set(&hwc->prev_count, period);
> +}

[ ... ]

> +static void riscv_iommu_pmu_update(struct perf_event *event)
> +{
> +	struct hw_perf_event *hwc = &event->hw;
> +	struct riscv_iommu_pmu *pmu = to_riscv_iommu_pmu(event->pmu);
> +	u64 delta, prev, now;
> +	u32 idx = hwc->idx;
> +	u64 counter_mask = pmu->cntr_mask[idx];
> +
> +	/*
> +	 * riscv_iommu_pmu_set_period() resets the hardware counter and
> +	 * prev_count as two separate writes.
> +	 * Pairing a "prev" read from before that reset with a "now" read from
> +	 * after it would produce a nonsensical, possibly huge, delta.
> +	 * Re-checking prev_count after the hardware read detects that torn
> +	 * pairing and retries.
> +	 */
> +	do {
> +		prev = local64_read(&hwc->prev_count);
> +		now = riscv_iommu_pmu_get_counter(pmu, idx);
> +	} while (prev != local64_read(&hwc->prev_count) ||
> +		 local64_cmpxchg(&hwc->prev_count, prev, now) != prev);
> +
> +	delta = (now - prev) & counter_mask;
> +	local64_add(delta, &event->count);
> +}

[ ... ]

> +static void riscv_iommu_pmu_read(struct perf_event *event)
> +{
> +	riscv_iommu_pmu_update(event);
> +}

[Severity: High]
Does this code allow a lockless read to calculate a torn update delta?

The commit message notes that IRQCHIP_MOVE_DEFERRED can cause the PMU interrupt
to be routed to a different CPU than the event is bound to.

If riscv_iommu_pmu_read() is called locklessly on CPU B, it can race with
riscv_iommu_pmu_irq_handler() running on CPU A. In riscv_iommu_pmu_set_period(),
the hardware counter is written before the shadow copy prev_count.

If CPU B reads the old near-overflow prev_count, and then CPU A resets the
hardware counter, CPU B could read the new reset hardware counter. Because
CPU A hasn't reset the software prev_count yet, CPU B's loop validation
(prev == local64_read()) would succeed, producing a massive wrap-around spike
in the event count.

[ ... ]

> +static int riscv_iommu_pmu_probe(struct auxiliary_device *auxdev,
> +				 const struct auxiliary_device_id *id)
> +{

[ ... ]

> +	ret = devm_add_action_or_reset(&auxdev->dev,
> +				       riscv_iommu_pmu_do_unregister,
> +				       iommu_pmu);
> +	if (ret) {
> +		/*
> +		 * The action above has already run, so the pmu is unregistered
> +		 * and pmu->cpu_pmu_context is freed. Drop the cpuhp instance
> +		 * with _nocalls() so the teardown callback cannot reach it, and
> +		 * so it does not outlive iommu_pmu on the global cpuhp list.
> +		 */
> +		cpuhp_state_remove_instance_nocalls(cpuhp_state, &iommu_pmu->node);
> +		return ret;
> +	}

[Severity: Medium]
Can this sequence result in a use-after-free during a concurrent CPU offline
event?

If devm_add_action_or_reset() fails to allocate a devres node, it will
immediately execute riscv_iommu_pmu_do_unregister(), which calls
perf_pmu_unregister() and frees the PMU per-CPU contexts.

Because the CPU hotplug instance is not removed before the teardown action
executes, a concurrent CPU hotplug callback (riscv_iommu_pmu_offline_cpu)
might be invoked in this window. Could this cause perf_pmu_migrate_context()
to be called on the already-freed contexts?

-- 
Sashiko AI review · https://sashiko.dev/#/patchset/20261002063315.1535817-1-zong.li@sifive.com?part=1

^ permalink raw reply	[flat|nested] 5+ messages in thread

end of thread, other threads:[~2026-10-02  9:14 UTC | newest]

Thread overview: 5+ messages (download: mbox.gz follow: Atom feed
-- links below jump to the message on this page --
2026-10-02  6:33 [PATCH v15 0/2] RISC-V IOMMU HPM support Zong Li
2026-10-02  6:33 ` [PATCH v15 1/2] drivers/perf: riscv-iommu: add risc-v iommu pmu driver Zong Li
2026-10-02  9:14   ` sashiko-bot
2026-10-02  6:33 ` [PATCH v15 2/2] iommu/riscv: create a auxiliary device for HPM Zong Li
2026-10-02  9:14   ` sashiko-bot

This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox