* [PATCH v3] perf: Add Arm Bus Monitor Unit driver
@ 2026-09-17 13:47 Robin Murphy
2026-09-17 14:02 ` sashiko-bot
2026-10-04 22:10 ` Will Deacon
0 siblings, 2 replies; 4+ messages in thread
From: Robin Murphy @ 2026-09-17 13:47 UTC (permalink / raw)
To: will; +Cc: mark.rutland, linux-perf-users, linux-arm-kernel
Arm's Bus Monitor Unit is a low-level performance analysis tool for
matching and counting transactions at interconnect interfaces. Each
BMU consists of a number of "interface monitoring units", with some
global controls to synchronise them and so permit accurate calculation
of cross-interface metrics. Each IMU contains, among other things, its
own PMU largely based on the CoreSight PMU architecture, but with some
functional incompatibilities due to the nature of how much per-event
filtering needs to be crammed in to categorise AXI/CHI transactions at
the protocol level. As such it gets an honorary home in arm_cspmu/ in
order to share the common register definitions, but for simplicity does
not try to bend the functional abstraction to fit.
Signed-off-by: Robin Murphy <robin.murphy@arm.com>
---
v3:
- Couple of inconsequential tweaks to shut Sashiko up about things
that would never actually happen anyway
- Drop unneeded DT patches that I can't be bothered to fix right now
---
drivers/perf/arm_cspmu/Kconfig | 6 +
drivers/perf/arm_cspmu/Makefile | 2 +
drivers/perf/arm_cspmu/arm-bmu.c | 528 +++++++++++++++++++++++++++++++
3 files changed, 536 insertions(+)
create mode 100644 drivers/perf/arm_cspmu/arm-bmu.c
diff --git a/drivers/perf/arm_cspmu/Kconfig b/drivers/perf/arm_cspmu/Kconfig
index 6f4e28fc84a2..0cb369f91023 100644
--- a/drivers/perf/arm_cspmu/Kconfig
+++ b/drivers/perf/arm_cspmu/Kconfig
@@ -27,3 +27,9 @@ config AMPERE_CORESIGHT_PMU_ARCH_SYSTEM_PMU
In the first phase, the driver enables support on MCU PMU used in
AmpereOne SoC family.
+
+config ARM_BMU
+ tristate "ARM Bus Monitor Unit"
+ depends on (ARM64 && ACPI) || COMPILE_TEST
+ help
+ Support perf event counting on Arm Bus Monitor Unit devices.
diff --git a/drivers/perf/arm_cspmu/Makefile b/drivers/perf/arm_cspmu/Makefile
index 220a734efd54..622b8fc3d5dc 100644
--- a/drivers/perf/arm_cspmu/Makefile
+++ b/drivers/perf/arm_cspmu/Makefile
@@ -8,3 +8,5 @@ arm_cspmu_module-y := arm_cspmu.o
obj-$(CONFIG_NVIDIA_CORESIGHT_PMU_ARCH_SYSTEM_PMU) += nvidia_cspmu.o
obj-$(CONFIG_AMPERE_CORESIGHT_PMU_ARCH_SYSTEM_PMU) += ampere_cspmu.o
+
+obj-$(CONFIG_ARM_BMU) += arm-bmu.o
diff --git a/drivers/perf/arm_cspmu/arm-bmu.c b/drivers/perf/arm_cspmu/arm-bmu.c
new file mode 100644
index 000000000000..fe25954ac836
--- /dev/null
+++ b/drivers/perf/arm_cspmu/arm-bmu.c
@@ -0,0 +1,528 @@
+// SPDX-License-Identifier: GPL-2.0
+// Copyright (C) 2025-2026 Arm Limited
+
+#include <linux/acpi.h>
+#include <linux/bitfield.h>
+#include <linux/interrupt.h>
+#include <linux/io-64-nonatomic-lo-hi.h>
+#include <linux/module.h>
+#include <linux/perf_event.h>
+#include <linux/platform_device.h>
+#include <linux/property.h>
+
+#include "arm_cspmu.h"
+
+#define MCU_CONFIG 0x0000
+#define MCU_PMU_GLOBAL_EN 0x0008
+#define MCU_PMU_INTERRUPT_STATUS 0x0090
+
+#define MCUCFG_NUM_IMU_MONITORS GENMASK_ULL(63, 60)
+#define MCUCFG_PMU_ELEMENT_SIZE GENMASK_ULL(11, 8)
+#define MCUCFG_PMU_ELEMENT_START GENMASK_ULL(7, 0)
+
+#define BMU_PMIMPDEF 0x0200
+#define BMU_PMCR 0x0e10
+
+/* These are reasonable expectations for now */
+#define PMU_MAX_COUNTERS 16
+#define MAX_IMUS 16
+
+/* Event attributes */
+#define BMU_CONFIG_IMU GENMASK_ULL(63, 56)
+#define BMU_CONFIG_EVENT GENMASK_ULL(55, 0)
+#define BMU_CONFIG_FILTER GENMASK_ULL(63, 0)
+#define BMU_CONFIG_MATCH GENMASK_ULL(63, 0)
+#define BMU_CONFIG_MASK GENMASK_ULL(63, 0)
+
+#define BMU_EVENT_IMU(e) FIELD_GET(BMU_CONFIG_IMU, (e)->attr.config)
+#define BMU_EVENT_EVENT(e) FIELD_GET(BMU_CONFIG_EVENT, (e)->attr.config)
+#define BMU_EVENT_FILTER(e) FIELD_GET(BMU_CONFIG_FILTER, (e)->attr.config1)
+#define BMU_EVENT_MATCH(e) FIELD_GET(BMU_CONFIG_MATCH, (e)->attr.config2)
+#define BMU_EVENT_MASK(e) FIELD_GET(BMU_CONFIG_MASK, (e)->attr.config3)
+
+static unsigned long arm_bmu_cpuhp_state;
+
+struct arm_bmu_pmu {
+ void __iomem *base;
+ struct perf_event *evcnt[PMU_MAX_COUNTERS];
+ int num_counters;
+};
+
+struct arm_bmu {
+ struct device *dev;
+ void __iomem *base;
+ struct pmu pmu;
+ struct hlist_node cpuhp_node;
+ int cpu;
+ int irq;
+ int num_imus;
+ struct arm_bmu_pmu imus[] __counted_by(num_imus);
+};
+
+#define to_bmu(p) container_of(p, struct arm_bmu, pmu)
+#define event_imu(e) (e)->hw.event_base
+#define to_bmu_pmu(e) &to_bmu(e->pmu)->imus[event_imu(e)]
+
+struct arm_bmu_format_attr {
+ struct device_attribute attr;
+ DECLARE_BITMAP(field, 64);
+ int config;
+};
+
+#define BMU_FORMAT_ATTR(_name, _cfg, _fld) \
+ (&((struct arm_bmu_format_attr[]) {{ \
+ .attr = __ATTR(_name, 0444, arm_bmu_format_show, NULL), \
+ .field = { BITMAP_FROM_U64(_fld) }, \
+ .config = _cfg, \
+ }})[0].attr.attr)
+
+static ssize_t arm_bmu_format_show(struct device *dev,
+ struct device_attribute *attr, char *buf)
+{
+ struct arm_bmu_format_attr *fmt = container_of(attr, typeof(*fmt), attr);
+
+ if (!fmt->config)
+ return sysfs_emit(buf, "config:%*pbl\n", 64, &fmt->field);
+
+ return sysfs_emit(buf, "config%d:%*pbl\n", fmt->config, 64, &fmt->field);
+}
+
+static struct attribute *arm_bmu_format_attrs[] = {
+ BMU_FORMAT_ATTR(imu, 0, BMU_CONFIG_IMU),
+ BMU_FORMAT_ATTR(event, 0, BMU_CONFIG_EVENT),
+ BMU_FORMAT_ATTR(filter, 1, BMU_CONFIG_FILTER),
+ BMU_FORMAT_ATTR(match, 2, BMU_CONFIG_MATCH),
+ BMU_FORMAT_ATTR(mask, 3, BMU_CONFIG_MASK),
+ NULL
+};
+
+static const struct attribute_group arm_bmu_format_attrs_group = {
+ .name = "format",
+ .attrs = arm_bmu_format_attrs,
+};
+
+static ssize_t arm_bmu_cpumask_show(struct device *dev,
+ struct device_attribute *attr, char *buf)
+{
+ struct arm_bmu *bmu = to_bmu(dev_get_drvdata(dev));
+
+ return sysfs_emit(buf, "%*pbl\n", cpumask_pr_args(cpumask_of(bmu->cpu)));
+}
+
+static struct device_attribute arm_bmu_cpumask_attr =
+ __ATTR(cpumask, 0444, arm_bmu_cpumask_show, NULL);
+
+static ssize_t arm_bmu_identifier_show(struct device *dev,
+ struct device_attribute *attr, char *buf)
+{
+ struct arm_bmu *bmu = to_bmu(dev_get_drvdata(dev));
+ u32 reg = readl_relaxed(bmu->imus[0].base + PMIIDR);
+
+ return sysfs_emit(buf, "%x\n", reg);
+}
+
+static struct device_attribute arm_bmu_identifier_attr =
+ __ATTR(identifier, 0444, arm_bmu_identifier_show, NULL);
+
+static struct attribute *arm_bmu_other_attrs[] = {
+ &arm_bmu_cpumask_attr.attr,
+ &arm_bmu_identifier_attr.attr,
+ NULL
+};
+
+static const struct attribute_group arm_bmu_other_attr_group = {
+ .attrs = arm_bmu_other_attrs,
+};
+
+static const struct attribute_group *arm_bmu_attr_groups[] = {
+ &arm_bmu_format_attrs_group,
+ &arm_bmu_other_attr_group,
+ NULL
+};
+
+static void arm_bmu_enable(struct pmu *pmu)
+{
+ writel_relaxed(1, to_bmu(pmu)->base + MCU_PMU_GLOBAL_EN);
+}
+
+static void arm_bmu_disable(struct pmu *pmu)
+{
+ writel_relaxed(0, to_bmu(pmu)->base + MCU_PMU_GLOBAL_EN);
+}
+
+static int arm_bmu_validate_group(struct perf_event *event)
+{
+ struct arm_bmu *bmu = to_bmu(event->pmu);
+ struct perf_event *sibling, *leader = event->group_leader;
+ int num[MAX_IMUS] = { 0 };
+
+ ++num[event_imu(event)];
+ if (leader != event) {
+ if (leader->pmu == &bmu->pmu)
+ ++num[event_imu(leader)];
+ for_each_sibling_event(sibling, leader) {
+ if (sibling->pmu == &bmu->pmu)
+ ++num[event_imu(sibling)];
+ }
+ }
+ for (int i = 0; i < bmu->num_imus; i++) {
+ if (num[i] > bmu->imus[i].num_counters)
+ return -EINVAL;
+ }
+ return 0;
+}
+
+static int arm_bmu_event_init(struct perf_event *event)
+{
+ struct arm_bmu *bmu = to_bmu(event->pmu);
+
+ if (event->attr.type != event->pmu->type)
+ return -ENOENT;
+
+ if (is_sampling_event(event))
+ return -EINVAL;
+
+ event->cpu = bmu->cpu;
+ event_imu(event) = BMU_EVENT_IMU(event);
+ if (event_imu(event) >= bmu->num_imus)
+ return -EINVAL;
+
+ return arm_bmu_validate_group(event);
+}
+
+static u64 arm_bmu_read_evcnt(struct perf_event *event)
+{
+ struct arm_bmu_pmu *pmu = to_bmu_pmu(event);
+ void __iomem *reg_base = pmu->base + event->hw.idx * sizeof(u64);
+ u64 lo, hi_old, hi_new;
+ int retries = 3; /* 1st time unlucky, 2nd improbable, 3rd just broken */
+
+ hi_new = readl_relaxed(reg_base + PMEVCNTR_HI);
+ do {
+ hi_old = hi_new;
+ lo = readl_relaxed(reg_base + PMEVCNTR_LO);
+ hi_new = readl_relaxed(reg_base + PMEVCNTR_HI);
+ } while (hi_new != hi_old && --retries);
+ WARN_ON(!retries);
+
+ return (hi_new << 32) | lo;
+}
+
+static void arm_bmu_event_read(struct perf_event *event)
+{
+ struct hw_perf_event *hw = &event->hw;
+ u64 count, prev;
+
+ do {
+ prev = local64_read(&hw->prev_count);
+ count = arm_bmu_read_evcnt(event);
+ } while (local64_cmpxchg(&hw->prev_count, prev, count) != prev);
+
+ count -= prev;
+ local64_add(count, &event->count);
+}
+
+static void arm_bmu_event_start(struct perf_event *event, int flags)
+{
+ struct arm_bmu_pmu *pmu = to_bmu_pmu(event);
+
+ writel_relaxed(1ULL << event->hw.idx, pmu->base + PMCNTENSET);
+}
+
+static void arm_bmu_event_stop(struct perf_event *event, int flags)
+{
+ struct arm_bmu_pmu *pmu = to_bmu_pmu(event);
+
+ writel_relaxed(1ULL << event->hw.idx, pmu->base + PMCNTENCLR);
+ if (flags & PERF_EF_UPDATE)
+ arm_bmu_event_read(event);
+}
+
+static int arm_bmu_event_add(struct perf_event *event, int flags)
+{
+ struct arm_bmu_pmu *pmu = to_bmu_pmu(event);
+ struct hw_perf_event *hw = &event->hw;
+ void __iomem *reg_base;
+
+ hw->idx = 0;
+ while (pmu->evcnt[hw->idx]) {
+ if (++hw->idx == pmu->num_counters)
+ return -ENOSPC;
+ }
+ pmu->evcnt[hw->idx] = event;
+
+ reg_base = pmu->base + hw->idx * sizeof(u64);
+ lo_hi_writeq_relaxed(BMU_EVENT_EVENT(event), reg_base + PMEVTYPER);
+ lo_hi_writeq_relaxed(BMU_EVENT_FILTER(event), reg_base + PMEVFILTR);
+ lo_hi_writeq_relaxed(BMU_EVENT_MATCH(event), reg_base + PMEVFILT2R);
+ lo_hi_writeq_relaxed(BMU_EVENT_MASK(event), reg_base + BMU_PMIMPDEF);
+
+ local64_set(&hw->prev_count, S64_MIN);
+ lo_hi_writeq_relaxed(S64_MIN, reg_base + PMEVCNTR_LO);
+
+ if (flags & PERF_EF_START)
+ arm_bmu_event_start(event, 0);
+ return 0;
+}
+
+static void arm_bmu_event_del(struct perf_event *event, int flags)
+{
+ struct arm_bmu_pmu *pmu = to_bmu_pmu(event);
+
+ arm_bmu_event_stop(event, PERF_EF_UPDATE);
+ pmu->evcnt[event->hw.idx] = NULL;
+}
+
+static void arm_bmu_pmu_irq(struct arm_bmu *bmu, int imu)
+{
+ struct arm_bmu_pmu *pmu = bmu->imus + imu;
+ u32 reg = readl_relaxed(pmu->base + PMOVSCLR);
+ u64 __iomem *pmevcnt = pmu->base + PMEVCNTR_LO;
+
+ for (int i = 0; i < PMU_MAX_COUNTERS; i++) {
+ if (!(reg & (1U << i)))
+ continue;
+ if (!pmu->evcnt[i]) {
+ dev_dbg(bmu->dev, "Spurious oveflow on IMU %d counter %d?\n", imu, i);
+ continue;
+ }
+ arm_bmu_event_read(pmu->evcnt[i]);
+ local64_set(&pmu->evcnt[i]->hw.prev_count, S64_MIN);
+ lo_hi_writeq_relaxed(S64_MIN, pmevcnt + i);
+ }
+ writel_relaxed(reg, pmu->base + PMOVSCLR);
+}
+
+static irqreturn_t arm_bmu_handle_irq(int irq, void *dev_id)
+{
+ struct arm_bmu *bmu = dev_id;
+ u32 reg = readl_relaxed(bmu->base + MCU_PMU_INTERRUPT_STATUS);
+ irqreturn_t ret = reg ? IRQ_HANDLED : IRQ_NONE;
+
+ for (int i = 0; i < bmu->num_imus && reg; i++, reg >>= 1) {
+ if (reg & 1)
+ arm_bmu_pmu_irq(bmu, i);
+ }
+ return ret;
+}
+
+static int arm_bmu_probe(struct platform_device *pdev)
+{
+ struct device *dev = &pdev->dev;
+ const struct resource *res;
+ struct arm_bmu *bmu;
+ const char *name = NULL;
+ void __iomem *base;
+ static atomic_t n;
+ int err, num, sz, off;
+ u64 cfg;
+ u32 reg;
+
+ res = platform_get_resource(pdev, IORESOURCE_MEM, 0);
+ if (!res)
+ return -EINVAL;
+
+ /* PMUs and MPAM monitors are intermingled so we can't claim the whole resource */
+ base = devm_ioremap(dev, res->start, resource_size(res));
+ if (!base)
+ return -ENOMEM;
+
+ /* Global and per-PMU-page security all defaults to Root/Secure only */
+ writel_relaxed(1, base + MCU_PMU_GLOBAL_EN);
+ reg = readl_relaxed(base + MCU_PMU_GLOBAL_EN);
+ if (reg != 1)
+ return dev_err_probe(dev, -ENODEV, "Non-Secure access to BMU not enabled?\n");
+ writel_relaxed(0, base + MCU_PMU_GLOBAL_EN);
+
+ cfg = lo_hi_readq_relaxed(base + MCU_CONFIG);
+ num = 1 + FIELD_GET(MCUCFG_NUM_IMU_MONITORS, cfg);
+ /* We don't expect to have dual-page complications to worry about */
+ sz = FIELD_GET(MCUCFG_PMU_ELEMENT_SIZE, cfg);
+ if (sz != 1)
+ return dev_err_probe(dev, -EINVAL, "PMU_ELEMENT_SIZE 0x%x not supported\n", sz);
+
+ /* The PMU pages *are* exclusively ours */
+ off = SZ_4K * FIELD_GET(MCUCFG_PMU_ELEMENT_START, cfg);
+ if (!devm_request_mem_region(dev, res->start + off, num * SZ_4K, dev_name(dev)))
+ return dev_err_probe(dev, -EADDRINUSE, "Unable to request PMU region\n");
+
+ bmu = devm_kzalloc(dev, struct_size(bmu, imus, num), GFP_KERNEL);
+ if (!bmu)
+ return -ENOMEM;
+
+ bmu->dev = dev;
+ bmu->base = base;
+ bmu->num_imus = num;
+ platform_set_drvdata(pdev, bmu);
+
+ base += off;
+ for (int i = 0; i < bmu->num_imus; i++, base += SZ_4K) {
+ /* At least PMCFGR.SIZE should always be nonzero if visible */
+ reg = readl_relaxed(base + PMCFGR);
+ if (!reg) {
+ dev_warn(dev, "Non-Secure access to PMU %d not enabled?\n", i);
+ num = 0;
+ } else {
+ num = 1 + FIELD_GET(PMCFGR_N, reg);
+ }
+ if (num > PMU_MAX_COUNTERS) {
+ dev_notice(dev, "PMU %d has %d counters, only using %d\n",
+ i, num, PMU_MAX_COUNTERS);
+ num = PMU_MAX_COUNTERS;
+ }
+ bmu->imus[i].base = base;
+ bmu->imus[i].num_counters = num;
+
+ writel_relaxed(PMCR_P | PMCR_E, base + BMU_PMCR);
+ writel_relaxed(U32_MAX, base + PMCNTENCLR);
+ writel_relaxed(U32_MAX, base + PMOVSCLR);
+ writel_relaxed(U32_MAX, base + PMINTENSET);
+ }
+
+ bmu->cpu = cpumask_local_spread(atomic_fetch_inc(&n), dev_to_node(dev));
+ bmu->irq = platform_get_irq(pdev, 0);
+ if (bmu->irq > 0) {
+ err = devm_request_irq(dev, bmu->irq, arm_bmu_handle_irq,
+ IRQF_NOBALANCING | IRQF_NO_THREAD,
+ dev_name(dev), bmu);
+ if (err)
+ bmu->irq = err;
+ else
+ irq_set_affinity(bmu->irq, cpumask_of(bmu->cpu));
+ }
+ if (bmu->irq < 0)
+ dev_info(dev, "Continuing without IRQ\n");
+
+ bmu->pmu = (struct pmu) {
+ .module = THIS_MODULE,
+ .parent = dev,
+ .attr_groups = arm_bmu_attr_groups,
+ .capabilities = PERF_PMU_CAP_NO_EXCLUDE,
+ .task_ctx_nr = perf_invalid_context,
+ .pmu_enable = arm_bmu_enable,
+ .pmu_disable = arm_bmu_disable,
+ .event_init = arm_bmu_event_init,
+ .add = arm_bmu_event_add,
+ .del = arm_bmu_event_del,
+ .start = arm_bmu_event_start,
+ .stop = arm_bmu_event_stop,
+ .read = arm_bmu_event_read,
+ };
+
+ name = acpi_device_uid(ACPI_COMPANION(dev));
+
+ if (name)
+ name = devm_kasprintf(dev, GFP_KERNEL, "arm_bmu_%s", name);
+ else
+ name = devm_kasprintf(dev, GFP_KERNEL, "arm_bmu_%llx", (u64)(res->start >> 12));
+ if (!name)
+ return -ENOMEM;
+
+ err = cpuhp_state_add_instance_nocalls(arm_bmu_cpuhp_state, &bmu->cpuhp_node);
+ if (err)
+ return err;
+
+ err = perf_pmu_register(&bmu->pmu, name, -1);
+ if (err)
+ cpuhp_state_remove_instance_nocalls(arm_bmu_cpuhp_state, &bmu->cpuhp_node);
+
+ return err;
+}
+
+static void arm_bmu_remove(struct platform_device *pdev)
+{
+ struct arm_bmu *bmu = platform_get_drvdata(pdev);
+
+ for (int i = 0; i < bmu->num_imus; i++)
+ writel_relaxed(U32_MAX, bmu->imus[i].base + PMINTENCLR);
+
+ perf_pmu_unregister(&bmu->pmu);
+ cpuhp_state_remove_instance_nocalls(arm_bmu_cpuhp_state, &bmu->cpuhp_node);
+}
+
+#ifdef CONFIG_ACPI
+static const struct acpi_device_id arm_bmu_acpi_match[] = {
+ { "ARMHB001" },
+ {}
+};
+MODULE_DEVICE_TABLE(acpi, arm_bmu_acpi_match);
+#endif
+
+static struct platform_driver arm_bmu_driver = {
+ .driver = {
+ .name = "arm-bmu",
+ .acpi_match_table = ACPI_PTR(arm_bmu_acpi_match),
+ .suppress_bind_attrs = true,
+ },
+ .probe = arm_bmu_probe,
+ .remove = arm_bmu_remove,
+};
+
+static void arm_bmu_migrate(struct arm_bmu *bmu, unsigned int cpu)
+{
+ perf_pmu_migrate_context(&bmu->pmu, bmu->cpu, cpu);
+ if (bmu->irq > 0)
+ irq_set_affinity(bmu->irq, cpumask_of(cpu));
+ bmu->cpu = cpu;
+}
+
+static int arm_bmu_online_cpu(unsigned int cpu, struct hlist_node *cpuhp_node)
+{
+ struct arm_bmu *bmu;
+ int node;
+
+ bmu = hlist_entry_safe(cpuhp_node, struct arm_bmu, cpuhp_node);
+ node = dev_to_node(bmu->dev);
+ if (cpu_to_node(bmu->cpu) != node && cpu_to_node(cpu) == node)
+ arm_bmu_migrate(bmu, cpu);
+ return 0;
+}
+
+static int arm_bmu_offline_cpu(unsigned int cpu, struct hlist_node *cpuhp_node)
+{
+ struct arm_bmu *bmu;
+ unsigned int target;
+ int node;
+
+ bmu = hlist_entry_safe(cpuhp_node, struct arm_bmu, cpuhp_node);
+ if (cpu != bmu->cpu)
+ return 0;
+
+ node = dev_to_node(bmu->dev);
+ target = cpumask_any_and_but(cpumask_of_node(node), cpu_online_mask, cpu);
+ if (target >= nr_cpu_ids)
+ target = cpumask_any_but(cpu_online_mask, cpu);
+
+ if (target < nr_cpu_ids)
+ arm_bmu_migrate(bmu, target);
+ return 0;
+}
+
+static int __init arm_bmu_init(void)
+{
+ int ret;
+
+ ret = cpuhp_setup_state_multi(CPUHP_AP_ONLINE_DYN, "perf/arm/bmu:online",
+ arm_bmu_online_cpu, arm_bmu_offline_cpu);
+ if (ret < 0)
+ return ret;
+
+ arm_bmu_cpuhp_state = ret;
+
+ ret = platform_driver_register(&arm_bmu_driver);
+ if (ret)
+ cpuhp_remove_multi_state(arm_bmu_cpuhp_state);
+ return ret;
+}
+
+static void __exit arm_bmu_exit(void)
+{
+ platform_driver_unregister(&arm_bmu_driver);
+ cpuhp_remove_multi_state(arm_bmu_cpuhp_state);
+}
+
+module_init(arm_bmu_init);
+module_exit(arm_bmu_exit);
+
+MODULE_DESCRIPTION("ARM Bus Monitor Unit Driver");
+MODULE_LICENSE("GPL v2");
--
2.54.0.dirty
^ permalink raw reply related [flat|nested] 4+ messages in thread* Re: [PATCH v3] perf: Add Arm Bus Monitor Unit driver
2026-09-17 13:47 [PATCH v3] perf: Add Arm Bus Monitor Unit driver Robin Murphy
@ 2026-09-17 14:02 ` sashiko-bot
2026-10-04 22:10 ` Will Deacon
1 sibling, 0 replies; 4+ messages in thread
From: sashiko-bot @ 2026-09-17 14:02 UTC (permalink / raw)
To: Robin Murphy; +Cc: linux-perf-users
Thank you for your contribution! Sashiko AI review found 6 potential issue(s) to consider:
- [High] arm-bmu: out-of-bounds MMIO access when calculating PMU base addresses
- [High] arm-bmu: race condition in CPU hotplug instance registration causes UAF/NULL deref
- [High] acpi: NULL pointer dereference when accessing ACPI_COMPANION(dev)
- [Medium] arm-bmu: hardware interrupts remain enabled on probe error paths and when missing IRQ
- [Medium] arm-bmu: hardware counter events are lost during overflow handling
- [Low] acpi: implicit declaration of acpi_device_uid() causes build failure when !CONFIG_ACPI
--
commit 7371a9609aa611a01ae74feaaad4d9c2b61e0cd7
Author: Robin Murphy <robin.murphy@arm.com>
perf: Add Arm Bus Monitor Unit driver
Arm's Bus Monitor Unit is a low-level performance analysis tool for matching
and counting transactions at interconnect interfaces. This commit adds a
driver for the BMU PMU architecture under arm_cspmu/.
> diff --git a/drivers/perf/arm_cspmu/arm-bmu.c b/drivers/perf/arm_cspmu/arm-bmu.c
> index 0000000000000..fe25954ac8361
> --- /dev/null
> +++ b/drivers/perf/arm_cspmu/arm-bmu.c
[ ... ]
> +static void arm_bmu_pmu_irq(struct arm_bmu *bmu, int imu)
> +{
> + struct arm_bmu_pmu *pmu = bmu->imus + imu;
> + u32 reg = readl_relaxed(pmu->base + PMOVSCLR);
> + u64 __iomem *pmevcnt = pmu->base + PMEVCNTR_LO;
> +
> + for (int i = 0; i < PMU_MAX_COUNTERS; i++) {
> + if (!(reg & (1U << i)))
> + continue;
> + if (!pmu->evcnt[i]) {
> + dev_dbg(bmu->dev, "Spurious oveflow on IMU %d counter %d?\n", imu, i);
> + continue;
> + }
> + arm_bmu_event_read(pmu->evcnt[i]);
> + local64_set(&pmu->evcnt[i]->hw.prev_count, S64_MIN);
> + lo_hi_writeq_relaxed(S64_MIN, pmevcnt + i);
[Severity: Medium]
Will hardware events occurring between arm_bmu_event_read() and this
lo_hi_writeq_relaxed() be lost since the actively running counter is
overwritten without being paused?
[ ... ]
> +static int arm_bmu_probe(struct platform_device *pdev)
> +{
> + struct device *dev = &pdev->dev;
> + const struct resource *res;
> + struct arm_bmu *bmu;
> + const char *name = NULL;
> + void __iomem *base;
> + static atomic_t n;
> + int err, num, sz, off;
> + u64 cfg;
> + u32 reg;
> +
> + res = platform_get_resource(pdev, IORESOURCE_MEM, 0);
> + if (!res)
> + return -EINVAL;
> +
> + /* PMUs and MPAM monitors are intermingled so we can't claim the whole resource */
> + base = devm_ioremap(dev, res->start, resource_size(res));
> + if (!base)
> + return -ENOMEM;
[ ... ]
> + cfg = lo_hi_readq_relaxed(base + MCU_CONFIG);
> + num = 1 + FIELD_GET(MCUCFG_NUM_IMU_MONITORS, cfg);
[ ... ]
> + /* The PMU pages *are* exclusively ours */
> + off = SZ_4K * FIELD_GET(MCUCFG_PMU_ELEMENT_START, cfg);
> + if (!devm_request_mem_region(dev, res->start + off, num * SZ_4K, dev_name(dev)))
> + return dev_err_probe(dev, -EADDRINUSE, "Unable to request PMU region\n");
[ ... ]
> + base += off;
> + for (int i = 0; i < bmu->num_imus; i++, base += SZ_4K) {
> + /* At least PMCFGR.SIZE should always be nonzero if visible */
> + reg = readl_relaxed(base + PMCFGR);
[Severity: High]
Could this result in an out-of-bounds MMIO read/write if the hardware reports
off and num values such that off + num * SZ_4K exceeds the mapped
resource_size(res)? It doesn't appear the calculated offset and size are
validated against the ioremap region size before being dereferenced.
[ ... ]
> + writel_relaxed(U32_MAX, base + PMCNTENCLR);
> + writel_relaxed(U32_MAX, base + PMOVSCLR);
> + writel_relaxed(U32_MAX, base + PMINTENSET);
> + }
> +
> + bmu->cpu = cpumask_local_spread(atomic_fetch_inc(&n), dev_to_node(dev));
> + bmu->irq = platform_get_irq(pdev, 0);
> + if (bmu->irq > 0) {
> + err = devm_request_irq(dev, bmu->irq, arm_bmu_handle_irq,
> + IRQF_NOBALANCING | IRQF_NO_THREAD,
> + dev_name(dev), bmu);
> + if (err)
> + bmu->irq = err;
> + else
> + irq_set_affinity(bmu->irq, cpumask_of(bmu->cpu));
> + }
> + if (bmu->irq < 0)
> + dev_info(dev, "Continuing without IRQ\n");
[Severity: Medium]
Do hardware interrupts remain enabled if the driver proceeds here without an
IRQ handler? It looks like PMINTENSET was written earlier in the loop, but
the flags aren't cleared via PMINTENCLR when an IRQ isn't provided.
[ ... ]
> + bmu->pmu = (struct pmu) {
> + .module = THIS_MODULE,
> + .parent = dev,
[ ... ]
> + };
> +
> + name = acpi_device_uid(ACPI_COMPANION(dev));
[Severity: High]
Can this cause a NULL pointer dereference if the driver binds to a device
without an ACPI companion? The pointer from ACPI_COMPANION(dev) appears to
be fed directly into acpi_device_uid() without being checked first.
[Severity: Low]
Additionally, since Kconfig allows building with CONFIG_COMPILE_TEST=y and
CONFIG_ACPI=n:
drivers/perf/arm_cspmu/Kconfig:ARM_BMU
depends on (ARM64 && ACPI) || COMPILE_TEST
Does this unconditional call to acpi_device_uid() cause a build failure due
to an implicit declaration when !CONFIG_ACPI?
> +
> + if (name)
> + name = devm_kasprintf(dev, GFP_KERNEL, "arm_bmu_%s", name);
> + else
> + name = devm_kasprintf(dev, GFP_KERNEL, "arm_bmu_%llx", (u64)(res->start >> 12));
> + if (!name)
> + return -ENOMEM;
> +
> + err = cpuhp_state_add_instance_nocalls(arm_bmu_cpuhp_state, &bmu->cpuhp_node);
> + if (err)
> + return err;
[Severity: Medium]
If this returns err, are hardware interrupts left enabled because PMINTENCLR
is not called on the error path?
> +
> + err = perf_pmu_register(&bmu->pmu, name, -1);
> + if (err)
> + cpuhp_state_remove_instance_nocalls(arm_bmu_cpuhp_state, &bmu->cpuhp_node);
> +
> + return err;
> +}
[Severity: High]
Can adding the hotplug instance before perf_pmu_register() completes cause a
race condition where a concurrent hotplug event triggers arm_bmu_offline_cpu()
and dereferences an uninitialized context?
> +
> +static void arm_bmu_remove(struct platform_device *pdev)
> +{
> + struct arm_bmu *bmu = platform_get_drvdata(pdev);
> +
> + for (int i = 0; i < bmu->num_imus; i++)
> + writel_relaxed(U32_MAX, bmu->imus[i].base + PMINTENCLR);
> +
> + perf_pmu_unregister(&bmu->pmu);
> + cpuhp_state_remove_instance_nocalls(arm_bmu_cpuhp_state, &bmu->cpuhp_node);
[Severity: High]
Does removing the instance after perf_pmu_unregister() leave a window open for
a use-after-free during teardown?
> +}
--
Sashiko AI review · https://sashiko.dev/#/patchset/c936505c43275f434bce924b7657e42bfb25579c.1789652832.git.robin.murphy@arm.com?part=1
^ permalink raw reply [flat|nested] 4+ messages in thread* Re: [PATCH v3] perf: Add Arm Bus Monitor Unit driver
2026-09-17 13:47 [PATCH v3] perf: Add Arm Bus Monitor Unit driver Robin Murphy
2026-09-17 14:02 ` sashiko-bot
@ 2026-10-04 22:10 ` Will Deacon
2026-10-05 13:15 ` Robin Murphy
1 sibling, 1 reply; 4+ messages in thread
From: Will Deacon @ 2026-10-04 22:10 UTC (permalink / raw)
To: Robin Murphy; +Cc: mark.rutland, linux-perf-users, linux-arm-kernel
On Thu, Sep 17, 2026 at 02:47:12PM +0100, Robin Murphy wrote:
> Arm's Bus Monitor Unit is a low-level performance analysis tool for
> matching and counting transactions at interconnect interfaces. Each
> BMU consists of a number of "interface monitoring units", with some
> global controls to synchronise them and so permit accurate calculation
> of cross-interface metrics. Each IMU contains, among other things, its
> own PMU largely based on the CoreSight PMU architecture, but with some
> functional incompatibilities due to the nature of how much per-event
> filtering needs to be crammed in to categorise AXI/CHI transactions at
> the protocol level. As such it gets an honorary home in arm_cspmu/ in
> order to share the common register definitions, but for simplicity does
> not try to bend the functional abstraction to fit.
>
> Signed-off-by: Robin Murphy <robin.murphy@arm.com>
>
> ---
> v3:
> - Couple of inconsequential tweaks to shut Sashiko up about things
> that would never actually happen anyway
> - Drop unneeded DT patches that I can't be bothered to fix right now
Please can you take a look at the Sashiko comments on this one?
Will
^ permalink raw reply [flat|nested] 4+ messages in thread
* Re: [PATCH v3] perf: Add Arm Bus Monitor Unit driver
2026-10-04 22:10 ` Will Deacon
@ 2026-10-05 13:15 ` Robin Murphy
0 siblings, 0 replies; 4+ messages in thread
From: Robin Murphy @ 2026-10-05 13:15 UTC (permalink / raw)
To: Will Deacon; +Cc: mark.rutland, linux-perf-users, linux-arm-kernel
On 04/10/2026 11:10 pm, Will Deacon wrote:
> On Thu, Sep 17, 2026 at 02:47:12PM +0100, Robin Murphy wrote:
>> Arm's Bus Monitor Unit is a low-level performance analysis tool for
>> matching and counting transactions at interconnect interfaces. Each
>> BMU consists of a number of "interface monitoring units", with some
>> global controls to synchronise them and so permit accurate calculation
>> of cross-interface metrics. Each IMU contains, among other things, its
>> own PMU largely based on the CoreSight PMU architecture, but with some
>> functional incompatibilities due to the nature of how much per-event
>> filtering needs to be crammed in to categorise AXI/CHI transactions at
>> the protocol level. As such it gets an honorary home in arm_cspmu/ in
>> order to share the common register definitions, but for simplicity does
>> not try to bend the functional abstraction to fit.
>>
>> Signed-off-by: Robin Murphy <robin.murphy@arm.com>
>>
>> ---
>> v3:
>> - Couple of inconsequential tweaks to shut Sashiko up about things
>> that would never actually happen anyway
>> - Drop unneeded DT patches that I can't be bothered to fix right now
>
> Please can you take a look at the Sashiko comments on this one?
I did start writing a reply at the time, but it was getting so sarcastic
that even I thought better of sending it...
In (polite) summary:
1 - Gibberish
2 - Yes if firmware blatantly mis-describes the hardware then Linux may
not work properly. Whoop de do.
3 - Yeah so?
4 - Pants-on-head gibberish, but the second part about COMPILE_TEST is
fair (an oversight from ripping out the DT support) - fix is already in
-next via the ACPI tree[1]
5 - Yeah so?
6,7 - Yes, welcome to perf...
Cheers,
Robin.
[1]
https://lore.kernel.org/all/CAJZ5v0g4k6FhyQMypkRyVJMkSBUsmeuetJuxRA826PcA7Mo5Hg@mail.gmail.com/
^ permalink raw reply [flat|nested] 4+ messages in thread
end of thread, other threads:[~2026-10-05 13:15 UTC | newest]
Thread overview: 4+ messages (download: mbox.gz follow: Atom feed
-- links below jump to the message on this page --
2026-09-17 13:47 [PATCH v3] perf: Add Arm Bus Monitor Unit driver Robin Murphy
2026-09-17 14:02 ` sashiko-bot
2026-10-04 22:10 ` Will Deacon
2026-10-05 13:15 ` Robin Murphy
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox