From mboxrd@z Thu Jan 1 00:00:00 1970 Return-Path: X-Spam-Checker-Version: SpamAssassin 3.4.0 (2014-02-07) on aws-us-west-2-korg-lkml-1.web.codeaurora.org Received: from bombadil.infradead.org (bombadil.infradead.org [198.137.202.133]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.lore.kernel.org (Postfix) with ESMTPS id 37967CA0ED3 for ; Mon, 2 Sep 2024 14:48:30 +0000 (UTC) DKIM-Signature: v=1; a=rsa-sha256; q=dns/txt; c=relaxed/relaxed; d=lists.infradead.org; s=bombadil.20210309; h=Sender:List-Subscribe:List-Help :List-Post:List-Archive:List-Unsubscribe:List-Id:In-Reply-To:Content-Type: MIME-Version:References:Message-ID:Subject:Cc:To:From:Date:Reply-To: Content-Transfer-Encoding:Content-ID:Content-Description:Resent-Date: Resent-From:Resent-Sender:Resent-To:Resent-Cc:Resent-Message-ID:List-Owner; bh=EG4K4tZ64shUUYntSjFiS9jCJSFIZ7Fiv4AIrxOFd7k=; b=X3dozrXV7qo53jREiWeuQftDqX xN8W/3Hto6Wz+sdwFT1PQsscXB8Nbo6TvvMhCYBb7JhKfNDwKT5dPr/ILNQfQEFvwVU40oOjrnRAP M0+pTPkbp2my3mckYGmCUIJ2Qsbe06MuTfDT0uF3z8mK6x2fyYsZDyYXE8GWaXj2lPx4bSiT1+ufS AIg7vOQsXtP1BqPE5ZboegNyXltYbWNli6wenuH4bYHGT+SK12Amla/ZhwNiN7DVc1soLcICZO+Fj UVVMit8A3EvKUkms4mVvVe1yNBx10AxVzifelF5ZfLl9DgMT9jbGiTvWqA3KmWazf9HhzGxVcn/Vq 2+nF/Ecw==; Received: from localhost ([::1] helo=bombadil.infradead.org) by bombadil.infradead.org with esmtp (Exim 4.97.1 #2 (Red Hat Linux)) id 1sl8Lq-0000000EjAf-1TsG; Mon, 02 Sep 2024 14:48:18 +0000 Received: from nyc.source.kernel.org ([147.75.193.91]) by bombadil.infradead.org with esmtps (Exim 4.97.1 #2 (Red Hat Linux)) id 1sl8Kv-0000000Ej57-0lvk for linux-arm-kernel@lists.infradead.org; Mon, 02 Sep 2024 14:47:22 +0000 Received: from smtp.kernel.org (transwarp.subspace.kernel.org [100.75.92.58]) by nyc.source.kernel.org (Postfix) with ESMTP id CCE0FA416AC; Mon, 2 Sep 2024 14:47:12 +0000 (UTC) Received: by smtp.kernel.org (Postfix) with ESMTPSA id 85DBCC4CEC2; Mon, 2 Sep 2024 14:47:18 +0000 (UTC) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/simple; d=kernel.org; s=k20201202; t=1725288439; bh=X45jWrJCTaBMiYV5408CS0aYMQPVZIj6Q9txx4dbO+o=; h=Date:From:To:Cc:Subject:References:In-Reply-To:From; b=ZwjrXm89et9gPbeEmry7VYwgv1sa69m/R0JFIUjqkxST8BgocZYVOED9qEhmrzDd9 jMOEqRiSnIJpk0RnlGTBJ6W7AGaOD+xKAGVZ3Rm6f3X/vZyOQRglwm1eqHk7mQMhtu JKLj5gRFVYKUnmKthNWL34ySPrxnXZTPjq8QQ5ng6DcarXQLdwUKqYTteUf3rWXasp 976bgykl8UEn0cU6BDS8WWgqWtHtpR+s9E57DRVaBuHqbvMhLjKc15scwzFHg9J0pv e3Gl/RrmW4/mbAlUs+EpeKsHvg2A3H2LG7tY2vFCKAlOn+Lui6F1yPerYA6OBZ9vG/ /aeWPuVccwT7A== Date: Mon, 2 Sep 2024 15:47:15 +0100 From: Will Deacon To: Robin Murphy Cc: mark.rutland@arm.com, linux-arm-kernel@lists.infradead.org, linux-kernel@vger.kernel.org, jialong.yang@shingroup.cn Subject: Re: [PATCH v3 2/3] perf: Add driver for Arm NI-700 interconnect PMU Message-ID: <20240902144714.GA11443@willie-the-truck> References: <275e8ef450eeaf837468ce34e2c6930d59091fbc.1725037424.git.robin.murphy@arm.com> MIME-Version: 1.0 Content-Type: text/plain; charset=us-ascii Content-Disposition: inline In-Reply-To: <275e8ef450eeaf837468ce34e2c6930d59091fbc.1725037424.git.robin.murphy@arm.com> User-Agent: Mutt/1.10.1 (2018-07-13) X-CRM114-Version: 20100106-BlameMichelson ( TRE 0.8.0 (BSD) ) MR-646709E3 X-CRM114-CacheID: sfid-20240902_074721_379629_4FCFC5C5 X-CRM114-Status: GOOD ( 39.30 ) X-BeenThere: linux-arm-kernel@lists.infradead.org X-Mailman-Version: 2.1.34 Precedence: list List-Id: List-Unsubscribe: , List-Archive: List-Post: List-Help: List-Subscribe: , Sender: "linux-arm-kernel" Errors-To: linux-arm-kernel-bounces+linux-arm-kernel=archiver.kernel.org@lists.infradead.org Hi Robin, On Fri, Aug 30, 2024 at 06:19:34PM +0100, Robin Murphy wrote: > The Arm NI-700 Network-on-Chip Interconnect has a relatively > straightforward design with a hierarchy of voltage, power, and clock > domains, where each clock domain then contains a number of interface > units and a PMU which can monitor events thereon. As such, it begets a > relatively straightforward driver to interface those PMUs with perf. > > Even more so than with arm-cmn, users will require detailed knowledge of > the wider system topology in order to meaningfully analyse anything, > since the interconnect itself cannot know what lies beyond the boundary > of each inscrutably-numbered interface. Given that, for now they are > also expected to refer to the NI-700 documentation for the relevant > event IDs to provide as well. An identifier is implemented so we can > come back and add jevents if anyone really wants to. > > Signed-off-by: Robin Murphy > > --- > v2: > - Add basic usage documentation > - Use __counted_by attribute > - Make group validation logic clearer (and drop PMU type check > which perf_event_open() already takes care of) > - Add retry limit to arm_ni_read_ccnt() > v3: > - Update .remove to return void > - Fix group leader validation and make the naming clearer > - Drop NUMA_NO_NODE check for CPU online (the only way that could > actually pass both other migration conditions is if the NUMA info > is so messed up that it's not worth worrying about anyway) Thanks, this is looking pretty good now. I just have a few random comments based on another read-through of the code. > diff --git a/Documentation/admin-guide/perf/arm-ni.rst b/Documentation/admin-guide/perf/arm-ni.rst > new file mode 100644 > index 000000000000..3cd7d0f75f0f > --- /dev/null > +++ b/Documentation/admin-guide/perf/arm-ni.rst > @@ -0,0 +1,17 @@ > +==================================== > +Arm Network-on Chip Interconnect PMU > +==================================== > + > +NI-700 and friends implement a distinct PMU for each clock domain within the > +interconnect. Correspondingly, the driver exposes multiple PMU devices named > +arm_ni__cd_, where is an (abritrary) instance identifier and is typo: abritrary > +the clock domain ID within that particular instance. If multiple NI instances > +exist within a system, the PMU devices can be correlated with the underlying > +hardware instance via sysfs parentage. > + > +Each PMU exposes base event aliases for the interface types present in its clock > +domain. These require qualifying with the "eventid" and "nodeid" parameters > +to specify the event code to count and the interface at which to count it > +(per the configured hardware ID as reflected in the xxNI_NODE_INFO register). > +The exception is the "cycles" alias for the PMU cycle counter, which is encoded > +with the PMU node type and needs no further qualification. [...] > +static ssize_t arm_ni_format_show(struct device *dev, > + struct device_attribute *attr, char *buf) > +{ > + struct arm_ni_format_attr *fmt = container_of(attr, typeof(*fmt), attr); > + int lo = __ffs(fmt->field), hi = __fls(fmt->field); > + > + return sysfs_emit(buf, "config:%d-%d\n", lo, hi); > +} Nit: if you end up adding single-bit config fields in the future, this will quietly do the wrong thing. Maybe safe-guard the 'lo==hi' case (even if you just warn once and return without doing anything). [...] > +static int arm_ni_init_cd(struct arm_ni *ni, struct arm_ni_node *node, u64 res_start) > +{ > + struct arm_ni_cd *cd = ni->cds + node->id; > + const char *name; > + int err; > + > + cd->id = node->id; > + cd->num_units = node->num_components; > + cd->units = devm_kcalloc(ni->dev, cd->num_units, sizeof(*(cd->units)), GFP_KERNEL); > + if (!cd->units) > + return -ENOMEM; > + > + for (int i = 0; i < cd->num_units; i++) { > + u32 reg = readl_relaxed(node->base + NI_CHILD_PTR(i)); > + void __iomem *unit_base = ni->base + reg; > + struct arm_ni_unit *unit = cd->units + i; > + > + reg = readl_relaxed(unit_base + NI_NODE_TYPE); > + unit->type = FIELD_GET(NI_NODE_TYPE_NODE_TYPE, reg); > + unit->id = FIELD_GET(NI_NODE_TYPE_NODE_ID, reg); > + > + switch (unit->type) { > + case NI_PMU: > + reg = readl_relaxed(unit_base + NI_PMCFGR); > + if (!reg) { > + dev_info(ni->dev, "No access to PMU %d\n", cd->id); > + devm_kfree(ni->dev, cd->units); > + return 0; > + } > + unit->ns = true; > + cd->pmu_base = unit_base; > + break; > + case NI_ASNI: > + case NI_AMNI: > + case NI_HSNI: > + case NI_HMNI: > + case NI_PMNI: > + unit->pmusela = unit_base + NI700_PMUSELA; > + writel_relaxed(1, unit->pmusela); > + if (readl_relaxed(unit->pmusela) != 1) > + dev_info(ni->dev, "No access to node 0x%04x%04x\n", unit->id, unit->type); > + else > + unit->ns = true; > + break; > + default: > + /* > + * e.g. FMU - thankfully bits 3:2 of FMU_ERR_FR0 are RES0 so > + * can't alias any of the leaf node types we're looking for. > + */ > + dev_dbg(ni->dev, "Mystery node 0x%04x%04x\n", unit->id, unit->type); > + break; > + } > + } > + > + res_start += cd->pmu_base - ni->base; > + if (!devm_request_mem_region(ni->dev, res_start, SZ_4K, dev_name(ni->dev))) { > + dev_err(ni->dev, "Failed to request PMU region 0x%llx\n", res_start); > + return -EBUSY; > + } > + > + writel_relaxed(NI_PMCR_RESET_CCNT | NI_PMCR_RESET_EVCNT, > + cd->pmu_base + NI_PMCR); > + writel_relaxed(U32_MAX, cd->pmu_base + NI_PMCNTENCLR); > + writel_relaxed(U32_MAX, cd->pmu_base + NI_PMOVSCLR); > + writel_relaxed(U32_MAX, cd->pmu_base + NI_PMINTENSET); > + > + cd->irq = platform_get_irq(to_platform_device(ni->dev), cd->id); > + if (cd->irq < 0) > + return cd->irq; > + > + err = devm_request_irq(ni->dev, cd->irq, arm_ni_handle_irq, > + IRQF_NOBALANCING | IRQF_NO_THREAD, > + dev_name(ni->dev), cd); > + if (err) > + return err; > + > + cd->cpu = cpumask_local_spread(0, dev_to_node(ni->dev)); > + cd->pmu = (struct pmu) { > + .module = THIS_MODULE, > + .parent = ni->dev, > + .attr_groups = arm_ni_attr_groups, > + .capabilities = PERF_PMU_CAP_NO_EXCLUDE, > + .task_ctx_nr = perf_invalid_context, > + .pmu_enable = arm_ni_pmu_enable, > + .pmu_disable = arm_ni_pmu_disable, > + .event_init = arm_ni_event_init, > + .add = arm_ni_event_add, > + .del = arm_ni_event_del, > + .start = arm_ni_event_start, > + .stop = arm_ni_event_stop, > + .read = arm_ni_event_read, > + }; > + > + name = devm_kasprintf(ni->dev, GFP_KERNEL, "arm_ni_%d_cd_%d", ni->id, cd->id); > + if (!name) > + return -ENOMEM; > + > + err = cpuhp_state_add_instance(arm_ni_hp_state, &cd->cpuhp_node); > + if (err) > + return err; What happens if there's a CPU hotplug operation here? Can we end up calling perf_pmu_migrate_context() concurrently with perf_pmu_register()? > + return perf_pmu_register(&cd->pmu, name, -1); Clean up the cpuhp instance if this fails? > +static void arm_ni_remove(struct platform_device *pdev) > +{ > + struct arm_ni *ni = platform_get_drvdata(pdev); > + > + for (int i = 0; i < ni->num_cds; i++) { > + struct arm_ni_cd *cd = ni->cds + i; > + > + if (!cd->pmu_base) > + continue; > + > + writel_relaxed(0, cd->pmu_base + NI_PMCR); > + writel_relaxed(U32_MAX, cd->pmu_base + NI_PMINTENCLR); > + perf_pmu_unregister(&cd->pmu); Similarly here, it feels like a CPU hotplug operation could cause problems here. > + cpuhp_state_remove_instance_nocalls(arm_ni_hp_state, &cd->cpuhp_node); > + } > +} Will