From mboxrd@z Thu Jan 1 00:00:00 1970 Return-Path: Received: (majordomo@vger.kernel.org) by vger.kernel.org via listexpand id S932301AbcCCIQn (ORCPT ); Thu, 3 Mar 2016 03:16:43 -0500 Received: from mail-bn1on0099.outbound.protection.outlook.com ([157.56.110.99]:27776 "EHLO na01-bn1-obe.outbound.protection.outlook.com" rhost-flags-OK-OK-OK-FAIL) by vger.kernel.org with ESMTP id S1756779AbcCCIEN (ORCPT ); Thu, 3 Mar 2016 03:04:13 -0500 Authentication-Results: spf=none (sender IP is 165.204.84.221) smtp.mailfrom=amd.com; alien8.de; dkim=none (message not signed) header.d=none;alien8.de; dmarc=permerror action=none header.from=amd.com; X-WSS-ID: 0O3GEEV-07-NVP-02 X-M-MSG: From: Huang Rui To: Borislav Petkov , Thomas Gleixner , "Peter Zijlstra" , Ingo Molnar , "Andy Lutomirski" , Robert Richter , "Jacob Shin" , Arnaldo Carvalho de Melo , Kan Liang CC: , , , Suravee Suthikulpanit , Aravind Gopalakrishnan , Borislav Petkov , Fengguang Wu , Huang Rui , Guenter Roeck Subject: [PATCH v6 2/2] perf/x86/amd/power: Add AMD accumulated power reporting mechanism Date: Thu, 3 Mar 2016 16:04:44 +0800 Message-ID: <1456992284-4808-3-git-send-email-ray.huang@amd.com> X-Mailer: git-send-email 1.9.1 In-Reply-To: <1456992284-4808-1-git-send-email-ray.huang@amd.com> References: <1456992284-4808-1-git-send-email-ray.huang@amd.com> MIME-Version: 1.0 Content-Type: text/plain X-EOPAttributedMessage: 0 X-Forefront-Antispam-Report: CIP:165.204.84.221;CTRY:US;IPV:NLI;EFV:NLI;SFV:NSPM;SFS:(10009020)(6009001)(2980300002)(428002)(189002)(199003)(164054003)(575784001)(92566002)(48376002)(189998001)(551984002)(86362001)(33646002)(101416001)(5001770100001)(36756003)(53416004)(50466002)(87936001)(5003940100001)(50226001)(19580405001)(19580395003)(2950100001)(5008740100001)(77096005)(11100500001)(47776003)(229853001)(76176999)(586003)(2906002)(4326007)(105586002)(5003600100002)(1096002)(1220700001)(106466001)(50986999)(5001960100004)(2004002);DIR:OUT;SFP:1101;SCL:1;SRVR:DM3PR12MB0858;H:atltwp01.amd.com;FPR:;SPF:None;MLV:sfv;MX:1;A:1;LANG:en; X-MS-Office365-Filtering-Correlation-Id: 61873b51-ee72-4797-0dd9-08d3433a66b8 X-Microsoft-Exchange-Diagnostics: 1;DM3PR12MB0858;2:Xzme7OZGu8se77Q1+/HRspehsKRW/gRYu/Nyr+8RSLw2l0qfy45le4xc/v0mXMmYuDfD6Zc1sp0ZL3PW/AqraxO79lhCYyNhqmS5QieppssCO+D3o7LNiclVHh+p59Nb3/sHBPMqgv5teOPECJ23cjaoo8c0jpcRiKw/lYj57XcyzD2Cu1QoJgiL7LVU8SZV;3:IQB/2NATigym0L5TBzsHs7eRMzKA5E1lacOEwYcZ3g5Oe4YGDOa1kUAxXdz8svb3C1vy2G76J+zB+67Lve/wddFjg5hDn8kDFwuao2N/T0Sh7Vciaa4V3IniBS0fuyfj2gFQUxlU+CDxnTQug5eL0+IAeQtFwIzadRPOZfVNfdsHLP/CAS5bXygdm078kmBgAZ/uTeCIQPr+t8BZrRD6p9yaEtJz0hRgW/Fd8BiQ8Ec=;25:cyJml/vUqelksq5085OMnVQXv/To3rWKNStc6x0SFvbgwtypP4/QkdyXyTH4f3gyiOMTs1eRDx7E/kmN+bHHFyFNHWNCupYA3KlVBcBkHcMzzv8rAS/RdBofEUQ2bQnySZtq7TFX0jGn4jfsQVK+LO7yIZe20sRyvZtU0GFEkrQWyGvd3Wp6/jYzoPAFai5dp2WBJ+09VQfRcIgeWZUo2cb3yyOkpuod3NOGOwzn4Q8N8FJtv3Rp4CaAEdOzDfF1amPk/DzfofsVduY1Kni1g8u/qeO+t/N24AHPZlmo1MqpsmHlVMwwHYXcoFrf2tiwJf4VN8RQaGQ8zZYAUDdkFw== X-Microsoft-Antispam: UriScan:;BCL:0;PCL:0;RULEID:;SRVR:DM3PR12MB0858; X-Microsoft-Exchange-Diagnostics: 1;DM3PR12MB0858;20:IUFun/eBk5wp8pKMso6fc7Lb0bFd6ct0cXHcotrdMcf2AGinY6xNqGVs8EzoAFYwR6JWtJ+tWjMy1QQng4LFlq31TkZDySfjpXbQTiw4aGiFaOuZ8EpkHsUnm1ZQ5ILx8whWluy8psxJiCtpkBetJhtOL+xSlCXGdwzUniFbzP8X2QESfnfbaOyMAf7yez5puAg/3z3fqH77rcIK+Y3zzsk/SNO/tgkIZz6V6gVNfQXvN6XVD61x2cqkppN/eMpEoIbBa+rsYRmIZmwAwnRayTPqld9eku8Et8YczOHIkK0HD0TMVzXBIhFxmp+lu5TE4BHfQ3fuyKVvleuzXw9HM9xVFyQlN0HIHT2J0Kef/D1LijK05ZZn60bQz3v10B6mWXndSE/HrRiXyuij8ceedHwCEPfw3jtN2mSmeqJIarQgmYxYjQENdarR8/OJ8TwUrPej9sPh/4scvj7tFXRErHNnPDX93ti/mk4AYLbdtFmeVDJud8/ImO2tHb4ameLa X-Microsoft-Antispam-PRVS: X-Exchange-Antispam-Report-Test: UriScan:; X-Exchange-Antispam-Report-CFA-Test: BCL:0;PCL:0;RULEID:(601004)(2401047)(5005006)(13018025)(13017025)(13015025)(13024025)(13023025)(8121501046)(3002001)(10201501046);SRVR:DM3PR12MB0858;BCL:0;PCL:0;RULEID:;SRVR:DM3PR12MB0858; X-Microsoft-Exchange-Diagnostics: 1;DM3PR12MB0858;4:TN5AyfcN+r+SBxc7v8X/6l6xPuQ71b5pnMR35Z4Gvvev2bgFBbrgaNoPdFWRrcWC83Hmy+RjK+kDmtY1fHJSAS6CT9Xa6lRBBa/2HpjKGl7e/96usK9eDOkYIIa1z9PmOSHS3dCOp+UaiuGZxEKyrkdJ2gHUNIyqF5BCDj/l+k1zi+PSSf+z97BbF9hORrISN2VdASS2D9UwiCYplTZmyfvvyyB2pzgrNVNZacDsNb1n9F8fhFgW8EvFm/7PZ+nbIUhjhOccIiGmBZ0X4d97BPBFEon58ea1FXR9+6YljbYDBiL5uErSiaIvWbnuo/sci4MajRNsSWOYwtaxyQEA444cfj0jzGzfikpl6cLFqedGERHhSMlwYjPTKVhhS9OXIC2TRQ2gaoVvyPY6+P++/swLs2cfnG7Pd6TE70REQVKSEXfMApHkfTrYHshgVPn6hNJNsw10Ui5KkXxgm84j2Q== X-Forefront-PRVS: 0870212862 X-Microsoft-Exchange-Diagnostics: =?us-ascii?Q?1;DM3PR12MB0858;23:N8EGrvLAP6gnNYChLACJZyIQWQ/iLzsfkoFckCC7N?= =?us-ascii?Q?fcxvrlyz3r3lmgkYwaVLOleuZAqns5I/lUCH3C5crVGIzF4Ljixd72Rcol8E?= =?us-ascii?Q?xiPUEqSs5DEYMVmH6lJrn+XCwUExpPq2DZklm7H2Bwn4LTbbiipV8gt2A3ke?= =?us-ascii?Q?2c1+HBDxLpKUU+9fO/jxpabdF//XKznToGU7bkOQIxb/Iec6ruXopx3mbcWW?= =?us-ascii?Q?1lyWyNfUcof7c1Rwlo/lQDsRUQGdKTS+MLLVANyW+9PgR1Bi91KZC5eUMTY8?= =?us-ascii?Q?KTHpb09Rx+pzHL4BXtG0KJeQLDMcYjOD1D4VEIDZfE85njfoAESHDfrebtWi?= =?us-ascii?Q?8Ki07oipG9ivuzLd7HgZkc1LyKCgsFdyRlVEpdlBbuUIMS+WADFcn8mpQed6?= =?us-ascii?Q?R+cQeVICSrmsEB5OOwEs7BHZWpLjcvyisvmz/hMQJTVq5Asd9zJuG+CIxoq8?= =?us-ascii?Q?COAvag1Vks6bJNQ/NVtqGK0clu23cqNUa1R/TXV1aeslDZVwwFLsE6+hjJw3?= =?us-ascii?Q?eT72Vo2odq5OmhHua88Bh866cHk6Fg4a7bSUxBvOwlZcZw+9JXmGjpAwGKCu?= =?us-ascii?Q?4TPr8gf9yjy+v5bzsjopLWC87raKF7snPwCux4EmWvhZytY9WTPP3g76Vq1M?= =?us-ascii?Q?f0meNE+LMXqux1qCzDeJqIh51p049y0dVQy+BFG7t7Y6qI0glBFwDN1Kr+1v?= =?us-ascii?Q?9FkVml6ANI0gGbQrhThTwMt55jaoD8HhuUeytiivMF14xR55AA3+OKyslXGl?= =?us-ascii?Q?jO+RS8Cd6Wv06oCgTiVm5tsXX9LppAES7/Y0+DnQzPDUyNYIFO3fGNL1Gu9i?= =?us-ascii?Q?CmyOUJqrdy7LiC9XcxjbHrSeZUC1Fcfy9a1Hz3LsOPC4AXlLwreAwtd9SmFG?= =?us-ascii?Q?3ZCa8ufmVYxNe7ehh56rGwMS8qqNGR7SbDHzVWKLvSaaA0e2/6Fl409FBmGS?= =?us-ascii?Q?7mdOPNdoyQ67S1HCeLY7p4Cga2+Uh6AFad3ztU5NXw8qK2K66Nw+uOVmonN3?= =?us-ascii?Q?mP9wX+Z2dfNIMGlVXGAsJrZPMj2MWyZzFfKAC+jqBlSIxV9QNpi4vS7w3KYa?= =?us-ascii?Q?WANDo/bqKoKAPnxw01/ryjG/UJhDB7ll/BdNnlN5bZe26TaOg=3D=3D?= X-Microsoft-Exchange-Diagnostics: 1;DM3PR12MB0858;5:/rzJP2fvoXAxUtRGVDC8vx3wDRDYH+dbB63nB0D1+tWuIUCqgyhxFajcxPP+f9t+ljIG8wcTigMIuH8OXi47mxcNzEEyuwigO2wAQvXKJDsdv9lx0U6ob0W3clkCZ9i2iA6hORhhEdMJ/E7CKZVleA==;24:Knc7WA4WTyLpdV7zvTYwPyZndvPScdAKOC3uxSB8nHhuPKehDHXSu7d3whSPrG95zqjv35r5BfQjDkzcAYD7J3DSKWYii8qXepkLcsj3WbU=;20:fhO+JRODJrg9k1MwDwVvJmn5po1q5+0P7FLcW8Gc5AxFmOghKHxaBXwV/T4b6FfDLCQfGIBFnYaTgRWRwHHuviZZUhZMaZ2DpnWTz5yYwDMzgGvrbTc03+ZBFVdRctonAtqz+saLyVIJmRZM3ZQb2kp6hJh1wA7UubORIFBiP8BbuBX/eS+7j8uTcK0ndgAyO1q+zdUDUvdgFl8Rgy8TCrOzTrPsKgFp/x+hpQUZFGHVSbNNgBAs3BDA5Ob3wp/X SpamDiagnosticOutput: 1:23 SpamDiagnosticMetadata: NSPM X-OriginatorOrg: amd.com X-MS-Exchange-CrossTenant-OriginalArrivalTime: 03 Mar 2016 08:04:08.8917 (UTC) X-MS-Exchange-CrossTenant-Id: 3dd8961f-e488-4e60-8e11-a82d994e183d X-MS-Exchange-CrossTenant-OriginalAttributedTenantConnectingIp: TenantId=3dd8961f-e488-4e60-8e11-a82d994e183d;Ip=[165.204.84.221];Helo=[atltwp01.amd.com] X-MS-Exchange-CrossTenant-FromEntityHeader: HybridOnPrem X-MS-Exchange-Transport-CrossTenantHeadersStamped: DM3PR12MB0858 Sender: linux-kernel-owner@vger.kernel.org List-ID: X-Mailing-List: linux-kernel@vger.kernel.org Introduce an AMD accumlated power reporting mechanism for the Family 15h, Model 60h processor that can be used to calculate the average power consumed by a processor during a measurement interval. The feature support is indicated by CPUID Fn8000_0007_EDX[12]. This feature will be implemented both in hwmon and perf. The current design provides one event to report per package/processor power consumption by counting each compute unit power value. Here the gory details of how the computation is done: --------------------------------------------------------------------- * Tsample: compute unit power accumulator sample period * Tref: the PTSC counter period (PTSC: performance timestamp counter) * N: the ratio of compute unit power accumulator sample period to the PTSC period * Jmax: max compute unit accumulated power which is indicated by MSR_C001007b[MaxCpuSwPwrAcc] * Jx/Jy: compute unit accumulated power which is indicated by MSR_C001007a[CpuSwPwrAcc] * Tx/Ty: the value of performance timestamp counter which is indicated by CU_PTSC MSR_C0010280[PTSC] * PwrCPUave: CPU average power i. Determine the ratio of Tsample to Tref by executing CPUID Fn8000_0007. N = value of CPUID Fn8000_0007_ECX[CpuPwrSampleTimeRatio[15:0]]. ii. Read the full range of the cumulative energy value from the new MSR MaxCpuSwPwrAcc. Jmax = value returned. iii. At time x, software reads CpuSwPwrAcc and samples the PTSC. Jx = value read from CpuSwPwrAcc and Tx = value read from PTSC. iv. At time y, software reads CpuSwPwrAcc and samples the PTSC. Jy = value read from CpuSwPwrAcc and Ty = value read from PTSC. v. Calculate the average power consumption for a compute unit over time period (y-x). Unit of result is uWatt: if (Jy < Jx) // Rollover has occurred Jdelta = (Jy + Jmax) - Jx else Jdelta = Jy - Jx PwrCPUave = N * Jdelta * 1000 / (Ty - Tx) ---------------------------------------------------------------------- Simple example: root@hr-zp:/home/ray/tip# ./tools/perf/perf stat -a -e 'power/power-pkg/' make -j4 CHK include/config/kernel.release CHK include/generated/uapi/linux/version.h CHK include/generated/utsrelease.h CHK include/generated/timeconst.h CHK include/generated/bounds.h CHK include/generated/asm-offsets.h CALL scripts/checksyscalls.sh CHK include/generated/compile.h SKIPPED include/generated/compile.h Building modules, stage 2. Kernel: arch/x86/boot/bzImage is ready (#40) MODPOST 4225 modules Performance counter stats for 'system wide': 183.44 mWatts power/power-pkg/ 341.837270111 seconds time elapsed root@hr-zp:/home/ray/tip# ./tools/perf/perf stat -a -e 'power/power-pkg/' sleep 10 Performance counter stats for 'system wide': 0.18 mWatts power/power-pkg/ 10.012551815 seconds time elapsed Suggested-by: Peter Zijlstra Suggested-by: Ingo Molnar Suggested-by: Borislav Petkov Signed-off-by: Huang Rui Cc: Guenter Roeck --- Hi Thomas, Thanks to suggest the updates for power_cpu_init(), but there are some minor issues with below codes. target = cpumask_any_and(&cpu_mask, topology_sibling_cpumask(cpu)); if (target < nr_cpu_ids) cpumask_set_cpu(cpu, &cpu_mask); For example: cores number is 4, and compute unit number is 2 (2 cores per compute units). When amd_power_pmu completes initialization, the cpumask will be set as 1111. Actually, I just would like to choose one core per compute unit. So the expected value should be 0101. when core 2 and core 3 are both online (they are in same compute unit) and cpu_mask is 1111, set core 2 offline, then 1111 -> 1011, set core 2 online again, then 1011 -> 1111. Actually, we don't expect two cores both set at cpu_mask in the same compute unit. So as inspired by you, I do below change for power_cpu_init and tested in my platform, please see my codes. :-) Thanks, Rui --- arch/x86/Kconfig | 9 + arch/x86/kernel/cpu/Makefile | 1 + arch/x86/kernel/cpu/perf_event_amd_power.c | 360 +++++++++++++++++++++++++++++ include/linux/perf_event.h | 4 + 4 files changed, 374 insertions(+) create mode 100644 arch/x86/kernel/cpu/perf_event_amd_power.c diff --git a/arch/x86/Kconfig b/arch/x86/Kconfig index 330e738..a0d3fb0 100644 --- a/arch/x86/Kconfig +++ b/arch/x86/Kconfig @@ -1202,6 +1202,15 @@ config MICROCODE_OLD_INTERFACE def_bool y depends on MICROCODE +config PERF_EVENTS_AMD_POWER + depends on PERF_EVENTS && CPU_SUP_AMD + tristate "AMD Processor Power Reporting Mechanism" + ---help--- + Provide power reporting mechanism support for AMD processors. + Currently, it leverages X86_FEATURE_ACC_POWER + (CPUID Fn8000_0007_EDX[12]) interface to calculate the + average power consumption on Family 15h processors. + config X86_MSR tristate "/dev/cpu/*/msr - Model-specific register support" ---help--- diff --git a/arch/x86/kernel/cpu/Makefile b/arch/x86/kernel/cpu/Makefile index faa7b52..95d8419 100644 --- a/arch/x86/kernel/cpu/Makefile +++ b/arch/x86/kernel/cpu/Makefile @@ -34,6 +34,7 @@ obj-$(CONFIG_PERF_EVENTS) += perf_event.o ifdef CONFIG_PERF_EVENTS obj-$(CONFIG_CPU_SUP_AMD) += perf_event_amd.o perf_event_amd_uncore.o +obj-$(CONFIG_PERF_EVENTS_AMD_POWER) += perf_event_amd_power.o ifdef CONFIG_AMD_IOMMU obj-$(CONFIG_CPU_SUP_AMD) += perf_event_amd_iommu.o endif diff --git a/arch/x86/kernel/cpu/perf_event_amd_power.c b/arch/x86/kernel/cpu/perf_event_amd_power.c new file mode 100644 index 0000000..42d5072 --- /dev/null +++ b/arch/x86/kernel/cpu/perf_event_amd_power.c @@ -0,0 +1,360 @@ +/* + * Performance events - AMD Processor Power Reporting Mechanism + * + * Copyright (C) 2016 Advanced Micro Devices, Inc. + * + * Author: Huang Rui + * + * This program is free software; you can redistribute it and/or modify + * it under the terms of the GNU General Public License version 2 as + * published by the Free Software Foundation. + */ + +#include +#include +#include +#include +#include "perf_event.h" + +#define MSR_F15H_CU_PWR_ACCUMULATOR 0xc001007a +#define MSR_F15H_CU_MAX_PWR_ACCUMULATOR 0xc001007b +#define MSR_F15H_PTSC 0xc0010280 + +/* Event code: LSB 8 bits, passed in attr->config any other bit is reserved. */ +#define AMD_POWER_EVENT_MASK 0xFFULL + +/* + * Accumulated power status counters. + */ +#define AMD_POWER_EVENTSEL_PKG 1 + +/* + * The ratio of compute unit power accumulator sample period to the + * PTSC period. + */ +static unsigned int cpu_pwr_sample_ratio; +static unsigned int cu_num; + +/* Maximum accumulated power of a compute unit. */ +static u64 max_cu_acc_power; + +static struct pmu pmu_class; + +/* + * Accumulated power represents the sum of each compute unit's (CU) power + * consumption. On any core of each CU we read the total accumulated power from + * MSR_F15H_CU_PWR_ACCUMULATOR. cpu_mask represents CPU bit map of all cores + * which are picked to measure the power for the CUs they belong to. + */ +static cpumask_t cpu_mask; + +static void event_update(struct perf_event *event) +{ + struct hw_perf_event *hwc = &event->hw; + u64 prev_pwr_acc, new_pwr_acc, prev_ptsc, new_ptsc; + u64 delta, tdelta; + + prev_pwr_acc = hwc->pwr_acc; + prev_ptsc = hwc->ptsc; + rdmsrl(MSR_F15H_CU_PWR_ACCUMULATOR, new_pwr_acc); + rdmsrl(MSR_F15H_PTSC, new_ptsc); + + /* + * Calculate the CU power consumption over a time period, the unit of + * final value (delta) is micro-Watts. Then add it to the event count. + */ + if (new_pwr_acc < prev_pwr_acc) { + delta = max_cu_acc_power + new_pwr_acc; + delta -= prev_pwr_acc; + } else + delta = new_pwr_acc - prev_pwr_acc; + + delta *= cpu_pwr_sample_ratio * 1000; + tdelta = new_ptsc - prev_ptsc; + + do_div(delta, tdelta); + local64_add(delta, &event->count); +} + +static void __pmu_event_start(struct perf_event *event) +{ + if (WARN_ON_ONCE(!(event->hw.state & PERF_HES_STOPPED))) + return; + + event->hw.state = 0; + + rdmsrl(MSR_F15H_PTSC, event->hw.ptsc); + rdmsrl(MSR_F15H_CU_PWR_ACCUMULATOR, event->hw.pwr_acc); +} + +static void pmu_event_start(struct perf_event *event, int mode) +{ + __pmu_event_start(event); +} + +static void pmu_event_stop(struct perf_event *event, int mode) +{ + struct hw_perf_event *hwc = &event->hw; + + /* Mark event as deactivated and stopped. */ + if (!(hwc->state & PERF_HES_STOPPED)) + hwc->state |= PERF_HES_STOPPED; + + /* Check if software counter update is necessary. */ + if ((mode & PERF_EF_UPDATE) && !(hwc->state & PERF_HES_UPTODATE)) { + /* + * Drain the remaining delta count out of an event + * that we are disabling: + */ + event_update(event); + hwc->state |= PERF_HES_UPTODATE; + } +} + +static int pmu_event_add(struct perf_event *event, int mode) +{ + struct hw_perf_event *hwc = &event->hw; + + hwc->state = PERF_HES_UPTODATE | PERF_HES_STOPPED; + + if (mode & PERF_EF_START) + __pmu_event_start(event); + + return 0; +} + +static void pmu_event_del(struct perf_event *event, int flags) +{ + pmu_event_stop(event, PERF_EF_UPDATE); +} + +static int pmu_event_init(struct perf_event *event) +{ + u64 cfg = event->attr.config & AMD_POWER_EVENT_MASK; + + /* Only look at AMD power events. */ + if (event->attr.type != pmu_class.type) + return -ENOENT; + + /* Unsupported modes and filters. */ + if (event->attr.exclude_user || + event->attr.exclude_kernel || + event->attr.exclude_hv || + event->attr.exclude_idle || + event->attr.exclude_host || + event->attr.exclude_guest || + /* no sampling */ + event->attr.sample_period) + return -EINVAL; + + if (cfg != AMD_POWER_EVENTSEL_PKG) + return -EINVAL; + + return 0; +} + +static void pmu_event_read(struct perf_event *event) +{ + event_update(event); +} + +static ssize_t +get_attr_cpumask(struct device *dev, struct device_attribute *attr, char *buf) +{ + return cpumap_print_to_pagebuf(true, buf, &cpu_mask); +} + +static DEVICE_ATTR(cpumask, S_IRUGO, get_attr_cpumask, NULL); + +static struct attribute *pmu_attrs[] = { + &dev_attr_cpumask.attr, + NULL, +}; + +static struct attribute_group pmu_attr_group = { + .attrs = pmu_attrs, +}; + +/* + * Currently it only supports to report the power of each + * processor/package. + */ +EVENT_ATTR_STR(power-pkg, power_pkg, "event=0x01"); + +EVENT_ATTR_STR(power-pkg.unit, power_pkg_unit, "mWatts"); + +/* Convert the count from micro-Watts to milli-Watts. */ +EVENT_ATTR_STR(power-pkg.scale, power_pkg_scale, "1.000000e-3"); + +static struct attribute *events_attr[] = { + EVENT_PTR(power_pkg), + EVENT_PTR(power_pkg_unit), + EVENT_PTR(power_pkg_scale), + NULL, +}; + +static struct attribute_group pmu_events_group = { + .name = "events", + .attrs = events_attr, +}; + +PMU_FORMAT_ATTR(event, "config:0-7"); + +static struct attribute *formats_attr[] = { + &format_attr_event.attr, + NULL, +}; + +static struct attribute_group pmu_format_group = { + .name = "format", + .attrs = formats_attr, +}; + +static const struct attribute_group *attr_groups[] = { + &pmu_attr_group, + &pmu_format_group, + &pmu_events_group, + NULL, +}; + +static struct pmu pmu_class = { + .attr_groups = attr_groups, + /* system-wide only */ + .task_ctx_nr = perf_invalid_context, + .event_init = pmu_event_init, + .add = pmu_event_add, + .del = pmu_event_del, + .start = pmu_event_start, + .stop = pmu_event_stop, + .read = pmu_event_read, +}; + +static void power_cpu_exit(int cpu) +{ + int target = nr_cpumask_bits; + + if (!cpumask_test_and_clear_cpu(cpu, &cpu_mask)) + return; + + /* + * Find a new CPU on the same compute unit, if was set in cpumask + * and still some CPUs on compute unit. Then migrate event and + * context to new CPU. + */ + target = cpumask_any_but(topology_sibling_cpumask(cpu), cpu); + if (target < nr_cpumask_bits) { + cpumask_set_cpu(target, &cpu_mask); + perf_pmu_migrate_context(&pmu_class, cpu, target); + } +} + +static void power_cpu_init(int cpu) +{ + /* + * 1) If any CPU is set at cpu_mask in the same compute unit, do + * nothing. + * 2) If no CPU is set at cpu_mask in the same compute unit, + * set current STARTING CPU. + * + * If cpu_mask and topology_sibling_cpumask has intersected + * bits, that means any CPU is set in the same compute unit. + * But cpumask_weight(topology_sibling_cpumask(cpu)) == 1 + * means no CPU is set on cpu_mask in the same compute unit + * before init current STARTING CPU. + */ + if (!cpumask_intersects(&cpu_mask, topology_sibling_cpumask(cpu)) && + cpumask_weight(topology_sibling_cpumask(cpu)) == 1) + cpumask_set_cpu(cpu, &cpu_mask); +} + +static int +power_cpu_notifier(struct notifier_block *self, unsigned long action, void *hcpu) +{ + unsigned int cpu = (long)hcpu; + + switch (action & ~CPU_TASKS_FROZEN) { + case CPU_STARTING: + power_cpu_init(cpu); + break; + case CPU_DOWN_PREPARE: + power_cpu_exit(cpu); + break; + default: + break; + } + + return NOTIFY_OK; +} + +static struct notifier_block power_cpu_notifier_nb = { + .notifier_call = power_cpu_notifier, + .priority = CPU_PRI_PERF, +}; + +static const struct x86_cpu_id cpu_match[] = { + { .vendor = X86_VENDOR_AMD, .family = 0x15 }, + {}, +}; + +static int __init amd_power_pmu_init(void) +{ + int i, ret; + u64 tmp; + + if (!x86_match_cpu(cpu_match)) + return 0; + + if (!boot_cpu_has(X86_FEATURE_ACC_POWER)) + return -ENODEV; + + cu_num = boot_cpu_data.x86_max_cores / smp_num_siblings; + + cpu_pwr_sample_ratio = cpuid_ecx(0x80000007); + + if (rdmsrl_safe(MSR_F15H_CU_MAX_PWR_ACCUMULATOR, &tmp)) { + pr_err("Failed to read max compute unit power accumulator MSR\n"); + return -ENODEV; + } + max_cu_acc_power = tmp; + + cpu_notifier_register_begin(); + + /* Choose one online core of each compute unit. */ + for (i = 0; i < boot_cpu_data.x86_max_cores; i += smp_num_siblings) { + WARN_ON(cpumask_empty(topology_sibling_cpumask(i))); + cpumask_set_cpu(cpumask_any(topology_sibling_cpumask(i)), &cpu_mask); + } + + for_each_online_cpu(i) + power_cpu_init(i); + + __register_cpu_notifier(&power_cpu_notifier_nb); + + ret = perf_pmu_register(&pmu_class, "power", -1); + if (WARN_ON(ret)) { + pr_warn("AMD Power PMU registration failed\n"); + goto out; + } + + pr_info("AMD Power PMU detected, %d compute units\n", cu_num); + +out: + cpu_notifier_register_done(); + + return ret; +} +module_init(amd_power_pmu_init); + +static void __exit amd_power_pmu_exit(void) +{ + cpu_notifier_register_begin(); + __unregister_cpu_notifier(&power_cpu_notifier_nb); + cpu_notifier_register_done(); + + perf_pmu_unregister(&pmu_class); +} +module_exit(amd_power_pmu_exit); + +MODULE_AUTHOR("Huang Rui "); +MODULE_DESCRIPTION("AMD Processor Power Reporting Mechanism"); +MODULE_LICENSE("GPL v2"); diff --git a/include/linux/perf_event.h b/include/linux/perf_event.h index f9828a4..01ea21c 100644 --- a/include/linux/perf_event.h +++ b/include/linux/perf_event.h @@ -128,6 +128,10 @@ struct hw_perf_event { struct { /* itrace */ int itrace_started; }; + struct { /* amd_power */ + u64 pwr_acc; + u64 ptsc; + }; #ifdef CONFIG_HAVE_HW_BREAKPOINT struct { /* breakpoint */ /* -- 1.9.1