From mboxrd@z Thu Jan 1 00:00:00 1970 Return-Path: X-Spam-Checker-Version: SpamAssassin 3.4.0 (2014-02-07) on aws-us-west-2-korg-lkml-1.web.codeaurora.org Received: from mail.kernel.org (mail.kernel.org [198.145.29.99]) by smtp.lore.kernel.org (Postfix) with ESMTP id D7CF5C433F5 for ; Mon, 25 Oct 2021 20:57:27 +0000 (UTC) Received: from lists.ozlabs.org (lists.ozlabs.org [112.213.38.117]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by mail.kernel.org (Postfix) with ESMTPS id C8C7F60F4F for ; Mon, 25 Oct 2021 20:57:26 +0000 (UTC) DMARC-Filter: OpenDMARC Filter v1.4.1 mail.kernel.org C8C7F60F4F Authentication-Results: mail.kernel.org; dmarc=fail (p=none dis=none) header.from=linux.vnet.ibm.com Authentication-Results: mail.kernel.org; spf=pass smtp.mailfrom=lists.ozlabs.org Received: from boromir.ozlabs.org (localhost [IPv6:::1]) by lists.ozlabs.org (Postfix) with ESMTP id 4HdS2F1TWKz3c7B for ; Tue, 26 Oct 2021 07:57:25 +1100 (AEDT) Authentication-Results: lists.ozlabs.org; dkim=fail reason="signature verification failed" (2048-bit key; unprotected) header.d=ibm.com header.i=@ibm.com header.a=rsa-sha256 header.s=pp1 header.b=KzmW0PYG; dkim-atps=neutral Authentication-Results: lists.ozlabs.org; spf=none (no SPF record) smtp.mailfrom=linux.vnet.ibm.com (client-ip=148.163.158.5; helo=mx0a-001b2d01.pphosted.com; envelope-from=atrajeev@linux.vnet.ibm.com; receiver=) Authentication-Results: lists.ozlabs.org; dkim=pass (2048-bit key; unprotected) header.d=ibm.com header.i=@ibm.com header.a=rsa-sha256 header.s=pp1 header.b=KzmW0PYG; dkim-atps=neutral Received: from mx0a-001b2d01.pphosted.com (mx0b-001b2d01.pphosted.com [148.163.158.5]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by lists.ozlabs.org (Postfix) with ESMTPS id 4HdLzt0WpHz2xCG for ; Tue, 26 Oct 2021 04:10:01 +1100 (AEDT) Received: from pps.filterd (m0098414.ppops.net [127.0.0.1]) by mx0b-001b2d01.pphosted.com (8.16.1.2/8.16.1.2) with SMTP id 19PEl3JJ002245 for ; Mon, 25 Oct 2021 17:09:59 GMT DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=ibm.com; h=from : message-id : content-type : subject : date : in-reply-to : cc : to : references : mime-version; s=pp1; bh=sPytxFxORP4YNSUcCEA78+fyeR6SW3Me322sCL2kyfk=; b=KzmW0PYG1PWFzQ7e9Kg6+WGTY6+Q3MCY8g0LvZc5gdt/sr3VMzplPbQ6QGAf29ugHo1C C2Vfy/gCIT8yLRzZMMJqmE5kve5ZADZ2TKZrPHmcdgsMjAP0MYFwWWmpUhwIgI5uLpO3 grQF9lR5gbnZpnO71J4dG1XIjuQ/pzAedHA9+/fIoq+/XytPoyMzRCIdGAzpajf3Ukyl /FNfyJJEED4fzMOlKlrlrv0xCdZuYAtnWu8nx1wnGq6eCiDT/DIKrVKOAU81Vzhq6iMZ o1SnZCAyVpoErNIt1ybs7q7fuXHpYiP89+nQZoONQVxKsGcZfGjnXvG+1GlREWVQzRSE IQ== Received: from pps.reinject (localhost [127.0.0.1]) by mx0b-001b2d01.pphosted.com with ESMTP id 3bwt36usuu-1 (version=TLSv1.2 cipher=ECDHE-RSA-AES256-GCM-SHA384 bits=256 verify=NOT) for ; Mon, 25 Oct 2021 17:09:58 +0000 Received: from m0098414.ppops.net (m0098414.ppops.net [127.0.0.1]) by pps.reinject (8.16.0.43/8.16.0.43) with SMTP id 19PGsw6r001800 for ; Mon, 25 Oct 2021 17:09:58 GMT Received: from ppma06ams.nl.ibm.com (66.31.33a9.ip4.static.sl-reverse.com [169.51.49.102]) by mx0b-001b2d01.pphosted.com with ESMTP id 3bwt36usu8-1 (version=TLSv1.2 cipher=ECDHE-RSA-AES256-GCM-SHA384 bits=256 verify=NOT); Mon, 25 Oct 2021 17:09:58 +0000 Received: from pps.filterd (ppma06ams.nl.ibm.com [127.0.0.1]) by ppma06ams.nl.ibm.com (8.16.1.2/8.16.1.2) with SMTP id 19PH2rPn008952; Mon, 25 Oct 2021 17:09:56 GMT Received: from b06cxnps4075.portsmouth.uk.ibm.com (d06relay12.portsmouth.uk.ibm.com [9.149.109.197]) by ppma06ams.nl.ibm.com with ESMTP id 3bv9njh4wh-1 (version=TLSv1.2 cipher=ECDHE-RSA-AES256-GCM-SHA384 bits=256 verify=NOT); Mon, 25 Oct 2021 17:09:56 +0000 Received: from b06wcsmtp001.portsmouth.uk.ibm.com (b06wcsmtp001.portsmouth.uk.ibm.com [9.149.105.160]) by b06cxnps4075.portsmouth.uk.ibm.com (8.14.9/8.14.9/NCO v10.0) with ESMTP id 19PH9qn46947448 (version=TLSv1/SSLv3 cipher=DHE-RSA-AES256-GCM-SHA384 bits=256 verify=OK); Mon, 25 Oct 2021 17:09:52 GMT Received: from b06wcsmtp001.portsmouth.uk.ibm.com (unknown [127.0.0.1]) by IMSVA (Postfix) with ESMTP id 950D6A4062; Mon, 25 Oct 2021 17:09:52 +0000 (GMT) Received: from b06wcsmtp001.portsmouth.uk.ibm.com (unknown [127.0.0.1]) by IMSVA (Postfix) with ESMTP id 37878A4054; Mon, 25 Oct 2021 17:09:50 +0000 (GMT) Received: from smtpclient.apple (unknown [9.195.47.16]) by b06wcsmtp001.portsmouth.uk.ibm.com (Postfix) with ESMTPS; Mon, 25 Oct 2021 17:09:49 +0000 (GMT) From: Athira Rajeev Message-Id: <12A145C3-5264-4D88-BCA5-66228D5523AD@linux.vnet.ibm.com> Content-Type: multipart/alternative; boundary="Apple-Mail=_7FCB717C-8178-4C20-844D-B04D366CCAB5" Subject: Re: [PATCH V2] powerpc/perf: Enable PMU counters post partition migration if PMU is active Date: Mon, 25 Oct 2021 22:39:48 +0530 In-Reply-To: <87lf2mxpov.fsf@linux.ibm.com> To: Nathan Lynch References: <1626006357-1611-1-git-send-email-atrajeev@linux.vnet.ibm.com> <87lf2mxpov.fsf@linux.ibm.com> X-Mailer: Apple Mail (2.3654.120.0.1.13) X-TM-AS-GCONF: 00 X-Proofpoint-ORIG-GUID: UUmlWU1IdPIvYtN9o66rBAjBMndaWmB_ X-Proofpoint-GUID: Kl7i19pjfy6Uws7XgjrmrqBvA7dVJvS1 X-Proofpoint-UnRewURL: 0 URL was un-rewritten MIME-Version: 1.0 X-Proofpoint-Virus-Version: vendor=baseguard engine=ICAP:2.0.182.1,Aquarius:18.0.790,Hydra:6.0.425,FMLib:17.0.607.475 definitions=2021-10-25_06,2021-10-25_02,2020-04-07_01 X-Proofpoint-Spam-Details: rule=outbound_notspam policy=outbound score=0 mlxscore=0 bulkscore=0 adultscore=0 mlxlogscore=999 impostorscore=0 lowpriorityscore=0 priorityscore=1501 malwarescore=0 clxscore=1015 spamscore=0 suspectscore=0 phishscore=0 classifier=spam adjust=0 reason=mlx scancount=1 engine=8.12.0-2109230001 definitions=main-2110250099 X-Mailman-Approved-At: Tue, 26 Oct 2021 07:56:00 +1100 X-BeenThere: linuxppc-dev@lists.ozlabs.org X-Mailman-Version: 2.1.29 Precedence: list List-Id: Linux on PowerPC Developers Mail List List-Unsubscribe: , List-Archive: List-Post: List-Help: List-Subscribe: , Cc: kjain@linux.ibm.com, maddy@linux.ibm.com, linuxppc-dev@lists.ozlabs.org, Nicholas Piggin , rnsastry@linux.ibm.com Errors-To: linuxppc-dev-bounces+linuxppc-dev=archiver.kernel.org@lists.ozlabs.org Sender: "Linuxppc-dev" --Apple-Mail=_7FCB717C-8178-4C20-844D-B04D366CCAB5 Content-Type: text/plain; charset=utf-8 Content-Transfer-Encoding: quoted-printable > On 21-Oct-2021, at 10:47 PM, Nathan Lynch wrote: >=20 > Athira Rajeev > writes: >> During Live Partition Migration (LPM), it is observed that perf >> counter values reports zero post migration completion. However >> 'perf stat' with workload continues to show counts post migration >> since PMU gets disabled/enabled during sched switches. But incase >> of system/cpu wide monitoring, zero counts were reported with 'perf >> stat' after migration completion. >>=20 >> Example: >> ./perf stat -e r1001e -I 1000 >> time counts unit events >> 1.001010437 22,137,414 r1001e >> 2.002495447 15,455,821 r1001e >> <<>> As seen in next below logs, the counter values shows zero >> after migration is completed. >> <<>> >> 86.142535370 129,392,333,440 r1001e >> 87.144714617 0 r1001e >> 88.146526636 0 r1001e >> 89.148085029 0 r1001e >=20 > Confirmed in my environment: >=20 > 51.099987985 300,338 cache-misses > 52.101839374 296,586 cache-misses > 53.116089796 263,150 cache-misses > 54.117949249 232,290 cache-misses > 55.602029375 68,700,421,711 cache-misses > 56.610073969 0 cache-misses > 57.614732000 0 cache-misses >=20 > I wonder what it means that there is a very unlikely huge value before > the counter stops working -- I believe your example has this phenomenon > too. >=20 >=20 >> diff --git a/arch/powerpc/platforms/pseries/mobility.c b/arch/powerpc/pl= atforms/pseries/mobility.c >> index e83e089..ff7a77c 100644 >> --- a/arch/powerpc/platforms/pseries/mobility.c >> +++ b/arch/powerpc/platforms/pseries/mobility.c >> @@ -476,6 +476,8 @@ static int do_join(void *arg) >> retry: >> /* Must ensure MSR.EE off for H_JOIN. */ >> hard_irq_disable(); >> + /* Disable PMU before suspend */ >> + mobility_pmu_disable(); >> hvrc =3D plpar_hcall_norets(H_JOIN); >>=20 >> switch (hvrc) { >> @@ -530,6 +532,8 @@ static int do_join(void *arg) >> * reset the watchdog. >> */ >> touch_nmi_watchdog(); >> + /* Enable PMU after resuming */ >> + mobility_pmu_enable(); >> return ret; >> } >=20 > We should minimize calls into other subsystems from this context (the > callback function we've passed to stop_machine); it's fairly sensitive. > Can this be moved out to pseries_migrate_partition() or similar? Hi Nathan Thanks for the review. I will move the callbacks to =E2=80=9Cpseries_migrate_partition=E2=80=9D in= next version Athira. --Apple-Mail=_7FCB717C-8178-4C20-844D-B04D366CCAB5 Content-Transfer-Encoding: quoted-printable Content-Type: text/html; charset=utf-8

On 21-Oct-2021, at 10:47 PM, Nathan Lynch <nathanl@linux.ibm.com> wrote:

Athira Rajeev <atrajeev@linux.vnet.ibm.com> writes:
During Live Partition Migration = (LPM), it is observed that perf
counter values reports = zero post migration completion. However
'perf stat' with = workload continues to show counts post migration
since PMU = gets disabled/enabled during sched switches. But incase
of = system/cpu wide monitoring, zero counts were reported with 'perf
stat' after migration completion.

Example:
./perf stat -e r1001e -I 1000
          tim= e =             co= unts unit events
    1.001010437 =         22,137,414 =      r1001e
    2.002495447 =         15,455,821 =      r1001e
<<>> As = seen in next below logs, the counter values shows zero
       after migration is = completed.
<<>>
   86.142535370 =    129,392,333,440 =      r1001e
   87.144714617 =             &n= bsp;    0      r1001e
   88.146526636 =             &n= bsp;    0      r1001e
   89.148085029 =             &n= bsp;    0      r1001e

Confirmed in my environment:

   51.099987985 =            300,338 =      cache-misses
   52.101839374 =            296,586 =      cache-misses
   53.116089796 =            263,150 =      cache-misses
   54.117949249 =            232,290 =      cache-misses
   55.602029375 =     68,700,421,711 =      cache-misses
   56.610073969 =             &n= bsp;    0 =      cache-misses
   57.614732000 =             &n= bsp;    0 =      cache-misses

I wonder what it means that there is a very unlikely huge = value before
the counter = stops working -- I believe your example has this phenomenon
too.


diff = --git a/arch/powerpc/platforms/pseries/mobility.c = b/arch/powerpc/platforms/pseries/mobility.c
index = e83e089..ff7a77c 100644
--- = a/arch/powerpc/platforms/pseries/mobility.c
+++ = b/arch/powerpc/platforms/pseries/mobility.c
@@ -476,6 = +476,8 @@ static int do_join(void *arg)
retry:
= /* Must ensure MSR.EE off for H_JOIN. */
= hard_irq_disable();
+ /* Disable PMU before suspend = */
+ mobility_pmu_disable();
hvrc =3D = plpar_hcall_norets(H_JOIN);

switch = (hvrc) {
@@ -530,6 +532,8 @@ static int do_join(void = *arg)
 * = reset the watchdog.
 */
= touch_nmi_watchdog();
+ /* Enable PMU after resuming = */
+ mobility_pmu_enable();
return = ret;
}

We should minimize calls into other subsystems from this = context (the
callback = function we've passed to stop_machine); it's fairly sensitive.
Can this be moved out to = pseries_migrate_partition() or similar?

Hi Nathan

Thanks= for the review.
I will move the callbacks to = =E2=80=9Cpseries_migrate_partition=E2=80=9D in next = version

Athira.

= --Apple-Mail=_7FCB717C-8178-4C20-844D-B04D366CCAB5--