From mboxrd@z Thu Jan 1 00:00:00 1970 Return-Path: X-Spam-Checker-Version: SpamAssassin 3.4.0 (2014-02-07) on aws-us-west-2-korg-lkml-1.web.codeaurora.org Received: from bombadil.infradead.org (bombadil.infradead.org [198.137.202.133]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.lore.kernel.org (Postfix) with ESMTPS id EA940C433EF for ; Sun, 27 Feb 2022 13:22:09 +0000 (UTC) DKIM-Signature: v=1; a=rsa-sha256; q=dns/txt; c=relaxed/relaxed; d=lists.infradead.org; s=bombadil.20210309; h=Sender: Content-Transfer-Encoding:Content-Type:List-Subscribe:List-Help:List-Post: List-Archive:List-Unsubscribe:List-Id:In-Reply-To:MIME-Version:References: Message-ID:Subject:Cc:To:From:Date:Reply-To:Content-ID:Content-Description: Resent-Date:Resent-From:Resent-Sender:Resent-To:Resent-Cc:Resent-Message-ID: List-Owner; bh=UmMTIcszojQml7n+lOTlN1xrzt00tKt000EREJLONcw=; b=saG3ptgwNGNx6T abVSoFjLwJB0oXMzuf4JMACWxLWli4F8361AdrnNJ21pQkx2E8Smb4/WY7c4yc2DILYbPiVG/4s+h kau0n+OA2YY8ueT/sMJr4rElD9ZJbs8AH9jftC3qDYOQp7G42XasIF0ixmWth0UOBurJqoIeuR9/H WU8whWf3lbNBO3HtjnKzbKV8ZRArCn2ZVqe4pqZ7MzLXjbCN+PQVlYwhtli8AMfFN2LhHZzoXavht f+Z4S+hO8bN8GlSs0dn04LbwAyVOfLcD4cLsgzHyD3/DSDc6ztnV0R8JO1onilAwSzEgGHa9WcLuW FAwedyLIRSSp9D6PXPMA==; Received: from localhost ([::1] helo=bombadil.infradead.org) by bombadil.infradead.org with esmtp (Exim 4.94.2 #2 (Red Hat Linux)) id 1nOJTd-009Fsq-4G; Sun, 27 Feb 2022 13:20:41 +0000 Received: from mail-pj1-x102c.google.com ([2607:f8b0:4864:20::102c]) by bombadil.infradead.org with esmtps (Exim 4.94.2 #2 (Red Hat Linux)) id 1nOJTY-009FsH-57 for linux-arm-kernel@lists.infradead.org; Sun, 27 Feb 2022 13:20:39 +0000 Received: by mail-pj1-x102c.google.com with SMTP id gb21so8810890pjb.5 for ; Sun, 27 Feb 2022 05:20:35 -0800 (PST) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=linaro.org; s=google; h=date:from:to:cc:subject:message-id:references:mime-version :content-disposition:content-transfer-encoding:in-reply-to; bh=eEz3Leq2u/W/C+1jbRePQ1kSHJwb+1+qvEjuj/ehSuA=; b=wjdw9ZdcHIC+9T5rjiGK0vRE7lu5K7AV99CFY8ovS93CTqbHL9xfz6AO+TukU7w9Pd H18N2jr3Q2o8Yf13Zfd2PrwBXMK7RRM0F2TToBUy7NjrU/6gqNghNmY+Cu+S4kNLiApb NAVK3OGBHpfAJwouCy9lHIbUfkQO6yMUiCOIpddjiJgXzwtJyD2VPOmdUVgFRYMtFaVH Nd5AK2hQaH/2oX9p+L3BMiHqmLR69dThYu0qypWQSJ/kogFHoZ4UWDb0jDGJeeMvO267 YDFrhmVv3St4L74L2OeuWz8uad2j981DkpN7yNAIiL1vEvXIYjYCAlLSRSfh6bolaa1/ GWRQ== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20210112; h=x-gm-message-state:date:from:to:cc:subject:message-id:references :mime-version:content-disposition:content-transfer-encoding :in-reply-to; bh=eEz3Leq2u/W/C+1jbRePQ1kSHJwb+1+qvEjuj/ehSuA=; b=BcJnqEgM5akhkGW/L1vYd+fxP4RlMz86TlBA2EfJiyY5A0aFqKmozaIJ8V5lfSJcgM uSahyZjRCiWQpuBcxg6XZfVfGSX2efK0Xq5e4f1MNjEWoP02Oh2XUWK72ky6xfue7SzX oTg8PI8KS570valAD8yNc1ed+bPwwXGU9PWlNEBfMWFUly/9ywyS23XOzUb8Zb7lxE9z mldLcFnVRcgBXKBdHeU/k63+L7nnVW/ANcludlQwq0JEx1mMu0+/bOQtunFqRXf75t1O wyQK/s9tPXu7AtqQ3Qa1VCJjvnUga0AWMeUgA9RQOx9pOYFX+0qodtaE419kk+Vrzo7E W2og== X-Gm-Message-State: AOAM53312sOCodkr4419Joc8d9M7jDv552uLO2NGzobuYgFiTsNs1gFX bCyT/+UOJFT2aKuJS6KzUG1Xlw== X-Google-Smtp-Source: ABdhPJyEHIQ8pREEDoGfnq9YgfGEEpmJI5qGuFyv/WDDBLW8midPhF54jJFsYqPNjXHwWj9/GRYpKg== X-Received: by 2002:a17:90b:216:b0:1bc:5d68:e7aa with SMTP id fy22-20020a17090b021600b001bc5d68e7aamr12130330pjb.57.1645968034367; Sun, 27 Feb 2022 05:20:34 -0800 (PST) Received: from leoy-ThinkPad-X240s ([204.124.180.219]) by smtp.gmail.com with ESMTPSA id mw7-20020a17090b4d0700b001b8baf6b6f5sm7790543pjb.50.2022.02.27.05.20.22 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Sun, 27 Feb 2022 05:20:33 -0800 (PST) Date: Sun, 27 Feb 2022 21:20:19 +0800 From: Leo Yan To: German Gomez Cc: Ali Saidi , acme@kernel.org, alexander.shishkin@linux.intel.com, andrew.kilroy@arm.com, benh@kernel.crashing.org, james.clark@arm.com, john.garry@huawei.com, jolsa@redhat.com, linux-arm-kernel@lists.infradead.org, linux-kernel@vger.kernel.org, linux-perf-users@vger.kernel.org, mark.rutland@arm.com, mathieu.poirier@linaro.org, mingo@redhat.com, namhyung@kernel.org, peterz@infradead.org, will@kernel.org Subject: Re: [PATCH 2/2] perf arm-spe: Parse more SPE fields and store source Message-ID: <20220227132019.GA107053@leoy-ThinkPad-X240s> References: <20220128210245.4628-1-alisaidi@amazon.com> <7eca7a1d-a5a2-2aab-b3cf-5d83cb8ccf4f@arm.com> <20220212041927.GA763461@leoy-ThinkPad-X240s> <9266bfb6-341c-1d9c-e96f-c9f856a5ffb6@arm.com> MIME-Version: 1.0 Content-Disposition: inline In-Reply-To: <9266bfb6-341c-1d9c-e96f-c9f856a5ffb6@arm.com> X-CRM114-Version: 20100106-BlameMichelson ( TRE 0.8.0 (BSD) ) MR-646709E3 X-CRM114-CacheID: sfid-20220227_052036_308694_AAC01BE6 X-CRM114-Status: GOOD ( 22.19 ) X-BeenThere: linux-arm-kernel@lists.infradead.org X-Mailman-Version: 2.1.34 Precedence: list List-Id: List-Unsubscribe: , List-Archive: List-Post: List-Help: List-Subscribe: , Content-Type: text/plain; charset="iso-8859-1" Content-Transfer-Encoding: quoted-printable Sender: "linux-arm-kernel" Errors-To: linux-arm-kernel-bounces+linux-arm-kernel=archiver.kernel.org@lists.infradead.org On Mon, Feb 21, 2022 at 08:41:43PM +0000, German Gomez wrote: [...] > Some comments: > = > # ARM_SPE_OP_ATOMIC > = > =A0 This might be a hack, but can we not represent it as both LD&SR as the > =A0 atomic op would combine both? > = > =A0 data_src.mem_op =3D PERF_MEM_OP_LOAD | PERF_MEM_OP_STORE; BTH, I don't understand well for this question, but let me explain a bit: We cannot use 'LOAD | STORE' to present the atomic operation. Please see Armv8 ARM section D10.2.7 Operation Type packet, it would give out more details. Atomic operation is an extra attribution for a load or store operations, it could be an atomic load or store, or load-acquire/store-release instructions, or load-exclusive/store-exclusive instructions. The function arm_spe_pkt_desc_op_type() in perf would also give more info about atomic operations. > # ARM_SPE_OP_EXCL (instructions ldxr/stxr) > = > =A0 x86 doesn't seem to have similar instructions with similar semantics > =A0 (please correct me if I'm wrong). For this arch, PERF_MEM_LOCK_LOCK > =A0 probably suffices. > = > =A0 PPC seems to have similar instructions to arm64 (lwarx/stwcx). I don't > =A0 know if they also have instructions with same semantics as x86. > = > =A0 I think it makes sense to have a PERF_MEM_LOCK_EXCL. If not, reusing > =A0 PERF_MEM_LOCK_LOCK is the quicker alternative. On Arm archs, I think OP_EXCL means Load-Exclusive and Store-Exclusive instructions. Different archs have different memory model and different atomic instructions, e.g. on Armv7, we uses Load-Exclusive and Store-Exclusive instructions for spinlock and on Armv8 we uses Load-Acquire and Store-Release instructions for spinlock. I have no any knowledge for x86 and PPC archs. Seems to me, x86 uses compare-and-swap instruction and PPC's lwarx/stwcx instructions "are primitive, or simple, instructions used to perform a read-modify-write operation to storage" [1]. So I personally think we can define PERF_MEM_LOCK_EXCL type for Arm arches and fill into the field perf_mem_data_src::lock: data_src.lock =3D PERF_MEM_LOCK_EXCL; ... or we can consider to introduce a field perf_mem_data_src::atomic and fill a new type PERF_MEM_ATOMIC_EXCL into this new field: data_src.atomic =3D PERF_MEM_ATOMIC_EXCL; [1] https://www.ibm.com/docs/en/aix/7.2?topic=3Dset-lwarx-load-word-reserve= -indexed-instruction > # ARM_SPE_OP_SVE_SG > = > =A0 (I'm sorry if this is too far out of scope of the original patch. Let > =A0 me know if you would prefer to discuss it on a separate channel) > = > =A0 On a separate note, I'm also looking at incorporating some of the SVE > =A0 bits in the perf samples. > =A0 > =A0 For this, do you think it makes sense to have two mem_* categories in > =A0 perf_mem_data_src: > = > =A0 mem_vector (2 bits) > =A0=A0=A0 - simd > =A0=A0=A0 - other (SVE in arm64) I think we can define below vector types: PERF_MEM_VECTOR_SIMD PERF_MEM_VECTOR_SVE The tricky thing is "other"... Based on the description for "Operation Type packet payload (Other)" in the Armv8 Arm, I think we even need to add an extra operation type PERF_MEM_OP_OTHER and assign it to data_src.mem_op field. > =A0 mem_src (1 bit) > =A0=A0=A0 - sparse (scatter/gather loads/stores in SVE, as well as simd) How about the naming "mem_attr" for new field and define two attributions: PERF_MEM_ATTR_SPARSE -> Gather/Scatter operation PERF_MEM_ATTR_PRED -> Predicated operation Just remind, we cannot only approve within Arm related developers, it's good to seek more wider review from other Arch developers when you send new patch set. Thanks, Leo _______________________________________________ linux-arm-kernel mailing list linux-arm-kernel@lists.infradead.org http://lists.infradead.org/mailman/listinfo/linux-arm-kernel