From mboxrd@z Thu Jan 1 00:00:00 1970 Return-Path: X-Spam-Checker-Version: SpamAssassin 3.4.0 (2014-02-07) on aws-us-west-2-korg-lkml-1.web.codeaurora.org Received: from bombadil.infradead.org (bombadil.infradead.org [198.137.202.133]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.lore.kernel.org (Postfix) with ESMTPS id 14585EE4993 for ; Mon, 21 Aug 2023 12:11:20 +0000 (UTC) DKIM-Signature: v=1; a=rsa-sha256; q=dns/txt; c=relaxed/relaxed; d=lists.infradead.org; s=bombadil.20210309; h=Sender: Content-Transfer-Encoding:Content-Type:List-Subscribe:List-Help:List-Post: List-Archive:List-Unsubscribe:List-Id:In-Reply-To:MIME-Version:References: Message-ID:Subject:Cc:To:From:Date:Reply-To:Content-ID:Content-Description: Resent-Date:Resent-From:Resent-Sender:Resent-To:Resent-Cc:Resent-Message-ID: List-Owner; bh=IuBC+4Z2f45r2C60Z7MEYEsq/vko2plFJTRi9x41VVE=; b=ZRZVk0SDA4KhB8 oHfZDb6XW8h5b4H7x6DJPgY8hKnjBAFw8z760BGtgrtBWDJaBTpr9/QmvFTT69z2mAU2duHUcVP15 HkEcXJtmejXSJJsgqg1suiziUDDIBmMj7LJ9mrcUSiKZttUvyL/8y9rybrck3dgytaZa7jwAh0VZl ShntvRSP6rDb1mV1w61R9hD7/ljDvW2jBrSu0hHIjWST1q4Gc2AL5MWtwsX4Q3kR97PT4IDFtiL3H IlBHkz4kM+g22cUy3AwhPGYo7teAkY+fh4P53b1ONO7JfM6vVVCV5vRVu4Dr+4HaSBO7wt2FkehVN LUofGy6S49XT7dJhVLsg==; Received: from localhost ([::1] helo=bombadil.infradead.org) by bombadil.infradead.org with esmtp (Exim 4.96 #2 (Red Hat Linux)) id 1qY3kL-00DwkC-0L; Mon, 21 Aug 2023 12:11:01 +0000 Received: from dfw.source.kernel.org ([139.178.84.217]) by bombadil.infradead.org with esmtps (Exim 4.96 #2 (Red Hat Linux)) id 1qY3kF-00Dwir-23 for linux-arm-kernel@lists.infradead.org; Mon, 21 Aug 2023 12:10:57 +0000 Received: from smtp.kernel.org (relay.kernel.org [52.25.139.140]) (using TLSv1.3 with cipher TLS_AES_256_GCM_SHA384 (256/256 bits) key-exchange X25519 server-signature RSA-PSS (2048 bits)) (No client certificate requested) by dfw.source.kernel.org (Postfix) with ESMTPS id E32B26334F; Mon, 21 Aug 2023 12:10:54 +0000 (UTC) Received: by smtp.kernel.org (Postfix) with ESMTPSA id 7D07CC433C8; Mon, 21 Aug 2023 12:10:53 +0000 (UTC) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/simple; d=kernel.org; s=k20201202; t=1692619854; bh=RjRb9BdupNx+xiC24F7DnlTGb1n6/oJFy2C9WHeWCko=; h=Date:From:To:Cc:Subject:References:In-Reply-To:From; b=PS3UUqO33/2+YxheK/j4VyAegBFzcAkRyUQ/YikAowF38UR/fIZZ5JG9H55fLiewN SwKiF3+xLM7ks+24jDKK8gQv/gvBDF7F+6Vn5vDocE1kmzF0vduFqnqVsAEFfTwLCb e0Bx1F3rC8Scj/v66LzsdDVQKOTYbkYSOzfClSKWbvJ6+xp5u/N26W6TmeHSMY/RSj Mu7FM3vx6tn5PViIonWe3ydBQt5eElXvC6PqRURTWzo90dnKRlgUmypq8ZvLx5G+9e a4kLNU55toDIpFuYyhp9AM1BP0E+NocAUVkkNU6sqT4MVEgNDOA/lknpxlZHaMQQId hqmq8WwnFEEKA== Date: Mon, 21 Aug 2023 13:10:50 +0100 From: Will Deacon To: Mark Brown Cc: Catalin Marinas , Benjamin Herrenschmidt , linux-arm-kernel@lists.infradead.org, linux-kernel@vger.kernel.org Subject: Re: [PATCH] arm64/fpsimd: Suppress SVE access traps when loading FPSIMD state Message-ID: <20230821121049.GA19670@willie-the-truck> References: <20230807-arm64-sve-trap-mitigation-v1-1-d92eed1d2855@kernel.org> MIME-Version: 1.0 Content-Disposition: inline In-Reply-To: <20230807-arm64-sve-trap-mitigation-v1-1-d92eed1d2855@kernel.org> User-Agent: Mutt/1.10.1 (2018-07-13) X-CRM114-Version: 20100106-BlameMichelson ( TRE 0.8.0 (BSD) ) MR-646709E3 X-CRM114-CacheID: sfid-20230821_051055_759008_EDD4C8E0 X-CRM114-Status: GOOD ( 32.48 ) X-BeenThere: linux-arm-kernel@lists.infradead.org X-Mailman-Version: 2.1.34 Precedence: list List-Id: List-Unsubscribe: , List-Archive: List-Post: List-Help: List-Subscribe: , Content-Type: text/plain; charset="us-ascii" Content-Transfer-Encoding: 7bit Sender: "linux-arm-kernel" Errors-To: linux-arm-kernel-bounces+linux-arm-kernel=archiver.kernel.org@lists.infradead.org On Mon, Aug 07, 2023 at 11:20:38PM +0100, Mark Brown wrote: > When we are in a syscall we take the opportunity to discard the SVE state, > saving only the FPSIMD subset of the register state. When we reload the > state from memory we reenable SVE access traps, stopping tracking SVE until > the task takes another SVE access trap. This means that for a task which is > actively using SVE many blocking system calls will have the additional > overhead of a SVE access trap. > > As SVE deployment is progressing we are seeing much wider use of the SVE > instruction set, including performance optimised implementations of > operations like memset() and memcpy(), which mean that even tasks which are > not obviously floating point based can end up with substantial SVE usage. > > It does not, however, make sense to just unconditionally use the full SVE > register state all the time since it is larger than the FPSIMD register > state so there is overhead saving and restoring it on context switch and > our requirement to flush the register state not shared with FPSIMD on > syscall also creates a noticeable overhead on system call. > > I did some instrumentation which counted the number of SVE access traps > and the number of times we loaded FPSIMD only register state for each task. > Testing with Debian Bookworm this showed that during boot the overwhelming > majority of tasks triggered another SVE access trap more than 50% of the > time after loading FPSIMD only state with a substantial number near 100%, > though some programs had a very small number of SVE accesses most likely > from startup. There were few tasks in the range 5-45%, most tasks either > used SVE frequently or used it only a tiny proportion of times. As expected > older distributions which do not have the SVE performance work available > showed no SVE usage in general applications. > > This indicates that there should be some useful benefit from reducing the > number of SVE access traps for blocking system calls like we did for non > blocking system calls in commit 8c845e273104 ("arm64/sve: Leave SVE enabled > on syscall if we don't context switch"). Let's do this by counting the > number of times we have loaded FPSIMD only register state for SVE tasks > and only disabling traps after some number of times, otherwise leaving > traps disabled and flushing the non-shared register state like we would on > trap. > > I pulled 64 out of thin air for the number of flushes to do, there is > doubtless room for tuning here. Ideally we would be able to tell if the > task is actually using SVE but without using performance counters (which > would be substantial work) we can't currently tell. I picked the number > because so many of the tasks using SVE used it so frequently. > > This means that for a task which is actively using SVE the number of SVE > access traps will be substantially reduced but applications which use SVE > only very infrequently will avoid the overheads associated with tracking > SVE state once the counter expires. > > There should be no functional change resulting from this, it is purely a > performance optimisation. Do you have any performance numbers to motivate this change? It would be interesting, for example, to see how changing the timeout value affects the results for some real workloads. Will _______________________________________________ linux-arm-kernel mailing list linux-arm-kernel@lists.infradead.org http://lists.infradead.org/mailman/listinfo/linux-arm-kernel