From mboxrd@z Thu Jan 1 00:00:00 1970 Return-Path: X-Spam-Checker-Version: SpamAssassin 3.4.0 (2014-02-07) on aws-us-west-2-korg-lkml-1.web.codeaurora.org Received: from bombadil.infradead.org (bombadil.infradead.org [198.137.202.133]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.lore.kernel.org (Postfix) with ESMTPS id EA480CCFA03 for ; Mon, 3 Nov 2025 17:08:47 +0000 (UTC) DKIM-Signature: v=1; a=rsa-sha256; q=dns/txt; c=relaxed/relaxed; d=lists.infradead.org; s=bombadil.20210309; h=Sender:Reply-To:List-Subscribe: List-Help:List-Post:List-Archive:List-Unsubscribe:List-Id:In-Reply-To: Content-Type:MIME-Version:References:Message-ID:Subject:Cc:To:From:Date: Content-Transfer-Encoding:Content-ID:Content-Description:Resent-Date: Resent-From:Resent-Sender:Resent-To:Resent-Cc:Resent-Message-ID:List-Owner; bh=YA8sXFA0OPROcJCL+f55Bq+qqdOX8UURJdXJr72IWQ4=; b=k0rz5K5TxgTkJl2NXKuFMWftzQ N3rCuMgrPuqc8CHds7LXwwE6Nc6aajra6ciUzNdm04VyRkXkgmrbr1muOXUNqaruRYRBTxQXW/T9n +8aJ0UZb68PQqtzPW4/Ru13VGQ/oolV9m5L2udpTBcu0EB1hFS08/ZBD/7shFyE04jdMQAMePyB4J OsgHQVgH7Rek+vWBViX/+nOGqNNjQjeXbreaDr0b5v6+1iF17A59pxRBLGp7VaPVwtG8EkHEbUBUU rbnpeMozWrO3izidNmYElaJCVBqBdat3w2TU0FVg2VS4skLSvlIMOSVQ2Yx6I1mgC/kmuf5g2EtOr 1LUpEJ4Q==; Received: from localhost ([::1] helo=bombadil.infradead.org) by bombadil.infradead.org with esmtp (Exim 4.98.2 #2 (Red Hat Linux)) id 1vFy2s-0000000ANkE-49a6; Mon, 03 Nov 2025 17:08:42 +0000 Received: from tor.source.kernel.org ([2600:3c04:e001:324:0:1991:8:25]) by bombadil.infradead.org with esmtps (Exim 4.98.2 #2 (Red Hat Linux)) id 1vFy2r-0000000ANjw-49Na for linux-arm-kernel@lists.infradead.org; Mon, 03 Nov 2025 17:08:42 +0000 Received: from smtp.kernel.org (transwarp.subspace.kernel.org [100.75.92.58]) by tor.source.kernel.org (Postfix) with ESMTP id 6C69D600B0; Mon, 3 Nov 2025 17:08:41 +0000 (UTC) Received: by smtp.kernel.org (Postfix) with ESMTPSA id 0C1FFC4CEFD; Mon, 3 Nov 2025 17:08:41 +0000 (UTC) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/simple; d=kernel.org; s=k20201202; t=1762189721; bh=nc9pwfo7tVhfkofi9WsCHyPsawYIqzmCCe2hZfm+ryk=; h=Date:From:To:Cc:Subject:Reply-To:References:In-Reply-To:From; b=NPLq8Qxjk5TF2XoJ3zCNnbZ8M1n51tjVIBvXcBq0DCmi75P9TiTP+gla/3+NQ6kIG ec171K/lqcI/7q3+JNfTlBvy5eFF+tqfh6FiJEithvtCOfxDJ1P9KFV7X9SQy6R7wm LLfCZGkKk5NNubTVqLPOHZI3QbET/7DHSJt/ayYOkswXKLJmkjEOD4SStXpJC/MMWX n1Ls3fQRTAY2MRktVlxsEqKZnTCVpFkB/jSGbVb2W8I5CUq0ydnb1IjUY71U/dLOBK wJUXvdfqTdj1ffixMOMeJlkD9CQSTvvjOYGgLd0zoO1u5ujBnxFQ/VLkINpSwZDcyA bTLK6d+XFNUhg== Received: by paulmck-ThinkPad-P17-Gen-1.home (Postfix, from userid 1000) id E6772CE0B94; Mon, 3 Nov 2025 09:08:39 -0800 (PST) Date: Mon, 3 Nov 2025 09:08:39 -0800 From: "Paul E. McKenney" To: Mathieu Desnoyers Cc: rcu@vger.kernel.org, linux-kernel@vger.kernel.org, kernel-team@meta.com, rostedt@goodmis.org, Catalin Marinas , Will Deacon , Mark Rutland , Sebastian Andrzej Siewior , linux-arm-kernel@lists.infradead.org, bpf@vger.kernel.org Subject: Re: [PATCH 17/19] srcu: Optimize SRCU-fast-updown for arm64 Message-ID: References: <082fb8ba-91b8-448e-a472-195eb7b282fd@paulmck-laptop> <20251102214436.3905633-17-paulmck@kernel.org> MIME-Version: 1.0 Content-Type: text/plain; charset=us-ascii Content-Disposition: inline In-Reply-To: X-BeenThere: linux-arm-kernel@lists.infradead.org X-Mailman-Version: 2.1.34 Precedence: list List-Id: List-Unsubscribe: , List-Archive: List-Post: List-Help: List-Subscribe: , Reply-To: paulmck@kernel.org Sender: "linux-arm-kernel" Errors-To: linux-arm-kernel-bounces+linux-arm-kernel=archiver.kernel.org@lists.infradead.org On Mon, Nov 03, 2025 at 08:34:10AM -0500, Mathieu Desnoyers wrote: > On 2025-11-02 16:44, Paul E. McKenney wrote: > > Some arm64 platforms have slow per-CPU atomic operations, for example, > > the Neoverse V2. This commit therefore moves SRCU-fast from per-CPU > > atomic operations to interrupt-disabled non-read-modify-write-atomic > > atomic_read()/atomic_set() operations. This works because > > SRCU-fast-updown is not invoked from read-side primitives, which > > means that if srcu_read_unlock_fast() NMI handlers. This means that > > srcu_read_lock_fast_updown() and srcu_read_unlock_fast_updown() can > > exclude themselves and each other > > > > This reduces the overhead of calls to srcu_read_lock_fast_updown() and > > srcu_read_unlock_fast_updown() from about 100ns to about 12ns on an ARM > > Neoverse V2. Although this is not excellent compared to about 2ns on x86, > > it sure beats 100ns. > > > > This command was used to measure the overhead: > > > > tools/testing/selftests/rcutorture/bin/kvm.sh --torture refscale --allcpus --duration 5 --configs NOPREEMPT --kconfig "CONFIG_NR_CPUS=64 CONFIG_TASKS_TRACE_RCU=y" --bootargs "refscale.loops=100000 refscale.guest_os_delay=5 refscale.nreaders=64 refscale.holdoff=30 torture.disable_onoff_at_boot refscale.scale_type=srcu-fast-updown refscale.verbose_batched=8 torture.verbose_sleep_frequency=8 torture.verbose_sleep_duration=8 refscale.nruns=100" --trust-make > > > Hi Paul, > > At a high level, what are you trying to achieve with this ? I am working around the high single-CPU cost of arm64 LSE instructions, as in about 50ns per compared non-LSE of about 5ns per. The 50ns rules them out for uretprobes, for example. But Catalin's later patch is in all ways better than mine, so I will be keeping this one only until Catalin's hits mainline. Once that happens, I will revert this one the following merge window. (It might be awhile because of the testing required on a wide range of platforms.) > AFAIU, you are trying to remove the cost of atomics on per-cpu > data from srcu-fast read lock/unlock for frequent calls for > CONFIG_NEED_SRCU_NMI_SAFE=y, am I on the right track ? > > [disclaimer: I've looked only briefly at your proposed patch.] > Then there are various other less specific approaches to consider > before introducing such architecture and use-case specific work-around. > > One example is the libside (user level) rcu implementation which uses > two counters per cpu [1]. One counter is the rseq fast path, and the > second counter is for atomics (as fallback). > > If the typical scenario we want to optimize for is thread context, we > can probably remove the atomic from the fast path with just preempt off > by partitioning the per-cpu counters further, one possibility being: > > struct percpu_srcu_fast_pair { > unsigned long lock, unlock; > }; > > struct percpu_srcu_fast { > struct percpu_srcu_fast_pair thread; > struct percpu_srcu_fast_pair irq; > }; > > And the grace period sums both thread and irq counters. > > Thoughts ? One complication here is that we need srcu_down_read() at task level and the matching srcu_up_read() at softirq and/or hardirq level. Or am I missing a trick in your proposed implementation? Thanx, Paul > Thanks, > > Mathieu > > [1] https://github.com/compudj/libside/blob/master/src/rcu.h#L71 > > -- > Mathieu Desnoyers > EfficiOS Inc. > https://www.efficios.com