From mboxrd@z Thu Jan 1 00:00:00 1970 Received: from smtp.kernel.org (aws-us-west-2-korg-mail-alma10-1.taild15c8.ts.net [100.103.45.18]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 6CF0351EDF1; Wed, 30 Sep 2026 17:45:27 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=100.103.45.18 ARC-Seal:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1790790328; cv=none; b=icAENfViqiZvR5I5lQSBIfkAxrflsw4rIH3yPIzJiIa+KO4C/BjQ+EDCaFbzXfRofkyjWZAdSI9lDSDh7wardtooT1MopQ6U7n2H5zKBd7zHPwgTF2b7o7qo2hUaJd9uVW4CgtEjoDuc/TWd6kIRYbZal2TTANuxRfZhefucw8M= ARC-Message-Signature:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1790790328; c=relaxed/simple; bh=rUf1OY+BGc/F9Irrs2O1t6mTkIKdozPlZb4oXGkta6E=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=EkcXfGqpHoq9g3LLjUfHmcvK6u6u5B8wI/eXvkGNVmAYLKt3JcoW66obMIA0YKuf6pAW/DkuN4tHANKD5YZN1VRsWXJDLRcuQv3l4PDxBdlsc62zAodo49ZtCpVRautkFeckOjJ6vjidOB8qgAwI/VBeSHunipZt2FyiANqN53c= ARC-Authentication-Results:i=1; smtp.subspace.kernel.org; dkim=pass (1024-bit key) header.d=linuxfoundation.org header.i=@linuxfoundation.org header.b=HTn5CUHH; arc=none smtp.client-ip=100.103.45.18 Authentication-Results: smtp.subspace.kernel.org; dkim=pass (1024-bit key) header.d=linuxfoundation.org header.i=@linuxfoundation.org header.b="HTn5CUHH" Received: by smtp.kernel.org (Postfix) with ESMTPSA id AF0F41F000FF; Wed, 30 Sep 2026 17:45:26 +0000 (UTC) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=linuxfoundation.org; s=korg; t=1790790327; bh=TdiKgBQnkLKkHoe970DAcVfXBLpY3tBkJnPLKfnq3EM=; h=From:To:Cc:Subject:Date:In-Reply-To:References; b=HTn5CUHHMalVLqKttN6xFza3m2hy5MYQWb+5k3ZM76QAyS9vJscSOiKIRLgK0YSUO 6paCRJuFzHUaFve+GtzoM89+im8tAZCd0xqpSxR+S32uwGa5vUoUBgMwdW64bRNaQI 5Coyw6IWeNfdHGZ7UJ2VZJpKznbNI8TmnWlxibSw= From: Greg Kroah-Hartman To: stable@vger.kernel.org Cc: Greg Kroah-Hartman , patches@lists.linux.dev, Catalin Marinas , "Paul E. McKenney" , Will Deacon , Palmer Dabbelt , Sasha Levin Subject: [PATCH 6.12 791/877] arm64: Use load LSE atomics for the non-return per-CPU atomic operations Date: Wed, 30 Sep 2026 17:28:22 +0200 Message-ID: <20260930152431.783514002@linuxfoundation.org> X-Mailer: git-send-email 2.55.0 In-Reply-To: <20260930152414.738996857@linuxfoundation.org> References: <20260930152414.738996857@linuxfoundation.org> User-Agent: quilt/0.69 X-stable: review X-Patchwork-Hint: ignore Precedence: bulk X-Mailing-List: patches@lists.linux.dev List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: 8bit 6.12-stable review patch. If anyone has any objections, please let me know. ------------------ From: Catalin Marinas [ Upstream commit 535fdfc5a228524552ee8810c9175e877e127c27 ] The non-return per-CPU this_cpu_*() atomic operations are implemented as STADD/STCLR/STSET when FEAT_LSE is available. On many microarchitecture implementations, these instructions tend to be executed "far" in the interconnect or memory subsystem (unless the data is already in the L1 cache). This is in general more efficient when there is contention as it avoids bouncing cache lines between CPUs. The load atomics (e.g. LDADD without XZR as destination), OTOH, tend to be executed "near" with the data loaded into the L1 cache. STADD executed back to back as in srcu_read_{lock,unlock}*() incur an additional overhead due to the default posting behaviour on several CPU implementations. Since the per-CPU atomics are unlikely to be used concurrently on the same memory location, encourage the hardware to to execute them "near" by issuing load atomics - LDADD/LDCLR/LDSET - with the destination register unused (but not XZR). Signed-off-by: Catalin Marinas Link: https://lore.kernel.org/r/e7d539ed-ced0-4b96-8ecd-048a5b803b85@paulmck-laptop Reported-by: Paul E. McKenney Tested-by: Paul E. McKenney Cc: Will Deacon Reviewed-by: Palmer Dabbelt [will: Add comment and link to the discussion thread] Signed-off-by: Will Deacon Stable-dep-of: 8cf2093f5372 ("arm64: percpu: Fix LSE operations on {8,16}-bit types") Signed-off-by: Sasha Levin Signed-off-by: Greg Kroah-Hartman --- arch/arm64/include/asm/percpu.h | 15 +++++++++++---- 1 file changed, 11 insertions(+), 4 deletions(-) --- a/arch/arm64/include/asm/percpu.h +++ b/arch/arm64/include/asm/percpu.h @@ -77,7 +77,7 @@ __percpu_##name##_case_##sz(void *ptr, u " stxr" #sfx "\t%w[loop], %" #w "[tmp], %[ptr]\n" \ " cbnz %w[loop], 1b", \ /* LSE atomics */ \ - #op_lse "\t%" #w "[val], %[ptr]\n" \ + #op_lse "\t%" #w "[val], %" #w "[tmp], %[ptr]\n" \ __nops(3)) \ : [loop] "=&r" (loop), [tmp] "=&r" (tmp), \ [ptr] "+Q"(*(u##sz *)ptr) \ @@ -124,9 +124,16 @@ PERCPU_RW_OPS(8) PERCPU_RW_OPS(16) PERCPU_RW_OPS(32) PERCPU_RW_OPS(64) -PERCPU_OP(add, add, stadd) -PERCPU_OP(andnot, bic, stclr) -PERCPU_OP(or, orr, stset) + +/* + * Use value-returning atomics for CPU-local ops as they are more likely + * to execute "near" to the CPU (e.g. in L1$). + * + * https://lore.kernel.org/r/e7d539ed-ced0-4b96-8ecd-048a5b803b85@paulmck-laptop + */ +PERCPU_OP(add, add, ldadd) +PERCPU_OP(andnot, bic, ldclr) +PERCPU_OP(or, orr, ldset) PERCPU_RET_OP(add, add, ldadd) #undef PERCPU_RW_OPS