Linux-ARM-Kernel Archive on lore.kernel.org
 help / color / mirror / Atom feed
From: Muhammad Usama Anjum <usama.anjum@arm.com>
To: Catalin Marinas <catalin.marinas@arm.com>,
	Leonardo Bras <leo.bras@arm.com>
Cc: usama.anjum@arm.com, Linus Walleij <linusw@kernel.org>,
	Will Deacon <will@kernel.org>, Marc Zyngier <maz@kernel.org>,
	Oliver Upton <oupton@kernel.org>, Joey Gouly <joey.gouly@arm.com>,
	Suzuki K Poulose <suzuki.poulose@arm.com>,
	Zenghui Yu <yuzenghui@huawei.com>,
	linux-arm-kernel@lists.infradead.org, kvmarm@lists.linux.dev
Subject: Re: [PATCH v2] arm64: clear_page[s] using memset
Date: Thu, 8 Oct 2026 14:30:20 +0100	[thread overview]
Message-ID: <ec298ee5-5080-494d-b735-7ba7bd20cdf0@arm.com> (raw)
In-Reply-To: <arP619hd9Rx8Vlrx@arm.com>

On 23/09/2026 5:14 pm, Catalin Marinas wrote:
> On Wed, Sep 16, 2026 at 04:54:09PM +0100, Leonardo Bras wrote:
>> On Wed, Sep 16, 2026 at 04:28:51PM +0100, Leonardo Bras wrote:
>>> On Wed, Sep 16, 2026 at 12:03:56PM +0200, Linus Walleij wrote:
>>>> There is no need to try to second-guess the compiler when
>>>> clearing memory. Just call memset() like everyone else.
>>>>
>>>> Since memset() already has an architecture-local MOPS
>>>> optimization, we do not need to do anything else to preserve
>>>> the MOPS optimization.
>>>>
>>>> While at it, implement the shorthand for directly calling
>>>> the new prototype clear_pages() for larger page chunks.
>>>>
>>>> No performance regressions can be seen, the fastpath
>>>> benchmarks differences are in the noise.
>>>>
>>>> Usama Anjum tested next-20260821 with one warm-up and three repeats
>>>> in four sessions, for a total of 12 measured runs. The commands were:
>>>>
>>>>   perf bench mem memset -k 1GB -f default -s 16GB
>>>>   perf bench mem mmap -p 1GB -f demand -s 32GB -l 5
>>>>   perf bench mem mmap -p 4KB -f demand -s 32GB -l 5
>>>>
>>>> The results were:
>>>>
>>>> aws-m7g.metal:
>>>>   Benchmark    Base bytes/sec   Change with patch
>>>>   memset 1GB   63932542232.12                1.92%
>>>>   mmap 1GB     63272579168.03               -0.57%
>>>>   mmap 4KB     49692830849.48               -1.11%
>>>>
>>>> cesw-aarch64-ampereone-1s-a192-32x:
>>>>   Benchmark    Base bytes/sec   Change with patch
>>>>   memset 1GB   33895562998.71                0.25%
>>>>   mmap 1GB     34338454210.17                1.16%
>>>>   mmap 4KB     25687107580.90               -0.85%
> [...]
>>> Looking on that, I see that the memset() implementation uses setp, setm, 
>>> sete, while the clear_page()'s uses setpn, setmn, setn for the case with 
>>> MOPS. But then, reading into the docs, the instructions seem pretty much 
>>> the same thing. 
> [...]
>> Oh, I browsed a bit here, and IIUC none of the tested machines have 
>> FEAT_MOPS, is that right?
>>
>> If that's the case, the tests are exactly to what is different between 
>> patched and current versions. There should be no impact on MOPS version as 
>> the instructions are basically the same.
> 
> Logically, yes, they are the same. From a performance perspective, there
> may be a difference between the temporal and non-temporal variants,
> depending on the usage.
> 
> I think it would be good to run the benchmarks with the current
> implementation without DC ZVA. I don't think we have an easy way to do
> this on the command line, so we can simply hard-code the DZP=1 check and
> fall back to the STP or STNP in both cases.
> 
> Otherwise I'm fine with the patch as well, good clean-up.


With DC ZVA bypassed, clear_page() STNP and patched memset() STP showed only
2 statistically significant regressions on AmpereOne machine.

Results for SUT Class aws-m7g.metal:

+-----------+------------------------------------------------+-----------------------+---------------------+--------------------+-------------------------+
| Benchmark | Result Class                                   | next-20260922-vanilla | next-20260922-patch | next-20260922-stnp | next-20260922-patch_stp |
+===========+================================================+=======================+=====================+====================+=========================+
| perf/mem  | memset -k 1GB -f default -s 16GB (bytes/sec)   |        65367617364.09 |              -0.32% |              0.70% |                  -0.80% |
|           | mmap -p 1GB -f demand -s 32GB -l 5 (bytes/sec) |        63750796288.65 |               0.59% |             -1.56% |                  -1.15% |
|           | mmap -p 4KB -f demand -s 32GB -l 5 (bytes/sec) |        49588573632.22 |              -0.39% |             -0.40% |                  -0.93% |
+-----------+------------------------------------------------+-----------------------+---------------------+--------------------+-------------------------+

Results for SUT Class cesw-aarch64-ampereone-1s-a192-32x:

+-----------+------------------------------------------------+-----------------------+---------------------+--------------------+-------------------------+
| Benchmark | Result Class                                   | next-20260922-vanilla | next-20260922-patch | next-20260922-stnp | next-20260922-patch_stp |
+===========+================================================+=======================+=====================+====================+=========================+
| perf/mem  | memset -k 1GB -f default -s 16GB (bytes/sec)   |        35882584735.80 |              -0.36% |              0.43% |                  -0.00% |
|           | mmap -p 1GB -f demand -s 32GB -l 5 (bytes/sec) |        36204572635.77 |              -1.17% |         (R) -2.07% |                   0.14% |
|           | mmap -p 4KB -f demand -s 32GB -l 5 (bytes/sec) |        27149223683.79 |               0.54% |         (R) -2.48% |                  -0.13% |
+-----------+------------------------------------------------+-----------------------+---------------------+--------------------+-------------------------+

Patches
-------

Forced STNP:

diff --git a/arch/arm64/lib/clear_page.S b/arch/arm64/lib/clear_page.S
index bd6f7d5eb6eb6..29796024b3254 100644
--- a/arch/arm64/lib/clear_page.S
+++ b/arch/arm64/lib/clear_page.S
@@ -28,8 +28,7 @@ alternative_else_nop_endif
 	ret
 .Lno_mops:
 #endif
-	mrs	x1, dczid_el0
-	tbnz	x1, #4, 2f	/* Branch if DC ZVA is prohibited */
+	b	2f		/* Test-only: disable DC ZVA. */
 	and	w1, w1, #0xf
 	mov	x2, #4
 	lsl	x1, x2, x1

Forced STP:

diff --git a/arch/arm64/lib/memset.S b/arch/arm64/lib/memset.S
index 97157da65ec6b..9a51511dc57f7 100644
--- a/arch/arm64/lib/memset.S
+++ b/arch/arm64/lib/memset.S
@@ -146,8 +146,7 @@ SYM_FUNC_START_LOCAL(__pi_memset_generic)
 	cmp	count, #128
 	b.lt	.Lnot_short /*count is at least  128 bytes*/
 
-	mrs	tmp1, dczid_el0
-	tbnz	tmp1, #4, .Lnot_short
+	b	.Lnot_short	/* Test-only: disable DC ZVA. */
 	mov	tmp3w, #4
 	and	zva_len, tmp1w, #15	/* Safety: other bits reserved.  */
 	lsl	zva_len, tmp3w, zva_len

-- 
Thanks,
Usama


      reply	other threads:[~2026-10-08 13:31 UTC|newest]

Thread overview: 5+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-09-16 10:03 [PATCH v2] arm64: clear_page[s] using memset Linus Walleij
2026-09-16 15:28 ` Leonardo Bras
2026-09-16 15:54   ` Leonardo Bras
2026-09-23 16:14     ` Catalin Marinas
2026-10-08 13:30       ` Muhammad Usama Anjum [this message]

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=ec298ee5-5080-494d-b735-7ba7bd20cdf0@arm.com \
    --to=usama.anjum@arm.com \
    --cc=catalin.marinas@arm.com \
    --cc=joey.gouly@arm.com \
    --cc=kvmarm@lists.linux.dev \
    --cc=leo.bras@arm.com \
    --cc=linusw@kernel.org \
    --cc=linux-arm-kernel@lists.infradead.org \
    --cc=maz@kernel.org \
    --cc=oupton@kernel.org \
    --cc=suzuki.poulose@arm.com \
    --cc=will@kernel.org \
    --cc=yuzenghui@huawei.com \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox