From: Andrew Morton <akpm@linux-foundation.org>
To: mm-commits@vger.kernel.org,ziy@nvidia.com,willy@infradead.org,will@kernel.org,ryan.roberts@arm.com,Markus.Elfring@web.de,ioworker0@gmail.com,gshan@redhat.com,dev.jain@arm.com,david@redhat.com,catalin.marinas@arm.com,21cnbao@gmail.com,xavier_qy@163.com,akpm@linux-foundation.org
Subject: [to-be-updated] mm-contpte-optimize-loop-to-reduce-redundant-operations.patch removed from -mm tree
Date: Wed, 16 Apr 2025 15:35:21 -0700 [thread overview]
Message-ID: <20250416223522.8175CC4CEE2@smtp.kernel.org> (raw)
The quilt patch titled
Subject: mm/contpte: optimize loop to reduce redundant operations
has been removed from the -mm tree. Its filename was
mm-contpte-optimize-loop-to-reduce-redundant-operations.patch
This patch was dropped because an updated version will be issued
------------------------------------------------------
From: Xavier <xavier_qy@163.com>
Subject: mm/contpte: optimize loop to reduce redundant operations
Date: Tue, 15 Apr 2025 16:22:05 +0800
Optimize contpte_ptep_get() by adding early termination logic. Check if
the dirty and young bits of orig_pte are already set and skip redundant
bit-setting operations during the loop. This reduces unnecessary
iterations and improves performance.
The function's execution time and instruction statistics have been traced
using perf, and the following are the operation results on a certain
Qualcomm mobile phone chip:
Instruction Statistics - Before Optimization
# count event_name # count / runtime
20,814,352 branch-load-misses # 662.244 K/sec
41,894,986,323 branch-loads # 1.333 G/sec
1,957,415 iTLB-load-misses # 62.278 K/sec
49,872,282,100 iTLB-loads # 1.587 G/sec
302,808,096 L1-icache-load-misses # 9.634 M/sec
49,872,282,100 L1-icache-loads # 1.587 G/sec
Total test time: 31.485237 seconds.
Instruction Statistics - After Optimization
# count event_name # count / runtime
19,340,524 branch-load-misses # 688.753 K/sec
38,510,185,183 branch-loads # 1.371 G/sec
1,812,716 iTLB-load-misses # 64.554 K/sec
47,673,923,151 iTLB-loads # 1.698 G/sec
675,853,661 L1-icache-load-misses # 24.068 M/sec
47,673,923,151 L1-icache-loads # 1.698 G/sec
Total test time: 28.108048 seconds.
Function Statistics - Before Optimization
Arch: arm64
Event: cpu-cycles (type 0, config 0)
Samples: 1419716
Event count: 99618088900
Overhead Symbol
21.42% lock_release
21.26% lock_acquire
20.88% arch_counter_get_cntvct
14.32% _raw_spin_unlock_irq
6.79% contpte_ptep_get
2.20% test_contpte_perf
1.82% follow_page_pte
0.97% lock_acquired
0.97% rcu_is_watching
0.89% mlock_pte_range
0.84% sched_clock_noinstr
0.70% handle_softirqs.llvm.8218488130471452153
0.58% test_preempt_disable_long
0.57% _raw_spin_unlock_irqrestore
0.54% arch_stack_walk
0.51% vm_normal_folio
0.48% check_preemption_disabled
0.47% stackinfo_get_task
0.36% try_grab_folio
0.34% preempt_count
0.32% trace_preempt_on
0.29% trace_preempt_off
0.24% debug_smp_processor_id
Function Statistics - After Optimization
Arch: arm64
Event: cpu-cycles (type 0, config 0)
Samples: 1431006
Event count: 118856425042
Overhead Symbol
22.59% lock_release
22.13% arch_counter_get_cntvct
22.08% lock_acquire
15.32% _raw_spin_unlock_irq
2.26% test_contpte_perf
1.50% follow_page_pte
1.49% arch_stack_walk
1.30% rcu_is_watching
1.09% lock_acquired
1.07% sched_clock_noinstr
0.88% handle_softirqs.llvm.12507768597002095717
0.88% trace_preempt_off
0.76% _raw_spin_unlock_irqrestore
0.61% check_preemption_disabled
0.52% trace_preempt_on
0.50% mlock_pte_range
0.43% try_grab_folio
0.41% folio_mark_accessed
0.40% vm_normal_folio
0.38% test_preempt_disable_long
0.28% contpte_ptep_get
0.27% __traceiter_android_rvh_preempt_disable
0.26% debug_smp_processor_id
0.24% return_address
0.20% __pte_offset_map_lock
0.19% unwind_next_frame_record
If there is no problem with my test program, it can be seen that there is a
significant performance improvement both in the overall number of instructions
and the execution time of contpte_ptep_get.
If any reviewers have time, you can also test it on your machines for comparison.
I have enabled THP and hugepages-64kB.
Test function:
#define PAGE_SIZE 4096
#define CONT_PTES 16
#define TEST_SIZE (4096* CONT_PTES * PAGE_SIZE)
void rwdata(char *buf)
{
for (size_t i = 0; i < TEST_SIZE; i += PAGE_SIZE) {
buf[i] = 'a';
volatile char c = buf[i];
}
}
void test_contpte_perf()
{
char *buf;
int ret = posix_memalign((void **)&buf, PAGE_SIZE, TEST_SIZE);
if (ret != 0) {
perror("posix_memalign failed");
exit(EXIT_FAILURE);
}
rwdata(buf);
for (int j = 0; j < 500; j++) {
mlock(buf, TEST_SIZE);
rwdata(buf);
munlock(buf, TEST_SIZE);
}
free(buf);
}
Link: https://lkml.kernel.org/r/20250415082205.2249918-2-xavier_qy@163.com
Signed-off-by: Xavier <xavier_qy@163.com>
Cc: Barry Song <21cnbao@gmail.com>
Cc: Catalin Marinas <catalin.marinas@arm.com>
Cc: David Hildenbrand <david@redhat.com>
Cc: Dev Jain <dev.jain@arm.com>
Cc: Gavin Shan <gshan@redhat.com>
Cc: Lance Yang <ioworker0@gmail.com>
Cc: Matthew Wilcox (Oracle) <willy@infradead.org>
Cc: Ryan Roberts <ryan.roberts@arm.com>
Cc: Will Deacon <will@kernel.org>
Cc: Zi Yan <ziy@nvidia.com>
Cc: Markus Elfring <Markus.Elfring@web.de>
Signed-off-by: Andrew Morton <akpm@linux-foundation.org>
---
arch/arm64/mm/contpte.c | 20 ++++++++++++++++++--
1 file changed, 18 insertions(+), 2 deletions(-)
--- a/arch/arm64/mm/contpte.c~mm-contpte-optimize-loop-to-reduce-redundant-operations
+++ a/arch/arm64/mm/contpte.c
@@ -152,6 +152,16 @@ void __contpte_try_unfold(struct mm_stru
}
EXPORT_SYMBOL_GPL(__contpte_try_unfold);
+/* Note: in order to improve efficiency, using this macro will modify the
+ * passed-in parameters.*/
+#define CHECK_CONTPTE_FLAG(start, ptep, orig_pte, flag) \
+ for (; (start) < CONT_PTES; (start)++, (ptep)++) { \
+ if (pte_##flag(__ptep_get(ptep))) { \
+ orig_pte = pte_mk##flag(orig_pte); \
+ break; \
+ } \
+ }
+
pte_t contpte_ptep_get(pte_t *ptep, pte_t orig_pte)
{
/*
@@ -169,11 +179,17 @@ pte_t contpte_ptep_get(pte_t *ptep, pte_
for (i = 0; i < CONT_PTES; i++, ptep++) {
pte = __ptep_get(ptep);
- if (pte_dirty(pte))
+ if (pte_dirty(pte)) {
orig_pte = pte_mkdirty(orig_pte);
+ CHECK_CONTPTE_FLAG(i, ptep, orig_pte, young);
+ break;
+ }
- if (pte_young(pte))
+ if (pte_young(pte)) {
orig_pte = pte_mkyoung(orig_pte);
+ CHECK_CONTPTE_FLAG(i, ptep, orig_pte, dirty);
+ break;
+ }
}
return orig_pte;
_
Patches currently in -mm which might be from xavier_qy@163.com are
reply other threads:[~2025-04-16 22:35 UTC|newest]
Thread overview: [no followups] expand[flat|nested] mbox.gz Atom feed
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=20250416223522.8175CC4CEE2@smtp.kernel.org \
--to=akpm@linux-foundation.org \
--cc=21cnbao@gmail.com \
--cc=Markus.Elfring@web.de \
--cc=catalin.marinas@arm.com \
--cc=david@redhat.com \
--cc=dev.jain@arm.com \
--cc=gshan@redhat.com \
--cc=ioworker0@gmail.com \
--cc=mm-commits@vger.kernel.org \
--cc=ryan.roberts@arm.com \
--cc=will@kernel.org \
--cc=willy@infradead.org \
--cc=xavier_qy@163.com \
--cc=ziy@nvidia.com \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is an external index of several public inboxes,
see mirroring instructions on how to clone and mirror
all data and code used by this external index.