Re: [mm/contpte v3 0/1] mm/contpte: Optimize loop to reduce redundant operations

All of lore.kernel.org
 help / color / mirror / Atom feed

From: Catalin Marinas <catalin.marinas@arm.com>
To: Xavier <xavier_qy@163.com>
Cc: ryan.roberts@arm.com, dev.jain@arm.com, ioworker0@gmail.com,
	21cnbao@gmail.com, akpm@linux-foundation.org, david@redhat.com,
	gshan@redhat.com, linux-arm-kernel@lists.infradead.org,
	linux-kernel@vger.kernel.org, will@kernel.org,
	willy@infradead.org, ziy@nvidia.com
Subject: Re: [mm/contpte v3 0/1] mm/contpte: Optimize loop to reduce redundant operations
Date: Wed, 16 Apr 2025 13:47:46 +0100	[thread overview]
Message-ID: <Z_-m8s5EUrL4DAME@arm.com> (raw)
In-Reply-To: <20250415082205.2249918-1-xavier_qy@163.com>

On Tue, Apr 15, 2025 at 04:22:04PM +0800, Xavier wrote:
> Patch V3 has changed the while loop to a for loop according to the suggestions
> of Dev.

For some reason, my email (office365) rejected all these patches (not
even quarantined), I only got the replies. Anyway, I can get them from
the lore archive.

> Meanwhile, to improve efficiency, the definition of local variables has
> been removed. This macro is only used within the current function and there
> will be no additional risks. In order to verify the optimization performance of
> Patch V3, a test function has been designed. By repeatedly calling mlock in a
> loop, the kernel is made to call contpte_ptep_get extensively to test the
> optimization effect of this function.
> The function's execution time and instruction statistics have been traced using
> perf, and the following are the operation results on a certain Qualcomm mobile
> phone chip:
> 
> Instruction Statistics - Before Optimization
> #          count  event_name              # count / runtime
>       20,814,352  branch-load-misses      # 662.244 K/sec
>   41,894,986,323  branch-loads            # 1.333 G/sec
>        1,957,415  iTLB-load-misses        # 62.278 K/sec
>   49,872,282,100  iTLB-loads              # 1.587 G/sec
>      302,808,096  L1-icache-load-misses   # 9.634 M/sec
>   49,872,282,100  L1-icache-loads         # 1.587 G/sec
> 
> Total test time: 31.485237 seconds.
> 
> Instruction Statistics - After Optimization
> #          count  event_name              # count / runtime
>       19,340,524  branch-load-misses      # 688.753 K/sec
>   38,510,185,183  branch-loads            # 1.371 G/sec
>        1,812,716  iTLB-load-misses        # 64.554 K/sec
>   47,673,923,151  iTLB-loads              # 1.698 G/sec
>      675,853,661  L1-icache-load-misses   # 24.068 M/sec
>   47,673,923,151  L1-icache-loads         # 1.698 G/sec
> 
> Total test time: 28.108048 seconds.

We'd need to reproduce these numbers on other platforms as well and with
different page sizes. I hope Ryan can do some tests next week.

Purely looking at the patch, I don't like the complexity. I'd rather go
with your v1 if the numbers are fairly similar (even if slightly slower).

However, I don't trust microbenchmarks like calling mlock() in a loop.
It was hand-crafted to dirty the whole buffer (making ptes young+dirty)
before mlock() to make the best out of the rewritten contpte_ptep_get().
Are there any real world workloads that would benefit from such change?

As it stands, I think this patch needs better justification.

Thanks.

-- 
Catalin

next prev parent reply	other threads:[~2025-04-16 12:49 UTC|newest]

Thread overview: 49+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2025-04-07  9:22 [PATCH v1] mm/contpte: Optimize loop to reduce redundant operations Xavier
2025-04-07 11:29 ` Lance Yang
2025-04-07 12:56   ` Xavier
2025-04-07 13:31     ` Lance Yang
2025-04-07 16:19       ` Dev Jain
2025-04-08  8:58         ` [PATCH v2 0/1] " Xavier
2025-04-08  8:58           ` [PATCH v2 1/1] " Xavier
2025-04-09  4:09             ` Gavin Shan
2025-04-09 15:10               ` Xavier
2025-04-10  0:58                 ` Gavin Shan
2025-04-10  2:53                   ` Xavier
2025-04-10  3:06                     ` Gavin Shan
2025-04-08  9:17         ` [PATCH v1] " Lance Yang
2025-04-09 15:15           ` Xavier
2025-04-10 21:25 ` Barry Song
2025-04-11 12:03   ` David Laight
2025-04-12  7:18     ` Barry Song
2025-04-11 17:30   ` Dev Jain
2025-04-12  5:05     ` Lance Yang
2025-04-12  5:27       ` Xavier
2025-04-14  8:06       ` Ryan Roberts
2025-04-14  8:51         ` Dev Jain
2025-04-14 12:11           ` Ryan Roberts
2025-04-15  8:22             ` [mm/contpte v3 0/1] " Xavier
2025-04-15  8:22               ` [mm/contpte v3 1/1] " Xavier
2025-04-15  9:01                 ` [mm/contpte v3] " Markus Elfring
2025-04-16  8:57                 ` [mm/contpte v3 1/1] " David Laight
2025-04-16 16:15                   ` Xavier
2025-04-16 12:54                 ` Ryan Roberts
2025-04-16 16:09                   ` Xavier
2025-04-22  9:33                   ` Xavier
2025-04-30 23:17                     ` Barry Song
2025-05-01 12:39                       ` Xavier
2025-05-01 21:19                         ` Barry Song
2025-05-01 21:32                           ` Barry Song
2025-05-04  2:39                             ` Xavier
2025-05-08  1:29                               ` Barry Song
2025-05-08  7:03                             ` [PATCH v4] arm64/mm: Optimize loop to reduce redundant operations of contpte_ptep_get Xavier Xia
2025-05-08  8:30                               ` David Hildenbrand
2025-05-09  9:17                                 ` Xavier
2025-05-09  9:25                                   ` David Hildenbrand
2025-05-09  2:09                               ` Barry Song
2025-05-09  9:20                                 ` Xavier
2025-04-16  2:10               ` [mm/contpte v3 0/1] mm/contpte: Optimize loop to reduce redundant operations Andrew Morton
2025-04-16  3:25                 ` Xavier
2025-04-16 12:47               ` Catalin Marinas [this message]
2025-04-16 15:08                 ` Xavier
2025-04-16 12:48               ` Ryan Roberts
2025-04-16 15:22                 ` Xavier

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=Z_-m8s5EUrL4DAME@arm.com \
    --to=catalin.marinas@arm.com \
    --cc=21cnbao@gmail.com \
    --cc=akpm@linux-foundation.org \
    --cc=david@redhat.com \
    --cc=dev.jain@arm.com \
    --cc=gshan@redhat.com \
    --cc=ioworker0@gmail.com \
    --cc=linux-arm-kernel@lists.infradead.org \
    --cc=linux-kernel@vger.kernel.org \
    --cc=ryan.roberts@arm.com \
    --cc=will@kernel.org \
    --cc=willy@infradead.org \
    --cc=xavier_qy@163.com \
    --cc=ziy@nvidia.com \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link

Be sure your reply has a Subject: header at the top and a blank line before the message body.

This is an external index of several public inboxes,
see mirroring instructions on how to clone and mirror
all data and code used by this external index.