From: "David Hildenbrand (Arm)" <david@kernel.org>
To: Pedro Demarchi Gomes <pedrodemargomes@gmail.com>,
Andrew Morton <akpm@linux-foundation.org>,
Xu Xin <xu.xin16@zte.com.cn>,
Chengming Zhou <chengming.zhou@linux.dev>
Cc: linux-mm@kvack.org, linux-kernel@vger.kernel.org
Subject: Re: [RFC PATCH] mm/ksm: use checksum to speed up page comparison
Date: Wed, 29 Jul 2026 11:38:38 +0200 [thread overview]
Message-ID: <98c91508-4aa7-4390-b2f0-e39f16bdedef@kernel.org> (raw)
In-Reply-To: <20260716122039.679173-1-pedrodemargomes@gmail.com>
On 7/16/26 14:20, Pedro Demarchi Gomes wrote:
> Use page checksums as the primary ordering key when traversing the stable
> and unstable trees and fall back to memcmp_pages() only when checksums
> match.
> Since struct ksm_stable_node does not have a checksum field, create one
> in a union with migration list fields, so when we encounter a migration
> page while scanning an address space we have to recalculate the page
> checksum. This avoids increasing the size of struct ksm_stable_node,
> which is maintained at 64 bytes, as show below.
>
> pedro@fedora:~/tmp/linux$ pahole -C ksm_stable_node ./vmlinux
> struct ksm_stable_node {
> union {
> struct {
> struct rb_node node __attribute__((__aligned__(8))); /* 0 24 */
> unsigned int checksum; /* 24 4 */
> } __attribute__((__aligned__(8))) __attribute__((__aligned__(8))); /* 0 32 */
> struct {
> struct list_head * head; /* 0 8 */
> struct {
> struct hlist_node hlist_dup; /* 8 16 */
> struct list_head list; /* 24 16 */
> }; /* 8 32 */
> }; /* 0 40 */
> } __attribute__((__aligned__(8))); /* 0 40 */
> struct hlist_head hlist; /* 40 8 */
> union {
> long unsigned int kpfn; /* 48 8 */
> long unsigned int chain_prune_time; /* 48 8 */
> }; /* 48 8 */
> int rmap_hlist_len; /* 56 4 */
> int nid; /* 60 4 */
>
> /* size: 64, cachelines: 1, members: 5 */
> /* forced alignments: 1 */
> } __attribute__((__aligned__(8)));
>
> To evaluate this change it was used two benchmarks, bench1.c and
> bench2.c. The first one allocates 8G of pages with different content,
> and the second one allocates 8G of pages where the first 4G are the same
> as the last 4G. The two benchmarks and the system ksm configuration are
> presented below.
>
> bench1.c:
>
> int main() {
> size_t size = 8ULL * 1024*1024*1024;
> unsigned long long int numpages = size/PAGESZ;
> char *pages = mmap(NULL, size, PROT_READ | PROT_WRITE, MAP_PRIVATE | MAP_ANONYMOUS, -1, 0);
>
> // Generate #numpages pages with different contents
> for (unsigned long long i = 0; i < numpages; i++) {
> *((unsigned long long *) &pages[i*PAGESZ]) = i;
> }
>
> if (madvise(pages, size, MADV_MERGEABLE) != 0) {
> perror("madvise MADV_MERGEABLE failed");
> return 1;
> }
> printf("Wait...\n");
> getchar();
> return 0;
> }
>
> bench2.c:
>
> int main() {
> size_t size = 8ULL * 1024*1024*1024;
> unsigned long long int numpages = size/PAGESZ;
> char *pages = mmap(NULL, size, PROT_READ | PROT_WRITE, MAP_PRIVATE | MAP_ANONYMOUS, -1, 0);
>
> // Generate #numpages pages with different contents
> for (unsigned long long i = 0; i < numpages/2; i++) {
> *((unsigned long long *) &pages[i*PAGESZ]) = i;
> *((unsigned long long *) &pages[(numpages-i-1)*PAGESZ]) = i;
> }
>
> if (madvise(pages, size, MADV_MERGEABLE) != 0) {
> perror("madvise MADV_MERGEABLE failed");
> return 1;
> }
> printf("Wait...\n");
> getchar();
> return 0;
> }
>
Hi!
> Configuration:
>
> echo never > /sys/kernel/mm/transparent_hugepage/enabled
> echo 1 > /sys/kernel/mm/ksm/sleep_millisecs
> echo 100000 > /sys/kernel/mm/ksm/pages_to_scan
Given that the default is 100, and sleep_millisecs is 200 ... and it is known
that frequent scanning is harmful for performance, what is the real world impact
in common setups?
IOW, do we even notice / care?
--
Cheers,
David
prev parent reply other threads:[~2026-07-29 9:38 UTC|newest]
Thread overview: 2+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-07-16 12:20 [RFC PATCH] mm/ksm: use checksum to speed up page comparison Pedro Demarchi Gomes
2026-07-29 9:38 ` David Hildenbrand (Arm) [this message]
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=98c91508-4aa7-4390-b2f0-e39f16bdedef@kernel.org \
--to=david@kernel.org \
--cc=akpm@linux-foundation.org \
--cc=chengming.zhou@linux.dev \
--cc=linux-kernel@vger.kernel.org \
--cc=linux-mm@kvack.org \
--cc=pedrodemargomes@gmail.com \
--cc=xu.xin16@zte.com.cn \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is an external index of several public inboxes,
see mirroring instructions on how to clone and mirror
all data and code used by this external index.