From mboxrd@z Thu Jan 1 00:00:00 1970 Received: from mail-ua1-f49.google.com (mail-ua1-f49.google.com [209.85.222.49]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 5BD833A5433 for ; Thu, 30 Jul 2026 14:23:18 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.222.49 ARC-Seal:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1785421399; cv=none; b=WhWbQHJOlWFMJV3oBoYArQc8o4A1iR4lib+mFESXMkvfKD9aW6qkcKwXCCzXNNZAwkyK/TNTlGaXILOnpl95agFD1sNv31GWqu2E7BOsLJlWX6jbkKrqeE86XPA5OKTffJqS33vQ1W33e/gNfq3vJjSu8EscMr8XKVu/RYWvuZ0= ARC-Message-Signature:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1785421399; c=relaxed/simple; bh=8TaOH++ZmXhN2HMt1hbHhOyLeyaiE5urweGmVtcx7jY=; h=Date:From:To:Cc:Subject:Message-ID:References:MIME-Version: Content-Type:Content-Disposition:In-Reply-To; b=dCVoADUyoNRq1ZgPmVkVwj1bcOnytqnfcdBDvbpiNgfggDMKEtmnayYikyPoXKa23DDY/dVGEqWLYzy5RfcgpqPAV2R6zMoALq3EXgiRjtu6QNEHJIzzqidURKJgfX1WUeoP8SFA/gavvGEDSe4/82FC9+jKQaZQEzzchz6hMno= ARC-Authentication-Results:i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com; spf=pass smtp.mailfrom=gmail.com; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b=YfTRZVer; arc=none smtp.client-ip=209.85.222.49 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=gmail.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b="YfTRZVer" Received: by mail-ua1-f49.google.com with SMTP id a1e0cc1a2514c-9693bbb962eso1220076241.2 for ; Thu, 30 Jul 2026 07:23:18 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=gmail.com; s=20251104; t=1785421397; x=1786026197; darn=vger.kernel.org; h=in-reply-to:content-disposition:content-type:mime-version :references:message-id:subject:cc:to:from:date:from:to:cc:subject :date:message-id:reply-to:content-type; bh=G74EfrGlVf7uX6WfKSZkQJO5gHzdmhMpULqAQUJlO/A=; b=YfTRZVerAGI7/ikkoy+ZnMVcln819x2Gku2KMS0rTFNPNJvhE7HjaE1GbI/Y3u9eVZ XbJtG7tW7YL68VczsRaLOKvRvZ5++OHX/Fkv48OxhKPXFXjrGkE20k6EyrfJdiunZlOC e98Wm9o56UPaUT0nSuuprxP5lTqxtElaZpiltqRD6P5DOwnSIBbgfjND2IPJ4KYECI4U gZxoYGeCwtD+Y6T+zhTa8JpJW8yE+IyW/6fvoKH+PgxLQVPDX7p0kOT1GW74DkdWXKDm YjWIS1rKR+oDDQMTWFxU6iJ+CqhpAXWwcBToiK/AbErCELY8sqK5CdzCt3Pg7g648MoB A6og== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1785421397; x=1786026197; h=in-reply-to:content-disposition:content-type:mime-version :references:message-id:subject:cc:to:from:date:x-gm-gg :x-gm-message-state:from:to:cc:subject:date:message-id:reply-to :content-type; bh=G74EfrGlVf7uX6WfKSZkQJO5gHzdmhMpULqAQUJlO/A=; b=aUU5o9Y8FGl7iO5N+JLkbGngpBsnPmwsu7SHXZTimUBleGZc267wVeHISdfKZ6Mj6A fspsPvEt8mTQQ+0cSFk3fJOU8zfQj7v3S9BzEMF55uBQn+crS9pIIHxgeolWJqPIaknd wnjJEfN6yibLR5UbURHUX65aBcg9CLjkPcJCW5UaixjIToqctKqNegL8xWgQgaUo6vtq nFJRoriNORAd8fX14pGDmJYtgnsVlXcYqfxNQh9Ro37MM1y+4qJXzTy9eGYmpjYwdkor f/JsmNpp5rbBxMjx1wtJyQzBPbeEr7M8xMlV51XLoghW0FFSN439wI88AgviVxv3M1uf ODug== X-Forwarded-Encrypted: i=1; AHgh+RoxE1bFDJkxCe4PhmAixj/nfpM0x5yG99ejOB3L1EI3CSx1ttJUJ1jzHg938mfmtURwif6jx2wgsCWvbz4=@vger.kernel.org X-Gm-Message-State: AOJu0YxyOyrZmBVYcIFbjXfrZYjyGnvyf69Oc/pdL3RTm2b7LVQ4VHyB lZeOA+He04jJFPBSyDK/mIqx20+6zW9qXJz4e4zAzVUpTkhzrqgrxw1T X-Gm-Gg: AR+sD12p90npmgAzqDhNAkIr4563dI2Tz2B/ctesfoXIetYsLOmbIsnJ/jOZEXJ7Ekz tW7WQI/szKc9Fd4wFEHxw3FXUYUMJ2y8TWsE3SfUf52kE4ey6yp5Awn2dQ3lxU9tYiK2/f+t2pi CWA9/YuH4J6SpCYUG3VnzNt0d4v3sJq6fYAjR05nGXZPG2L4p6edBVR291tJjFqqiCfnSLN/rxW 2B5El9dzUSHq79FhVM8WSiltN5IB9mZHS0ENN8tttno20qknVYXKlGNeG5Ma+e9xZRVkydCRPyw WKkEHLdF5k/Nf54j1WZfhZFypXDiXg62/RMkDZAGnFZmfs8TfPa1EjunrMGPzHFUN36NC5KgxZq t9Xc08EijqJqW8K6SnhPsvKTIDvXDvNqT93eIHvrg52c340ZzlSLr7TA6vNzyFc75Qw2z4xg8Pi U84SUK/0U3G/B+dr3JTIe+dCDTBPIckjG+SUw1MuIabzgNaP4HBuB0dX6dNPzCHv9eo4WGnOvRW oiD5Wp8UJTV+fXf6arCWN7wLJ7LnTUFUTKH X-Received: by 2002:a05:6102:304f:b0:737:4cac:52f7 with SMTP id ada2fe7eead31-75750171354mr1067997137.19.1785421397113; Thu, 30 Jul 2026 07:23:17 -0700 (PDT) Received: from fedora ([2804:30c:1f53:aa00:1495:7d0a:c9fe:c63a]) by smtp.gmail.com with ESMTPSA id a1e0cc1a2514c-977cddf7ba2sm1262117241.13.2026.07.30.07.23.14 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Thu, 30 Jul 2026 07:23:16 -0700 (PDT) Date: Thu, 30 Jul 2026 11:22:25 -0300 From: Pedro Demarchi Gomes To: "David Hildenbrand (Arm)" Cc: Andrew Morton , Xu Xin , Chengming Zhou , linux-mm@kvack.org, linux-kernel@vger.kernel.org Subject: Re: [RFC PATCH] mm/ksm: use checksum to speed up page comparison Message-ID: References: <20260716122039.679173-1-pedrodemargomes@gmail.com> <98c91508-4aa7-4390-b2f0-e39f16bdedef@kernel.org> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Type: text/plain; charset=utf-8 Content-Disposition: inline In-Reply-To: <98c91508-4aa7-4390-b2f0-e39f16bdedef@kernel.org> On Wed, Jul 29, 2026 at 11:38:38AM +0200, David Hildenbrand (Arm) wrote: > On 7/16/26 14:20, Pedro Demarchi Gomes wrote: > > Use page checksums as the primary ordering key when traversing the stable > > and unstable trees and fall back to memcmp_pages() only when checksums > > match. > > Since struct ksm_stable_node does not have a checksum field, create one > > in a union with migration list fields, so when we encounter a migration > > page while scanning an address space we have to recalculate the page > > checksum. This avoids increasing the size of struct ksm_stable_node, > > which is maintained at 64 bytes, as show below. > > > > pedro@fedora:~/tmp/linux$ pahole -C ksm_stable_node ./vmlinux > > struct ksm_stable_node { > > union { > > struct { > > struct rb_node node __attribute__((__aligned__(8))); /* 0 24 */ > > unsigned int checksum; /* 24 4 */ > > } __attribute__((__aligned__(8))) __attribute__((__aligned__(8))); /* 0 32 */ > > struct { > > struct list_head * head; /* 0 8 */ > > struct { > > struct hlist_node hlist_dup; /* 8 16 */ > > struct list_head list; /* 24 16 */ > > }; /* 8 32 */ > > }; /* 0 40 */ > > } __attribute__((__aligned__(8))); /* 0 40 */ > > struct hlist_head hlist; /* 40 8 */ > > union { > > long unsigned int kpfn; /* 48 8 */ > > long unsigned int chain_prune_time; /* 48 8 */ > > }; /* 48 8 */ > > int rmap_hlist_len; /* 56 4 */ > > int nid; /* 60 4 */ > > > > /* size: 64, cachelines: 1, members: 5 */ > > /* forced alignments: 1 */ > > } __attribute__((__aligned__(8))); > > > > To evaluate this change it was used two benchmarks, bench1.c and > > bench2.c. The first one allocates 8G of pages with different content, > > and the second one allocates 8G of pages where the first 4G are the same > > as the last 4G. The two benchmarks and the system ksm configuration are > > presented below. > > > > bench1.c: > > > > int main() { > > size_t size = 8ULL * 1024*1024*1024; > > unsigned long long int numpages = size/PAGESZ; > > char *pages = mmap(NULL, size, PROT_READ | PROT_WRITE, MAP_PRIVATE | MAP_ANONYMOUS, -1, 0); > > > > // Generate #numpages pages with different contents > > for (unsigned long long i = 0; i < numpages; i++) { > > *((unsigned long long *) &pages[i*PAGESZ]) = i; > > } > > > > if (madvise(pages, size, MADV_MERGEABLE) != 0) { > > perror("madvise MADV_MERGEABLE failed"); > > return 1; > > } > > printf("Wait...\n"); > > getchar(); > > return 0; > > } > > > > bench2.c: > > > > int main() { > > size_t size = 8ULL * 1024*1024*1024; > > unsigned long long int numpages = size/PAGESZ; > > char *pages = mmap(NULL, size, PROT_READ | PROT_WRITE, MAP_PRIVATE | MAP_ANONYMOUS, -1, 0); > > > > // Generate #numpages pages with different contents > > for (unsigned long long i = 0; i < numpages/2; i++) { > > *((unsigned long long *) &pages[i*PAGESZ]) = i; > > *((unsigned long long *) &pages[(numpages-i-1)*PAGESZ]) = i; > > } > > > > if (madvise(pages, size, MADV_MERGEABLE) != 0) { > > perror("madvise MADV_MERGEABLE failed"); > > return 1; > > } > > printf("Wait...\n"); > > getchar(); > > return 0; > > } > > > > Hi! > Hi! > > Configuration: > > > > echo never > /sys/kernel/mm/transparent_hugepage/enabled > > echo 1 > /sys/kernel/mm/ksm/sleep_millisecs > > echo 100000 > /sys/kernel/mm/ksm/pages_to_scan > > Given that the default is 100, and sleep_millisecs is 200 ... and it is known > that frequent scanning is harmful for performance, what is the real world impact > in common setups? > > IOW, do we even notice / care? > As described in the KSM admin guide [1], the default values for pages_to_scan and sleep_millisecs are intended for demonstration purposes rather than production use. At LPC 2023, Stefan Roesch described the use of KSM in a production workload at Meta [2] and presented optimizations such as Smart Scan and Advisor Mode, both of which have since been merged to reduce KSM's scanning overhead. This RFC proposes a complementary optimization. KSM already computes a checksum for every scanned page to identify frequently changing pages and avoid unnecessary stable and unstable tree lookups. Reusing that checksum as an index into those trees can further reduce lookup costs, lowering CPU usage without changing KSM's behavior. [1] https://docs.kernel.org/admin-guide/mm/ksm.html [2] https://lpc.events/event/17/contributions/1625/ > -- > Cheers, > > David >