From mboxrd@z Thu Jan 1 00:00:00 1970 Received: from mail-vk1-f175.google.com (mail-vk1-f175.google.com [209.85.221.175]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id A43FF3A5E67 for ; Thu, 30 Jul 2026 20:07:22 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.221.175 ARC-Seal:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1785442044; cv=none; b=O0fSM5+d50N7ElkIV66iuMoaMGE8wF/kaNAOMPhzc38HK8iE3GKSeDzjycDpr3eqW3f5czh11MK98lufZkxk9nkY87HG6DgRNMspEVDMKValjyKGtqAzRkYtDiZIA0oJxmIX96vONAx8xz1hAwBZFDKD8QH3Y65s3B00pAc/WZk= ARC-Message-Signature:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1785442044; c=relaxed/simple; bh=OQ9pvmLESpXnnDiRajB4/N7WuQHvGk6Y2n//ngFkT6U=; h=Date:From:To:Cc:Subject:Message-ID:References:MIME-Version: Content-Type:Content-Disposition:In-Reply-To; b=Tbypfl9UywSYYN6b3oPStbXAmJZWbypwlhNd0Htw1w6KZ4x+0GlVWYJj7pEMilnBU0bXX7YNBc4pPFBy2fnhvfcVyvJdfn1NdDlkW37/GNz7nwzEfSGBW6zEl29TcyM6PE9CBNdZKy4TIQWf5NsSZ9ycvibQO3W9fZUueSpAUdk= ARC-Authentication-Results:i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com; spf=pass smtp.mailfrom=gmail.com; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b=p0TwxNyI; arc=none smtp.client-ip=209.85.221.175 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=gmail.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b="p0TwxNyI" Received: by mail-vk1-f175.google.com with SMTP id 71dfb90a1353d-5bfdbc68bbeso94535e0c.3 for ; Thu, 30 Jul 2026 13:07:22 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=gmail.com; s=20251104; t=1785442041; x=1786046841; darn=vger.kernel.org; h=in-reply-to:content-disposition:content-type:mime-version :references:message-id:subject:cc:to:from:date:from:to:cc:subject :date:message-id:reply-to:content-type; bh=I/UMoM94HTH0CifsnpKfjDYBiZszucZn3J+xZcqQgUs=; b=p0TwxNyICvLXfv1VUGgTx1V1t2TNKwoG/2weL/NxagfFMKd7wzT8BNjdYIJpO1k/+9 81KNfzGNjdI8HHPcyI6CJrnBcoinutV4rSAnhL5StiRfqgCh8MDHnTC19UjmaSB1YpQt IauhAgeITqIKEoqrtAbgWrVmc4aXzngmzTUi/aGs6Ey7LxBYVXDn34dPTQZ8Ose7ylCC G4lXr+Odmo/puf6143ECCiXClJtaySNw0ncYzb3wK4VmlNdgx7y1YCV8O2Cjss6WCfyk HT6rdOu4/tZb00CNgDAidDC8FE42YmC1M71WOOCYXUt+2rXtrIhahzvsTkxtmYnt0+wY 16eQ== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1785442041; x=1786046841; h=in-reply-to:content-disposition:content-type:mime-version :references:message-id:subject:cc:to:from:date:x-gm-gg :x-gm-message-state:from:to:cc:subject:date:message-id:reply-to :content-type; bh=I/UMoM94HTH0CifsnpKfjDYBiZszucZn3J+xZcqQgUs=; b=n5AExB45fYoaGYU/9Vu1YnhZFhnQEt4YIMwwpuOBYZ4IUA6CThzWOEb3+8Yn1Yod6D qoyR7inzWYvaptUaRmsnxo5k5UmkNThIHvWw+9qKyUXsxDIOZb0UjJdmD7uU89A0tl66 8MM1fRWqH0qaKdmRiVHFKR7ZosV7N9tG47SDlZCvagCDs03CRGl4xIEKM8A0Le5wCScJ 7noXXRXBjfMKmkgEK1wZwWrzWHJ6STgM8SiRowHEZ163a61C92AeWTDZSKmc38aCB22U YTBrBwMmagYUylWKgNiBk7ILir5vZ2V03bCXFnUbjltQV8iVyUquXQBL3S9UKB2CSvC9 j5IQ== X-Forwarded-Encrypted: i=1; AHgh+RqidrN5xfM5g3QC78/MiGNxta2r4YnwThLLEmQYk+E4ihYL1K+BSXVSuq2gEgYxdgYX0ZtwY5jhBoL3LJs=@vger.kernel.org X-Gm-Message-State: AOJu0YwtoqB2sen7Ei3czLD7imJMo942LIs2SSycMCDEs37v0YJp+a/K RMSLSceblElmQF35tkUtxvVKFDpxo+gIvSfahoU6/9Ek+A3SR272PmxVwlCpNoaD X-Gm-Gg: AR+sD10vlfr8OITaqETDCSMDPl7MwrIacS/dFnDuka/+6xeN6g4fmpbEBS81d+FEJye RzZOGGE5JIfVyezeqoNz/gT/RLmwD1/xJYmJvRsz//P2Qaf5ypbWHjYEmNiWfnNpNsuSFYDHGGc rhYCtt33yWcX/u1+uwUsQhd0jx3hajayywxhqENm1OWq01y3ryEn8siwhiUO/32ZRY3Wxa9eV5y xnHxLrWwRl7dk2/fqk+/bnoLSXAvBSKhABBJLFK2+lwgIt7WesIyj+cT8in5z+fJfs460KEwd9H E1kF+TfnYbTzMqkEZ+GkIikABc+q7TFfc2k9ZXJrlzKidXfg1aO8cG4FT+oCGq0cncN3FlD0FWn P05ZgG/2Osxdrksnvw3lE/dwQnvqpPFdxfrvp68RLgQJaXrgOMyHFTw6/Hs2VXl5avT4qgGyZ1h sqiralksK+yJbMQqGc1pM9VhLSO8vQs100V2bBiMhiJ5zDEJWHC6mwSus574yB17tqNqZMf7Dd2 lcnngX7IOzeSgEZdV5FtL1r X-Received: by 2002:a05:6122:e46d:b0:5bd:faa8:74e1 with SMTP id 71dfb90a1353d-5c366276510mr1279552e0c.13.1785442041197; Thu, 30 Jul 2026 13:07:21 -0700 (PDT) Received: from fedora ([2804:30c:1f53:aa00:1495:7d0a:c9fe:c63a]) by smtp.gmail.com with ESMTPSA id 71dfb90a1353d-5c36674e7dbsm2178151e0c.16.2026.07.30.13.07.18 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Thu, 30 Jul 2026 13:07:20 -0700 (PDT) Date: Thu, 30 Jul 2026 17:06:28 -0300 From: Pedro Demarchi Gomes To: "David Hildenbrand (Arm)" Cc: Andrew Morton , Xu Xin , Chengming Zhou , linux-mm@kvack.org, linux-kernel@vger.kernel.org, shr@devkernel.io Subject: Re: [RFC PATCH] mm/ksm: use checksum to speed up page comparison Message-ID: References: <20260716122039.679173-1-pedrodemargomes@gmail.com> <98c91508-4aa7-4390-b2f0-e39f16bdedef@kernel.org> <89e64c5d-f317-413e-9fa3-9b7e0a7f0f5f@kernel.org> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Type: text/plain; charset=utf-8 Content-Disposition: inline In-Reply-To: <89e64c5d-f317-413e-9fa3-9b7e0a7f0f5f@kernel.org> On Thu, Jul 30, 2026 at 04:47:14PM +0200, David Hildenbrand (Arm) wrote: > On 7/30/26 16:22, Pedro Demarchi Gomes wrote: > > On Wed, Jul 29, 2026 at 11:38:38AM +0200, David Hildenbrand (Arm) wrote: > >> On 7/16/26 14:20, Pedro Demarchi Gomes wrote: > >>> Use page checksums as the primary ordering key when traversing the stable > >>> and unstable trees and fall back to memcmp_pages() only when checksums > >>> match. > >>> Since struct ksm_stable_node does not have a checksum field, create one > >>> in a union with migration list fields, so when we encounter a migration > >>> page while scanning an address space we have to recalculate the page > >>> checksum. This avoids increasing the size of struct ksm_stable_node, > >>> which is maintained at 64 bytes, as show below. > >>> > >>> pedro@fedora:~/tmp/linux$ pahole -C ksm_stable_node ./vmlinux > >>> struct ksm_stable_node { > >>> union { > >>> struct { > >>> struct rb_node node __attribute__((__aligned__(8))); /* 0 24 */ > >>> unsigned int checksum; /* 24 4 */ > >>> } __attribute__((__aligned__(8))) __attribute__((__aligned__(8))); /* 0 32 */ > >>> struct { > >>> struct list_head * head; /* 0 8 */ > >>> struct { > >>> struct hlist_node hlist_dup; /* 8 16 */ > >>> struct list_head list; /* 24 16 */ > >>> }; /* 8 32 */ > >>> }; /* 0 40 */ > >>> } __attribute__((__aligned__(8))); /* 0 40 */ > >>> struct hlist_head hlist; /* 40 8 */ > >>> union { > >>> long unsigned int kpfn; /* 48 8 */ > >>> long unsigned int chain_prune_time; /* 48 8 */ > >>> }; /* 48 8 */ > >>> int rmap_hlist_len; /* 56 4 */ > >>> int nid; /* 60 4 */ > >>> > >>> /* size: 64, cachelines: 1, members: 5 */ > >>> /* forced alignments: 1 */ > >>> } __attribute__((__aligned__(8))); > >>> > >>> To evaluate this change it was used two benchmarks, bench1.c and > >>> bench2.c. The first one allocates 8G of pages with different content, > >>> and the second one allocates 8G of pages where the first 4G are the same > >>> as the last 4G. The two benchmarks and the system ksm configuration are > >>> presented below. > >>> > >>> bench1.c: > >>> > >>> int main() { > >>> size_t size = 8ULL * 1024*1024*1024; > >>> unsigned long long int numpages = size/PAGESZ; > >>> char *pages = mmap(NULL, size, PROT_READ | PROT_WRITE, MAP_PRIVATE | MAP_ANONYMOUS, -1, 0); > >>> > >>> // Generate #numpages pages with different contents > >>> for (unsigned long long i = 0; i < numpages; i++) { > >>> *((unsigned long long *) &pages[i*PAGESZ]) = i; > >>> } > >>> > >>> if (madvise(pages, size, MADV_MERGEABLE) != 0) { > >>> perror("madvise MADV_MERGEABLE failed"); > >>> return 1; > >>> } > >>> printf("Wait...\n"); > >>> getchar(); > >>> return 0; > >>> } > >>> > >>> bench2.c: > >>> > >>> int main() { > >>> size_t size = 8ULL * 1024*1024*1024; > >>> unsigned long long int numpages = size/PAGESZ; > >>> char *pages = mmap(NULL, size, PROT_READ | PROT_WRITE, MAP_PRIVATE | MAP_ANONYMOUS, -1, 0); > >>> > >>> // Generate #numpages pages with different contents > >>> for (unsigned long long i = 0; i < numpages/2; i++) { > >>> *((unsigned long long *) &pages[i*PAGESZ]) = i; > >>> *((unsigned long long *) &pages[(numpages-i-1)*PAGESZ]) = i; > >>> } > >>> > >>> if (madvise(pages, size, MADV_MERGEABLE) != 0) { > >>> perror("madvise MADV_MERGEABLE failed"); > >>> return 1; > >>> } > >>> printf("Wait...\n"); > >>> getchar(); > >>> return 0; > >>> } > >>> > >> > >> Hi! > >> > > > > Hi! > > > >>> Configuration: > >>> > >>> echo never > /sys/kernel/mm/transparent_hugepage/enabled > >>> echo 1 > /sys/kernel/mm/ksm/sleep_millisecs > >>> echo 100000 > /sys/kernel/mm/ksm/pages_to_scan > >> > >> Given that the default is 100, and sleep_millisecs is 200 ... and it is known > >> that frequent scanning is harmful for performance, what is the real world impact > >> in common setups? > >> > >> IOW, do we even notice / care? > >> > > > > As described in the KSM admin guide [1], the default values for pages_to_scan > > and sleep_millisecs are intended for demonstration purposes rather than > > production use. > > Well, I argue that > > echo 1 > /sys/kernel/mm/ksm/sleep_millisecs > echo 100000 > /sys/kernel/mm/ksm/pages_to_scan > > is not for production use either? > Yes, this configuration was used only for a quick benchmark to test the RFC idea. > > > > At LPC 2023, Stefan Roesch described the use of KSM in a production workload at > > Meta [2] and presented optimizations such as Smart Scan and Advisor Mode, both > > of which have since been merged to reduce KSM's scanning overhead. > > > > This RFC proposes a complementary optimization. KSM already computes a checksum > > for every scanned page to identify frequently changing pages and avoid > > unnecessary stable and unstable tree lookups. Reusing that checksum as an index > > into those trees can further reduce lookup costs, lowering CPU usage without > > changing KSM's behavior. > > Yes, but the change is not completely trivial, so we better make sure that the > change is actually worth it in practice. > Ok. I am CCing Stefan Roesch to get his opinion on this. Thanks! > -- > Cheers, > > David >