From mboxrd@z Thu Jan 1 00:00:00 1970 Received: from foss.arm.com (foss.arm.com [217.140.110.172]) by smtp.subspace.kernel.org (Postfix) with ESMTP id 88682313293 for ; Thu, 23 Jul 2026 07:09:20 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=217.140.110.172 ARC-Seal:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1784790562; cv=none; b=rYEFPe+MW3sDukM+NWeW6jsiKBpz4MHSwA0mGu3/ayhGI7rjU+s7A7qCPKRUGBkE4gzlnxqfvxjFK7ZrdCx/spu4QBsr2lL2orGricyvtZ2hcVxBX/f30tfQkvGW54Py/IQlsy5UHFGr0JJnO0kBiJB/3SpST1BYjhy3ngtDvkY= ARC-Message-Signature:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1784790562; c=relaxed/simple; bh=r9EfMEmAfxVooZvdxkOc/5sBU7O5MNGAs6jqD165yJM=; h=From:To:Cc:Subject:Date:Message-ID:MIME-Version; b=k5NN7CMA5OkHk5ACVbG3+FgGBtIITnvBsAew75W2c8RCCcdibvfHW/6+6HcxImsz+4ylxQAE3Ns75m6ZyVXLCHB2KaA1MWe0BWgcTUBKDNqwlCU86Srxj4WwZPGzUyR/oWhd0ifK2J6TALeYpJe8a4pQ9q0DitNlbvIFyk0xwXI= ARC-Authentication-Results:i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=arm.com; spf=pass smtp.mailfrom=arm.com; dkim=pass (1024-bit key) header.d=arm.com header.i=@arm.com header.b=ou8rjE6e; arc=none smtp.client-ip=217.140.110.172 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=arm.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=arm.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (1024-bit key) header.d=arm.com header.i=@arm.com header.b="ou8rjE6e" Received: from usa-sjc-imap-foss1.foss.arm.com (unknown [10.121.207.14]) by usa-sjc-mx-foss1.foss.arm.com (Postfix) with ESMTP id 6030A1477; Thu, 23 Jul 2026 00:09:15 -0700 (PDT) Received: from cesw-amp-gbt-1s-m12830-01.blr.arm.com (cesw-amp-gbt-1s-m12830-01.blr.arm.com [10.164.195.33]) by usa-sjc-imap-foss1.foss.arm.com (Postfix) with ESMTPA id A2C123F66F; Thu, 23 Jul 2026 00:09:12 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=simple/simple; d=arm.com; s=foss; t=1784790559; bh=r9EfMEmAfxVooZvdxkOc/5sBU7O5MNGAs6jqD165yJM=; h=From:To:Cc:Subject:Date:From; b=ou8rjE6eBVg9yglueWD2KRr0100k416BiZ9gZo3ZXKNeqUtt4hov0ULKqVZqKepcG OnpuRIl/wgswEOFImGUU8U3ygWlG/Y2WL+OEvwP4qycfCrQSeT/LyklTGQE5yVABNC 7WAwtFWqAIo1tTscUAiy7qK3Ma8dfOKwV3lH3/qs= From: Dev Jain To: akpm@linux-foundation.org, david@kernel.org, ljs@kernel.org, hughd@google.com, chrisl@kernel.org, kasong@tencent.com Cc: Dev Jain , riel@surriel.com, liam@infradead.org, vbabka@kernel.org, harry@kernel.org, jannh@google.com, lance.yang@linux.dev, baolin.wang@linux.alibaba.com, shikemeng@huaweicloud.com, nphamcs@gmail.com, baoquan.he@linux.dev, baohua@kernel.org, youngjun.park@lge.com, linux-mm@kvack.org, linux-kernel@vger.kernel.org, rppt@kernel.org, surenb@google.com, mhocko@suse.com, pfalcato@suse.de, ryan.roberts@arm.com, anshuman.khandual@arm.com Subject: [PATCH 0/8] Optimize anonymous swapbacked large folio unmapping Date: Thu, 23 Jul 2026 07:08:56 +0000 Message-ID: <20260723070905.3422276-1-dev.jain@arm.com> X-Mailer: git-send-email 2.43.0 Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: 8bit Speed up unmapping of anonymous swapbacked large folios by clearing the ptes, and setting swap ptes, in one go. The following benchmark (stolen from Barry) is used to measure the time taken to swapout 256M worth of memory backed by 64K large folios: #define _GNU_SOURCE #include #include #include #include #include #include #include #define SIZE_MB 256 #define SIZE_BYTES (SIZE_MB * 1024 * 1024) int main() { void *addr = mmap(NULL, SIZE_BYTES, PROT_READ | PROT_WRITE, MAP_PRIVATE | MAP_ANONYMOUS, -1, 0); if (addr == MAP_FAILED) { perror("mmap failed"); return 1; } memset(addr, 0, SIZE_BYTES); struct timespec start, end; clock_gettime(CLOCK_MONOTONIC, &start); if (madvise(addr, SIZE_BYTES, MADV_PAGEOUT) != 0) { perror("madvise(MADV_PAGEOUT) failed"); munmap(addr, SIZE_BYTES); return 1; } clock_gettime(CLOCK_MONOTONIC, &end); long duration_ns = (end.tv_sec - start.tv_sec) * 1e9 + (end.tv_nsec - start.tv_nsec); printf("madvise(MADV_PAGEOUT) took %ld ns (%.3f ms)\n", duration_ns, duration_ns / 1e6); munmap(addr, SIZE_BYTES); return 0; } Performance as measured on a Linux VM on Apple M3 (arm64): Vanilla - Mean: 37401913 ns, std dev: 12% Patched - Mean: 17420282 ns, std dev: 11% resulting in more than 2x speedup. No regression observed on 4K folios. Performance as measured on bare metal x86: Vanilla - mean: 54986286 ns, std dev: 1.5% Patched - mean: 51930795 ns, std dev: 3% I tried magnifying the difference on x86 by using 1M large folios, but can't spot an obvious improvement (looks like my system is too fast to benefit from batched atomic operations!), hinting that the benefit lies mainly in the reduction of ptep_get() calls and the reduction of TLB flushes during contpte-unfolding, on arm64. No regression is observed on 4K folios on x86 too. --- Applies on mm-new (3d18f3499c48). mm-selftests pass. This breakout patchset is the final completion of https://lore.kernel.org/all/20260526063635.61721-1-dev.jain@arm.com/ The extra addition is the generic set_softleaf_ptes() helper. Dev Jain (8): mm/swapfile: add batched version of folio_dup_swap mm/swapfile: add batched version of folio_put_swap mm/rmap: mm/rmap: Add batched version of folio_try_share_anon_rmap_pte mm/internal: rename swap offset helpers to softleaf offset mm/internal: add set_softleaf_ptes mm/memory: use set_softleaf_ptes for uffd-wp markers mm: move anon-exclusive batch helper to internal.h mm/rmap: batch unmap anonymous swap-backed large folios include/linux/rmap.h | 51 +++++++++++++------- mm/internal.h | 79 +++++++++++++++++++++++++------ mm/memory.c | 20 +++----- mm/mprotect.c | 17 ------- mm/rmap.c | 109 +++++++++++++++++++++++++++++++------------ mm/shmem.c | 8 ++-- mm/swap.h | 35 ++++++++++++-- mm/swapfile.c | 31 ++++++------ 8 files changed, 235 insertions(+), 115 deletions(-) -- 2.43.0