From mboxrd@z Thu Jan 1 00:00:00 1970 Received: from smtp.kernel.org (aws-us-west-2-korg-mail-alma10-1.taild15c8.ts.net [100.103.45.18]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id CC4253E40E7; Tue, 8 Sep 2026 20:13:24 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=100.103.45.18 ARC-Seal:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1788898406; cv=none; b=a3wiqJX2rJnPEG2m46Gpu24Dbn9RpdZhIhBnC0VT7DYgqjiHX6Mf6wYU4XB9vtKf6G//19vwTjrhNGcAkGgyT7L2C3x/v1dwKUbiyBivN4suFCZY0vE3iag9CfHNAQo3uoTeKzftGcDplqB2RVktLwC2XE+ndOfI6UR6iHu57cU= ARC-Message-Signature:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1788898406; c=relaxed/simple; bh=vEqZNj/orhd3A72a0KyhZ8lk/+VHtX68aC52ZbOqwFo=; h=From:Date:Subject:MIME-Version:Content-Type:Message-Id:References: In-Reply-To:To:Cc; b=NcXJB6L0DmMJ5GgsiRQIuriX2jZxJcMNt95RpjuAOcy3RxwZ0bhTGbBlwRcZa+lJl8gwJNZ5gjaJWpwfgSoabbl4IVCkKh0haqyXVRWA3a9FXZNFFqfn9crifID0UnBrsayVVecNY6D3SA1aL7P07o/kuSanUIBl43OdSF0Gzpc= ARC-Authentication-Results:i=1; smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b=Fl/YCDPj; arc=none smtp.client-ip=100.103.45.18 Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b="Fl/YCDPj" Received: by smtp.kernel.org (Postfix) with ESMTPSA id 799EB1F00A3A; Tue, 8 Sep 2026 20:12:56 +0000 (UTC) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=kernel.org; s=k20260515; t=1788898404; bh=5K16GlzsQoy0h7+0CkyHUMkjHsarATEzCzpeYhQPaCY=; h=From:Date:Subject:References:In-Reply-To:To:Cc; b=Fl/YCDPjRM8WBFFd2e8aQAOkcjORUNEDvBkYPZ0czOwSoKB9JUkzPuLgU9/uPT+28 HH9oU0LFFmBARTFN9q3OvmwkkSEllMdPfkkJ9Y0U4lcd9dKI721He/knHyzy2TvflO /ou7DeEyJdwsk2wtsr6GfTfLEqwo/CfTcJ8EWEmtspNdXFBue7A1s/ztnSI4zxOZgM Nwfu2DysRSUinri8xi4SNPc7I2bJJhyUcL7w8lvFQJemkasfX86nje5c0yC+QmHl2H B9q5ltTeHhkkn7x4cw0Ot5AdEHPEyJ29dkPxUp/lKTqSNFoO8AzSgt/f+h+Kc06hgL bU7IDnReUhc5g== From: "Lorenzo Stoakes (ARM)" Date: Tue, 08 Sep 2026 21:01:27 +0100 Subject: [PATCH 23/39] mm/mlock: eliminate weird VMA_IO_BIT abuse and simplify Precedence: bulk X-Mailing-List: linux-trace-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Type: text/plain; charset="utf-8" Content-Transfer-Encoding: 7bit Message-Id: <20260908-b4-mmap-prepare-vma-flag-sanify-v1-23-dacf19cce22b@kernel.org> References: <20260908-b4-mmap-prepare-vma-flag-sanify-v1-0-dacf19cce22b@kernel.org> In-Reply-To: <20260908-b4-mmap-prepare-vma-flag-sanify-v1-0-dacf19cce22b@kernel.org> To: Andrew Morton , "Liam R. Howlett" , Vlastimil Babka , Jann Horn , Pedro Falcato , David Hildenbrand , Mike Rapoport , Suren Baghdasaryan , Michal Hocko , Jonathan Corbet , Greg Kroah-Hartman , Dennis Dalessandro , Jason Gunthorpe , Leon Romanovsky , Paul Moore , Stephen Smalley , Jaroslav Kysela , Takashi Iwai , Alexei Starovoitov , Daniel Borkmann , Andrii Nakryiko , Eduard Zingerman , Kumar Kartikeya Dwivedi , Zi Yan , Baolin Wang , Nico Pache , Ryan Roberts , Dev Jain , Barry Song , Lance Yang , Usama Arif , Kiryl Shutsemau , Doug Gilbert , "James E.J. Bottomley" , "Martin K. Petersen" , Jaya Kumar , Simona Vetter , Helge Deller , Sebastian Reichel , John Hubbard , Peter Xu , Masami Hiramatsu , Oleg Nesterov , Peter Zijlstra , Thomas Gleixner , Ingo Molnar , Borislav Petkov , Dave Hansen , x86@kernel.org, Arnaldo Carvalho de Melo , Namhyung Kim , Mark Rutland , Rik van Riel , Harry Yoo , Juri Lelli , Vincent Guittot , Maarten Lankhorst , Maxime Ripard , Thomas Zimmermann , David Airlie , Will Deacon , "Aneesh Kumar K.V" , Nick Piggin , Arnd Bergmann , Muchun Song , Oscar Salvador , "Matthew Wilcox (Oracle)" , Jan Kara , Marc Zyngier , Oliver Upton , Catalin Marinas , Madhavan Srinivasan , Anup Patel , Paul Walmsley , Palmer Dabbelt , Albert Ou , Christian Borntraeger , Janosch Frank , Claudio Imbrenda , Alexander Gordeev , Gerald Schaefer , Heiko Carstens , Vasily Gorbik , "David S. Miller" , Andreas Larsson , Alexander Viro , Christian Brauner , Matthew Brost , Joshua Hahn , Rakie Kim , Byungchul Park , Gregory Price , Ying Huang , Alistair Popple , Chris Li , Kairui Song , Kemeng Shi , Nhat Pham , Baoquan He , Youngjun Park , Johannes Weiner , Qi Zheng , Shakeel Butt , Axel Rasmussen , Yuanchu Xie , Wei Xu , Xu Xin , Chengming Zhou , Michal Hocko , Miklos Szeredi Cc: linux-mm@kvack.org, linux-kernel@vger.kernel.org, linux-doc@vger.kernel.org, linux-usb@vger.kernel.org, linux-rdma@vger.kernel.org, selinux@vger.kernel.org, linux-sound@vger.kernel.org, bpf@vger.kernel.org, linux-scsi@vger.kernel.org, linux-fbdev@vger.kernel.org, dri-devel@lists.freedesktop.org, linux-trace-kernel@vger.kernel.org, linux-perf-users@vger.kernel.org, linux-arch@vger.kernel.org, linux-fsdevel@vger.kernel.org, linux-arm-kernel@lists.infradead.org, kvmarm@lists.linux.dev, linuxppc-dev@lists.ozlabs.org, kvm@vger.kernel.org, kvm-riscv@lists.infradead.org, linux-riscv@lists.infradead.org, linux-s390@vger.kernel.org, sparclinux@vger.kernel.org, fuse-devel@lists.linux.dev, "Lorenzo Stoakes (ARM)" X-Mailer: b4 0.14.3 X-Developer-Signature: v=1; a=openpgp-sha256; l=7044; i=ljs@kernel.org; h=from:subject:message-id; bh=vEqZNj/orhd3A72a0KyhZ8lk/+VHtX68aC52ZbOqwFo=; b=owGbwMvMwCV2fu7ZrsZH9SKMp9WSGLIWZG7RL7u2o4dd6mTohG0XpLh3TvjX9Gf1FYOdaqsvH GXtOFR7paOUhUGMi0FWTJHl+Rfx/UEiYfM6L/i7wcxhZQIZwsDFKQATyUlmZOgIPuqrHpV47LaA aNfVw+2Lwhi9tDrPpU6+6l1edqHjhRDD/xxpMcYXGXueC8Z/OXa2u74xNjtI2T7XKerKxJOTr6f v5AEA X-Developer-Key: i=ljs@kernel.org; a=openpgp; fpr=E7F417BF5214569E89D04F46CF9DCD8A81E27F14 When performing mlock() or munlock() otherwise normal VMAs have VMA_IO_BIT solely to fix a race with migration which might otherwise double-count mlock VMAs. This is unnecessary - at the point of applying folio mlock state, whether setting or clearing PG_mlocked, we know whether or not we are locking. Solve this in two ways - thread a boolean through the page table walk indicating whether a lock or unlock is being performed, and run a locking walk with VMA_LOCKONFAULT_BIT set and VMA_LOCKED_BIT cleared. This state never occurs otherwise, as VMA_LOCKONFAULT_BIT always implies VMA_LOCKED_BIT. These are also always cleared together. Then, update folio_add_lru_vma() and mlock_folio() to check only for VMA_LOCKED_BIT, and update try_to_unmap_one() to check for VMA_LOCKED_MASK instead. Also remove the useless invocation of allow_mlock_munlock() which simply returns true if unlocking and instead rename it to allow_mlock() and only call it when locking. Finally, with the other mlock abuse of VMA_IO_BIT addressed, update mlock_vma_folio(), munlock_vma_folio() and folio_add_lru_vma() to simply test for VMA_LOCKED_BIT. While here, also replace some deprecated VMA flag predicates. Signed-off-by: Lorenzo Stoakes (ARM) --- mm/folio.c | 2 +- mm/internal.h | 5 ++--- mm/mlock.c | 51 +++++++++++++++++++-------------------------------- mm/rmap.c | 4 +++- 4 files changed, 25 insertions(+), 37 deletions(-) diff --git a/mm/folio.c b/mm/folio.c index 50a6dbe55998..a3f5c463f665 100644 --- a/mm/folio.c +++ b/mm/folio.c @@ -502,7 +502,7 @@ void folio_add_lru_vma(struct folio *folio, struct vm_area_struct *vma) { VM_BUG_ON_FOLIO(folio_test_lru(folio), folio); - if (unlikely((vma->vm_flags & (VM_LOCKED | VM_SPECIAL)) == VM_LOCKED)) + if (vma_test(vma, VMA_LOCKED_BIT)) mlock_new_folio(folio); else folio_add_lru(folio); diff --git a/mm/internal.h b/mm/internal.h index 6e27d3b10c01..04b1f1d3d960 100644 --- a/mm/internal.h +++ b/mm/internal.h @@ -975,8 +975,7 @@ void mlock_folio(struct folio *folio); static inline void mlock_vma_folio(struct folio *folio, struct vm_area_struct *vma) { - /* The VM_IO check prevents migration from double-counting during mlock. */ - if (unlikely((vma->vm_flags & (VM_LOCKED|VM_SPECIAL)) == VM_LOCKED)) + if (vma_test(vma, VMA_LOCKED_BIT)) mlock_folio(folio); } @@ -993,7 +992,7 @@ static inline void munlock_vma_folio(struct folio *folio, * always munlock the folio and page reclaim will correct it * if it's wrong. */ - if (unlikely(vma->vm_flags & VM_LOCKED)) + if (unlikely(vma_test(vma, VMA_LOCKED_BIT))) munlock_folio(folio); } diff --git a/mm/mlock.c b/mm/mlock.c index 39215a3eab1f..4235a1518fc9 100644 --- a/mm/mlock.c +++ b/mm/mlock.c @@ -316,22 +316,10 @@ static inline unsigned int folio_mlock_step(struct folio *folio, return folio_pte_batch(folio, pte, ptent, count); } -static inline bool allow_mlock_munlock(struct folio *folio, +static inline bool allow_mlock(struct folio *folio, struct vm_area_struct *vma, unsigned long start, unsigned long end, unsigned int step) { - /* - * For unlock, allow munlock large folio which is partially - * mapped to VMA. As it's possible that large folio is - * mlocked and VMA is split later. - * - * During memory pressure, such kind of large folio can - * be split. And the pages are not in VM_LOCKed VMA - * can be reclaimed. - */ - if (!vma_test(vma, VMA_LOCKED_BIT)) - return true; - /* folio_within_range() cannot take KSM, but any small folio is OK */ if (!folio_test_large(folio)) return true; @@ -352,6 +340,7 @@ static int mlock_pte_range(pmd_t *pmd, unsigned long addr, { struct vm_area_struct *vma = walk->vma; + const bool lock = walk->private; spinlock_t *ptl; pte_t *start_pte, *pte; pte_t ptent; @@ -368,7 +357,7 @@ static int mlock_pte_range(pmd_t *pmd, unsigned long addr, folio = pmd_folio(*pmd); if (folio_is_zone_device(folio)) goto out; - if (vma_test(vma, VMA_LOCKED_BIT)) + if (lock) mlock_folio(folio); else munlock_folio(folio); @@ -390,10 +379,10 @@ static int mlock_pte_range(pmd_t *pmd, unsigned long addr, continue; step = folio_mlock_step(folio, pte, addr, end); - if (!allow_mlock_munlock(folio, vma, start, end, step)) + if (lock && !allow_mlock(folio, vma, start, end, step)) goto next_entry; - if (vma_test(vma, VMA_LOCKED_BIT)) + if (lock) mlock_folio(folio); else munlock_folio(folio); @@ -428,31 +417,29 @@ static void mlock_vma_pages_range(struct vm_area_struct *vma, .pmd_entry = mlock_pte_range, .walk_lock = PGWALK_WRLOCK_VERIFY, }; + const bool lock = vma_flags_test(new_vma_flags, VMA_LOCKED_BIT); + vma_flags_t walk_flags = *new_vma_flags; /* - * There is a slight chance that concurrent page migration, - * or page reclaim finding a page of this now-VMA_LOCKED_BIT vma, - * will call mlock_vma_folio() and raise page's mlock_count: - * double counting, leaving the page unevictable indefinitely. - * Communicate this danger to mlock_vma_folio() with VMA_IO_BIT, - * which is a VMA_SPECIAL_FLAGS flag not allowed on VMA_LOCKED_BIT vmas. - * mmap_lock is held in write mode here, so this weird - * combination should not be visible to other mmap_lock users; - * but WRITE_ONCE so rmap walkers must see VMA_IO_BIT if VMA_LOCKED_BIT. + * LOCKONFAULT without LOCKED never otherwise occurs: it marks a walk in + * progress so that rmap-side callers, which test VMA_LOCKED_BIT, do not + * count folios, while try_to_unmap_one(), which tests VMA_LOCKED_MASK, + * still refuses to unmap them. */ - if (vma_flags_test(new_vma_flags, VMA_LOCKED_BIT)) - vma_flags_set(new_vma_flags, VMA_IO_BIT); + if (lock) { + vma_flags_clear(&walk_flags, VMA_LOCKED_BIT); + vma_flags_set(&walk_flags, VMA_LOCKONFAULT_BIT); + } + vma_start_write(vma); - vma_flags_reset_once(vma, new_vma_flags); + vma_flags_reset_once(vma, &walk_flags); lru_add_drain(); - walk_page_range_vma(vma, start, end, &mlock_walk_ops, NULL); + walk_page_range_vma(vma, start, end, &mlock_walk_ops, (void *)lock); lru_add_drain(); - if (vma_flags_test(new_vma_flags, VMA_IO_BIT)) { - vma_flags_clear(new_vma_flags, VMA_IO_BIT); + if (lock) vma_flags_reset_once(vma, new_vma_flags); - } } /* diff --git a/mm/rmap.c b/mm/rmap.c index 5fefe5b060b1..120c894d2dde 100644 --- a/mm/rmap.c +++ b/mm/rmap.c @@ -2239,9 +2239,11 @@ static bool try_to_unmap_one(struct folio *folio, struct vm_area_struct *vma, /* * If the folio is in an mlock()d vma, we must not swap it out. + * VMA_LOCKONFAULT_BIT alone marks an mlock walk in progress, see + * mlock_vma_pages_range(). */ if (!(flags & TTU_IGNORE_MLOCK) && - (vma->vm_flags & VM_LOCKED)) { + vma_test_any_mask(vma, VMA_LOCKED_MASK)) { ptes++; /* -- 2.55.0