From mboxrd@z Thu Jan 1 00:00:00 1970 Return-Path: X-Spam-Checker-Version: SpamAssassin 3.4.0 (2014-02-07) on aws-us-west-2-korg-lkml-1.web.codeaurora.org Received: from kanga.kvack.org (kanga.kvack.org [205.233.56.17]) (using TLSv1 with cipher DHE-RSA-AES256-SHA (256/256 bits)) (No client certificate requested) by smtp.lore.kernel.org (Postfix) with ESMTPS id 7CCF5C9830E for ; Fri, 25 Sep 2026 22:28:17 +0000 (UTC) Received: by kanga.kvack.org (Postfix) id 3E47A6B0088; Fri, 25 Sep 2026 18:28:16 -0400 (EDT) Received: by kanga.kvack.org (Postfix, from userid 40) id 396A36B008A; Fri, 25 Sep 2026 18:28:16 -0400 (EDT) X-Delivered-To: int-list-linux-mm@kvack.org Received: by kanga.kvack.org (Postfix, from userid 63042) id 25E6C6B008C; Fri, 25 Sep 2026 18:28:16 -0400 (EDT) X-Delivered-To: linux-mm@kvack.org Received: from relay.hostedemail.com (smtprelay0017.hostedemail.com [216.40.44.17]) by kanga.kvack.org (Postfix) with ESMTP id E6E0B6B0088 for ; Fri, 25 Sep 2026 18:28:15 -0400 (EDT) Received: from smtpin20.hostedemail.com (lb01a-stub [10.200.18.249]) by unirelay07.hostedemail.com (Postfix) with ESMTP id 71AA11607DB for ; Fri, 25 Sep 2026 22:28:15 +0000 (UTC) X-FDA: 85253724150.20.4BD7AF7 Received: from tor.source.kernel.org (tor.source.kernel.org [172.105.4.254]) by imf12.hostedemail.com (Postfix) with ESMTP id 9F36D40003 for ; Fri, 25 Sep 2026 22:28:13 +0000 (UTC) Authentication-Results: imf12.hostedemail.com; dkim=pass header.d=linux-foundation.org header.s=korg header.b=S6120kZC; spf=pass (imf12.hostedemail.com: domain of akpm@linux-foundation.org designates 172.105.4.254 as permitted sender) smtp.mailfrom=akpm@linux-foundation.org; dmarc=none ARC-Message-Signature: i=1; a=rsa-sha256; c=relaxed/relaxed; d=hostedemail.com; s=arc-20220608; t=1790375293; h=from:from:sender:reply-to:subject:subject:date:date: message-id:message-id:to:to:cc:cc:mime-version:mime-version: content-type:content-type: content-transfer-encoding:content-transfer-encoding: in-reply-to:in-reply-to:references:references:dkim-signature; bh=/FeqC6NsreHCzW9fFs+vzkIExlhoQtzmbN9AT65QI1Y=; b=qXzNU8p8plG77ECwZfgurfjeWJZX2LRdCrwoo5N0vZX0B+gXUcDVevyZIhdaTDtlRoBCGd svXI3R0gMpPAhEHPuaXrwD3fcnChTnifa+pR67jQRac59UghvcBs4nzpL2JxqqohOCKWeA KNnTERowi3ROMWZ1gsYB2JBlX21J7b0= ARC-Seal: i=1; a=rsa-sha256; d=hostedemail.com; s=arc-20220608; cv=none; t=1790375293; b=HL41fM4txJYg/Z63gxVcJpMDgTnuGMV/eUA/z8cqXDZ/XMZwhvHvIIDUM5BCtfFMaus60P OrnkYud4pyM58+5i8VHtBaTqUSzMQ7HZv/54m8laDYi9TpPDcTS0GTBgUwTC0piPkmY2QQ +jKZlIwhaaCHS7zNqMTTtO653aktZCc= ARC-Authentication-Results: i=1; imf12.hostedemail.com; dkim=pass header.d=linux-foundation.org header.s=korg header.b=S6120kZC; spf=pass (imf12.hostedemail.com: domain of akpm@linux-foundation.org designates 172.105.4.254 as permitted sender) smtp.mailfrom=akpm@linux-foundation.org; dmarc=none Received: from smtp.kernel.org (quasi.space.kernel.org [100.103.45.18]) by tor.source.kernel.org (Postfix) with ESMTP id 70FF960136; Fri, 25 Sep 2026 22:28:12 +0000 (UTC) Received: by smtp.kernel.org (Postfix) with ESMTPSA id 5D2B91F000FF; Fri, 25 Sep 2026 22:28:09 +0000 (UTC) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=linux-foundation.org; s=korg; t=1790375292; bh=/FeqC6NsreHCzW9fFs+vzkIExlhoQtzmbN9AT65QI1Y=; h=Date:From:To:Cc:Subject:In-Reply-To:References; b=S6120kZCR12iXPZt0xcZL0+Wfn1gU1RZB7NUKDNkZNJefGSfG9P6RA8Epfy+rvbkU xVs4fjUj0UHw+k+Gz9ZkgCGkot9F+dI1a3SWmLuTTxquEIyiCktMR74MXAiWNxdK6p IPoozz5ITxp6Mrzfo2l82/2c0yVPGzL/vW8rZJG0= Date: Fri, 25 Sep 2026 15:28:08 -0700 From: Andrew Morton To: "Lorenzo Stoakes (ARM)" Cc: David Hildenbrand , Zi Yan , Baolin Wang , "Liam R. Howlett" , Nico Pache , Ryan Roberts , Dev Jain , Barry Song , Lance Yang , Usama Arif , Kiryl Shutsemau , Guo Ren , Brian Cain , Geert Uytterhoeven , Dinh Nguyen , Simon Schuster , Jonas Bonn , Stefan Kristiansson , Stafford Horne , Rich Felker , John Paul Adrian Glaubitz , Paul Walmsley , Palmer Dabbelt , Albert Ou , Alexandre Ghiti , Russell King , Vineet Gupta , Michal Simek , Chris Zankel , Max Filippov , Will Deacon , "Aneesh Kumar K.V" , Nick Piggin , Peter Zijlstra , "David S. Miller" , Andreas Larsson , Richard Henderson , Matt Turner , Magnus Lindholm , Catalin Marinas , Mark Rutland , Huacai Chen , WANG Xuerui , Thomas Bogendoerfer , "James E.J. Bottomley" , Helge Deller , Madhavan Srinivasan , Michael Ellerman , "Christophe Leroy (CS GROUP)" , Heiko Carstens , Vasily Gorbik , Alexander Gordeev , Christian Borntraeger , Sven Schnelle , Richard Weinberger , Anton Ivanov , Johannes Berg , Thomas Gleixner , Ingo Molnar , Borislav Petkov , Dave Hansen , x86@kernel.org, "H. Peter Anvin" , Arnd Bergmann , Vlastimil Babka , Mike Rapoport , Suren Baghdasaryan , Michal Hocko , Jason Gunthorpe , John Hubbard , Peter Xu , Yoshinori Sato , Shakeel Butt , Jonathan Corbet , Randy Dunlap , linux-mm@kvack.org, linux-kernel@vger.kernel.org, linux-csky@vger.kernel.org, linux-hexagon@vger.kernel.org, linux-m68k@lists.linux-m68k.org, linux-openrisc@vger.kernel.org, linux-sh@vger.kernel.org, linux-riscv@lists.infradead.org, linux-arm-kernel@lists.infradead.org, linux-snps-arc@lists.infradead.org, linux-arch@vger.kernel.org, sparclinux@vger.kernel.org, linux-alpha@vger.kernel.org, loongarch@lists.linux.dev, linux-mips@vger.kernel.org, linux-parisc@vger.kernel.org, linuxppc-dev@lists.ozlabs.org, linux-s390@vger.kernel.org, linux-um@lists.infradead.org, Hugh Dickins , Qi Zheng , linux-doc@vger.kernel.org, Greg Ungerer Subject: Re: [PATCH v5 00/12] mm: make userland page table freeing RCU-safe Message-Id: <20260925152808.d4cae19cc002b1bba26502af@linux-foundation.org> In-Reply-To: <20260925-rcu-pagetable-freeing-v5-0-31e91065fea4@kernel.org> References: <20260925-rcu-pagetable-freeing-v5-0-31e91065fea4@kernel.org> X-Mailer: Sylpheed 3.8.0beta1 (GTK+ 2.24.33; x86_64-pc-linux-gnu) Mime-Version: 1.0 Content-Type: text/plain; charset=US-ASCII Content-Transfer-Encoding: 7bit X-Rspamd-Server: rspam12 X-Rspamd-Queue-Id: 9F36D40003 X-Stat-Signature: o88e5zjqd8mr9cy6b9zf5nqz8hn3rusz X-Rspam-User: X-HE-Tag: 1790375293-26871 X-HE-Meta: U2FsdGVkX1/JLuQLnoZ+1DKXwyh2Tg0ApkqlyzqgUjAUZGTTjoYsniP63N7IAIKlfaK/SeIJCzwAca00OgM9Y+cTwdpc0ItTBHjxfggXn9rBeRc6gPx3U7HnfM6xGX4m8G08A6bgT+VwDp0HzddMY1id7XPc0lerSTZ2KsvfkZywljgQGQT4oXUnDb+Qv8bzC8OTeKLLpW0AIlMJmndgtXSN5IlmfafXff5pzU6CNWhrUvSEUy9ch/vlziCqZvQQG9FrLM8aJNkGlBhRwqnfS14g+Pd4F6B7VqbzGvUoRdj7T8cgjm4RZzm6O9nJEmZciCHSkLC9VnT1e2L2IbqyKnCgwIij7zcjIy+GZZOHs9ENiOW9+Gx5DZ08EgIoE5khu2DbNmeV3sWtRGbTuDjora87Q2P111FqIy5y0P6SQyMfBJR/gAgpBj/N34okExm8lDxsQ1UNjUNOdXaybDvIBmFZdaohtSd8dnpN3NdnB2PJcR4Ueroz1+JnIvHFjJy0frC5GQiI4CZNF08Zl/gHU6R4eiBTlmy/qLf7fpARL2/uFmYJw9mZ8miNeyNe8GPJAa3d2ZrnVSALGoZhPSockeKBe1hQa8z/iRoJtw8QqEaNue218FixPBWiNjpIlGV6/D1o9UGO1WitVvK9s8lAzgkpUZf1W5PN3Z8xwmcXRa14fjjZlueFQ7RWS5y6vZvhYHef8xB2bTYwQ055lNoNTWCPZ2W/8mkSCJwSkibJ6zpy6N15NZfmxi8W/m+watTIgJs8sd+ewtlC7iGgtCOP+zfy10GpmLMW8/Cy4kPwCxeVuG6VxLZJFsgBUzPwQvZD0TariZ4i6WS81BpfBRMGdJjZiryImVTUeJQlVQTge35deoD0rKqXDpvUfpPc6W1lwuglzDpPErvPyk3cl/KaxarsRiZLLSZkT1MAtPrTvvE1gWv1Okg+CGYWaZtO85/W+zY8L+uj4g4hOyqUiLv eaS2PvYU 9hTxzNgPmTtGFvr45f8xG/tALcfRYbqRhekBlVUt7KwPHUijyc2PZxJnnER9CRZY1QvfhwznZ+SOC2J9peBOf9PrkKTSENPEuH2BbY5DZyNFQGMyo2dVXFVvMvCnNJu79gtP4Im3/ihbY3P1uFaMHuFl4XQOXWgGyhdqMsQgS4sc6Wuz563Qtdoi/AiHPfnwxrYo4F4oPjfp/dQ+EOO5FDilfuqxGrj2QCGfRKqU9m7hGQvO+HkZstdTuRsuqIcn6a8rki+bdoEM5p6j+dJYqcXhjGy8giAz248YZISEpaYH8O+oZJJV7ht0Y+LbFVdEH60oDrq+alCV3EHy6BA4YHqRYIRnJ+Yp////L Sender: owner-linux-mm@kvack.org Precedence: bulk X-Loop: owner-majordomo@kvack.org List-ID: List-Subscribe: List-Unsubscribe: On Fri, 25 Sep 2026 21:09:35 +0100 "Lorenzo Stoakes (ARM)" wrote: > The majority of architectures in the kernel defer page table freeing until > an RCU grace period has elapsed, this series converts all remaining > architectures to do so too and eliminates CONFIG_MMU_GATHER_RCU_TABLE_FREE > altogether. > > This is important because it enables safe lockless page table walking under > RCU alone. > > Doing so allows for reduced lock contention, avoids lock ordering concerns > and enables fast, efficient and correct page table walking as a result. Thanks, I've updated mm-unstable to this version. > v5: > * Collected tags (thanks all! :) > * Updated 1/12 to correctly set result to SCAN_ALLOC_HUGE_PAGE_FAIL should > the page table allocation fail, as per Lance. > * Updated 1/12 to rename alloc_deposit_pte() to alloc_deposit_pte_table() > and dropped the powerpc hash MMU paragraph from the commit message, as > per David. > * Updated 9/12 to use a BH-disabling spinlock rather than an IRQ-safe one, > as per Matthew. > * Updated the subject line of 11/12 to clarify that the patch removes > now-unneeded code, as per David. > * Updated 12/12 documentation as per David, Lance. Here's how v5 altered mm.git: Documentation/mm/process_addrs.rst | 9 +++++++-- arch/m68k/mm/motorola.c | 16 +++++++--------- mm/khugepaged.c | 4 ++-- 3 files changed, 16 insertions(+), 13 deletions(-) --- a/arch/m68k/mm/motorola.c~b +++ a/arch/m68k/mm/motorola.c @@ -181,7 +181,7 @@ static void *add_pointer_table(struct mm new = PD_PTABLE(pt_addr); PD_MARKBITS(new) = ptable_mask(type) - 1; - scoped_guard(spinlock_irqsave, &ptable_lock) + scoped_guard(spinlock_bh, &ptable_lock) list_add(new, &ptable_list[type]); return (pmd_t *)pt_addr; @@ -191,16 +191,15 @@ void *get_pointer_table(struct mm_struct { unsigned int tmp, off; unsigned long mask; - unsigned long flags; ptable_desc *dp; void *ret; - spin_lock_irqsave(&ptable_lock, flags); + spin_lock_bh(&ptable_lock); dp = ptable_list[type].next; mask = list_empty(&ptable_list[type]) ? 0 : PD_MARKBITS(dp); if (mask == 0) { - spin_unlock_irqrestore(&ptable_lock, flags); + spin_unlock_bh(&ptable_lock); return add_pointer_table(mm, type); } @@ -213,7 +212,7 @@ void *get_pointer_table(struct mm_struct } ret = ptdesc_address(PD_PTDESC(dp)) + off; - spin_unlock_irqrestore(&ptable_lock, flags); + spin_unlock_bh(&ptable_lock); return ret; } @@ -223,9 +222,8 @@ int free_pointer_table(void *table, int unsigned long ptable = (unsigned long)table; unsigned long pt_addr = ptable & PAGE_MASK; unsigned int mask = 1U << ((ptable - pt_addr)/ptable_size(type)); - unsigned long flags; - spin_lock_irqsave(&ptable_lock, flags); + spin_lock_bh(&ptable_lock); dp = PD_PTABLE(pt_addr); if (PD_MARKBITS (dp) & mask) @@ -236,7 +234,7 @@ int free_pointer_table(void *table, int if (PD_MARKBITS(dp) == ptable_mask(type)) { /* all tables in ptdesc are free, free ptdesc */ list_del(dp); - spin_unlock_irqrestore(&ptable_lock, flags); + spin_unlock_bh(&ptable_lock); mmu_page_dtor((void *)pt_addr); pagetable_dtor_free(virt_to_ptdesc((void *)pt_addr)); @@ -249,7 +247,7 @@ int free_pointer_table(void *table, int list_move(dp, &ptable_list[type]); } - spin_unlock_irqrestore(&ptable_lock, flags); + spin_unlock_bh(&ptable_lock); return 0; } --- a/Documentation/mm/process_addrs.rst~b +++ a/Documentation/mm/process_addrs.rst @@ -541,8 +541,13 @@ We establish basic locking rules when in after an RCU grace period has elapsed. However, any entry found must be revalidated after the page table lock is taken (such as the :c:func:`!pmd_same` recheck performed by :c:func:`!pte_offset_map_lock`) - before it is acted upon. Changing an entry requires the page table - lock and one of the locks that excludes teardown (mmap or VMA lock). + before it is acted upon. Changing an entry requires the page table lock + and one of the locks that excludes teardown (any one of the mmap, VMA or + rmap locks). +* When traversing page tables under RCU alone it is important to take care + when operating upon leaf entries - if the value is operated upon (for + instance getting the folio associated with a PTE) an appropriate lock must + be taken to prevent concurrent modification. * Reads from and writes to page table entries must be *appropriately* atomic. See the section on atomicity below for details. * Populating previously empty entries requires that the mmap or VMA locks are --- a/mm/khugepaged.c~b +++ a/mm/khugepaged.c @@ -1278,7 +1278,7 @@ static enum scan_result alloc_charge_fol return SCAN_SUCCEED; } -static pgtable_t alloc_deposit_pte(struct mm_struct *mm) +static pgtable_t alloc_deposit_pte_table(struct mm_struct *mm) { /* * khugepaged is run from a kernel thread, so need to manually set the @@ -1328,7 +1328,7 @@ static enum scan_result collapse_huge_pa } if (is_pmd_order(order)) { - pgtable = alloc_deposit_pte(mm); + pgtable = alloc_deposit_pte_table(mm); if (!pgtable) { result = SCAN_ALLOC_HUGE_PAGE_FAIL; goto out_nolock; _