From mboxrd@z Thu Jan 1 00:00:00 1970 Received: from tor.source.kernel.org (tor.source.kernel.org [172.105.4.254]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 22BE439A7E0 for ; Tue, 1 Sep 2026 14:41:41 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=172.105.4.254 ARC-Seal:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1788273703; cv=none; b=QKkDUFPaIZ/0vijh/tBmkchRBNNsXhK2/BmcDtQANJo3sqkZtDyuguHySEtZxEJDMFdKIoQ55asI6BjS7wbGMey/UqaMDKsJSrPo+mtHQvIRrz2OxZosU5DTgVEFeuPy1WwyOY4sn0sOtOXgw/LGhtzKfeQRSscWUtkXugrF3BY= ARC-Message-Signature:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1788273703; c=relaxed/simple; bh=aD1aeCXoZUTVkYkw8DYLgpTcX29H0YL7tFV55GqLOUs=; h=Date:From:To:Cc:Subject:Message-ID:References:MIME-Version: Content-Type:Content-Disposition:In-Reply-To; b=Kg0YJIQSpcC7pslO8T05kZVO6WNAjlu4sEqMno5Bp7RX0IkKToEwDNv9GUgd144LbmZ1/Ot7zlQ5++IjuFxeYkEUIecj3643ni74lTjLXUd0Wp5hLxgVMUiW85a11OXfwkK5K5i93EdZurCmqFdY9/8cUvdvbhOo1mYtHhzRG5w= ARC-Authentication-Results:i=1; smtp.subspace.kernel.org; dmarc=pass (p=quarantine dis=none) header.from=kernel.org; spf=pass smtp.mailfrom=kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b=DCkYCgJH; arc=none smtp.client-ip=172.105.4.254 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=quarantine dis=none) header.from=kernel.org Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=kernel.org Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b="DCkYCgJH" Received: from smtp.kernel.org (quasi.space.kernel.org [100.103.45.18]) by tor.source.kernel.org (Postfix) with ESMTP id EEA18602BE; Tue, 1 Sep 2026 14:41:40 +0000 (UTC) Received: by smtp.kernel.org (Postfix) with ESMTPSA id 5A26C1F000E9; Tue, 1 Sep 2026 14:41:20 +0000 (UTC) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=kernel.org; s=k20260515; t=1788273700; bh=aD1aeCXoZUTVkYkw8DYLgpTcX29H0YL7tFV55GqLOUs=; h=Date:From:To:Cc:Subject:References:In-Reply-To; b=DCkYCgJHTMeXpcfI011IFyZWluRGgMafgbwW8u29qgQBVZ9Dxsno0gDXnl7hHZ0m+ lkqnWV4wHDwKSWKnoQda11OIQ06fvWLfa32Ye05qYmX1Z7boJPT6aJZEl4EVv8r+G6 Q8c5h5Eq9o8u4IVkNkDvCB17dlRL52JTqjsLsbOVCb03G2bti6Yh9u4sRvw0rWJ1Ca qwHtYbfIXyVj5uLcNFlMoaGF0xHllKs6NACCcVCPrJufvl5Fed3xFkkaUWs7j7W7ZV siJKVu6dFavxSGvLfYUXvuVBZcE7zllPLV8WeZFb1IHv1rO4zZ/Mqb/Gua1p6L6cXY pVvsocfBPnW+w== Date: Tue, 1 Sep 2026 15:41:17 +0100 From: "Lorenzo Stoakes (ARM)" To: Jason Gunthorpe Cc: Kiryl Shutsemau , Andrew Morton , David Hildenbrand , Zi Yan , Baolin Wang , "Liam R. Howlett" , Nico Pache , Ryan Roberts , Dev Jain , Barry Song , Lance Yang , Usama Arif , Guo Ren , Brian Cain , Geert Uytterhoeven , Dinh Nguyen , Simon Schuster , Jonas Bonn , Stefan Kristiansson , Stafford Horne , Yoshinori Sato , Rich Felker , John Paul Adrian Glaubitz , Paul Walmsley , Palmer Dabbelt , Albert Ou , Alexandre Ghiti , Russell King , Vineet Gupta , Michal Simek , Chris Zankel , Max Filippov , Will Deacon , "Aneesh Kumar K.V" , Nick Piggin , Peter Zijlstra , "David S. Miller" , Andreas Larsson , Richard Henderson , Matt Turner , Magnus Lindholm , Catalin Marinas , Mark Rutland , Huacai Chen , WANG Xuerui , Thomas Bogendoerfer , "James E.J. Bottomley" , Helge Deller , Madhavan Srinivasan , Michael Ellerman , "Christophe Leroy (CS GROUP)" , Heiko Carstens , Vasily Gorbik , Alexander Gordeev , Christian Borntraeger , Sven Schnelle , Richard Weinberger , Anton Ivanov , Johannes Berg , Thomas Gleixner , Ingo Molnar , Borislav Petkov , Dave Hansen , x86@kernel.org, "H. Peter Anvin" , Arnd Bergmann , Vlastimil Babka , Mike Rapoport , Suren Baghdasaryan , Michal Hocko , John Hubbard , Peter Xu , linux-mm@kvack.org, linux-kernel@vger.kernel.org, linux-csky@vger.kernel.org, linux-hexagon@vger.kernel.org, linux-m68k@lists.linux-m68k.org, linux-openrisc@vger.kernel.org, linux-sh@vger.kernel.org, linux-riscv@lists.infradead.org, linux-arm-kernel@lists.infradead.org, linux-snps-arc@lists.infradead.org, linux-arch@vger.kernel.org, sparclinux@vger.kernel.org, linux-alpha@vger.kernel.org, loongarch@lists.linux.dev, linux-mips@vger.kernel.org, linux-parisc@vger.kernel.org, linuxppc-dev@lists.ozlabs.org, linux-s390@vger.kernel.org, linux-um@lists.infradead.org, Hugh Dickins , Qi Zheng Subject: Re: [PATCH 01/12] mm/huge_memory: zap deposited page tables after an RCU grace period Message-ID: References: <20260901-rcu-pagetable-freeing-v1-0-5456a81c8212@kernel.org> <20260901-rcu-pagetable-freeing-v1-1-5456a81c8212@kernel.org> <20260901142408.GA56830@ziepe.ca> Precedence: bulk X-Mailing-List: linux-m68k@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Type: text/plain; charset=us-ascii Content-Disposition: inline In-Reply-To: <20260901142408.GA56830@ziepe.ca> On Tue, Sep 01, 2026 at 11:24:08AM -0300, Jason Gunthorpe wrote: > On Tue, Sep 01, 2026 at 03:12:45PM +0100, Lorenzo Stoakes (ARM) wrote: > > > It won't be costly at the time of the calls obviously as its deferred. Maybe > > increase some time spent in softirq but again is 512x that big of a deal? > > > > I'm not sure how you'd both defer the free and somehow utilise mmu_gather here > > either really, certainly not without it becoming extremely messy. > > The less costly version is to thread the page to be freed onto the > mmu_gather through a linked list in the struct page memory. This is > super cheap since it is just a singly linked list operation. > > Then when the mmu_gather is flushed it does a single call_rcu using > the rcu head of the struct page of the head of the list. The callback > clears the entire linked list of pages. > > Since you have to tlb flush anyhow, it makes sense to always use the > mmu_gather. For example the design I ended up with for iommupt > accumulates all the invalidations and all the free-able memory into a > gather then invalidates and frees. > > This allows maximizing the tlbi efficiency too. You can't do call_srcu > until you flush the tlb and if you call once per table then you are > also tlb flushing once per table too. > > So if the kernel really does want to clear out 512 leaf tables the > optimal implementation is one range tlbi for 512 entries followed by > one call_rcu to free the memory. Hence the gather.. I think there's some confusion here. This isn't the path in which a page table is being freed, the _deposited_ table is zapped, in zap_deposited_table(). That is, the page table kept in reserve for THP split, that is not currently mapped. It amounts to a __free_pages() call. The TLB operations are in e.g. zap_huge_pmd() etc. and nobody has complained about inefficiencies there. So, unless I'm missing something here, TLB flushes play no role in this whatsoever. The issue Kiryl raised was that instead of immediately freeing page tables, they are now batched up individually by call_rcu(). I personally find it difficult to imagine the numbers here would be problematic or certainly cause anything observable beyond what is observable now. So I'm going to have to say, unless it can be clearly demonstrated this is problematic, I don't think there's any reason to add additional complexity here. > > Jason -- Cheers, Lorenzo