All of lore.kernel.org
 help / color / mirror / Atom feed
From: "Lorenzo Stoakes (ARM)" <ljs@kernel.org>
To: Kiryl Shutsemau <kas@kernel.org>
Cc: Jason Gunthorpe <jgg@ziepe.ca>,
	 Andrew Morton <akpm@linux-foundation.org>,
	David Hildenbrand <david@kernel.org>, Zi Yan <ziy@nvidia.com>,
	 Baolin Wang <baolin.wang@linux.alibaba.com>,
	"Liam R. Howlett" <liam@infradead.org>,
	 Nico Pache <nico.pache@linux.dev>,
	Ryan Roberts <ryan.roberts@arm.com>, Dev Jain <dev.jain@arm.com>,
	 Barry Song <baohua@kernel.org>,
	Lance Yang <lance.yang@linux.dev>,
	 Usama Arif <usama.arif@linux.dev>, Guo Ren <guoren@kernel.org>,
	Brian Cain <bcain@kernel.org>,
	 Geert Uytterhoeven <geert@linux-m68k.org>,
	Dinh Nguyen <dinguyen@kernel.org>,
	 Simon Schuster <schuster.simon@siemens-energy.com>,
	Jonas Bonn <jonas@southpole.se>,
	 Stefan Kristiansson <stefan.kristiansson@saunalahti.fi>,
	Stafford Horne <shorne@gmail.com>,
	 Yoshinori Sato <ysato@users.sourceforge.jp>,
	Rich Felker <dalias@libc.org>,
	 John Paul Adrian Glaubitz <glaubitz@physik.fu-berlin.de>,
	Paul Walmsley <pjw@kernel.org>,
	 Palmer Dabbelt <palmer@dabbelt.com>,
	Albert Ou <aou@eecs.berkeley.edu>,
	 Alexandre Ghiti <alex@ghiti.fr>,
	Russell King <linux@armlinux.org.uk>,
	 Vineet Gupta <vgupta@kernel.org>,
	Michal Simek <monstr@monstr.eu>, Chris Zankel <chris@zankel.net>,
	 Max Filippov <jcmvbkbc@gmail.com>, Will Deacon <will@kernel.org>,
	 "Aneesh Kumar K.V" <aneesh.kumar@kernel.org>,
	Nick Piggin <npiggin@gmail.com>,
	 Peter Zijlstra <peterz@infradead.org>,
	"David S. Miller" <davem@davemloft.net>,
	 Andreas Larsson <andreas@gaisler.com>,
	Richard Henderson <richard.henderson@linaro.org>,
	 Matt Turner <mattst88@gmail.com>,
	Magnus Lindholm <linmag7@gmail.com>,
	 Catalin Marinas <catalin.marinas@arm.com>,
	Mark Rutland <mark.rutland@arm.com>,
	 Huacai Chen <chenhuacai@kernel.org>,
	WANG Xuerui <kernel@xen0n.name>,
	 Thomas Bogendoerfer <tsbogend@alpha.franken.de>,
	"James E.J. Bottomley" <James.Bottomley@hansenpartnership.com>,
	 Helge Deller <deller@gmx.de>,
	Madhavan Srinivasan <maddy@linux.ibm.com>,
	 Michael Ellerman <mpe@ellerman.id.au>,
	"Christophe Leroy (CS GROUP)" <chleroy@kernel.org>,
	 Heiko Carstens <hca@linux.ibm.com>,
	Vasily Gorbik <gor@linux.ibm.com>,
	 Alexander Gordeev <agordeev@linux.ibm.com>,
	Christian Borntraeger <borntraeger@linux.ibm.com>,
	 Sven Schnelle <svens@linux.ibm.com>,
	Richard Weinberger <richard@nod.at>,
	 Anton Ivanov <anton.ivanov@cambridgegreys.com>,
	Johannes Berg <johannes@sipsolutions.net>,
	 Thomas Gleixner <tglx@kernel.org>,
	Ingo Molnar <mingo@redhat.com>, Borislav Petkov <bp@alien8.de>,
	 Dave Hansen <dave.hansen@linux.intel.com>,
	x86@kernel.org, "H. Peter Anvin" <hpa@zytor.com>,
	 Arnd Bergmann <arnd@arndb.de>,
	Vlastimil Babka <vbabka@kernel.org>,
	 Mike Rapoport <rppt@kernel.org>,
	Suren Baghdasaryan <surenb@google.com>,
	 Michal Hocko <mhocko@suse.com>,
	John Hubbard <jhubbard@nvidia.com>, Peter Xu <peterx@redhat.com>,
	 linux-mm@kvack.org, linux-kernel@vger.kernel.org,
	linux-csky@vger.kernel.org,  linux-hexagon@vger.kernel.org,
	linux-m68k@lists.linux-m68k.org, linux-openrisc@vger.kernel.org,
	 linux-sh@vger.kernel.org, linux-riscv@lists.infradead.org,
	 linux-arm-kernel@lists.infradead.org,
	linux-snps-arc@lists.infradead.org, linux-arch@vger.kernel.org,
	 sparclinux@vger.kernel.org, linux-alpha@vger.kernel.org,
	loongarch@lists.linux.dev,  linux-mips@vger.kernel.org,
	linux-parisc@vger.kernel.org, linuxppc-dev@lists.ozlabs.org,
	 linux-s390@vger.kernel.org, linux-um@lists.infradead.org,
	Hugh Dickins <hughd@google.com>,  Qi Zheng <qi.zheng@linux.dev>
Subject: Re: [PATCH 01/12] mm/huge_memory: zap deposited page tables after an RCU grace period
Date: Tue, 1 Sep 2026 16:45:14 +0100	[thread overview]
Message-ID: <apbvoN6ttgHfI_xg@gremlin> (raw)
In-Reply-To: <apbuO_QveZNAz3ns@thinkstation>

On Tue, Sep 01, 2026 at 04:28:48PM +0100, Kiryl Shutsemau wrote:
> On Tue, Sep 01, 2026 at 03:41:17PM +0100, Lorenzo Stoakes (ARM) wrote:
> > On Tue, Sep 01, 2026 at 11:24:08AM -0300, Jason Gunthorpe wrote:
> > > On Tue, Sep 01, 2026 at 03:12:45PM +0100, Lorenzo Stoakes (ARM) wrote:
> > >
> > > > It won't be costly at the time of the calls obviously as its deferred. Maybe
> > > > increase some time spent in softirq but again is 512x that big of a deal?
> > > >
> > > > I'm not sure how you'd both defer the free and somehow utilise mmu_gather here
> > > > either really, certainly not without it becoming extremely messy.
> > >
> > > The less costly version is to thread the page to be freed onto the
> > > mmu_gather through a linked list in the struct page memory. This is
> > > super cheap since it is just a singly linked list operation.
> > >
> > > Then when the mmu_gather is flushed it does a single call_rcu using
> > > the rcu head of the struct page of the head of the list. The callback
> > > clears the entire linked list of pages.
> > >
> > > Since you have to tlb flush anyhow, it makes sense to always use the
> > > mmu_gather. For example the design I ended up with for iommupt
> > > accumulates all the invalidations and all the free-able memory into a
> > > gather then invalidates and frees.
> > >
> > > This allows maximizing the tlbi efficiency too. You can't do call_srcu
> > > until you flush the tlb and if you call once per table then you are
> > > also tlb flushing once per table too.
> > >
> > > So if the kernel really does want to clear out 512 leaf tables the
> > > optimal implementation is one range tlbi for 512 entries followed by
> > > one call_rcu to free the memory. Hence the gather..
> >
> > I think there's some confusion here.
> >
> > This isn't the path in which a page table is being freed, the _deposited_
> > table is zapped, in zap_deposited_table().
> >
> > That is, the page table kept in reserve for THP split, that is not
> > currently mapped.
> >
> > It amounts to a __free_pages() call.
> >
> > The TLB operations are in e.g. zap_huge_pmd() etc. and nobody has
> > complained about inefficiencies there.
> >
> > So, unless I'm missing something here, TLB flushes play no role in this
> > whatsoever.
> >
> > The issue Kiryl raised was that instead of immediately freeing page tables,
> > they are now batched up individually by call_rcu().
> >
> > I personally find it difficult to imagine the numbers here would be
> > problematic or certainly cause anything observable beyond what is
> > observable now.
> >
> > So I'm going to have to say, unless it can be clearly demonstrated this is
> > problematic, I don't think there's any reason to add additional complexity
> > here.
>
> It would be nice to measure munmap() overhead here.
>
> I am worried about hitting DEFAULT_MAX_RCU_BLIMIT and trigger
> rcu_force_quiescent_state() which can be disruptive to the system.
>
> DEFAULT_MAX_RCU_BLIMIT is 10K, so it is ~20G of THP unmapped on x86.
>
> munmap() of 64G worth of THP should be enough to demonstrate the
> problem.

I mean you're going to hit that from RCU freeing page tables already, which
most architectures already do right?

So if RCU saturation is a problem, that problem already exists, but I've
not heard of that being a problem at all?

So you're going to have to demonstrate why this situation is markedly
different from that. And it's the same scale.

Overall I think freeing 64 GiB of mapped memory all at once will inevitably
be a slow operation, freeing them directly will also be a lengthily process.

And also it seems to me that RCU mishandling heavy load to the point of
causing system instability should a bug filed with RCU no?

Also note pte_free_defer() is already used in retract_page_tables() so a large
collapse could also hit this problem?

I'm not sure I'm convinced there's an issue here.

>
> --
>   Kiryl Shutsemau / Kirill A. Shutemov

--
Cheers, Lorenzo


WARNING: multiple messages have this Message-ID (diff)
From: "Lorenzo Stoakes (ARM)" <ljs@kernel.org>
To: Kiryl Shutsemau <kas@kernel.org>
Cc: Jason Gunthorpe <jgg@ziepe.ca>,
	 Andrew Morton <akpm@linux-foundation.org>,
	David Hildenbrand <david@kernel.org>, Zi Yan <ziy@nvidia.com>,
	 Baolin Wang <baolin.wang@linux.alibaba.com>,
	"Liam R. Howlett" <liam@infradead.org>,
	 Nico Pache <nico.pache@linux.dev>,
	Ryan Roberts <ryan.roberts@arm.com>, Dev Jain <dev.jain@arm.com>,
	 Barry Song <baohua@kernel.org>,
	Lance Yang <lance.yang@linux.dev>,
	 Usama Arif <usama.arif@linux.dev>, Guo Ren <guoren@kernel.org>,
	Brian Cain <bcain@kernel.org>,
	 Geert Uytterhoeven <geert@linux-m68k.org>,
	Dinh Nguyen <dinguyen@kernel.org>,
	 Simon Schuster <schuster.simon@siemens-energy.com>,
	Jonas Bonn <jonas@southpole.se>,
	 Stefan Kristiansson <stefan.kristiansson@saunalahti.fi>,
	Stafford Horne <shorne@gmail.com>,
	 Yoshinori Sato <ysato@users.sourceforge.jp>,
	Rich Felker <dalias@libc.org>,
	 John Paul Adrian Glaubitz <glaubitz@physik.fu-berlin.de>,
	Paul Walmsley <pjw@kernel.org>,
	 Palmer Dabbelt <palmer@dabbelt.com>,
	Albert Ou <aou@eecs.berkeley.edu>,
	 Alexandre Ghiti <alex@ghiti.fr>,
	Russell King <linux@armlinux.org.uk>,
	 Vineet Gupta <vgupta@kernel.org>,
	Michal Simek <monstr@monstr.eu>, Chris Zankel <chris@zankel.net>,
	 Max Filippov <jcmvbkbc@gmail.com>, Will Deacon <will@kernel.org>,
	 "Aneesh Kumar K.V" <aneesh.kumar@kernel.org>,
	Nick Piggin <npiggin@gmail.com>,
	 Peter Zijlstra <peterz@infradead.org>,
	"David S. Miller" <davem@davemloft.net>,
	 Andreas Larsson <andreas@gaisler.com>,
	Richard Henderson <richard.henderson@linaro.org>,
	 Matt Turner <mattst88@gmail.com>,
	Magnus Lindholm <linmag7@gmail.com>,
	 Catalin Marinas <catalin.marinas@arm.com>,
	Mark Rutland <mark.rutland@arm.com>,
	 Huacai Chen <chenhuacai@kernel.org>,
	WANG Xuerui <kernel@xen0n.name>,
	 Thomas Bogendoerfer <tsbogend@alpha.franken.de>,
	"James E.J. Bottomley" <James.Bottomley@hansenpartnership.com>,
	 Helge Deller <deller@gmx.de>,
	Madhavan Srinivasan <maddy@linux.ibm.com>,
	 Michael Ellerman <mpe@ellerman.id.au>,
	"Christophe Leroy (CS GROUP)" <chleroy@kernel.org>,
	 Heiko Carstens <hca@linux.ibm.com>,
	Vasily Gorbik <gor@linux.ibm.com>,
	 Alexander Gordeev <agordeev@linux.ibm.com>,
	Christian Borntraeger <borntraeger@linux.ibm.com>,
	 Sven Schnelle <svens@linux.ibm.com>,
	Richard Weinberger <richard@nod.at>,
	 Anton Ivanov <anton.ivanov@cambridgegreys.com>,
	Johannes Berg <johannes@sipsolutions.net>,
	 Thomas Gleixner <tglx@kernel.org>,
	Ingo Molnar <mingo@redhat.com>, Borislav Petkov <bp@alien8.de>,
	 Dave Hansen <dave.hansen@linux.intel.com>,
	x86@kernel.org, "H. Peter Anvin" <hpa@zytor.com>,
	 Arnd Bergmann <arnd@arndb.de>,
	Vlastimil Babka <vbabka@kernel.org>,
	 Mike Rapoport <rppt@kernel.org>,
	Suren Baghdasaryan <surenb@google.com>,
	 Michal Hocko <mhocko@suse.com>,
	John Hubbard <jhubbard@nvidia.com>, Peter Xu <peterx@redhat.com>,
	 linux-mm@kvack.org, linux-kernel@vger.kernel.org,
	linux-csky@vger.kernel.org,  linux-hexagon@vger.kernel.org,
	linux-m68k@lists.linux-m68k.org, linux-openrisc@vger.kernel.org,
	 linux-sh@vger.kernel.org, linux-riscv@lists.infradead.org,
	 linux-arm-kernel@lists.infradead.org,
	linux-snps-arc@lists.infradead.org, linux-arch@vger.kernel.org,
	 sparclinux@vger.kernel.org, linux-alpha@vger.kernel.org,
	loongarch@lists.linux.dev,  linux-mips@vger.kernel.org,
	linux-parisc@vger.kernel.org, linuxppc-dev@lists.ozlabs.org,
	 linux-s390@vger.kernel.org, linux-um@lists.infradead.org,
	Hugh Dickins <hughd@google.com>,  Qi Zheng <qi.zheng@linux.dev>
Subject: Re: [PATCH 01/12] mm/huge_memory: zap deposited page tables after an RCU grace period
Date: Tue, 1 Sep 2026 16:45:14 +0100	[thread overview]
Message-ID: <apbvoN6ttgHfI_xg@gremlin> (raw)
In-Reply-To: <apbuO_QveZNAz3ns@thinkstation>

On Tue, Sep 01, 2026 at 04:28:48PM +0100, Kiryl Shutsemau wrote:
> On Tue, Sep 01, 2026 at 03:41:17PM +0100, Lorenzo Stoakes (ARM) wrote:
> > On Tue, Sep 01, 2026 at 11:24:08AM -0300, Jason Gunthorpe wrote:
> > > On Tue, Sep 01, 2026 at 03:12:45PM +0100, Lorenzo Stoakes (ARM) wrote:
> > >
> > > > It won't be costly at the time of the calls obviously as its deferred. Maybe
> > > > increase some time spent in softirq but again is 512x that big of a deal?
> > > >
> > > > I'm not sure how you'd both defer the free and somehow utilise mmu_gather here
> > > > either really, certainly not without it becoming extremely messy.
> > >
> > > The less costly version is to thread the page to be freed onto the
> > > mmu_gather through a linked list in the struct page memory. This is
> > > super cheap since it is just a singly linked list operation.
> > >
> > > Then when the mmu_gather is flushed it does a single call_rcu using
> > > the rcu head of the struct page of the head of the list. The callback
> > > clears the entire linked list of pages.
> > >
> > > Since you have to tlb flush anyhow, it makes sense to always use the
> > > mmu_gather. For example the design I ended up with for iommupt
> > > accumulates all the invalidations and all the free-able memory into a
> > > gather then invalidates and frees.
> > >
> > > This allows maximizing the tlbi efficiency too. You can't do call_srcu
> > > until you flush the tlb and if you call once per table then you are
> > > also tlb flushing once per table too.
> > >
> > > So if the kernel really does want to clear out 512 leaf tables the
> > > optimal implementation is one range tlbi for 512 entries followed by
> > > one call_rcu to free the memory. Hence the gather..
> >
> > I think there's some confusion here.
> >
> > This isn't the path in which a page table is being freed, the _deposited_
> > table is zapped, in zap_deposited_table().
> >
> > That is, the page table kept in reserve for THP split, that is not
> > currently mapped.
> >
> > It amounts to a __free_pages() call.
> >
> > The TLB operations are in e.g. zap_huge_pmd() etc. and nobody has
> > complained about inefficiencies there.
> >
> > So, unless I'm missing something here, TLB flushes play no role in this
> > whatsoever.
> >
> > The issue Kiryl raised was that instead of immediately freeing page tables,
> > they are now batched up individually by call_rcu().
> >
> > I personally find it difficult to imagine the numbers here would be
> > problematic or certainly cause anything observable beyond what is
> > observable now.
> >
> > So I'm going to have to say, unless it can be clearly demonstrated this is
> > problematic, I don't think there's any reason to add additional complexity
> > here.
>
> It would be nice to measure munmap() overhead here.
>
> I am worried about hitting DEFAULT_MAX_RCU_BLIMIT and trigger
> rcu_force_quiescent_state() which can be disruptive to the system.
>
> DEFAULT_MAX_RCU_BLIMIT is 10K, so it is ~20G of THP unmapped on x86.
>
> munmap() of 64G worth of THP should be enough to demonstrate the
> problem.

I mean you're going to hit that from RCU freeing page tables already, which
most architectures already do right?

So if RCU saturation is a problem, that problem already exists, but I've
not heard of that being a problem at all?

So you're going to have to demonstrate why this situation is markedly
different from that. And it's the same scale.

Overall I think freeing 64 GiB of mapped memory all at once will inevitably
be a slow operation, freeing them directly will also be a lengthily process.

And also it seems to me that RCU mishandling heavy load to the point of
causing system instability should a bug filed with RCU no?

Also note pte_free_defer() is already used in retract_page_tables() so a large
collapse could also hit this problem?

I'm not sure I'm convinced there's an issue here.

>
> --
>   Kiryl Shutsemau / Kirill A. Shutemov

--
Cheers, Lorenzo

_______________________________________________
linux-snps-arc mailing list
linux-snps-arc@lists.infradead.org
http://lists.infradead.org/mailman/listinfo/linux-snps-arc

WARNING: multiple messages have this Message-ID (diff)
From: "Lorenzo Stoakes (ARM)" <ljs@kernel.org>
To: Kiryl Shutsemau <kas@kernel.org>
Cc: Jason Gunthorpe <jgg@ziepe.ca>,
	 Andrew Morton <akpm@linux-foundation.org>,
	David Hildenbrand <david@kernel.org>, Zi Yan <ziy@nvidia.com>,
	 Baolin Wang <baolin.wang@linux.alibaba.com>,
	"Liam R. Howlett" <liam@infradead.org>,
	 Nico Pache <nico.pache@linux.dev>,
	Ryan Roberts <ryan.roberts@arm.com>, Dev Jain <dev.jain@arm.com>,
	 Barry Song <baohua@kernel.org>,
	Lance Yang <lance.yang@linux.dev>,
	 Usama Arif <usama.arif@linux.dev>, Guo Ren <guoren@kernel.org>,
	Brian Cain <bcain@kernel.org>,
	 Geert Uytterhoeven <geert@linux-m68k.org>,
	Dinh Nguyen <dinguyen@kernel.org>,
	 Simon Schuster <schuster.simon@siemens-energy.com>,
	Jonas Bonn <jonas@southpole.se>,
	 Stefan Kristiansson <stefan.kristiansson@saunalahti.fi>,
	Stafford Horne <shorne@gmail.com>,
	 Yoshinori Sato <ysato@users.sourceforge.jp>,
	Rich Felker <dalias@libc.org>,
	 John Paul Adrian Glaubitz <glaubitz@physik.fu-berlin.de>,
	Paul Walmsley <pjw@kernel.org>,
	 Palmer Dabbelt <palmer@dabbelt.com>,
	Albert Ou <aou@eecs.berkeley.edu>,
	 Alexandre Ghiti <alex@ghiti.fr>,
	Russell King <linux@armlinux.org.uk>,
	 Vineet Gupta <vgupta@kernel.org>,
	Michal Simek <monstr@monstr.eu>, Chris Zankel <chris@zankel.net>,
	 Max Filippov <jcmvbkbc@gmail.com>, Will Deacon <will@kernel.org>,
	 "Aneesh Kumar K.V" <aneesh.kumar@kernel.org>,
	Nick Piggin <npiggin@gmail.com>,
	 Peter Zijlstra <peterz@infradead.org>,
	"David S. Miller" <davem@davemloft.net>,
	 Andreas Larsson <andreas@gaisler.com>,
	Richard Henderson <richard.henderson@linaro.org>,
	 Matt Turner <mattst88@gmail.com>,
	Magnus Lindholm <linmag7@gmail.com>,
	 Catalin Marinas <catalin.marinas@arm.com>,
	Mark Rutland <mark.rutland@arm.com>,
	 Huacai Chen <chenhuacai@kernel.org>,
	WANG Xuerui <kernel@xen0n.name>,
	 Thomas Bogendoerfer <tsbogend@alpha.franken.de>,
	"James E.J. Bottomley" <James.Bottomley@hansenpartnership.com>,
	 Helge Deller <deller@gmx.de>,
	Madhavan Srinivasan <maddy@linux.ibm.com>,
	 Michael Ellerman <mpe@ellerman.id.au>,
	"Christophe Leroy (CS GROUP)" <chleroy@kernel.org>,
	 Heiko Carstens <hca@linux.ibm.com>,
	Vasily Gorbik <gor@linux.ibm.com>,
	 Alexander Gordeev <agordeev@linux.ibm.com>,
	Christian Borntraeger <borntraeger@linux.ibm.com>,
	 Sven Schnelle <svens@linux.ibm.com>,
	Richard Weinberger <richard@nod.at>,
	 Anton Ivanov <anton.ivanov@cambridgegreys.com>,
	Johannes Berg <johannes@sipsolutions.net>,
	 Thomas Gleixner <tglx@kernel.org>,
	Ingo Molnar <mingo@redhat.com>, Borislav Petkov <bp@alien8.de>,
	 Dave Hansen <dave.hansen@linux.intel.com>,
	x86@kernel.org, "H. Peter Anvin" <hpa@zytor.com>,
	 Arnd Bergmann <arnd@arndb.de>,
	Vlastimil Babka <vbabka@kernel.org>,
	 Mike Rapoport <rppt@kernel.org>,
	Suren Baghdasaryan <surenb@google.com>,
	 Michal Hocko <mhocko@suse.com>,
	John Hubbard <jhubbard@nvidia.com>, Peter Xu <peterx@redhat.com>,
	 linux-mm@kvack.org, linux-kernel@vger.kernel.org,
	linux-csky@vger.kernel.org,  linux-hexagon@vger.kernel.org,
	linux-m68k@lists.linux-m68k.org, linux-openrisc@vger.kernel.org,
	 linux-sh@vger.kernel.org, linux-riscv@lists.infradead.org,
	 linux-arm-kernel@lists.infradead.org,
	linux-snps-arc@lists.infradead.org, linux-arch@vger.kernel.org,
	 sparclinux@vger.kernel.org, linux-alpha@vger.kernel.org,
	loongarch@lists.linux.dev,  linux-mips@vger.kernel.org,
	linux-parisc@vger.kernel.org, linuxppc-dev@lists.ozlabs.org,
	 linux-s390@vger.kernel.org, linux-um@lists.infradead.org,
	Hugh Dickins <hughd@google.com>,  Qi Zheng <qi.zheng@linux.dev>
Subject: Re: [PATCH 01/12] mm/huge_memory: zap deposited page tables after an RCU grace period
Date: Tue, 1 Sep 2026 16:45:14 +0100	[thread overview]
Message-ID: <apbvoN6ttgHfI_xg@gremlin> (raw)
In-Reply-To: <apbuO_QveZNAz3ns@thinkstation>

On Tue, Sep 01, 2026 at 04:28:48PM +0100, Kiryl Shutsemau wrote:
> On Tue, Sep 01, 2026 at 03:41:17PM +0100, Lorenzo Stoakes (ARM) wrote:
> > On Tue, Sep 01, 2026 at 11:24:08AM -0300, Jason Gunthorpe wrote:
> > > On Tue, Sep 01, 2026 at 03:12:45PM +0100, Lorenzo Stoakes (ARM) wrote:
> > >
> > > > It won't be costly at the time of the calls obviously as its deferred. Maybe
> > > > increase some time spent in softirq but again is 512x that big of a deal?
> > > >
> > > > I'm not sure how you'd both defer the free and somehow utilise mmu_gather here
> > > > either really, certainly not without it becoming extremely messy.
> > >
> > > The less costly version is to thread the page to be freed onto the
> > > mmu_gather through a linked list in the struct page memory. This is
> > > super cheap since it is just a singly linked list operation.
> > >
> > > Then when the mmu_gather is flushed it does a single call_rcu using
> > > the rcu head of the struct page of the head of the list. The callback
> > > clears the entire linked list of pages.
> > >
> > > Since you have to tlb flush anyhow, it makes sense to always use the
> > > mmu_gather. For example the design I ended up with for iommupt
> > > accumulates all the invalidations and all the free-able memory into a
> > > gather then invalidates and frees.
> > >
> > > This allows maximizing the tlbi efficiency too. You can't do call_srcu
> > > until you flush the tlb and if you call once per table then you are
> > > also tlb flushing once per table too.
> > >
> > > So if the kernel really does want to clear out 512 leaf tables the
> > > optimal implementation is one range tlbi for 512 entries followed by
> > > one call_rcu to free the memory. Hence the gather..
> >
> > I think there's some confusion here.
> >
> > This isn't the path in which a page table is being freed, the _deposited_
> > table is zapped, in zap_deposited_table().
> >
> > That is, the page table kept in reserve for THP split, that is not
> > currently mapped.
> >
> > It amounts to a __free_pages() call.
> >
> > The TLB operations are in e.g. zap_huge_pmd() etc. and nobody has
> > complained about inefficiencies there.
> >
> > So, unless I'm missing something here, TLB flushes play no role in this
> > whatsoever.
> >
> > The issue Kiryl raised was that instead of immediately freeing page tables,
> > they are now batched up individually by call_rcu().
> >
> > I personally find it difficult to imagine the numbers here would be
> > problematic or certainly cause anything observable beyond what is
> > observable now.
> >
> > So I'm going to have to say, unless it can be clearly demonstrated this is
> > problematic, I don't think there's any reason to add additional complexity
> > here.
>
> It would be nice to measure munmap() overhead here.
>
> I am worried about hitting DEFAULT_MAX_RCU_BLIMIT and trigger
> rcu_force_quiescent_state() which can be disruptive to the system.
>
> DEFAULT_MAX_RCU_BLIMIT is 10K, so it is ~20G of THP unmapped on x86.
>
> munmap() of 64G worth of THP should be enough to demonstrate the
> problem.

I mean you're going to hit that from RCU freeing page tables already, which
most architectures already do right?

So if RCU saturation is a problem, that problem already exists, but I've
not heard of that being a problem at all?

So you're going to have to demonstrate why this situation is markedly
different from that. And it's the same scale.

Overall I think freeing 64 GiB of mapped memory all at once will inevitably
be a slow operation, freeing them directly will also be a lengthily process.

And also it seems to me that RCU mishandling heavy load to the point of
causing system instability should a bug filed with RCU no?

Also note pte_free_defer() is already used in retract_page_tables() so a large
collapse could also hit this problem?

I'm not sure I'm convinced there's an issue here.

>
> --
>   Kiryl Shutsemau / Kirill A. Shutemov

--
Cheers, Lorenzo

_______________________________________________
linux-riscv mailing list
linux-riscv@lists.infradead.org
http://lists.infradead.org/mailman/listinfo/linux-riscv

  reply	other threads:[~2026-09-01 15:45 UTC|newest]

Thread overview: 117+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-09-01 11:01 [PATCH 00/12] mm: make userland page table freeing RCU-safe Lorenzo Stoakes (ARM)
2026-09-01 11:01 ` Lorenzo Stoakes (ARM)
2026-09-01 11:01 ` Lorenzo Stoakes (ARM)
2026-09-01 11:01 ` [PATCH 01/12] mm/huge_memory: zap deposited page tables after an RCU grace period Lorenzo Stoakes (ARM)
2026-09-01 11:01   ` Lorenzo Stoakes (ARM)
2026-09-01 11:01   ` Lorenzo Stoakes (ARM)
2026-09-01 11:18   ` sashiko-bot
2026-09-01 13:19   ` Kiryl Shutsemau
2026-09-01 13:19     ` Kiryl Shutsemau
2026-09-01 13:19     ` Kiryl Shutsemau
2026-09-01 14:12     ` Lorenzo Stoakes (ARM)
2026-09-01 14:12       ` Lorenzo Stoakes (ARM)
2026-09-01 14:12       ` Lorenzo Stoakes (ARM)
2026-09-01 14:24       ` Jason Gunthorpe
2026-09-01 14:24         ` Jason Gunthorpe
2026-09-01 14:24         ` Jason Gunthorpe
2026-09-01 14:41         ` Lorenzo Stoakes (ARM)
2026-09-01 14:41           ` Lorenzo Stoakes (ARM)
2026-09-01 14:41           ` Lorenzo Stoakes (ARM)
2026-09-01 15:28           ` Kiryl Shutsemau
2026-09-01 15:28             ` Kiryl Shutsemau
2026-09-01 15:28             ` Kiryl Shutsemau
2026-09-01 15:45             ` Lorenzo Stoakes (ARM) [this message]
2026-09-01 15:45               ` Lorenzo Stoakes (ARM)
2026-09-01 15:45               ` Lorenzo Stoakes (ARM)
2026-09-01 17:11               ` Kiryl Shutsemau
2026-09-01 17:11                 ` Kiryl Shutsemau
2026-09-01 17:11                 ` Kiryl Shutsemau
2026-09-01 17:14                 ` Lorenzo Stoakes (ARM)
2026-09-01 17:14                   ` Lorenzo Stoakes (ARM)
2026-09-01 17:14                   ` Lorenzo Stoakes (ARM)
2026-09-01 15:54             ` Liam R. Howlett
2026-09-01 15:54               ` Liam R. Howlett
2026-09-01 15:54               ` Liam R. Howlett
2026-09-01 16:06               ` Jason Gunthorpe
2026-09-01 16:06                 ` Jason Gunthorpe
2026-09-01 16:06                 ` Jason Gunthorpe
2026-09-01 17:13               ` Kiryl Shutsemau
2026-09-01 17:13                 ` Kiryl Shutsemau
2026-09-01 17:13                 ` Kiryl Shutsemau
2026-09-01 17:47                 ` Liam R. Howlett
2026-09-01 17:47                   ` Liam R. Howlett
2026-09-01 17:47                   ` Liam R. Howlett
2026-09-01 11:01 ` [PATCH 02/12] mm: enable MMU_GATHER_RCU_TABLE_FREE for most 2-level architectures Lorenzo Stoakes (ARM)
2026-09-01 11:01   ` Lorenzo Stoakes (ARM)
2026-09-01 11:01   ` Lorenzo Stoakes (ARM)
2026-09-01 11:22   ` sashiko-bot
2026-09-01 11:01 ` [PATCH 03/12] mm: enable MMU_GATHER_RCU_TABLE_FREE for MMU riscv Lorenzo Stoakes (ARM)
2026-09-01 11:01   ` Lorenzo Stoakes (ARM)
2026-09-01 11:01   ` Lorenzo Stoakes (ARM)
2026-09-01 11:14   ` sashiko-bot
2026-09-01 11:01 ` [PATCH 04/12] mm: enable MMU_GATHER_RCU_TABLE_FREE for MMU arm Lorenzo Stoakes (ARM)
2026-09-01 11:01   ` Lorenzo Stoakes (ARM)
2026-09-01 11:01   ` Lorenzo Stoakes (ARM)
2026-09-01 11:18   ` sashiko-bot
2026-09-01 11:01 ` [PATCH 05/12] mm: enable MMU_GATHER_RCU_TABLE_FREE for arc, microblaze, xtensa Lorenzo Stoakes (ARM)
2026-09-01 11:01   ` Lorenzo Stoakes (ARM)
2026-09-01 11:01   ` Lorenzo Stoakes (ARM)
2026-09-01 11:28   ` sashiko-bot
2026-09-07 21:07   ` Suren Baghdasaryan
2026-09-07 21:07     ` Suren Baghdasaryan
2026-09-07 21:07     ` Suren Baghdasaryan
2026-09-08 11:28     ` Lorenzo Stoakes (ARM)
2026-09-08 11:28       ` Lorenzo Stoakes (ARM)
2026-09-08 11:28       ` Lorenzo Stoakes (ARM)
2026-09-01 11:01 ` [PATCH 06/12] mm: enable MMU_GATHER_RCU_TABLE_FREE for sparc64 Lorenzo Stoakes (ARM)
2026-09-01 11:01   ` Lorenzo Stoakes (ARM)
2026-09-01 11:01   ` Lorenzo Stoakes (ARM)
2026-09-01 11:23   ` sashiko-bot
2026-09-01 11:01 ` [PATCH 07/12] mm: enable MMU_GATHER_RCU_TABLE_FREE for m68k-coldfire Lorenzo Stoakes (ARM)
2026-09-01 11:01   ` Lorenzo Stoakes (ARM)
2026-09-01 11:01   ` Lorenzo Stoakes (ARM)
2026-09-01 11:23   ` sashiko-bot
2026-09-01 11:01 ` [PATCH 08/12] mm: enable MMU_GATHER_RCU_TABLE_FREE for sh-X2 Lorenzo Stoakes (ARM)
2026-09-01 11:01   ` Lorenzo Stoakes (ARM)
2026-09-01 11:01   ` Lorenzo Stoakes (ARM)
2026-09-01 11:30   ` sashiko-bot
2026-09-01 11:01 ` [PATCH 09/12] mm: enable MMU_GATHER_RCU_TABLE_FREE for m68k-motorola Lorenzo Stoakes (ARM)
2026-09-01 11:01   ` Lorenzo Stoakes (ARM)
2026-09-01 11:01   ` Lorenzo Stoakes (ARM)
2026-09-01 11:30   ` sashiko-bot
2026-09-01 11:01 ` [PATCH 10/12] mm: enable MMU_GATHER_RCU_TABLE_FREE for sparc32 Lorenzo Stoakes (ARM)
2026-09-01 11:01   ` Lorenzo Stoakes (ARM)
2026-09-01 11:01   ` Lorenzo Stoakes (ARM)
2026-09-01 11:37   ` sashiko-bot
2026-09-01 11:01 ` [PATCH 11/12] mm: make userland page table freeing RCU-safe Lorenzo Stoakes (ARM)
2026-09-01 11:01   ` Lorenzo Stoakes (ARM)
2026-09-01 11:01   ` Lorenzo Stoakes (ARM)
2026-09-01 11:52   ` sashiko-bot
2026-09-01 13:33   ` Kiryl Shutsemau
2026-09-01 13:33     ` Kiryl Shutsemau
2026-09-01 13:33     ` Kiryl Shutsemau
2026-09-01 14:03     ` Lorenzo Stoakes (ARM)
2026-09-01 14:03       ` Lorenzo Stoakes (ARM)
2026-09-01 14:03       ` Lorenzo Stoakes (ARM)
2026-09-01 11:01 ` [PATCH 12/12] mm: change the contract for free_pgtables(), update docs Lorenzo Stoakes (ARM)
2026-09-01 11:01   ` Lorenzo Stoakes (ARM)
2026-09-01 11:01   ` Lorenzo Stoakes (ARM)
2026-09-01 11:37   ` sashiko-bot
2026-09-01 13:57   ` Kiryl Shutsemau
2026-09-01 13:57     ` Kiryl Shutsemau
2026-09-01 13:57     ` Kiryl Shutsemau
2026-09-01 14:31     ` Lorenzo Stoakes (ARM)
2026-09-01 14:31       ` Lorenzo Stoakes (ARM)
2026-09-01 14:31       ` Lorenzo Stoakes (ARM)
2026-09-01 17:15       ` Kiryl Shutsemau
2026-09-01 17:15         ` Kiryl Shutsemau
2026-09-01 17:15         ` Kiryl Shutsemau
2026-09-01 17:25         ` Lorenzo Stoakes (ARM)
2026-09-01 17:25           ` Lorenzo Stoakes (ARM)
2026-09-01 17:25           ` Lorenzo Stoakes (ARM)
2026-09-07 16:53           ` Suren Baghdasaryan
2026-09-07 16:53             ` Suren Baghdasaryan
2026-09-07 16:53             ` Suren Baghdasaryan
2026-09-08 11:29             ` Lorenzo Stoakes (ARM)
2026-09-08 11:29               ` Lorenzo Stoakes (ARM)
2026-09-08 11:29               ` Lorenzo Stoakes (ARM)

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=apbvoN6ttgHfI_xg@gremlin \
    --to=ljs@kernel.org \
    --cc=James.Bottomley@hansenpartnership.com \
    --cc=agordeev@linux.ibm.com \
    --cc=akpm@linux-foundation.org \
    --cc=alex@ghiti.fr \
    --cc=andreas@gaisler.com \
    --cc=aneesh.kumar@kernel.org \
    --cc=anton.ivanov@cambridgegreys.com \
    --cc=aou@eecs.berkeley.edu \
    --cc=arnd@arndb.de \
    --cc=baohua@kernel.org \
    --cc=baolin.wang@linux.alibaba.com \
    --cc=bcain@kernel.org \
    --cc=borntraeger@linux.ibm.com \
    --cc=bp@alien8.de \
    --cc=catalin.marinas@arm.com \
    --cc=chenhuacai@kernel.org \
    --cc=chleroy@kernel.org \
    --cc=chris@zankel.net \
    --cc=dalias@libc.org \
    --cc=dave.hansen@linux.intel.com \
    --cc=davem@davemloft.net \
    --cc=david@kernel.org \
    --cc=deller@gmx.de \
    --cc=dev.jain@arm.com \
    --cc=dinguyen@kernel.org \
    --cc=geert@linux-m68k.org \
    --cc=glaubitz@physik.fu-berlin.de \
    --cc=gor@linux.ibm.com \
    --cc=guoren@kernel.org \
    --cc=hca@linux.ibm.com \
    --cc=hpa@zytor.com \
    --cc=hughd@google.com \
    --cc=jcmvbkbc@gmail.com \
    --cc=jgg@ziepe.ca \
    --cc=jhubbard@nvidia.com \
    --cc=johannes@sipsolutions.net \
    --cc=jonas@southpole.se \
    --cc=kas@kernel.org \
    --cc=kernel@xen0n.name \
    --cc=lance.yang@linux.dev \
    --cc=liam@infradead.org \
    --cc=linmag7@gmail.com \
    --cc=linux-alpha@vger.kernel.org \
    --cc=linux-arch@vger.kernel.org \
    --cc=linux-arm-kernel@lists.infradead.org \
    --cc=linux-csky@vger.kernel.org \
    --cc=linux-hexagon@vger.kernel.org \
    --cc=linux-kernel@vger.kernel.org \
    --cc=linux-m68k@lists.linux-m68k.org \
    --cc=linux-mips@vger.kernel.org \
    --cc=linux-mm@kvack.org \
    --cc=linux-openrisc@vger.kernel.org \
    --cc=linux-parisc@vger.kernel.org \
    --cc=linux-riscv@lists.infradead.org \
    --cc=linux-s390@vger.kernel.org \
    --cc=linux-sh@vger.kernel.org \
    --cc=linux-snps-arc@lists.infradead.org \
    --cc=linux-um@lists.infradead.org \
    --cc=linux@armlinux.org.uk \
    --cc=linuxppc-dev@lists.ozlabs.org \
    --cc=loongarch@lists.linux.dev \
    --cc=maddy@linux.ibm.com \
    --cc=mark.rutland@arm.com \
    --cc=mattst88@gmail.com \
    --cc=mhocko@suse.com \
    --cc=mingo@redhat.com \
    --cc=monstr@monstr.eu \
    --cc=mpe@ellerman.id.au \
    --cc=nico.pache@linux.dev \
    --cc=npiggin@gmail.com \
    --cc=palmer@dabbelt.com \
    --cc=peterx@redhat.com \
    --cc=peterz@infradead.org \
    --cc=pjw@kernel.org \
    --cc=qi.zheng@linux.dev \
    --cc=richard.henderson@linaro.org \
    --cc=richard@nod.at \
    --cc=rppt@kernel.org \
    --cc=ryan.roberts@arm.com \
    --cc=schuster.simon@siemens-energy.com \
    --cc=shorne@gmail.com \
    --cc=sparclinux@vger.kernel.org \
    --cc=stefan.kristiansson@saunalahti.fi \
    --cc=surenb@google.com \
    --cc=svens@linux.ibm.com \
    --cc=tglx@kernel.org \
    --cc=tsbogend@alpha.franken.de \
    --cc=usama.arif@linux.dev \
    --cc=vbabka@kernel.org \
    --cc=vgupta@kernel.org \
    --cc=will@kernel.org \
    --cc=x86@kernel.org \
    --cc=ysato@users.sourceforge.jp \
    --cc=ziy@nvidia.com \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is an external index of several public inboxes,
see mirroring instructions on how to clone and mirror
all data and code used by this external index.