From: Yeoreum Yun <yeoreum.yun@arm.com>
To: Dave Hansen <dave.hansen@intel.com>
Cc: Yeoreum Yun <yeoreum.yun@arm.com>,
Russell King <linux@armlinux.org.uk>,
Huacai Chen <chenhuacai@kernel.org>,
WANG Xuerui <kernel@xen0n.name>,
Thomas Bogendoerfer <tsbogend@alpha.franken.de>,
Catalin Marinas <catalin.marinas@arm.com>,
Will Deacon <will@kernel.org>, Arnd Bergmann <arnd@arndb.de>,
Andrew Morton <akpm@linux-foundation.org>,
Kairui Song <kasong@tencent.com>, Qi Zheng <qi.zheng@linux.dev>,
Shakeel Butt <shakeel.butt@linux.dev>,
Barry Song <baohua@kernel.org>,
Axel Rasmussen <axelrasmussen@google.com>,
Yuanchu Xie <yuanchu@google.com>, Wei Xu <weixugc@google.com>,
Johannes Weiner <hannes@cmpxchg.org>,
David Hildenbrand <david@kernel.org>,
Michal Hocko <mhocko@kernel.org>,
Lorenzo Stoakes <ljs@kernel.org>,
Tianrui Zhao <zhaotianrui@loongson.cn>,
Bibo Mao <maobibo@loongson.cn>, Anup Patel <anup@brainfault.org>,
Atish Patra <atish.patra@linux.dev>,
Paul Walmsley <pjw@kernel.org>,
Palmer Dabbelt <palmer@dabbelt.com>,
Albert Ou <aou@eecs.berkeley.edu>,
Alexandre Ghiti <alex@ghiti.fr>,
Dave Hansen <dave.hansen@linux.intel.com>,
Andy Lutomirski <luto@kernel.org>,
Peter Zijlstra <peterz@infradead.org>,
Thomas Gleixner <tglx@kernel.org>, Ingo Molnar <mingo@redhat.com>,
Borislav Petkov <bp@alien8.de>,
x86@kernel.org, "H. Peter Anvin" <hpa@zytor.com>,
"Liam R. Howlett" <liam@infradead.org>,
Vlastimil Babka <vbabka@kernel.org>,
Mike Rapoport <rppt@kernel.org>,
Suren Baghdasaryan <surenb@google.com>,
Michal Hocko <mhocko@suse.com>, Jonas Bonn <jonas@southpole.se>,
Stefan Kristiansson <stefan.kristiansson@saunalahti.fi>,
Stafford Horne <shorne@gmail.com>,
linux-arm-kernel@lists.infradead.org,
linux-kernel@vger.kernel.org, loongarch@lists.linux.dev,
linux-mips@vger.kernel.org, linux-arch@vger.kernel.org,
linux-mm@kvack.org, kvm@vger.kernel.org,
kvm-riscv@lists.infradead.org, linux-riscv@lists.infradead.org,
linux-openrisc@vger.kernel.org
Subject: Re: [PATCH RFC v2 14/20] x86: mm: skip pud setup when using generic compile-time folded pagetable
Date: Wed, 22 Jul 2026 18:18:33 +0100 [thread overview]
Message-ID: <amD7ab_KuJPdBQBh@e129823.arm.com> (raw)
In-Reply-To: <0ea931f1-4801-4d32-b6ce-4787a8d3daa9@intel.com>
On Wed, Jul 22, 2026 at 09:33:55AM -0700, Dave Hansen wrote:
> On 7/22/26 08:30, Yeoreum Yun wrote:
> > We want to rework how set_pXd() behaves for generic compile-time folded
>
> Please move this all to imperative voice. I think I asked for this
> before, but perhaps I didn't. Either way, please fix the whole series.
Oh. Sorry, I've forgotten to edit this commit message.
I'll fix in next around.
>
> > page tables by disallowing its use and triggering a compile-time error
> > when it is used improperly, ensuring that the actual first-level set_pXd()
> > function is used instead.
> >
> > The behaviour of pXd_page() will change with generic compile-time folded
> > page tables by disallowing its use and triggering a compile-time error
> > when it's used improperly, ensuring that the actual.
> >
> > To prepare fot that, skip collapse_pud_page() and populate_pud() according
> > to CONFIG_PGTABLE_LEVELS.
> >
> > There should be no functional change.
>
> If this wasn't done, what would happen? There would be a compile error,
> right?
>
> Shouldn't that be said out loud somewhere?
Yes compile error for using pud_page(). I'll shout out detail in the
commit message.
>
> > diff --git a/arch/x86/mm/pat/set_memory.c b/arch/x86/mm/pat/set_memory.c
> > index 078689aa7206..91b33a265a43 100644
> > --- a/arch/x86/mm/pat/set_memory.c
> > +++ b/arch/x86/mm/pat/set_memory.c
> > @@ -1346,7 +1346,7 @@ static int collapse_pud_page(pud_t *pud, unsigned long addr,
> > pmd_t *pmd, first;
> > int i;
> >
> > - if (!direct_gbpages)
> > + if (CONFIG_PGTABLE_LEVELS == 2 || !direct_gbpages)
> > return 0;
>
> This is subtly wrong. It probably shuts the compiler up, but it's subtly
> wrong.
>
> x86 has 4 paging modes. I'll use the SDM/Intel terminology:
>
> 32-bit:
> * 32-bit paging: 2-level
> * PAE: 3-level
> 64-bit:
> * 4-level paging
> * 5-level paging
>
> But "gbpages" are not available in PAE paging. It's a hardware
> limitation. This doesn't cause a functional problem because the pud
> level is not folded. But it is confusing and not really an accurate way
> to write this check.
>
> The *right* way to do this is probably something like this in a header:
>
> static inline bool direct_gbpages(void)
> {
> if (!IS_ENABLED(CONFIG_X86_DIRECT_GBPAGES))
> return false;
>
> return __direct_gbpages;
> }
>
> Which just happens to point readers over to this Kconfig nugget:
>
> config X86_DIRECT_GBPAGES
> def_bool y
> depends on X86_64
>
> which, combined with the knowledge that 32-bit has a 3-level mode, would
> make the proposed check obviously wrong.
>
> So I'll take this as a suitable alternative:
>
> /* Avoid compiling the below code if PUD-level mappings are impossible: */
> if (IS_ENABLED(CONFIG_X86_DIRECT_GBPAGES))
> return 0;
>
> if (!direct_gbpages)
> return 0;
>
> That will compile-time optimize the code away in the exact right
> conditions *and* fix the (assumed by me) compile error that you were
> chasing.
I thought the same. but It seemed odd at the time since direct_gbpages is
always 0 when !CONFIG_X86_DIRECT_GBPAGES.
Howver, I've missed the X86_DIRECT_GBPAGES is for X86_64 only and It was
wrong totally.
I'll follow your suggestion.
> > @@ -1730,7 +1730,8 @@ static int populate_pud(struct cpa_data *cpa, unsigned long start, p4d_t *p4d,
> > /*
> > * Map everything starting from the Gb boundary, possibly with 1G pages
> > */
> > - while (boot_cpu_has(X86_FEATURE_GBPAGES) && end - start >= PUD_SIZE) {
> > + while (CONFIG_PGTABLE_LEVELS > 3 && boot_cpu_has(X86_FEATURE_GBPAGES) &&
> > + end - start >= PUD_SIZE) {
> > set_pud(pud, pud_mkhuge(pfn_pud(cpa->pfn,
> > canon_pgprot(pud_pgprot))));
>
> This is an OK approach. But there's a way to fix this site *and*
> optimize a non-zero amount of other code at the same time. Add this hunk
> to arch/x86/Kconfig.cpufeatures:
>
> config X86_DISABLED_FEATURE_GBPAGES
> def_bool y
> depends on X86_32
>
> That will turn the boot_cpu_has() check in to something that can be
> resolved at compile time. It has the added advantage of compiling out
> all of the code under X86_FEATURE_GBPAGES everywhere else in the tree.
Does it? when I glimpse check, this wouldn't be compiled since
there is no bit for X86_DISABLED_FEATURE_GBPAGES and defining the
DISABLED bit for FEATURE_GBPAGES seems odd since bit X86_FEATURE_GBPAGES
is already defined.
In stead of CONFIG_PGTABLE_LEVEL > 3, as above, would it be better to
add check IS_ENABLED(CONFIG_X86_DIRECT_GBPAGES)?
Thanks.
--
Sincerely,
Yeoreum Yun
next prev parent reply other threads:[~2026-07-22 17:18 UTC|newest]
Thread overview: 50+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-07-22 15:30 [PATCH RFC v2 00/20] mm: optimize unnecessary loads due to ptep_get() and friends out Yeoreum Yun
2026-07-22 15:30 ` [PATCH RFC v2 01/20] ARM: mm: make nommu pgd_t a scalar Yeoreum Yun
2026-07-22 15:30 ` [PATCH RFC v2 02/20] ARM: mm: make 2-level " Yeoreum Yun
2026-07-22 15:30 ` [PATCH RFC v2 03/20] ARM: mm: remove custom pgdp_get() Yeoreum Yun
2026-07-22 15:30 ` [PATCH RFC v2 04/20] LoongArch: mm: define pud_leaf() only when PUD exists Yeoreum Yun
2026-07-22 15:30 ` [PATCH RFC v2 05/20] MIPS: " Yeoreum Yun
2026-07-22 15:30 ` [PATCH RFC v2 06/20] mm/pgtable: define (pgd|p4d|pud)_leaf() for folded page tables Yeoreum Yun
2026-07-22 15:51 ` sashiko-bot
2026-07-22 15:30 ` [PATCH RFC v2 07/20] mm/pgtable: define (pgd|p4d|pud)_offset_lockless() " Yeoreum Yun
2026-07-22 15:58 ` sashiko-bot
2026-07-22 19:25 ` Yeoreum Yun
2026-07-22 15:30 ` [PATCH RFC v2 08/20] mm: vmscan: remove stack copy address of pud pass in wallk_pud_range() Yeoreum Yun
2026-07-22 15:56 ` sashiko-bot
2026-07-22 19:31 ` Yeoreum Yun
2026-07-22 15:30 ` [PATCH RFC v2 09/20] loongarch: kvm: remove stack copy address of pXd in pXd_offset() Yeoreum Yun
2026-07-22 15:53 ` sashiko-bot
2026-07-22 19:51 ` Yeoreum Yun
2026-07-22 15:30 ` [PATCH RFC v2 10/20] riscv: " Yeoreum Yun
2026-07-22 15:30 ` [PATCH RFC v2 11/20] riscv: mm: use proper set_pXd() for generic compile-time folded patable in vmalloc_fault() Yeoreum Yun
2026-07-22 15:30 ` [PATCH RFC v2 12/20] x86: mm: define pudp_set_access_flags() when CONFIG_HAVE_ARCH_TRANSPARENT_HUGEPAGE_PUD is enabled only Yeoreum Yun
2026-07-22 16:02 ` Dave Hansen
2026-07-22 17:27 ` Yeoreum Yun
2026-07-22 15:30 ` [PATCH RFC v2 13/20] x86: mm: carve out the generic compile-time folded pgtable case in effective_prot() Yeoreum Yun
2026-07-22 16:07 ` sashiko-bot
2026-07-22 16:23 ` Yeoreum Yun
2026-07-22 16:11 ` Dave Hansen
2026-07-22 16:28 ` Yeoreum Yun
2026-07-22 17:37 ` Yeoreum Yun
2026-07-22 20:20 ` Dave Hansen
2026-07-22 21:00 ` Yeoreum Yun
2026-07-22 15:30 ` [PATCH RFC v2 14/20] x86: mm: skip pud setup when using generic compile-time folded pagetable Yeoreum Yun
2026-07-22 16:18 ` sashiko-bot
2026-07-22 19:08 ` Yeoreum Yun
2026-07-22 16:33 ` Dave Hansen
2026-07-22 17:18 ` Yeoreum Yun [this message]
2026-07-22 20:02 ` Dave Hansen
2026-07-22 20:18 ` Yeoreum Yun
2026-07-22 20:22 ` Dave Hansen
2026-07-22 20:28 ` H. Peter Anvin
2026-07-22 20:40 ` Yeoreum Yun
2026-07-22 15:30 ` [PATCH RFC v2 15/20] mm/pgtable: optimize pmdp_get() and friends for folded pagetable levels Yeoreum Yun
2026-07-22 16:15 ` sashiko-bot
2026-07-22 19:44 ` Yeoreum Yun
2026-07-22 15:30 ` [PATCH RFC v2 16/20] mm/pgtable: catch abuse of folded dummy pgd_t/p4d_t/pud_t Yeoreum Yun
2026-07-22 15:30 ` [PATCH RFC v2 17/20] mm/pgtable: disallow calling (pgd|p4d|pud)_page, pgd_page_vaddr() and (p4d|pud)_pgtable with dummy Yeoreum Yun
2026-07-22 15:30 ` [PATCH RFC v2 18/20] mm/pgtable: disallow calling folded set_pgd/set_p4d/set_pud " Yeoreum Yun
2026-07-22 15:30 ` [PATCH RFC v2 19/20] openrisc/pgtable: drop __pmd_offset() Yeoreum Yun
2026-07-22 15:30 ` [PATCH RFC v2 20/20] arm64: pgtable: convert pte_present() from macro to static inline Yeoreum Yun
2026-07-22 16:40 ` [PATCH RFC v2 00/20] mm: optimize unnecessary loads due to ptep_get() and friends out Dave Hansen
2026-07-22 17:30 ` Yeoreum Yun
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=amD7ab_KuJPdBQBh@e129823.arm.com \
--to=yeoreum.yun@arm.com \
--cc=akpm@linux-foundation.org \
--cc=alex@ghiti.fr \
--cc=anup@brainfault.org \
--cc=aou@eecs.berkeley.edu \
--cc=arnd@arndb.de \
--cc=atish.patra@linux.dev \
--cc=axelrasmussen@google.com \
--cc=baohua@kernel.org \
--cc=bp@alien8.de \
--cc=catalin.marinas@arm.com \
--cc=chenhuacai@kernel.org \
--cc=dave.hansen@intel.com \
--cc=dave.hansen@linux.intel.com \
--cc=david@kernel.org \
--cc=hannes@cmpxchg.org \
--cc=hpa@zytor.com \
--cc=jonas@southpole.se \
--cc=kasong@tencent.com \
--cc=kernel@xen0n.name \
--cc=kvm-riscv@lists.infradead.org \
--cc=kvm@vger.kernel.org \
--cc=liam@infradead.org \
--cc=linux-arch@vger.kernel.org \
--cc=linux-arm-kernel@lists.infradead.org \
--cc=linux-kernel@vger.kernel.org \
--cc=linux-mips@vger.kernel.org \
--cc=linux-mm@kvack.org \
--cc=linux-openrisc@vger.kernel.org \
--cc=linux-riscv@lists.infradead.org \
--cc=linux@armlinux.org.uk \
--cc=ljs@kernel.org \
--cc=loongarch@lists.linux.dev \
--cc=luto@kernel.org \
--cc=maobibo@loongson.cn \
--cc=mhocko@kernel.org \
--cc=mhocko@suse.com \
--cc=mingo@redhat.com \
--cc=palmer@dabbelt.com \
--cc=peterz@infradead.org \
--cc=pjw@kernel.org \
--cc=qi.zheng@linux.dev \
--cc=rppt@kernel.org \
--cc=shakeel.butt@linux.dev \
--cc=shorne@gmail.com \
--cc=stefan.kristiansson@saunalahti.fi \
--cc=surenb@google.com \
--cc=tglx@kernel.org \
--cc=tsbogend@alpha.franken.de \
--cc=vbabka@kernel.org \
--cc=weixugc@google.com \
--cc=will@kernel.org \
--cc=x86@kernel.org \
--cc=yuanchu@google.com \
--cc=zhaotianrui@loongson.cn \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox