From: Nadia Chambers <nadia.yvette.chambers@ik.me>
To: "David Hildenbrand (Arm)" <david@kernel.org>
Cc: Kiryl Shutsemau <kas@kernel.org>,
lsf-pc@lists.linux-foundation.org, linux-mm@kvack.org,
x86@kernel.org, linux-kernel@vger.kernel.org,
Andrew Morton <akpm@linux-foundation.org>,
Thomas Gleixner <tglx@linutronix.de>,
Ingo Molnar <mingo@redhat.com>, Borislav Petkov <bp@alien8.de>,
Dave Hansen <dave.hansen@linux.intel.com>,
Lorenzo Stoakes <lorenzo.stoakes@oracle.com>,
"Liam R. Howlett" <Liam.Howlett@oracle.com>,
Mike Rapoport <rppt@kernel.org>,
Matthew Wilcox <willy@infradead.org>,
Johannes Weiner <hannes@cmpxchg.org>,
Usama Arif <usama.arif@linux.dev>
Subject: Re: [LSF/MM/BPF TOPIC] 64k (or 16k) base page size on x86
Date: Mon, 20 Jul 2026 00:24:12 +0200 [thread overview]
Message-ID: <al0xogLqubAZh6SD@ik.me> (raw)
In-Reply-To: <17c5708d-3859-49a5-814e-bc3564bc3ac6@kernel.org>
Am Fr, Feb 20, 2026 um 11:24:37 +0100, David Hildenbrand (Arm) schrieb:
> Right, see the proposal from Dev on the list.
> From user-space POV, the pagesize would be 64K for these emulated processes.
> That is, VMAs must be suitable aligned etc.
> One key thing I think is that you could run such emulated-64k process (that
> actually support it!) with 4k processes on the same machine, like Arm is
> considering.
> You would have no weird "vma crosses base pages" handling, which is just
> rather nasty and makes my head hurt.
While it had to be written, it was a straightforward extension of
converting ->vm_pgoff to being in MMUPAGE_SIZE (Kiryl's PTE_SIZE) units,
recovering offsets into pages usw. Nothing seemed terribly hard to
understand about it. Really, the code as it stands rarely approaches
what it's doing in such a way that it broadly surveys the layout of vmas
and pages and be conscious of whether pages spanned vmas or multiple
vmas were representing maps of of fragments of pages, so I'm not
convinced there's anything particularly unique to be afraid of here.
Am Fr, Feb 20, 2026 um 11:24:37 +0100, David Hildenbrand (Arm) schrieb:
> Well, yes, like Willy says, there are already similar custom solutions for
> s390x and ppc.
> Pasha talked recently about the memory waste of 16k kernel stacks and how we
> would want to reduce that to 4k. In your proposal, it would be 64k, unless
> you somehow manage to allocate multiple kernel stacks from the same 64k
> page. My head hurts thinking about whether that could work, maybe it could
> (no idea about guard pages in there, though).
Never mind wasting memory; just getting it booting in 2003 needed
adjusting things so the THREAD_SIZE stayed constant and there were
things in the core kernel that got a divide by zero or some such when
THREAD_SIZE was too large. The effort for the fix in 2003 was just
getting adequate diagnostics, not the stack allocation code itself.
The mechanically assisted forward port did something that got it booting
that, upon critical examination, I don't like and uses a lot of space.
vmalloc and guard pages don't have anything special about them either,
apart from taking work to implement, which I've done before as part of
what appears to be a now non-extant patch kit, stack_paranoia.
Am Fr, Feb 20, 2026 um 11:24:37 +0100, David Hildenbrand (Arm) schrieb:
> Let's take a look at the history of page size usage on Arm (people can feel
> free to correct me):
> (1) Most distros were using 64k on Arm.
> (2) People realized that 64k was suboptimal many use cases (memory
> waste for stacks, pagecache, etc) and started to switch to 4k. I
> remember that mostly HPC-centric users sticked to 64k, but there was
> also demand from others to be able to stay on 64k.
> (3) Arm improved performance on a 4k kernel by adding cont-pte support,
> trying to get closer to 64k native performance.
> (4) Achieving 64k native performance is hard, which is why per-process
> page sizes are being explored to get the best out of both worlds
> (use 64k page size only where it really matters for performance).
There are ways to deal with internal fragmentation, like tail packing
for the pagecache, switching to slab for a lot of things, potentially
even pagetables, so they don't grow with PAGE_SIZE usw. Stacks might
take more work than the average kernel data structure, but I've
implemented changes to stack allocation before at points when the
vmalloc and size varying usw. usf. stack bits weren't in mainline.
Bitblitting latencies are also issues to try to keep bounded that I've
seen little discussion of here yet.
I suppose the bad news is that I didn't get the chance to actually work
on the internal fragmentation mitigation strategies I wrote out plans
for 23 years ago. They weren't that involved, though.
Am Fr, Feb 20, 2026 um 11:24:37 +0100, David Hildenbrand (Arm) schrieb:
> Arm clearly has the added benefit of actually benefiting from hardware
> support for 64k.
> IIUC, what you are proposing feels a bit like traveling back in time when it
> comes to the memory waste problem that Arm users encountered.
> Where do you see the big difference to 64k on Arm in your proposal? Would
> you currently also be running 64k Arm in production and the memory waste etc
> is acceptable?
My original 2003 intentions for the use of Hugh's ABI compatibility
technique were to make page clustering as a solution for 32-bit large
memory „forward-looking to 64-bit“, where at the time, 64-bit meant
IA64, by using it to provide a guarantee that small superpage
allocations would never fail due to external fragmentation, which design
goal came from above. While some colouring benefits usw. were
anticipated as byproducts, this idea about how to handle small
superpages was the primary way it was conceived of as addressing 64-bit
concerns, perhaps in part by dint of IA64-centrism, or potentially a
kind of inheritance of a requirement for an equivalent of the
very-differently-implemented mechanism for the same from HP-UX. It also
seemed rather safe to set the goal of „raising the floor“ so the very
smallest TLB entries wouldn't proliferate, and the smallest page sizes
consume a lesser fraction of available TLB entries.
Proposals for trying to use the technique as some kind of fully general
method for reaching for 1 GiB if not larger pages seem ill-conceived to
me. Those kinds of issues demand changes akin to transforming the
problem so it's in an entirely lesser complexity class, not adding a
small constant to the methods already in place attempting to deal with
large-scale external fragmentation. I liked Rik's partitioning memory
into arenas of pre-constructed 1 GiB or otherwise very large superpages
as a strategy for that, for instance.
I generally care enough about architectural coverage that my recent
mechanically assisted efforts maintained an 80+ -cell regression testing
matrix of 20+ architectures with PAGE_MMUSHIFT values of 0, 2, 4 and 6.
ARM notably represented an extra architecture variant via LPAE. The arch
code I cry about was not having got the chance to use the recent MIPS
PageGrain extension to provide a 1 KiB MMUPAGE_SIZE and a PAGE_MMUSHIFT
of 8 for demonstration purposes. If I had got that done, I might have
even emailed Babaoğlu, Joy, and possibly even McKusick about it, though
Karels is reputedly the most likely suspect for who outside of Babaoğlu
and Joy might have committed CLSIZE u.a. to 3BSD's VAX VM code in 1979.
I made at least earnest attempts to cover every architecture in the
kernel. There was one I couldn't find a toolchain for that just wired
all of the MMUPAGE macros to be identical to their PAGE analogues and
hoped at least built (arc?), but otherwise, I was trying to cover them
all, even if at the moment only in QEMU. I hope I can put a personal
many-architecture testing lab like I had in the '00s together for
myself again at some point, given the chance, but they'll have to be
small things like for embedded this time with how flats in Berlin are.
-- nyc
next prev parent reply other threads:[~2026-07-19 22:24 UTC|newest]
Thread overview: 60+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-02-19 15:08 [LSF/MM/BPF TOPIC] 64k (or 16k) base page size on x86 Kiryl Shutsemau
2026-02-19 15:17 ` Peter Zijlstra
2026-02-19 15:20 ` Peter Zijlstra
2026-02-19 15:27 ` Kiryl Shutsemau
2026-02-19 15:33 ` Pedro Falcato
2026-02-19 15:50 ` Kiryl Shutsemau
2026-02-19 15:53 ` David Hildenbrand (Arm)
2026-02-19 19:31 ` Pedro Falcato
2026-02-19 15:39 ` David Hildenbrand (Arm)
2026-02-19 15:54 ` Kiryl Shutsemau
2026-02-19 16:09 ` David Hildenbrand (Arm)
2026-02-20 2:55 ` Zi Yan
2026-07-19 17:30 ` Nadia Chambers
2026-02-19 17:09 ` Kiryl Shutsemau
2026-02-20 10:24 ` David Hildenbrand (Arm)
2026-02-20 12:07 ` Kiryl Shutsemau
2026-02-20 16:30 ` David Hildenbrand (Arm)
2026-02-20 19:33 ` Kalesh Singh
2026-02-23 11:04 ` David Hildenbrand (Arm)
2026-02-23 11:13 ` Kiryl Shutsemau
2026-02-23 11:27 ` David Hildenbrand (Arm)
2026-02-23 12:16 ` Kiryl Shutsemau
2026-02-23 15:14 ` Dave Hansen
2026-02-23 15:31 ` David Hildenbrand (Arm)
2026-02-23 15:45 ` Kiryl Shutsemau
2026-02-23 15:49 ` David Hildenbrand (Arm)
2026-02-23 16:22 ` Lorenzo Stoakes
2026-02-23 16:34 ` David Laight
2026-07-19 22:24 ` Nadia Chambers [this message]
2026-07-20 8:02 ` David Hildenbrand (Arm)
2026-07-20 8:49 ` Nadia Chambers
2026-02-19 23:24 ` Kalesh Singh
2026-02-20 12:10 ` Kiryl Shutsemau
2026-02-20 19:21 ` Kalesh Singh
2026-02-19 17:08 ` Dave Hansen
2026-02-19 22:05 ` Kiryl Shutsemau
2026-02-20 3:28 ` Liam R. Howlett
2026-02-20 12:33 ` Kiryl Shutsemau
2026-02-20 15:17 ` Liam R. Howlett
2026-02-20 15:50 ` Kiryl Shutsemau
2026-07-19 3:42 ` Nadia Chambers
2026-07-19 5:23 ` Hillf Danton
2026-07-19 6:09 ` Nadia Chambers
2026-02-19 17:30 ` Dave Hansen
2026-02-19 22:14 ` Kiryl Shutsemau
2026-02-19 22:21 ` Dave Hansen
2026-02-19 17:47 ` Matthew Wilcox
2026-02-19 22:26 ` Kiryl Shutsemau
2026-07-19 23:00 ` Nadia Chambers
2026-02-20 9:04 ` David Laight
2026-02-20 12:12 ` Kiryl Shutsemau
2026-07-19 22:31 ` Nadia Chambers
2026-04-29 14:39 ` Matthew Wilcox
2026-04-29 15:26 ` Kiryl Shutsemau
2026-05-01 18:05 ` David Hildenbrand (Arm)
2026-05-01 18:00 ` Kiryl Shutsemau
2026-05-01 18:02 ` David Hildenbrand (Arm)
2026-05-01 18:12 ` Kiryl Shutsemau
2026-05-01 18:31 ` David Hildenbrand (Arm)
2026-07-19 0:51 ` Nadia Chambers
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=al0xogLqubAZh6SD@ik.me \
--to=nadia.yvette.chambers@ik.me \
--cc=Liam.Howlett@oracle.com \
--cc=akpm@linux-foundation.org \
--cc=bp@alien8.de \
--cc=dave.hansen@linux.intel.com \
--cc=david@kernel.org \
--cc=hannes@cmpxchg.org \
--cc=kas@kernel.org \
--cc=linux-kernel@vger.kernel.org \
--cc=linux-mm@kvack.org \
--cc=lorenzo.stoakes@oracle.com \
--cc=lsf-pc@lists.linux-foundation.org \
--cc=mingo@redhat.com \
--cc=rppt@kernel.org \
--cc=tglx@linutronix.de \
--cc=usama.arif@linux.dev \
--cc=willy@infradead.org \
--cc=x86@kernel.org \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is an external index of several public inboxes,
see mirroring instructions on how to clone and mirror
all data and code used by this external index.