From: Matthew Wilcox <willy@infradead.org>
To: Artem Lytkin <iprintercanon@gmail.com>
Cc: Ackerley Tng <ackerleytng@google.com>,
David Hildenbrand <david@kernel.org>, Zi Yan <ziy@nvidia.com>,
Kiryl Shutsemau <kas@kernel.org>, Hugh Dickins <hughd@google.com>,
Baolin Wang <baolin.wang@linux.alibaba.com>,
linux-mm@kvack.org, linux-fsdevel@vger.kernel.org
Subject: Re: [RFC PATCH 0/4] Extend xas_split* to support splitting arbitrarily large entries
Date: Mon, 24 Aug 2026 21:31:28 +0100 [thread overview]
Message-ID: <aoyqIBz1EoY0R8qw@casper.infradead.org> (raw)
In-Reply-To: <20260824194755.216000-1-iprintercanon@gmail.com>
On Mon, Aug 24, 2026 at 10:47:55PM +0300, Artem Lytkin wrote:
> I'm asking because there is a second user for it that has nothing to do
> with guest_memfd. On arm64 with 64K pages HPAGE_PMD_ORDER is 13 and
> MAX_XAS_ORDER is 11, so MAX_PAGECACHE_ORDER clamps to 11 and shmem never
> offers a PMD order. MADV_COLLAPSE on tmpfs returns -EINVAL there, and
> huge=always hands out order-11 folios at best, so ShmemPmdMapped never
> leaves zero. Every other page cache file system tops out at order 11 as
> well, which is 128M and not PMD-mappable. pagemap.h admits it outright:
> "xas_split_alloc() does not support arbitrary orders. This implies no
> 512MB THP on ARM64 with 64KB base page size." Kiryl's khugepaged
> selftest series skips the shmem and file cases altogether for exactly
> this reason, and on 6 August he confirmed in that thread that uniform
> splits still go through xas_split_alloc(), so xas_try_split() on its
> own does not lift the cap.
But why would you want it? With 64KiB pages on arm64, you need 512MiB
for a PMD-sized page, and that's just too big for current hardware.
What's your bandwidth? 40GB per DIMM? With 8 DIMMs, that's an
aggregate of 320GB bandwidth, so that restricts you to 640 page faults
per second. That's too few for almost all workloads.
The page allocator also isn't great at doing order-13 allocations.
Since arm64 does support 2MiB TLBs, even with a 512MiB PMD size, I'm
not sure there's a huge performance win waiting for you on the other
side of all this work.
IOW is this driven by a real need, or just some completionist thing?
> That case needs two new levels of nodes rather than one: 130 nodes,
> 73 KiB, a long way from the 1G to 4K split Matthew was worried about.
> Raising MAX_XAS_ORDER to 3 * XA_CHUNK_SHIFT - 1 covers it, and
> MAX_PAGECACHE_ORDER stays min()'d with HPAGE_PMD_ORDER, so no page cache
> caller can ask for the order-17 worst case that cap admits.
130 nodes is still a lot to ask the slab allocator for! With 64KiB
pages it's not too bad -- you get 112 per page. But still ...
next prev parent reply other threads:[~2026-08-24 20:31 UTC|newest]
Thread overview: 16+ messages / expand[flat|nested] mbox.gz Atom feed top
2025-11-17 22:46 [RFC PATCH 0/4] Extend xas_split* to support splitting arbitrarily large entries Ackerley Tng
2025-11-17 22:46 ` [RFC PATCH 1/4] XArray: Initialize nodes while splitting instead of while allocating Ackerley Tng
2025-11-17 22:46 ` [RFC PATCH 2/4] XArray: Update xas_split_alloc() to allocate enough nodes to split large entries Ackerley Tng
2025-11-17 22:47 ` [RFC PATCH 3/4] XArray: Support splitting for arbitrarily " Ackerley Tng
2025-11-17 22:47 ` [RFC PATCH 4/4] XArray: test: Increase split order test range in check_split() Ackerley Tng
2025-12-08 3:32 ` Zi Yan
2025-12-08 16:47 ` Ackerley Tng
2025-11-17 23:22 ` [RFC PATCH 0/4] Extend xas_split* to support splitting arbitrarily large entries Matthew Wilcox
2025-11-17 23:43 ` Ackerley Tng
2025-11-18 8:51 ` David Hildenbrand (Red Hat)
2025-12-05 0:38 ` Ackerley Tng
2025-11-18 8:46 ` [syzbot ci] " syzbot ci
2026-08-24 19:47 ` [RFC PATCH 0/4] " Artem Lytkin
2026-08-24 20:31 ` Matthew Wilcox [this message]
2026-08-25 8:31 ` David Hildenbrand (Arm)
2026-08-25 9:08 ` Artem Lytkin
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=aoyqIBz1EoY0R8qw@casper.infradead.org \
--to=willy@infradead.org \
--cc=ackerleytng@google.com \
--cc=baolin.wang@linux.alibaba.com \
--cc=david@kernel.org \
--cc=hughd@google.com \
--cc=iprintercanon@gmail.com \
--cc=kas@kernel.org \
--cc=linux-fsdevel@vger.kernel.org \
--cc=linux-mm@kvack.org \
--cc=ziy@nvidia.com \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox