* + mm-mempolicy-use-vm_normal_folio_pmd-in-queue_folios_pmd.patch added to mm-new branch
@ 2026-09-12 8:21 Andrew Morton
0 siblings, 0 replies; only message in thread
From: Andrew Morton @ 2026-09-12 8:21 UTC (permalink / raw)
To: mm-commits, ziy, ying.huang, vbabka, stable, sashiko-bot,
rakie.kim, peterx, matthew.brost, ljs, liam, joshua.hahnjy, jgg,
jannh, david, byungchul, apopple, gourry, akpm
The patch titled
Subject: mm/mempolicy: use vm_normal_folio_pmd() in queue_folios_pmd()
has been added to the -mm mm-new branch. Its filename is
mm-mempolicy-use-vm_normal_folio_pmd-in-queue_folios_pmd.patch
This patch will shortly appear at
https://git.kernel.org/pub/scm/linux/kernel/git/akpm/25-new.git/tree/patches/mm-mempolicy-use-vm_normal_folio_pmd-in-queue_folios_pmd.patch
This patch will later appear in the mm-new branch at
git://git.kernel.org/pub/scm/linux/kernel/git/akpm/mm
Note, mm-new is a provisional staging ground for work-in-progress
patches, and acceptance into mm-new is a notification for others take
notice and to finish up reviews. Please do not hesitate to respond to
review feedback and post updated versions to replace or incrementally
fixup patches in mm-new.
The mm-new branch of mm.git is not included in linux-next
If a few days of testing in mm-new is successful, the patch will me moved
into mm.git's mm-unstable branch, which is included in linux-next
Before you just go and hit "reply", please:
a) Consider who else should be cc'ed
b) Prefer to cc a suitable mailing list as well
c) Ideally: find the original patch on the mailing list and do a
reply-to-all to that, adding suitable additional cc's
*** Remember to use Documentation/process/submit-checklist.rst when testing your code ***
The -mm tree is included into linux-next via various
branches at git://git.kernel.org/pub/scm/linux/kernel/git/akpm/mm
and is updated there most days
------------------------------------------------------
From: Gregory Price <gourry@gourry.net>
Subject: mm/mempolicy: use vm_normal_folio_pmd() in queue_folios_pmd()
Date: Fri, 11 Sep 2026 23:48:32 -0400
Patch series "mm: stop calling pmd_folio() on special PMDs", v2.
Two page table walkers resolve the folio behind a PMD with pmd_folio(),
which is only valid for a PMD mapping a refcounted struct page:
madvise_cold_or_pageout_pte_range() mm/madvise.c
queue_folios_pmd() mm/mempolicy.c
vmf_insert_pfn_pmd() installs special PMDs holding a raw pfn that need not
have a memmap entry at all. Both walkers can reach one and fault on the
first folio field read. The PTE halves of both already use
vm_normal_folio(); these two patches make the PMD halves match.
The four callers of vmf_insert_pfn_pmd(), and which walker each reaches:
drivers/vfio/pci/vfio_pci_core.c VM_PFNMAP mempolicy
drivers/gpu/drm/drm_gem_shmem_helper.c VM_PFNMAP mempolicy
drivers/gpu/drm/panthor/panthor_gem.c VM_PFNMAP mempolicy
drivers/hv/mshv_vtl_main.c VM_MIXEDMAP both
can_madv_lru_vma() rejects VM_PFNMAP, so only mshv_vtl_low reaches the
madvise walker, and that needs CAP_SYS_ADMIN. queue_pages_walk_ops
supplies its own ->test_walk, so walk_page_test()'s generic VM_PFNMAP skip
never runs and vfio-pci is reachable by any process holding the device fd.
Hence the different stable tags.
One behaviour change: mbind(MPOL_MF_STRICT) over a PMD mapped VM_PFNMAP
region now returns 0 rather than -EIO. The PTE loop already returned 0
there. drm_gem_shmem and panthor are where this is observable, since they
PMD map pages that do have a memmap entry and so never faulted.
Reproducer
==========
No hardware needed. An out of tree module stands in for the drivers
above: three misc devices, each with a ->huge_fault calling
vmf_insert_pfn_pmd(), plus VM_HUGEPAGE so the fault path takes the PMD
branch.
/dev/pmdspec_mixed VM_MIXEDMAP, pfn at the 1 TiB mark, no memmap
/dev/pmdspec_pfnmap VM_PFNMAP, pfn at the 1 TiB mark, no memmap
/dev/pmdspec_real VM_PFNMAP, real alloc_pages(PMD_ORDER) on node 0
Userspace maps the device into a PMD aligned window, reads one byte to
fault the PMD in, checks a module parameter to confirm it went in, then
issues the operation.
vng --run <bzImage> --user root --memory 4G --verbose \
--append "numa=fake=2" \
--exec "insmod pmdspec.ko && ./pmdspec_test <subtest>"
numa=fake=2 gives a node 1 to bind to; the module allocates its real page
on node 0, which is what makes queue_folio_required() true.
subtest operation parent series
--------------------------------------------------------------------
madv_cold madvise(MADV_COLD) oops ret=0
madv_pageout madvise(MADV_PAGEOUT) oops ret=0
mbind_mixed mbind(MPOL_BIND, n1, MPOL_MF_MOVE) oops ret=0
mbind_pfnmap mbind(MPOL_BIND, n1, MPOL_MF_STRICT) oops ret=0
mbind_real mbind(MPOL_BIND, n1, MPOL_MF_STRICT) -EIO ret=0
Two things the table shows that are easy to miss in the code:
- mbind_mixed passes only MPOL_MF_MOVE. MPOL_MF_STRICT is not needed for
a VM_MIXEDMAP vma: walk_page_test() only skips VM_PFNMAP, and
vma_migratable() is true for VM_MIXEDMAP.
- mbind_real demonstrates the user visible change (-EIO -> 0)
This patch (of 2):
mmap a VM_PFNMAP region whose ->huge_fault installs a PMD through
vmf_insert_pfn_pmd() - a vfio-pci MMIO BAR does this - then
mbind(p, len, MPOL_BIND, &mask, maxnode, MPOL_MF_STRICT);
With a stand-in module for the driver:
BUG: unable to handle page fault for address: fffff96dc0000008
RIP: 0010:queue_folios_pte_range+0xaf/0x440
walk_pgd_range+0x52b/0xaf0
__walk_page_range+0x6a/0x1d0
walk_page_range_mm_unsafe+0x193/0x230
queue_pages_range+0x64/0xa0
do_mbind+0x25e/0x640
queue_folios_pmd(), inlined above, calls pmd_folio() on that PMD. The pfn
is raw MMIO with no memmap entry, so the folio lands in unpopulated
vmemmap. Neither guard stops the walk:
walk_page_test() skips VM_PFNMAP, but queue_pages_walk_ops
supplies ->test_walk, so it never runs
queue_pages_test_walk() honours vma_migratable(), but only while
MPOL_MF_STRICT is clear
A VM_MIXEDMAP vma needs neither flag, being vma_migratable(), so plain
mbind(MPOL_MF_MOVE) reaches this too - and there the bad folio carries on
into migrate_folio_add() and folio_isolate_lru(). mshv_vtl_low is such a
mapping.
Use vm_normal_folio_pmd() and skip on NULL, as the PTE loop in
queue_folios_pte_range() already does with vm_normal_folio(). On the NULL
path, retain ACTION_CONTINUE handling for the huge zero PMD.
mbind(MPOL_MF_STRICT) over a PMD mapped VM_PFNMAP region now returns 0
rather than -EIO. The PTE loop already returned 0 there.
Link: https://lore.kernel.org/20260912034833.2952750-1-gourry@gourry.net
Link: https://lore.kernel.org/20260912034833.2952750-2-gourry@gourry.net
Signed-off-by: Gregory Price (Meta) <gourry@gourry.net>
Signed-off-by: Andrew Morton <akpm@linux-foundation.org>
Fixes: 3c8e44c9b369 ("mm: mark special bits for huge pfn mappings when inject")
Reported-by: sashiko-bot <sashiko-bot@kernel.org>
Closes: https://sashiko.dev/#/patchset/20260817220810.1175596-1-gourry%40gourry.net
Assisted-by: LLM
Acked-by: David Hildenbrand (Arm) <david@kernel.org>
Cc: Alistair Popple <apopple@nvidia.com>
Cc: Byungchul Park <byungchul@sk.com>
Cc: "Huang, Ying" <ying.huang@linux.alibaba.com>
Cc: Jann Horn <jannh@google.com>
Cc: Jason Gunthorpe <jgg@ziepe.ca>
Cc: Joshua Hahn <joshua.hahnjy@gmail.com>
Cc: Liam R. Howlett <liam@infradead.org>
Cc: Lorenzo Stoakes <ljs@kernel.org>
Cc: Matthew Brost <matthew.brost@intel.com>
Cc: Peter Xu <peterx@redhat.com>
Cc: Rakie Kim <rakie.kim@sk.com>
Cc: Vlastimil Babka <vbabka@kernel.org>
Cc: Zi Yan <ziy@nvidia.com>
Cc: <stable@vger.kernel.org>
---
mm/mempolicy.c | 16 +++++++++-------
1 file changed, 9 insertions(+), 7 deletions(-)
--- a/mm/mempolicy.c~mm-mempolicy-use-vm_normal_folio_pmd-in-queue_folios_pmd
+++ a/mm/mempolicy.c
@@ -667,7 +667,8 @@ static inline bool queue_folio_required(
return node_isset(nid, *qp->nmask) == !(flags & MPOL_MF_INVERT);
}
-static void queue_folios_pmd(pmd_t *pmd, struct mm_walk *walk)
+static void queue_folios_pmd(pmd_t *pmd, unsigned long addr,
+ struct mm_walk *walk)
{
struct folio *folio;
struct queue_pages *qp = walk->private;
@@ -678,13 +679,14 @@ static void queue_folios_pmd(pmd_t *pmd,
qp->nr_failed++;
return;
}
- folio = pmd_folio(pmdval);
- if (folio_is_zone_device(folio))
- return;
- if (is_huge_zero_folio(folio)) {
- walk->action = ACTION_CONTINUE;
+ folio = vm_normal_folio_pmd(walk->vma, addr, pmdval);
+ if (!folio) {
+ if (is_huge_zero_pmd(pmdval))
+ walk->action = ACTION_CONTINUE;
return;
}
+ if (folio_is_zone_device(folio))
+ return;
if (!queue_folio_required(folio, qp))
return;
if (!(qp->flags & (MPOL_MF_MOVE | MPOL_MF_MOVE_ALL)) ||
@@ -717,7 +719,7 @@ static int queue_folios_pte_range(pmd_t
ptl = pmd_trans_huge_lock(pmd, vma);
if (ptl) {
- queue_folios_pmd(pmd, walk);
+ queue_folios_pmd(pmd, addr, walk);
spin_unlock(ptl);
goto out;
}
_
Patches currently in -mm which might be from gourry@gourry.net are
mm-mempolicy-take-a-cpuset-cookie-for-the-interleave-node-count.patch
mm-mempolicy-use-srcu-for-the-weighted-interleave-state.patch
mm-mempolicy-stop-copying-the-nodemask-in-the-interleave-paths.patch
mm-huge_memory-skip-zone-device-folios-in-madvise_free_huge_pmd.patch
mm-madvise-skip-zone-device-folios-in-cold-pageout-pmd-range.patch
mm-mempolicy-skip-zone-device-folios-when-queueing-folios.patch
mm-memory_hotplug-factor-out-node_is_memoryless.patch
mm-mempolicy-use-vm_normal_folio_pmd-in-queue_folios_pmd.patch
mm-madvise-use-vm_normal_folio_pmd-in-cold-pageout-pmd-range.patch
mm-refactor-find_next_best_node-to-find_next_best_node_in.patch
mm-page_alloc-refactor-build_node_zonelist-out-of-build_zonelists.patch
^ permalink raw reply [flat|nested] only message in thread
only message in thread, other threads:[~2026-09-12 8:21 UTC | newest]
Thread overview: (only message) (download: mbox.gz follow: Atom feed
-- links below jump to the message on this page --
2026-09-12 8:21 + mm-mempolicy-use-vm_normal_folio_pmd-in-queue_folios_pmd.patch added to mm-new branch Andrew Morton
This is an external index of several public inboxes,
see mirroring instructions on how to clone and mirror
all data and code used by this external index.