Linux-mm Archive on lore.kernel.org
 help / color / mirror / Atom feed
* [RFC PATCH 0/9] mm: Memory providers for guest_memfd and iommufd
       [not found] <20260720111259.122911-1-dwmw2@infradead.org>
@ 2026-10-06 18:32 ` Fred Griffoul
  2026-10-06 18:32   ` [PATCH 1/9] KVM: guest_memfd: Add a writable result to get_pfn() Fred Griffoul
                     ` (8 more replies)
  0 siblings, 9 replies; 10+ messages in thread
From: Fred Griffoul @ 2026-10-06 18:32 UTC (permalink / raw)
  To: Paolo Bonzini, Sean Christopherson, Marc Zyngier, Oliver Upton,
	Andrew Morton, David Hildenbrand, Alexander Viro,
	Christian Brauner, Jan Kara, Jason Gunthorpe, Kevin Tian,
	Joerg Roedel, Will Deacon, Robin Murphy, Thomas Gleixner,
	Ingo Molnar, Borislav Petkov, Dave Hansen, x86, H . Peter Anvin,
	Jonathan Corbet, Shuah Khan
  Cc: David Woodhouse, Ackerley Tng, Lorenzo Stoakes, Liam R . Howlett,
	Vlastimil Babka, Mike Rapoport, Suren Baghdasaryan, Michal Hocko,
	Joey Gouly, Suzuki K Poulose, Zenghui Yu, Steffen Eiden,
	linux-kernel, kvm, kvmarm, iommu, linux-fsdevel, linux-mm,
	linux-kselftest

From: Fred Griffoul <fgriffo@amazon.co.uk>

Some drivers manage RAM outside the page allocator: a carve-out, device
memory, or memory that a host component moves between VMs while they
run. This memory often has no struct page. The driver wants to lend
pages to a VM and to its devices, and to take any of them back later.

David Woodhouse's series "KVM: Allow alternative providers of
guest_memfd backed by PFNMAP memory" lets such a driver back a
guest_memfd through struct kvm_gmem_ops:

  https://lore.kernel.org/kvm/20260720111259.122911-1-dwmw2@infradead.org/

That covers the guest, but not the VM's devices, which his cover
letter left as an open question. This series adds a small interface in
mm: the driver that owns the memory becomes a provider, and the code
that maps it attaches as a consumer. There are two consumers, a
guest_memfd backend and IOMMU_IOAS_MAP_FILE in iommufd.

The series is on top of David's v2 (base 0e35b9b6ec0f). Patch 1 was
part of the earlier dma-buf RFC. The provider backend depends on it, so
this series carries it.

Why not dma-buf
===============

The earlier RFC shared the memory as a dma-buf. Christian König
rejected that: an importer must not build its own page tables from a
dma-buf, and the use case should stay out of drivers/dma-buf. I
acknowledged that and dropped it; this series does not touch dma-buf.

Why a new interface
===================

kvm_gmem_ops gives KVM a way in. For devices, the only option today is
a dma-buf plus a private call that iommufd looks up with symbol_get(),
which is how vfio-pci works. Each new owner would need another such
lookup, and iommufd would end up knowing each owner by name. With a
common interface, a provider module only calls into mm, and nothing
needs symbol_get().

It also means the provider revokes a range once. The VMM hands the
same file to KVM and to iommufd; on a revoke the core forwards it to
every consumer, so KVM clears the range from stage-2, iommufd from the
IOMMU page tables, and guest_memfd from the VMM's own mapping, all
before the revoke returns.

The contract
============

For a page of a provider file, get_page() returns the frame, its type
(RAM or MMIO, read-only or not), and the largest aligned block of frames
of the same type, so that consumers can use large mappings. A page
without a frame is a hole.

Consumers take no references. A frame stays valid until the revoke that
removes it returns. To guarantee that, a consumer either holds a lock
across get_page() and the mapping that its revoke callback also takes,
or detects a racing revoke and retries.

The interface does not deal with folios. The native guest_memfd backend
and memfd pinning in iommufd handle pages from the page allocator, and
pages that can move or swap would need references. A backend that needs
full control over the file can still write its own kvm_gmem_ops.

Backends
========

The motivating user is a host-side memory manager. It reserves a region
of host RAM, lends parts of it to VMs and their devices, moves pages
between VMs, and has to survive a live update of the host kernel. The
sample module in this series models it.

Device-DAX could also be a provider: its range never moves, get_page()
is short, and it only revokes on unbind.  iommufd can map MMIO, but KVM
takes only RAM, because it picks the guest memory type itself.

Neither consumer keeps state about the provider's frames, which helps
with live update.  guest_memfd has nothing to save across kexec, since
KVM refills stage-2 on faults. The provider saves its frames and who
owns them. Devices cannot fault, so their mappings stay part of the
existing iommufd live update work, and the provider only has to give
back the same frames.

Confidential VMs
================

This series does not support them yet. For now only x86 VMs of type
KVM_X86_DEFAULT_VM can use a provider.

I don't think a provider should back private pages. Revoking a private
page loses its contents, and only guest_memfd knows which pages are
private. The interface could still be useful there later, with
guest_memfd acting as a provider for its own files so that iommufd only
maps shared pages. I left that for a later series.

Patches
=======

  Patch 1 lets a kvm_gmem_ops backend map a page read-only for the
  guest.  get_pfn() gets a writable output, and a guest write to such a
  page exits with KVM_EXIT_MEMORY_FAULT. KVM_MEM_READONLY can't be
  used for this, because guest_memfd slots don't allow it.

  Patch 2 adds the interface and the core, and the FOP_MEM_PROVIDER
  flag in struct file_operations.

  Patch 3 adds the provider backend to guest_memfd, with the
  GUEST_MEMFD_FLAG_USE_PROVIDER flag and a provider_fd field.
  guest_memfd also handles the userspace mapping of the file, if the
  provider allows it, so providers don't each have to get the fault and
  revoke handling right.

  Patch 4 prepares iommufd for tracking pages it does not pin. No
  functional change.

  Patch 5 lets IOMMU_IOAS_MAP_FILE take a provider file. iommufd maps
  one PAGE_SIZE entry per page and pins nothing. It leaves holes
  unmapped and honours read-only and MMIO pages. On a revoke, it
  unmaps the range and maps whatever the provider backs now.

  The PAGE_SIZE entries are deliberate for now: a partial revoke then
  never has to split a large IOMMU page. Using the block size the
  provider reports would be better, and needs the revoke to unmap and
  remap whole blocks. A device that accesses the range between the
  unmap and the map faults; replacing the entries in place would avoid
  that, but needs new support in the IOMMU drivers. Both are left for
  later.

  Patches 6 and 7 add iommufd selftests: two mock-domain queries and a
  mock provider.

  Patch 8 adds a sample provider. It gives each VM a child file, can
  move, donate and reclaim pages, and supports read-only pages. A child
  file is read-only to its holder; the owner changes it through the
  control device. It only uses the mm interface.

  Patch 9 adds a KVM selftest with one provider and two VMMs. After
  each change it checks what the guest, the device and the VMM's mapping
  see.

Testing
=======

x86_64, in QEMU with KASAN, PROVE_LOCKING and DEBUG_ATOMIC_SLEEP, with
two ranges hidden from the host with memmap=, one for each sample. No
KASAN report, lockdep splat or warning in any run.

  - mem_provider_test, the selftest in patch 9: pass.
  - guest_memfd_test, including the USE_PROVIDER flag check: pass.
  - iommufd_selftest, iommufd_ioas: the 12 provider tests pass in all
    four variants. The 4 failures are access_domain_destory, which the
    series does not touch: it needs hugetlb pages, which this VM has
    none of.
  - David's gmem_provider tests, which this series does not change:
    hugepage, revoke, iommufd and readonly pass; the SNP and vfio-pci
    tests skip for lack of hardware.

Not tested: arm64, where guest_memfd refuses providers; a real IOMMU,
since only the mock domain was used.

I wrote most of this series with AI help, and I am posting it as an RFC
to get feedback on the interface.

Fred Griffoul (9):
  KVM: guest_memfd: Add a writable result to get_pfn()
  mm: Add memory providers
  KVM: guest_memfd: Add a memory provider backing
  iommufd: Track the domains of pages that are not pinned
  iommufd: Map memory provider files
  iommufd/selftest: Add mock-domain IOVA queries
  iommufd/selftest: Add a mock memory provider
  samples/kvm: Add a memory provider sample
  KVM: selftests: Test a memory provider shared by KVM and iommufd

 Documentation/virt/kvm/api.rst                |  45 +-
 MAINTAINERS                                   |   8 +
 arch/arm64/kvm/mmu.c                          |  13 +-
 arch/arm64/kvm/nested.c                       |  16 +-
 arch/x86/kvm/mmu/mmu.c                        |  18 +-
 arch/x86/kvm/svm/sev.c                        |  13 +-
 arch/x86/kvm/x86.c                            |  10 +
 drivers/iommu/iommufd/Kconfig                 |   1 +
 drivers/iommu/iommufd/io_pagetable.c          |  37 +-
 drivers/iommu/iommufd/io_pagetable.h          |  54 +-
 drivers/iommu/iommufd/iommufd_test.h          |  36 +
 drivers/iommu/iommufd/pages.c                 | 340 ++++++-
 drivers/iommu/iommufd/selftest.c              | 286 ++++++
 include/linux/fs.h                            |   2 +
 include/linux/kvm_host.h                      |  20 +-
 include/linux/mem_provider.h                  | 211 +++++
 include/uapi/linux/iommufd.h                  |  11 +-
 include/uapi/linux/kvm.h                      |  11 +-
 mm/Kconfig                                    |   7 +
 mm/Makefile                                   |   1 +
 mm/mem_provider.c                             | 177 ++++
 samples/Kconfig                               |  17 +
 samples/Makefile                              |   1 +
 samples/kvm/Makefile                          |   1 +
 samples/kvm/gmem_provider.c                   |  67 +-
 samples/kvm/gmem_provider.h                   |  17 +
 samples/kvm/mem_provider_sample.c             | 855 ++++++++++++++++++
 samples/kvm/mem_provider_sample.h             | 135 +++
 tools/include/uapi/linux/kvm.h                |  11 +-
 tools/testing/selftests/iommu/iommufd.c       | 119 +++
 tools/testing/selftests/iommu/iommufd_utils.h |  64 ++
 tools/testing/selftests/kvm/Makefile.kvm      |   2 +
 .../testing/selftests/kvm/guest_memfd_test.c  |   6 +
 .../kvm/x86/gmem_provider_readonly_test.c     | 150 +++
 .../selftests/kvm/x86/mem_provider_test.c     | 582 ++++++++++++
 virt/kvm/Kconfig                              |   1 +
 virt/kvm/guest_memfd.c                        | 286 +++++-
 37 files changed, 3531 insertions(+), 100 deletions(-)
 create mode 100644 include/linux/mem_provider.h
 create mode 100644 mm/mem_provider.c
 create mode 100644 samples/kvm/mem_provider_sample.c
 create mode 100644 samples/kvm/mem_provider_sample.h
 create mode 100644 tools/testing/selftests/kvm/x86/gmem_provider_readonly_test.c
 create mode 100644 tools/testing/selftests/kvm/x86/mem_provider_test.c


base-commit: 0e35b9b6ec0ffcc5e23cbdec09f5c622ad532b53
prerequisite-patch-id: 2a0016e90f0690baef841a8c34c07b96e3f1b941
prerequisite-patch-id: c497c1ff1e9de8e9c470c58b163ccbfce4827ae7
prerequisite-patch-id: a16b61afd172662bde725e5c12805758f76b89fa
prerequisite-patch-id: 0ddee6c1b48fa853e6a94f2746340e525477d287
prerequisite-patch-id: 4dc9a395a2eb73283b20b75e2ba763ec72b2e22d
prerequisite-patch-id: a9758d7f8f6959dae6ad154902fa1269234dbc3c
prerequisite-patch-id: 0537bcc3ca5e3bb9b86fe7c99a245a9ff9d67831
prerequisite-patch-id: d5feb18b6630c99973259804418243d856c37ce7
prerequisite-patch-id: 2a08a105b6b5d46240a69f5bbf646211d44f8412
prerequisite-patch-id: 760b34834d5d8dcdf1ae1a09cc6aae267dd7760c
prerequisite-patch-id: dd94ecc7ae9472c2bf005c9cb1c483994eacd163
-- 
2.47.3



^ permalink raw reply	[flat|nested] 10+ messages in thread

* [PATCH 1/9] KVM: guest_memfd: Add a writable result to get_pfn()
  2026-10-06 18:32 ` [RFC PATCH 0/9] mm: Memory providers for guest_memfd and iommufd Fred Griffoul
@ 2026-10-06 18:32   ` Fred Griffoul
  2026-10-06 18:32   ` [PATCH 2/9] mm: Add memory providers Fred Griffoul
                     ` (7 subsequent siblings)
  8 siblings, 0 replies; 10+ messages in thread
From: Fred Griffoul @ 2026-10-06 18:32 UTC (permalink / raw)
  To: Paolo Bonzini, Sean Christopherson, Marc Zyngier, Oliver Upton,
	Andrew Morton, David Hildenbrand, Alexander Viro,
	Christian Brauner, Jan Kara, Jason Gunthorpe, Kevin Tian,
	Joerg Roedel, Will Deacon, Robin Murphy, Thomas Gleixner,
	Ingo Molnar, Borislav Petkov, Dave Hansen, x86, H . Peter Anvin,
	Jonathan Corbet, Shuah Khan
  Cc: David Woodhouse, Ackerley Tng, Lorenzo Stoakes, Liam R . Howlett,
	Vlastimil Babka, Mike Rapoport, Suren Baghdasaryan, Michal Hocko,
	Joey Gouly, Suzuki K Poulose, Zenghui Yu, Steffen Eiden,
	linux-kernel, kvm, kvmarm, iommu, linux-fsdevel, linux-mm,
	linux-kselftest

From: Fred Griffoul <fgriffo@amazon.co.uk>

A guest_memfd backing cannot map a page read-only for the guest:
get_pfn() returns only a frame, and KVM_MEM_READONLY is refused on
guest_memfd slots. A hypervisor therefore cannot give a guest a page it
may read but not write.

Let get_pfn() clear a writable result. The x86 and arm64 fault paths
then map the page without write permission, and a guest write exits
with KVM_EXIT_MEMORY_FAULT. The result covers the whole block that
max_order allows.

Callers that hand the page to something that writes it must refuse a
read-only page: the arm64 VNCR page and the SEV-SNP VMSA. populate() has
no writable result, so a backing cannot report read-only state through
it. The result does not cover KVM's own writes through the slot's host
address.

David's gmem_provider sample gains a read-only ioctl, and a selftest
checks that a guest write to such a page exits and lands once the page
is writable again.

Signed-off-by: Fred Griffoul <fgriffo@amazon.co.uk>
---
 arch/arm64/kvm/mmu.c                          |  13 +-
 arch/arm64/kvm/nested.c                       |  16 +-
 arch/x86/kvm/mmu/mmu.c                        |  18 ++-
 arch/x86/kvm/svm/sev.c                        |  13 +-
 include/linux/kvm_host.h                      |  16 +-
 samples/kvm/gmem_provider.c                   |  67 +++++++-
 samples/kvm/gmem_provider.h                   |  17 ++
 tools/testing/selftests/kvm/Makefile.kvm      |   1 +
 .../kvm/x86/gmem_provider_readonly_test.c     | 150 ++++++++++++++++++
 virt/kvm/guest_memfd.c                        |  21 ++-
 10 files changed, 309 insertions(+), 23 deletions(-)
 create mode 100644 tools/testing/selftests/kvm/x86/gmem_provider_readonly_test.c

diff --git a/arch/arm64/kvm/mmu.c b/arch/arm64/kvm/mmu.c
index 6c941aaa10c6..32e591edc69d 100644
--- a/arch/arm64/kvm/mmu.c
+++ b/arch/arm64/kvm/mmu.c
@@ -1607,7 +1607,7 @@ struct kvm_s2_fault_desc {
 
 static int gmem_abort(const struct kvm_s2_fault_desc *s2fd)
 {
-	bool write_fault, exec_fault;
+	bool write_fault, exec_fault, writable;
 	bool perm_fault = kvm_vcpu_trap_is_permission_fault(s2fd->vcpu);
 	enum kvm_pgtable_walk_flags flags = KVM_PGTABLE_WALK_SHARED;
 	enum kvm_pgtable_prot prot = KVM_PGTABLE_PROT_R;
@@ -1641,14 +1641,17 @@ static int gmem_abort(const struct kvm_s2_fault_desc *s2fd)
 	/* Pairs with the smp_wmb() in kvm_mmu_invalidate_end(). */
 	smp_rmb();
 
-	ret = kvm_gmem_get_pfn(kvm, s2fd->memslot, gfn, &pfn, &page, NULL);
-	if (ret) {
+	ret = kvm_gmem_get_pfn(kvm, s2fd->memslot, gfn, &pfn, &page, NULL,
+			       &writable);
+	if (ret || (write_fault && !writable)) {
 		kvm_prepare_memory_fault_exit(s2fd->vcpu, s2fd->fault_ipa, PAGE_SIZE,
 					      write_fault, exec_fault, false);
-		return ret;
+		if (!ret)
+			kvm_release_faultin_page(kvm, page, true, false);
+		return ret ?: -EFAULT;
 	}
 
-	if (!(s2fd->memslot->flags & KVM_MEM_READONLY))
+	if (!(s2fd->memslot->flags & KVM_MEM_READONLY) && writable)
 		prot |= KVM_PGTABLE_PROT_W;
 
 	if (s2fd->nested)
diff --git a/arch/arm64/kvm/nested.c b/arch/arm64/kvm/nested.c
index fb54f6dad995..fa86705df949 100644
--- a/arch/arm64/kvm/nested.c
+++ b/arch/arm64/kvm/nested.c
@@ -1411,11 +1411,19 @@ static int kvm_translate_vncr(struct kvm_vcpu *vcpu, bool *is_gmem)
 		if (is_error_noslot_pfn(pfn) || (write_fault && !writable))
 			return -EFAULT;
 	} else {
-		ret = kvm_gmem_get_pfn(vcpu->kvm, memslot, gfn, &pfn, &page, NULL);
-		if (ret) {
+		ret = kvm_gmem_get_pfn(vcpu->kvm, memslot, gfn, &pfn, &page, NULL,
+				       &writable);
+		/*
+		 * The VNCR page is always written, so a read-only page cannot
+		 * back it.  Exit on the first access, even a read, rather than
+		 * map it read-only and fault on the next write.
+		 */
+		if (ret || !writable) {
 			kvm_prepare_memory_fault_exit(vcpu, vt->wr.pa, PAGE_SIZE,
-					      write_fault, false, false);
-			return ret;
+					      true, false, false);
+			if (!ret)
+				kvm_release_faultin_page(vcpu->kvm, page, true, false);
+			return ret ?: -EFAULT;
 		}
 	}
 
diff --git a/arch/x86/kvm/mmu/mmu.c b/arch/x86/kvm/mmu/mmu.c
index 234d0a95abf5..f689ef5c2b46 100644
--- a/arch/x86/kvm/mmu/mmu.c
+++ b/arch/x86/kvm/mmu/mmu.c
@@ -4612,6 +4612,7 @@ static void kvm_mmu_finish_page_fault(struct kvm_vcpu *vcpu,
 static int kvm_mmu_faultin_pfn_gmem(struct kvm_vcpu *vcpu,
 				    struct kvm_page_fault *fault)
 {
+	bool writable;
 	int max_order, r;
 
 	if (!kvm_slot_has_gmem(fault->slot)) {
@@ -4620,13 +4621,26 @@ static int kvm_mmu_faultin_pfn_gmem(struct kvm_vcpu *vcpu,
 	}
 
 	r = kvm_gmem_get_pfn(vcpu->kvm, fault->slot, fault->gfn, &fault->pfn,
-			     &fault->refcounted_page, &max_order);
+			     &fault->refcounted_page, &max_order, &writable);
 	if (r) {
 		kvm_mmu_prepare_memory_fault_exit(vcpu, fault);
 		return r;
 	}
 
-	fault->map_writable = !(fault->slot->flags & KVM_MEM_READONLY);
+	/*
+	 * get_pfn() may clear writable, on top of the memslot's read-only flag:
+	 * a read-only page is mapped read-only, and a guest write to it exits to
+	 * userspace rather than being installed.
+	 */
+	fault->map_writable = !(fault->slot->flags & KVM_MEM_READONLY) &&
+			      writable;
+	if (fault->write && !fault->map_writable) {
+		kvm_mmu_prepare_memory_fault_exit(vcpu, fault);
+		kvm_release_faultin_page(vcpu->kvm, fault->refcounted_page,
+				 true, false);
+		fault->refcounted_page = NULL;
+		return -EFAULT;
+	}
 	fault->max_level = kvm_max_level_for_order(max_order);
 
 	return RET_PF_CONTINUE;
diff --git a/arch/x86/kvm/svm/sev.c b/arch/x86/kvm/svm/sev.c
index 125779c82bc4..f4f944c7da10 100644
--- a/arch/x86/kvm/svm/sev.c
+++ b/arch/x86/kvm/svm/sev.c
@@ -4026,6 +4026,7 @@ static void sev_snp_init_protected_guest_state(struct kvm_vcpu *vcpu)
 	struct vcpu_svm *svm = to_svm(vcpu);
 	struct kvm_memory_slot *slot;
 	struct page *page;
+	bool writable;
 	kvm_pfn_t pfn;
 	gfn_t gfn;
 
@@ -4063,9 +4064,17 @@ static void sev_snp_init_protected_guest_state(struct kvm_vcpu *vcpu)
 	 * The new VMSA will be private memory guest memory, so retrieve the
 	 * PFN from the gmem backend.
 	 */
-	if (kvm_gmem_get_pfn(vcpu->kvm, slot, gfn, &pfn, &page, NULL))
+	if (kvm_gmem_get_pfn(vcpu->kvm, slot, gfn, &pfn, &page, NULL,
+			     &writable))
 		return;
 
+	/* The CPU writes the VMSA on every VMRUN: a read-only page cannot be one. */
+	if (!writable) {
+		if (page)
+			kvm_release_page_clean(page);
+		return;
+	}
+
 	/*
 	 * From this point forward, the VMSA will always be a guest-mapped page
 	 * rather than the initial one allocated by KVM in svm->sev_es.vmsa. In
@@ -4996,7 +5005,7 @@ void sev_handle_rmp_fault(struct kvm_vcpu *vcpu, gpa_t gpa, u64 error_code)
 		return;
 	}
 
-	ret = kvm_gmem_get_pfn(kvm, slot, gfn, &pfn, &page, &order);
+	ret = kvm_gmem_get_pfn(kvm, slot, gfn, &pfn, &page, &order, NULL);
 	if (ret) {
 		pr_warn_ratelimited("SEV: Unexpected RMP fault, no backing page for private GPA 0x%llx\n",
 				    gpa);
diff --git a/include/linux/kvm_host.h b/include/linux/kvm_host.h
index 04fa0cb126f6..7281d0e94121 100644
--- a/include/linux/kvm_host.h
+++ b/include/linux/kvm_host.h
@@ -649,6 +649,15 @@ static inline bool kvm_slot_has_gmem(const struct kvm_memory_slot *slot)
  *    is responsible for put_page() after use.  If *page is left NULL the PFN
  *    is treated as non-refcounted, and its lifetime is owned by the
  *    implementation across bind()/unbind().
+ *  - *writable, when the caller passes it, is true on entry.  Clear it to
+ *    have KVM map the page read-only; a guest write then exits as a memory
+ *    fault.  It covers the whole block that *max_order allows, so that block
+ *    must be all writable or all read-only.  It applies to stage-2 mappings
+ *    only, not to KVM's own writes through the memslot's host address.  A
+ *    caller that hands the page to hardware that writes it passes @writable
+ *    and refuses a read-only page.  Only a caller that never writes the page
+ *    itself passes NULL.  populate() has no writable result, so KVM cannot
+ *    learn read-only state through it; it is used only to fill a page.
  *
  * Memory intended to back guest RAM MUST be reported as E820_TYPE_RAM by the
  * host so KVM maps it write-back (and applies the memory-encryption bit on
@@ -662,7 +671,8 @@ struct kvm_gmem_ops {
 		       struct kvm_memory_slot *slot);
 	int (*get_pfn)(struct file *file, struct kvm *kvm,
 		       struct kvm_memory_slot *slot, gfn_t gfn,
-		       kvm_pfn_t *pfn, struct page **page, int *max_order);
+		       kvm_pfn_t *pfn, struct page **page, int *max_order,
+		       bool *writable);
 	int (*populate)(struct file *file, struct kvm *kvm,
 			struct kvm_memory_slot *slot, gfn_t gfn,
 			kvm_pfn_t *pfn, struct page *src_page, int order);
@@ -2651,12 +2661,12 @@ static inline bool kvm_mem_is_private(struct kvm *kvm, gfn_t gfn)
 #ifdef CONFIG_KVM_GUEST_MEMFD
 int kvm_gmem_get_pfn(struct kvm *kvm, struct kvm_memory_slot *slot,
 		     gfn_t gfn, kvm_pfn_t *pfn, struct page **page,
-		     int *max_order);
+		     int *max_order, bool *writable);
 #else
 static inline int kvm_gmem_get_pfn(struct kvm *kvm,
 				   struct kvm_memory_slot *slot, gfn_t gfn,
 				   kvm_pfn_t *pfn, struct page **page,
-				   int *max_order)
+				   int *max_order, bool *writable)
 {
 	KVM_BUG_ON(1, kvm);
 	return -EIO;
diff --git a/samples/kvm/gmem_provider.c b/samples/kvm/gmem_provider.c
index 9728f5a8029b..99b288302d98 100644
--- a/samples/kvm/gmem_provider.c
+++ b/samples/kvm/gmem_provider.c
@@ -95,6 +95,7 @@ struct gmem_info {
 	gfn_t base_gfn;				/* recorded at bind, for revoke */
 	pgoff_t pgoff;				/* provider offset (pages) of the slot */
 	unsigned long *absent;			/* bitmap of currently-revoked pages */
+	unsigned long *readonly;		/* bitmap of pages the guest may not write */
 	struct list_head dmabufs;		/* struct gmem_dmabuf entries */
 	struct mutex dmabufs_lock;
 };
@@ -142,6 +143,22 @@ static int gmem_max_order(struct gmem_info *info, gfn_t gfn, unsigned long index
 		remaining = min(remaining, absent_next - index);
 	}
 
+	/*
+	 * Likewise a hugepage must be uniformly writable or uniformly
+	 * read-only: clamp at the next page whose read-only bit differs.
+	 */
+	if (info->readonly) {
+		unsigned long next;
+
+		if (test_bit(index, info->readonly))
+			next = find_next_zero_bit(info->readonly, info->npages,
+						  index + 1);
+		else
+			next = find_next_bit(info->readonly, info->npages,
+					     index + 1);
+		remaining = min(remaining, next - index);
+	}
+
 	if (IS_ALIGNED(pfn, 1UL << pud_order) &&
 	    IS_ALIGNED(gfn, 1UL << pud_order) &&
 	    remaining >= (1UL << pud_order))
@@ -157,7 +174,8 @@ static int gmem_max_order(struct gmem_info *info, gfn_t gfn, unsigned long index
 
 static int gmem_get_pfn(struct file *file, struct kvm *kvm,
 			struct kvm_memory_slot *slot, gfn_t gfn,
-			kvm_pfn_t *pfn, struct page **page, int *max_order)
+			kvm_pfn_t *pfn, struct page **page, int *max_order,
+			bool *writable)
 {
 	struct gmem_info *info = to_gmem_info(file);
 	pgoff_t index = gfn - slot->base_gfn + slot->gmem.pgoff;
@@ -172,6 +190,8 @@ static int gmem_get_pfn(struct file *file, struct kvm *kvm,
 	*pfn = info->base_pfn + index;
 	if (max_order)
 		*max_order = gmem_max_order(info, gfn, index);
+	if (writable)
+		*writable = !(info->readonly && test_bit(index, info->readonly));
 	return 0;
 }
 
@@ -371,6 +391,7 @@ static void gmem_release(struct file *file)
 	if (info->cma_pages)
 		free_contig_range(info->base_pfn, info->npages);
 	kvfree(info->absent);
+	kvfree(info->readonly);
 	kfree(info);
 	module_put(THIS_MODULE);
 }
@@ -581,6 +602,46 @@ static long gmem_fd_ioctl(struct file *file, unsigned int cmd, unsigned long arg
 	if (cmd == GMEM_PROVIDER_GET_DMABUF)
 		return gmem_provider_get_dmabuf(file);
 
+	if (cmd == GMEM_PROVIDER_SET_READONLY) {
+		struct gmem_provider_readonly r;
+		unsigned long clamped_start;
+
+		if (copy_from_user(&r, (void __user *)arg, sizeof(r)))
+			return -EFAULT;
+		if (!r.len || !PAGE_ALIGNED(r.offset) || !PAGE_ALIGNED(r.len) ||
+		    r.pad)
+			return -EINVAL;
+		start_index = r.offset >> PAGE_SHIFT;
+		end_index = start_index + (r.len >> PAGE_SHIFT);
+		if (end_index > info->npages || end_index < start_index)
+			return -EINVAL;
+
+		/*
+		 * Flip the bits, then drop the guest's existing mappings of the
+		 * range so the next access re-faults through get_pfn() and
+		 * picks up the new permission.  Making a range read-only must
+		 * tear down writable mappings; making it writable again is
+		 * also invalidated so a stale read-only mapping does not keep
+		 * exiting.
+		 */
+		mutex_lock(&info->lock);
+		if (r.readonly)
+			bitmap_set(info->readonly, start_index,
+				   end_index - start_index);
+		else
+			bitmap_clear(info->readonly, start_index,
+				     end_index - start_index);
+		clamped_start = max_t(unsigned long, start_index, info->pgoff);
+		if (info->kvm && end_index > clamped_start)
+			kvm_gmem_invalidate_range(info->kvm,
+						  info->base_gfn + clamped_start -
+							  info->pgoff,
+						  info->base_gfn + end_index -
+							  info->pgoff);
+		mutex_unlock(&info->lock);
+		return 0;
+	}
+
 	if (cmd != GMEM_PROVIDER_SET_PRESENT)
 		return -ENOTTY;
 	if (copy_from_user(&p, (void __user *)arg, sizeof(p)))
@@ -694,7 +755,8 @@ static long gmem_ctl_ioctl(struct file *file, unsigned int cmd, unsigned long ar
 
 	info->absent = kvzalloc(BITS_TO_LONGS(info->npages) * sizeof(unsigned long),
 				GFP_KERNEL);
-	if (!info->absent) {
+	info->readonly = kvzalloc_objs(unsigned long, BITS_TO_LONGS(info->npages));
+	if (!info->absent || !info->readonly) {
 		ret = -ENOMEM;
 		goto err_free_pages;
 	}
@@ -733,6 +795,7 @@ static long gmem_ctl_ioctl(struct file *file, unsigned int cmd, unsigned long ar
 		free_contig_range(page_to_pfn(pages), npages);
 err_free_info:
 	kvfree(info->absent);
+	kvfree(info->readonly);
 	kfree(info);
 err_put_kvm:
 	kvm_put_kvm(kvm);
diff --git a/samples/kvm/gmem_provider.h b/samples/kvm/gmem_provider.h
index 45f1b8257f60..51b49ef9a3c5 100644
--- a/samples/kvm/gmem_provider.h
+++ b/samples/kvm/gmem_provider.h
@@ -41,6 +41,23 @@ struct gmem_provider_present {
 
 #define GMEM_PROVIDER_SET_PRESENT	_IOW(GMEM_PROVIDER_IOCTL_BASE, 2, struct gmem_provider_present)
 
+/*
+ * ioctl on a provider fd: make a byte range read-only for the guest, or
+ * writable again.  KVM maps a read-only page without write permission and a
+ * guest write to it exits to userspace with KVM_EXIT_MEMORY_FAULT.  Existing
+ * mappings of the range are dropped so the change takes effect on the next
+ * access.
+ */
+struct gmem_provider_readonly {
+	__u64 offset;	/* byte offset into the provider region, page aligned */
+	__u64 len;	/* byte length, page aligned */
+	__u32 readonly;	/* 1 = guest may not write, 0 = guest may write */
+	__u32 pad;
+};
+
+#define GMEM_PROVIDER_SET_READONLY \
+	_IOW(GMEM_PROVIDER_IOCTL_BASE, 4, struct gmem_provider_readonly)
+
 /*
  * ioctl on a provider fd (returned by SETUP): export the backing region as a
  * dynamic dma-buf and return an fd for it, suitable for
diff --git a/tools/testing/selftests/kvm/Makefile.kvm b/tools/testing/selftests/kvm/Makefile.kvm
index 12004a487c32..4d7082448cea 100644
--- a/tools/testing/selftests/kvm/Makefile.kvm
+++ b/tools/testing/selftests/kvm/Makefile.kvm
@@ -81,6 +81,7 @@ TEST_GEN_PROGS_x86 += x86/fix_hypercall_test
 TEST_GEN_PROGS_x86 += x86/gmem_provider_test
 TEST_GEN_PROGS_x86 += x86/gmem_provider_hugepage_test
 TEST_GEN_PROGS_x86 += x86/gmem_provider_revoke_test
+TEST_GEN_PROGS_x86 += x86/gmem_provider_readonly_test
 TEST_GEN_PROGS_x86 += x86/gmem_provider_iommufd_test
 TEST_GEN_PROGS_x86 += x86/gmem_provider_vfio_test
 TEST_GEN_PROGS_x86 += x86/hwcr_msr_test
diff --git a/tools/testing/selftests/kvm/x86/gmem_provider_readonly_test.c b/tools/testing/selftests/kvm/x86/gmem_provider_readonly_test.c
new file mode 100644
index 000000000000..c3eedba7a16d
--- /dev/null
+++ b/tools/testing/selftests/kvm/x86/gmem_provider_readonly_test.c
@@ -0,0 +1,150 @@
+// SPDX-License-Identifier: GPL-2.0
+/*
+ * gmem_provider_readonly_test - exercise per-range read-only from a provider.
+ *
+ * Marks a provider-backed page read-only via an ioctl on the provider fd and
+ * checks that KVM honours the provider's answer: the guest can still read the
+ * page, a guest write exits to userspace with KVM_EXIT_MEMORY_FAULT rather than
+ * landing, and clearing the bit lets the write through. This is the mechanism
+ * a hypervisor uses to protect a page it shares with the guest, such as a
+ * sidecar's info page, without giving up the mapping.
+ *
+ * The test opens the provider with GMEM_PROVIDER_FLAG_MMAP_CAPABLE at SETUP time
+ * (gmem-only). Load the module with a backing region of at least DATA_SIZE.
+ */
+#include <fcntl.h>
+#include <errno.h>
+#include <stdint.h>
+#include <stdio.h>
+#include <string.h>
+#include <unistd.h>
+#include <sys/ioctl.h>
+#include <sys/mman.h>
+
+#include "test_util.h"
+#include "kvm_util.h"
+#include "processor.h"
+
+/* Mirrors samples/kvm/gmem_provider.h */
+struct gmem_provider_setup {
+	__s32 kvm_fd;
+	__u32 flags;
+	__u64 size;
+};
+
+struct gmem_provider_readonly {
+	__u64 offset;
+	__u64 len;
+	__u32 readonly;
+	__u32 pad;
+};
+
+#define GMEM_PROVIDER_SETUP		_IOW('G', 1, struct gmem_provider_setup)
+#define GMEM_PROVIDER_FLAG_MMAP_CAPABLE	(1u << 0)
+#define GMEM_PROVIDER_SET_READONLY	_IOW('G', 4, struct gmem_provider_readonly)
+
+#define DATA_SLOT	10
+#define DATA_GPA	(1ULL << 32)
+#define DATA_SIZE	0x200000ULL		/* 2 MiB region */
+#define MAGIC		0x1234abcdULL
+#define MAGIC2		0xfeedf00dULL
+
+/*
+ * Phase 1: read the page and report it.
+ * Phase 2: write to it.  With the page read-only this never returns to the
+ *          guest until userspace clears the bit; then it completes and the
+ *          guest reports what it wrote.
+ */
+static void guest_code(void)
+{
+	GUEST_SYNC(*(volatile uint64_t *)DATA_GPA);
+	*(volatile uint64_t *)DATA_GPA = MAGIC2;
+	GUEST_SYNC(*(volatile uint64_t *)DATA_GPA);
+	GUEST_DONE();
+}
+
+int main(void)
+{
+	struct vm_shape shape = {
+		.mode = VM_MODE_DEFAULT,
+		.type = KVM_X86_SW_PROTECTED_VM,
+	};
+	struct gmem_provider_setup setup = { .flags = GMEM_PROVIDER_FLAG_MMAP_CAPABLE };
+	struct gmem_provider_readonly req;
+	struct kvm_vcpu *vcpu;
+	struct kvm_vm *vm;
+	struct ucall uc;
+	int gmem_ctl, gmem_fd, r;
+	void *hva;
+
+	TEST_REQUIRE(kvm_check_cap(KVM_CAP_VM_TYPES) & BIT(KVM_X86_SW_PROTECTED_VM));
+
+	gmem_ctl = open("/dev/gmem_provider", O_RDWR);
+	__TEST_REQUIRE(gmem_ctl >= 0,
+		       "gmem_provider module not loaded (/dev/gmem_provider absent)");
+
+	vm = vm_create_shape_with_one_vcpu(shape, &vcpu, guest_code);
+
+	setup.kvm_fd = vm->fd;
+	setup.size = DATA_SIZE;
+	gmem_fd = ioctl(gmem_ctl, GMEM_PROVIDER_SETUP, &setup);
+	TEST_ASSERT(gmem_fd >= 0, "GMEM_PROVIDER_SETUP failed, errno %d", errno);
+
+	hva = mmap(NULL, DATA_SIZE, PROT_READ | PROT_WRITE, MAP_SHARED, gmem_fd, 0);
+	TEST_ASSERT(hva != MAP_FAILED, "mmap(provider) failed, errno %d", errno);
+
+	r = __vm_set_user_memory_region2(vm, DATA_SLOT, KVM_MEM_GUEST_MEMFD,
+					 DATA_GPA, DATA_SIZE, hva, gmem_fd, 0);
+	TEST_ASSERT(!r, "KVM_SET_USER_MEMORY_REGION2 failed: %d errno %d", r, errno);
+	virt_map(vm, DATA_GPA, DATA_GPA, 1);
+
+	/* Seed the page from the host before the guest ever touches it. */
+	*(volatile uint64_t *)hva = MAGIC;
+
+	/* 1) Make the page read-only for the guest. */
+	req = (struct gmem_provider_readonly){ .offset = 0, .len = 4096, .readonly = 1 };
+	r = ioctl(gmem_fd, GMEM_PROVIDER_SET_READONLY, &req);
+	TEST_ASSERT(!r, "set readonly ioctl failed, errno %d", errno);
+
+	/* 2) Guest read must still work and see the host's value. */
+	vcpu_run(vcpu);
+	TEST_ASSERT(get_ucall(vcpu, &uc) == UCALL_SYNC, "expected UCALL_SYNC");
+	TEST_ASSERT(uc.args[1] == MAGIC, "guest read 0x%lx, want MAGIC",
+		    (unsigned long)uc.args[1]);
+	pr_info("read-only: guest read 0x%llx\n", MAGIC);
+
+	/* 3) Guest write must exit to userspace, not land. */
+	r = _vcpu_run(vcpu);
+	TEST_ASSERT(r == -1 && errno == EFAULT &&
+		    vcpu->run->exit_reason == KVM_EXIT_MEMORY_FAULT,
+		    "read-only write: expected KVM_EXIT_MEMORY_FAULT (r=%d errno=%d exit_reason=%u %s)",
+		    r, errno, vcpu->run->exit_reason,
+		    exit_reason_str(vcpu->run->exit_reason));
+	TEST_ASSERT(*(volatile uint64_t *)hva == MAGIC,
+		    "guest write landed on a read-only page: host sees 0x%lx",
+		    (unsigned long)*(volatile uint64_t *)hva);
+	pr_info("read-only: guest write exited with KVM_EXIT_MEMORY_FAULT, page unchanged\n");
+
+	/* 4) Make it writable again; the retried write must complete. */
+	req.readonly = 0;
+	r = ioctl(gmem_fd, GMEM_PROVIDER_SET_READONLY, &req);
+	TEST_ASSERT(!r, "clear readonly ioctl failed, errno %d", errno);
+
+	vcpu_run(vcpu);
+	TEST_ASSERT(get_ucall(vcpu, &uc) == UCALL_SYNC, "expected UCALL_SYNC after clear");
+	TEST_ASSERT(uc.args[1] == MAGIC2, "after clear guest read 0x%lx, want MAGIC2",
+		    (unsigned long)uc.args[1]);
+	TEST_ASSERT(*(volatile uint64_t *)hva == MAGIC2,
+		    "host sees 0x%lx after guest write, want MAGIC2",
+		    (unsigned long)*(volatile uint64_t *)hva);
+	pr_info("writable: guest write 0x%llx landed -- read-only path works\n", MAGIC2);
+
+	vcpu_run(vcpu);
+	TEST_ASSERT(get_ucall(vcpu, &uc) == UCALL_DONE, "expected UCALL_DONE");
+
+	kvm_vm_free(vm);
+	munmap(hva, DATA_SIZE);
+	close(gmem_fd);
+	close(gmem_ctl);
+	return 0;
+}
diff --git a/virt/kvm/guest_memfd.c b/virt/kvm/guest_memfd.c
index d284bb70fe05..a509f1a96c0b 100644
--- a/virt/kvm/guest_memfd.c
+++ b/virt/kvm/guest_memfd.c
@@ -624,7 +624,7 @@ static void kvm_gmem_native_unbind(struct file *slot_file, struct kvm *kvm,
 static int  kvm_gmem_native_get_pfn(struct file *file, struct kvm *kvm,
 				    struct kvm_memory_slot *slot, gfn_t gfn,
 				    kvm_pfn_t *pfn, struct page **page,
-				    int *max_order);
+				    int *max_order, bool *writable);
 static void kvm_gmem_native_release(struct file *file);
 static int  kvm_gmem_native_mmap(struct file *file,
 				 struct vm_area_struct *vma);
@@ -988,7 +988,7 @@ static struct folio *__kvm_gmem_get_pfn(struct file *file,
 static int kvm_gmem_native_get_pfn(struct file *file, struct kvm *kvm,
 				   struct kvm_memory_slot *slot, gfn_t gfn,
 				   kvm_pfn_t *pfn, struct page **page,
-				   int *max_order)
+				   int *max_order, bool *writable)
 {
 	pgoff_t index = kvm_gmem_get_index(slot, gfn);
 	struct folio *folio;
@@ -1016,7 +1016,7 @@ static int kvm_gmem_native_get_pfn(struct file *file, struct kvm *kvm,
 
 int kvm_gmem_get_pfn(struct kvm *kvm, struct kvm_memory_slot *slot,
 		     gfn_t gfn, kvm_pfn_t *pfn, struct page **page,
-		     int *max_order)
+		     int *max_order, bool *writable)
 {
 	const struct kvm_gmem_ops *ops;
 
@@ -1029,7 +1029,10 @@ int kvm_gmem_get_pfn(struct kvm *kvm, struct kvm_memory_slot *slot,
 		return -EFAULT;
 
 	*page = NULL;
-	return ops->get_pfn(file, kvm, slot, gfn, pfn, page, max_order);
+	if (writable)
+		*writable = true;
+	return ops->get_pfn(file, kvm, slot, gfn, pfn, page, max_order,
+			    writable);
 }
 EXPORT_SYMBOL_FOR_KVM_INTERNAL(kvm_gmem_get_pfn);
 
@@ -1080,6 +1083,7 @@ static int kvm_gmem_populate_one(const struct kvm_gmem_ops *ops,
 				      void *opaque)
 {
 	struct page *ignored_page = NULL;
+	bool writable = true;
 	kvm_pfn_t pfn;
 	int ret;
 
@@ -1093,10 +1097,17 @@ static int kvm_gmem_populate_one(const struct kvm_gmem_ops *ops,
 				    src_page, 0);
 	else
 		ret = ops->get_pfn(file, kvm, slot, gfn, &pfn,
-				   &ignored_page, NULL);
+				   &ignored_page, NULL, &writable);
 	if (ret)
 		return ret;
 
+	/* post_populate() writes the page, so it cannot be read-only. */
+	if (!writable) {
+		if (ignored_page)
+			put_page(ignored_page);
+		return -EPERM;
+	}
+
 	ret = post_populate(kvm, gfn, pfn, src_page, opaque);
 
 	/*


^ permalink raw reply related	[flat|nested] 10+ messages in thread

* [PATCH 2/9] mm: Add memory providers
  2026-10-06 18:32 ` [RFC PATCH 0/9] mm: Memory providers for guest_memfd and iommufd Fred Griffoul
  2026-10-06 18:32   ` [PATCH 1/9] KVM: guest_memfd: Add a writable result to get_pfn() Fred Griffoul
@ 2026-10-06 18:32   ` Fred Griffoul
  2026-10-06 18:32   ` [PATCH 3/9] KVM: guest_memfd: Add a memory provider backing Fred Griffoul
                     ` (6 subsequent siblings)
  8 siblings, 0 replies; 10+ messages in thread
From: Fred Griffoul @ 2026-10-06 18:32 UTC (permalink / raw)
  To: Paolo Bonzini, Sean Christopherson, Marc Zyngier, Oliver Upton,
	Andrew Morton, David Hildenbrand, Alexander Viro,
	Christian Brauner, Jan Kara, Jason Gunthorpe, Kevin Tian,
	Joerg Roedel, Will Deacon, Robin Murphy, Thomas Gleixner,
	Ingo Molnar, Borislav Petkov, Dave Hansen, x86, H . Peter Anvin,
	Jonathan Corbet, Shuah Khan
  Cc: David Woodhouse, Ackerley Tng, Lorenzo Stoakes, Liam R . Howlett,
	Vlastimil Babka, Mike Rapoport, Suren Baghdasaryan, Michal Hocko,
	Joey Gouly, Suzuki K Poulose, Zenghui Yu, Steffen Eiden,
	linux-kernel, kvm, kvmarm, iommu, linux-fsdevel, linux-mm,
	linux-kselftest

From: Fred Griffoul <fgriffo@amazon.co.uk>

A driver that owns memory outside the page allocator, often without
struct page, has no common way to lend it to the code that maps it.
guest_memfd and iommufd each need their own hooks, and the owner must
keep them consistent when it takes memory back.

Add an interface between such an owner, the provider, and its
consumers. A consumer attaches to a provider file, asks for the frame
and attributes behind each page, and makes every mapping itself. When
the provider changes a range, it revokes it, and every consumer of the
file removes its mappings before the revoke returns.

A frame is therefore valid until its revoke returns, and consumers hold
no references. A revoke reaches every consumer of the file, so a
provider cannot take a frame from one consumer while another still maps
it.

A provider is found from its file: its file_operations sit in a struct
mem_provider_fops with the provider's operations, and FOP_MEM_PROVIDER
is set. Each provider file embeds a struct mem_provider_file that holds
its consumers, and the provider revokes through it. The core has no
registry and no global state, and struct file_operations does not grow.

The provider also owns the host memory type of its frames. Memory that
is not System RAM must have its type reserved, or PAT makes it uncached
on x86.

Suggested-by: David Woodhouse <dwmw@amazon.co.uk>
Signed-off-by: Fred Griffoul <fgriffo@amazon.co.uk>
---
 MAINTAINERS                  |   7 ++
 include/linux/fs.h           |   2 +
 include/linux/mem_provider.h | 211 +++++++++++++++++++++++++++++++++++
 mm/Kconfig                   |   7 ++
 mm/Makefile                  |   1 +
 mm/mem_provider.c            | 177 +++++++++++++++++++++++++++++
 6 files changed, 405 insertions(+)
 create mode 100644 include/linux/mem_provider.h
 create mode 100644 mm/mem_provider.c

diff --git a/MAINTAINERS b/MAINTAINERS
index f37a81950e25..6cab075a3ff7 100644
--- a/MAINTAINERS
+++ b/MAINTAINERS
@@ -17360,6 +17360,13 @@ T:	git git://git.kernel.org/pub/scm/linux/kernel/git/akpm/mm
 F:	include/uapi/asm-generic/mman-common.h
 F:	mm/madvise.c
 
+MEMORY PROVIDERS
+M:	Fred Griffoul <fgriffo@amazon.co.uk>
+L:	linux-mm@kvack.org
+S:	Maintained
+F:	include/linux/mem_provider.h
+F:	mm/mem_provider.c
+
 MEMORY TECHNOLOGY DEVICES (MTD)
 M:	Miquel Raynal <miquel.raynal@bootlin.com>
 M:	Richard Weinberger <richard@nod.at>
diff --git a/include/linux/fs.h b/include/linux/fs.h
index 50ce731a2b78..f459f0e34ca5 100644
--- a/include/linux/fs.h
+++ b/include/linux/fs.h
@@ -1980,6 +1980,8 @@ struct file_operations {
 #define FOP_ASYNC_LOCK		((__force fop_flags_t)(1 << 6))
 /* File system supports uncached read/write buffered IO */
 #define FOP_DONTCACHE		((__force fop_flags_t)(1 << 7))
+/* Lends memory to consumers: embedded in a struct mem_provider_fops */
+#define FOP_MEM_PROVIDER	((__force fop_flags_t)(1 << 8))
 
 /* Wrap a directory iterator that needs exclusive inode access */
 int wrap_directory_iterator(struct file *, struct dir_context *,
diff --git a/include/linux/mem_provider.h b/include/linux/mem_provider.h
new file mode 100644
index 000000000000..94e9352ebeb0
--- /dev/null
+++ b/include/linux/mem_provider.h
@@ -0,0 +1,211 @@
+/* SPDX-License-Identifier: GPL-2.0 */
+#ifndef _LINUX_MEM_PROVIDER_H
+#define _LINUX_MEM_PROVIDER_H
+
+#include <linux/bits.h>
+#include <linux/fs.h>
+#include <linux/list.h>
+#include <linux/rwsem.h>
+#include <linux/types.h>
+
+/*
+ * Memory providers
+ *
+ * A provider owns physical memory that the page allocator does not manage,
+ * and lends it to consumers through files that it hands to userspace.  A
+ * consumer is given such a file, attaches to it, and asks the provider for
+ * the frame behind each page.  The consumer makes every mapping of the
+ * frames itself, including the mappings into userspace; the provider makes
+ * none of them.
+ *
+ * When the frames or attributes behind a range change, the provider revokes
+ * the range.  The core calls every consumer attached to the file, and each
+ * one removes all its mappings of the range before the revoke returns.  A
+ * frame therefore stays valid for a consumer until the revoke that removes
+ * it has returned, and consumers hold no references.
+ *
+ * A consumer must not install a mapping of a frame after the revoke that
+ * removes it has returned.  It either holds a lock across get_page() and
+ * the map that its revoke callback also takes, or detects a revoke that ran
+ * in between and retries, as KVM does with its invalidation sequence.
+ *
+ * A revoke may also change the contents behind the range, for example when
+ * the provider moves a frame from one file to another.  A consumer that
+ * reports the pages written through its mappings, for live migration, must
+ * therefore treat the whole revoked range as written.
+ *
+ * A provider cannot take a frame back from one consumer only: a revoke
+ * reaches every consumer of the file.
+ *
+ * A provider gives its files a struct mem_provider_fops, which holds their
+ * file_operations, with FOP_MEM_PROVIDER and an owner set, and the provider's
+ * operations.  It embeds a struct mem_provider_file in the state of each
+ * file, returns it from attach(), and passes it to mem_provider_revoke().
+ * An attachment holds a reference to the file, so neither the file's state
+ * nor the module can go away while a consumer is attached.  MEM_PROVIDER
+ * has no prompt: a provider selects it, or depends on a consumer that does.
+ */
+
+/*
+ * Memory type of a frame, in the low bits of the attributes.  The provider
+ * makes the host's memory type for its frames match what it reports: for
+ * memory that is not System RAM, it reserves the type of the range, for
+ * example with memremap() or ioremap_wc(), before it hands frames out.  The
+ * architecture otherwise chooses the type of an unknown range, uncached on
+ * x86, and the VMM's mapping of a guest_memfd would differ from the guest's.
+ * On x86 the guest's type also depends on the host: KVM maps a frame uncached
+ * for the guest unless it is RAM in the firmware's E820 map, or a struct-page
+ * frame (DAX, ZONE_DEVICE) that PAT does not force uncached.  So the
+ * reservation matters for the guest too, not only for the VMM's mapping.
+ */
+#define MEM_PROVIDER_ATTR_TYPE		GENMASK(3, 0)
+#define MEM_PROVIDER_TYPE_RAM		0	/* cacheable system memory */
+#define MEM_PROVIDER_TYPE_MMIO		1	/* uncached device memory */
+#define MEM_PROVIDER_TYPE_MMIO_WC	2	/* write-combining device memory */
+						/* 3 to 15 are reserved */
+
+/*
+ * No consumer may map the frame writable: not for a guest, a device or
+ * userspace.
+ */
+#define MEM_PROVIDER_ATTR_READONLY	BIT(4)
+
+/*
+ * No consumer may map the frame into userspace.  Guest and device mappings
+ * are not affected.  The provider decides this for each page.
+ */
+#define MEM_PROVIDER_ATTR_NO_USER_MAP	BIT(5)
+
+#define MEM_PROVIDER_ATTR_VALID		(MEM_PROVIDER_ATTR_TYPE | \
+					 MEM_PROVIDER_ATTR_READONLY | \
+					 MEM_PROVIDER_ATTR_NO_USER_MAP)
+
+static inline unsigned int mem_provider_type(u32 attrs)
+{
+	return attrs & MEM_PROVIDER_ATTR_TYPE;
+}
+
+/*
+ * The provider side of one provider file, embedded in the provider's state
+ * for the file and set up with mem_provider_file_init().  It must stay until
+ * the file is released.  Its fields are private to the core.
+ */
+struct mem_provider_file {
+	struct rw_semaphore lock;
+	struct list_head consumers;
+};
+
+/**
+ * struct mem_provider_ops - What a provider implements.
+ */
+struct mem_provider_ops {
+	/**
+	 * @attach: A consumer starts to use @file.
+	 *
+	 * Called once for each consumer, and possibly for several at once.
+	 *
+	 * @size is the end of the range of the file that the consumer will
+	 * use, from offset 0, as the consumer's user asked for it: the size
+	 * of a guest_memfd, or the end of an iommufd mapping.  Return the
+	 * file's struct mem_provider_file, which is passed to the other
+	 * operations, or an ERR_PTR().  Return ERR_PTR(-EINVAL) if the file
+	 * is smaller than @size, so that the consumer fails the request at
+	 * once.  May sleep.
+	 */
+	struct mem_provider_file *(*attach)(struct file *file, loff_t size);
+
+	/**
+	 * @detach: A consumer is gone.
+	 *
+	 * Called once for each successful attach().  The consumer has removed
+	 * every mapping of the frames, and the core no longer calls its
+	 * revoke callback.  May sleep.
+	 */
+	void (*detach)(struct mem_provider_file *mpf);
+
+	/**
+	 * @get_page: Return the frame behind page @index.
+	 *
+	 * @pfn:       [out] The frame.
+	 * @max_order: [in, out] On entry, the largest order the consumer
+	 *             can use.  On return, the largest order for which the
+	 *             aligned block that contains @index is physically
+	 *             contiguous and has the same attributes.  It must not
+	 *             be larger than on entry.
+	 * @attrs:     [out] MEM_PROVIDER_ATTR_* for the block.
+	 *
+	 * Return 0, -EFAULT if the provider does not back the page now, or
+	 * another negative errno.
+	 *
+	 * May sleep, and may run at the same time as the provider changes
+	 * its state.  A result that is out of date is harmless, because the
+	 * provider revokes the range after the change.  Must not call into
+	 * the consumer, must not call mem_provider_revoke(), and must not take
+	 * a lock that the provider holds across mem_provider_revoke().
+	 */
+	int (*get_page)(struct mem_provider_file *mpf, pgoff_t index,
+			unsigned long *pfn, int *max_order, u32 *attrs);
+};
+
+/*
+ * The file_operations of a provider's files, with FOP_MEM_PROVIDER set in
+ * fops.fop_flags, and the provider's operations.
+ */
+struct mem_provider_fops {
+	struct file_operations fops;
+	const struct mem_provider_ops *ops;
+};
+
+/**
+ * struct mem_provider_attachment - One consumer's attachment to a provider
+ * file.  Embedded in the consumer's object, set by mem_provider_attach() and
+ * owned by the consumer until mem_provider_detach().  @ops must be NULL before
+ * the first mem_provider_attach(), for example by zeroing the attachment, so
+ * that mem_provider_detach() on an attachment that was never made does
+ * nothing.
+ * @ops: The provider's operations.  Read-only to the consumer.
+ * @mpf: The provider's per-file state.  Read-only to the consumer.
+ * @file: The provider file.  Read-only to the consumer.
+ * @size: The end of the range the consumer uses.  Read-only to the consumer.
+ * @revoke: Called when the frames behind a range change; see below.
+ * @node: Private to the core.
+ */
+struct mem_provider_attachment {
+	const struct mem_provider_ops *ops;
+	struct mem_provider_file *mpf;
+	struct file *file;
+	loff_t size;
+
+	/*
+	 * @revoke: The frames behind [@offset, @offset + @len) changed.
+	 *
+	 * The consumer removes every mapping it made of the range and asks
+	 * get_page() again when it next needs a page.  It must not return
+	 * while a mapping of the old frames remains, including one being
+	 * installed from an earlier get_page(): it excludes such a mapping or
+	 * makes it retry.  Revokes of the same file may call this at the same
+	 * time, and may call it before mem_provider_attach() returns.  May
+	 * sleep, and may call get_page().  Must not call
+	 * mem_provider_revoke(), and must not attach or detach any provider.
+	 */
+	void (*revoke)(struct mem_provider_attachment *att, loff_t offset,
+		       loff_t len);
+
+	struct list_head node;
+};
+
+/* Provider side. */
+void mem_provider_file_init(struct mem_provider_file *mpf);
+void mem_provider_revoke(struct mem_provider_file *mpf, loff_t offset,
+			 loff_t len);
+
+/* Consumer side. */
+int mem_provider_attach(struct mem_provider_attachment *att, struct file *file,
+			loff_t size,
+			void (*revoke)(struct mem_provider_attachment *att,
+				       loff_t offset, loff_t len));
+void mem_provider_detach(struct mem_provider_attachment *att);
+int mem_provider_get_page(struct mem_provider_attachment *att, pgoff_t index,
+			  unsigned long *pfn, int *max_order, u32 *attrs);
+
+#endif /* _LINUX_MEM_PROVIDER_H */
diff --git a/mm/Kconfig b/mm/Kconfig
index 9e0ca4824905..7085bf6666e5 100644
--- a/mm/Kconfig
+++ b/mm/Kconfig
@@ -774,6 +774,13 @@ config DEFAULT_MMAP_MIN_ADDR
 config ARCH_SUPPORTS_MEMORY_FAILURE
 	bool
 
+config MEM_PROVIDER
+	bool
+	help
+	  An interface that lets a driver lend memory that the page allocator
+	  does not manage to consumers such as guest_memfd and iommufd, and
+	  take it back.  See include/linux/mem_provider.h.
+
 config MEMORY_FAILURE
 	depends on MMU
 	depends on ARCH_SUPPORTS_MEMORY_FAILURE
diff --git a/mm/Makefile b/mm/Makefile
index eff9f9e7e061..35420b587f89 100644
--- a/mm/Makefile
+++ b/mm/Makefile
@@ -110,6 +110,7 @@ obj-$(CONFIG_CGROUP_HUGETLB) += hugetlb_cgroup.o
 obj-$(CONFIG_GUP_TEST) += gup_test.o
 obj-$(CONFIG_DMAPOOL_TEST) += dmapool_test.o
 obj-$(CONFIG_MEMORY_FAILURE) += memory-failure.o
+obj-$(CONFIG_MEM_PROVIDER) += mem_provider.o
 obj-$(CONFIG_HWPOISON_INJECT) += hwpoison-inject.o
 obj-$(CONFIG_DEBUG_KMEMLEAK) += kmemleak.o
 obj-$(CONFIG_DEBUG_RODATA_TEST) += rodata_test.o
diff --git a/mm/mem_provider.c b/mm/mem_provider.c
new file mode 100644
index 000000000000..5a2984afde10
--- /dev/null
+++ b/mm/mem_provider.c
@@ -0,0 +1,177 @@
+// SPDX-License-Identifier: GPL-2.0
+/*
+ * Memory providers: lend memory that the page allocator does not manage to
+ * consumers that map it.  See include/linux/mem_provider.h.
+ *
+ * Locking: the @lock of a struct mem_provider_file protects its list of
+ * consumers.  mem_provider_revoke() holds it for reading while it calls the
+ * consumers, so a consumer cannot detach until every revoke of its file has
+ * returned.  It follows that a revoke callback must not attach, detach or
+ * revoke, and that a consumer must not detach while it holds a lock that its
+ * revoke callback takes.  A revoke callback may call get_page(), which takes
+ * no lock of the core.
+ */
+#include <linux/export.h>
+#include <linux/file.h>
+#include <linux/fs.h>
+#include <linux/mem_provider.h>
+#include <linux/mm.h>
+#include <linux/rwsem.h>
+
+/**
+ * mem_provider_file_init() - Set up the provider side of a provider file.
+ * @mpf: Embedded in the provider's state for the file.
+ */
+void mem_provider_file_init(struct mem_provider_file *mpf)
+{
+	init_rwsem(&mpf->lock);
+	INIT_LIST_HEAD(&mpf->consumers);
+}
+EXPORT_SYMBOL_GPL(mem_provider_file_init);
+
+/**
+ * mem_provider_attach() - Attach a consumer to a provider file.
+ * @att: The consumer's attachment, filled in on success.
+ * @file: A file handed out by a provider.
+ * @size: The end of the range of @file that the consumer will use, in bytes
+ *        from offset 0.  The provider refuses a file that is smaller.
+ * @revoke: Called when the frames behind a range of @file change.
+ *
+ * Holds a reference to @file, and so to the provider's module, until
+ * mem_provider_detach().
+ *
+ * Return: 0, -EINVAL if @size or @revoke is invalid or the provider lacks an
+ * operation, -ENODEV if @file is not a provider file, or the error returned
+ * by the provider.
+ */
+int mem_provider_attach(struct mem_provider_attachment *att, struct file *file,
+			loff_t size,
+			void (*revoke)(struct mem_provider_attachment *att,
+				       loff_t offset, loff_t len))
+{
+	const struct mem_provider_ops *ops;
+	struct mem_provider_file *mpf;
+
+	if (size <= 0 || !revoke)
+		return -EINVAL;
+
+	if (!(file->f_op->fop_flags & FOP_MEM_PROVIDER))
+		return -ENODEV;
+	ops = container_of(file->f_op, struct mem_provider_fops, fops)->ops;
+	if (WARN_ON_ONCE(!ops->attach || !ops->detach || !ops->get_page))
+		return -EINVAL;
+
+	mpf = ops->attach(file, size);
+	if (IS_ERR(mpf))
+		return PTR_ERR(mpf);
+
+	att->ops = ops;
+	att->mpf = mpf;
+	att->file = get_file(file);
+	att->size = size;
+	att->revoke = revoke;
+
+	down_write(&mpf->lock);
+	list_add_tail(&att->node, &mpf->consumers);
+	up_write(&mpf->lock);
+	return 0;
+}
+EXPORT_SYMBOL_GPL(mem_provider_attach);
+
+/**
+ * mem_provider_detach() - Detach a consumer from a provider file.
+ * @att: An attachment from mem_provider_attach().
+ *
+ * The consumer must have removed every mapping of the provider's frames.
+ * Waits for revokes of the file that are in progress.  Does nothing if @att
+ * is not attached.
+ */
+void mem_provider_detach(struct mem_provider_attachment *att)
+{
+	struct mem_provider_file *mpf = att->mpf;
+
+	if (!att->ops)
+		return;
+
+	down_write(&mpf->lock);
+	list_del(&att->node);
+	up_write(&mpf->lock);
+
+	att->ops->detach(mpf);
+	fput(att->file);
+
+	att->ops = NULL;
+	att->mpf = NULL;
+	att->file = NULL;
+}
+EXPORT_SYMBOL_GPL(mem_provider_detach);
+
+/**
+ * mem_provider_get_page() - Get the frame behind a page of an attachment.
+ * @att: The attachment.
+ * @index: The page, in units of PAGE_SIZE from offset 0.
+ * @pfn: [out] The frame.
+ * @max_order: [in, out] See &mem_provider_ops.get_page.
+ * @attrs: [out] MEM_PROVIDER_ATTR_* for the frame.
+ *
+ * On return, @max_order is also limited so that the block is aligned in
+ * physical memory.  The frame stays valid until the consumer's revoke
+ * callback for the page has returned.
+ *
+ * Return: 0, -EFAULT if the provider does not back the page now, -EINVAL if
+ * @index is beyond the attachment, -EIO if the provider returned an invalid
+ * result, or another negative errno from the provider.
+ */
+int mem_provider_get_page(struct mem_provider_attachment *att, pgoff_t index,
+			  unsigned long *pfn, int *max_order, u32 *attrs)
+{
+	int order = *max_order;
+	int ret;
+
+	if (index >= DIV_ROUND_UP(att->size, PAGE_SIZE))
+		return -EINVAL;
+
+	*attrs = 0;
+	ret = att->ops->get_page(att->mpf, index, pfn, max_order, attrs);
+	if (ret)
+		return ret;
+
+	if (WARN_ON_ONCE(*max_order < 0 || *max_order > order ||
+			 (*attrs & ~MEM_PROVIDER_ATTR_VALID) ||
+			 mem_provider_type(*attrs) > MEM_PROVIDER_TYPE_MMIO_WC))
+		return -EIO;
+
+	/* A large mapping also needs the first frame of the block aligned. */
+	if (*max_order && ((*pfn ^ index) & ((1UL << *max_order) - 1)))
+		*max_order = __ffs(*pfn ^ index);
+	return 0;
+}
+EXPORT_SYMBOL_GPL(mem_provider_get_page);
+
+/**
+ * mem_provider_revoke() - Tell the consumers that frames changed.
+ * @mpf: The provider side of the file.
+ * @offset: The start of the range, in bytes.
+ * @len: The length of the range, in bytes.
+ *
+ * Call this after the provider has changed the frames or attributes behind
+ * the range.  Returns when every consumer of the file has removed its
+ * mappings of the range.  The provider must not hold a lock that its
+ * get_page() takes, because a consumer may call get_page() before it
+ * returns.  May sleep.
+ */
+void mem_provider_revoke(struct mem_provider_file *mpf, loff_t offset,
+			 loff_t len)
+{
+	struct mem_provider_attachment *att;
+
+	might_sleep();
+	if (len <= 0)
+		return;
+
+	down_read(&mpf->lock);
+	list_for_each_entry(att, &mpf->consumers, node)
+		att->revoke(att, offset, len);
+	up_read(&mpf->lock);
+}
+EXPORT_SYMBOL_GPL(mem_provider_revoke);


^ permalink raw reply related	[flat|nested] 10+ messages in thread

* [PATCH 3/9] KVM: guest_memfd: Add a memory provider backing
  2026-10-06 18:32 ` [RFC PATCH 0/9] mm: Memory providers for guest_memfd and iommufd Fred Griffoul
  2026-10-06 18:32   ` [PATCH 1/9] KVM: guest_memfd: Add a writable result to get_pfn() Fred Griffoul
  2026-10-06 18:32   ` [PATCH 2/9] mm: Add memory providers Fred Griffoul
@ 2026-10-06 18:32   ` Fred Griffoul
  2026-10-06 18:32   ` [PATCH 4/9] iommufd: Track the domains of pages that are not pinned Fred Griffoul
                     ` (5 subsequent siblings)
  8 siblings, 0 replies; 10+ messages in thread
From: Fred Griffoul @ 2026-10-06 18:32 UTC (permalink / raw)
  To: Paolo Bonzini, Sean Christopherson, Marc Zyngier, Oliver Upton,
	Andrew Morton, David Hildenbrand, Alexander Viro,
	Christian Brauner, Jan Kara, Jason Gunthorpe, Kevin Tian,
	Joerg Roedel, Will Deacon, Robin Murphy, Thomas Gleixner,
	Ingo Molnar, Borislav Petkov, Dave Hansen, x86, H . Peter Anvin,
	Jonathan Corbet, Shuah Khan
  Cc: David Woodhouse, Ackerley Tng, Lorenzo Stoakes, Liam R . Howlett,
	Vlastimil Babka, Mike Rapoport, Suren Baghdasaryan, Michal Hocko,
	Joey Gouly, Suzuki K Poulose, Zenghui Yu, Steffen Eiden,
	linux-kernel, kvm, kvmarm, iommu, linux-fsdevel, linux-mm,
	linux-kselftest

From: Fred Griffoul <fgriffo@amazon.co.uk>

A driver can back a guest_memfd today only by implementing
kvm_gmem_ops itself, and must then keep any device mapping of the same
memory consistent on its own.

Add GUEST_MEMFD_FLAG_USE_PROVIDER. KVM_CREATE_GUEST_MEMFD then takes a
memory provider file in provider_fd, which uses the first reserved word
and must be zero without the flag. Guest faults ask the provider for
each frame, and a revoke removes the range from the guest and from the
VMM's mapping. A page that is not backed, or that is not RAM, exits with
KVM_EXIT_MEMORY_FAULT. Read-only pages are mapped read only, and
fallocate() is not supported.

guest_memfd maps itself into userspace from the same frames, so that
a provider needs no fault handler. Pages marked NO_USER_MAP raise
SIGBUS there.

USE_PROVIDER requires GUEST_MEMFD_FLAG_MMAP, so that the memory slot
is gmem-only. An architecture opts in; x86 does so for VMs without
private or encrypted memory, and other architectures refuse the flag
for now. Native guest_memfd files now take the invalidate lock when
a file is added and when a closing file is unbound.

Signed-off-by: Fred Griffoul <fgriffo@amazon.co.uk>
---
 Documentation/virt/kvm/api.rst                |  45 ++-
 arch/x86/kvm/x86.c                            |  10 +
 include/linux/kvm_host.h                      |   4 +
 include/uapi/linux/kvm.h                      |  11 +-
 tools/include/uapi/linux/kvm.h                |  11 +-
 .../testing/selftests/kvm/guest_memfd_test.c  |   6 +
 virt/kvm/Kconfig                              |   1 +
 virt/kvm/guest_memfd.c                        | 265 +++++++++++++++++-
 8 files changed, 338 insertions(+), 15 deletions(-)

diff --git a/Documentation/virt/kvm/api.rst b/Documentation/virt/kvm/api.rst
index a5f9ee92f43e..d9aaf8a76b59 100644
--- a/Documentation/virt/kvm/api.rst
+++ b/Documentation/virt/kvm/api.rst
@@ -6431,7 +6431,9 @@ and cannot be resized  (guest_memfd files do however support PUNCH_HOLE).
   struct kvm_create_guest_memfd {
 	__u64 size;
 	__u64 flags;
-	__u64 reserved[6];
+	__s32 provider_fd;
+	__u32 pad;
+	__u64 reserved[5];
   };
 
 Conceptually, the inode backing a guest_memfd file represents physical memory,
@@ -6453,15 +6455,38 @@ a single guest_memfd file, but the bound ranges must not overlap).
 The capability KVM_CAP_GUEST_MEMFD_FLAGS enumerates the `flags` that can be
 specified via KVM_CREATE_GUEST_MEMFD.  Currently defined flags:
 
-  ============================ ================================================
-  GUEST_MEMFD_FLAG_MMAP        Enable using mmap() on the guest_memfd file
-                               descriptor.
-  GUEST_MEMFD_FLAG_INIT_SHARED Make all memory in the file shared during
-                               KVM_CREATE_GUEST_MEMFD (memory files created
-                               without INIT_SHARED will be marked private).
-                               Shared memory can be faulted into host userspace
-                               page tables. Private memory cannot.
-  ============================ ================================================
+  ============================= ================================================
+  GUEST_MEMFD_FLAG_MMAP         Enable using mmap() on the guest_memfd file
+                                descriptor.
+  GUEST_MEMFD_FLAG_INIT_SHARED  Make all memory in the file shared during
+                                KVM_CREATE_GUEST_MEMFD (memory files created
+                                without INIT_SHARED will be marked private).
+                                Shared memory can be faulted into host userspace
+                                page tables. Private memory cannot.
+  GUEST_MEMFD_FLAG_USE_PROVIDER Take the memory of the file from the memory
+                                provider behind `provider_fd`, instead of
+                                allocating it.  Requires
+                                GUEST_MEMFD_FLAG_MMAP.  Not supported for VMs
+                                with private or encrypted memory.
+  ============================= ================================================
+
+`pad` must be zero.  With GUEST_MEMFD_FLAG_USE_PROVIDER, `provider_fd` is a file
+handed out by a memory provider (see include/linux/mem_provider.h); without it,
+`provider_fd` must be zero.
+The provider decides which pages exist and what backs them, and can change
+this at any time.  The guest and any host mapping of the guest_memfd follow
+the change.  A guest access to a page that the provider does not back exits
+to userspace with KVM_EXIT_MEMORY_FAULT, and so does an access to a page that
+is not RAM.  A page that the provider marks read only is mapped read only for
+the guest.  fallocate() is not supported.  Only x86 VMs of type
+KVM_X86_DEFAULT_VM support GUEST_MEMFD_FLAG_USE_PROVIDER.
+
+mmap() of the guest_memfd maps the provider's pages into userspace on fault.
+An access raises SIGBUS if the provider does not back the page, if the page is
+not RAM, if the provider marks it as not to be mapped into userspace, or if it
+is a write to a read-only page.
+KVM's own accesses through the memslot's userspace address fail in the same
+cases.
 
 When the KVM MMU performs a PFN lookup to service a guest fault and the backing
 guest_memfd has the GUEST_MEMFD_FLAG_MMAP set, then the fault will always be
diff --git a/arch/x86/kvm/x86.c b/arch/x86/kvm/x86.c
index afcac1042947..3669316f5712 100644
--- a/arch/x86/kvm/x86.c
+++ b/arch/x86/kvm/x86.c
@@ -14124,6 +14124,16 @@ bool kvm_arch_supports_gmem_init_shared(struct kvm *kvm)
 	return !kvm_arch_has_private_mem(kvm);
 }
 
+/*
+ * Only VMs without encrypted memory: SEV and SEV-ES guests have no private
+ * memory in KVM's sense, but their memory would need reclaiming when a
+ * provider takes it back.
+ */
+bool kvm_arch_gmem_supports_provider(struct kvm *kvm)
+{
+	return !kvm || kvm->arch.vm_type == KVM_X86_DEFAULT_VM;
+}
+
 #ifdef CONFIG_HAVE_KVM_ARCH_GMEM_PREPARE
 int kvm_arch_gmem_prepare(struct kvm *kvm, gfn_t gfn, kvm_pfn_t pfn, int max_order)
 {
diff --git a/include/linux/kvm_host.h b/include/linux/kvm_host.h
index 7281d0e94121..3ea591642542 100644
--- a/include/linux/kvm_host.h
+++ b/include/linux/kvm_host.h
@@ -832,11 +832,15 @@ static inline bool kvm_arch_has_private_mem(struct kvm *kvm)
 
 #ifdef CONFIG_KVM_GUEST_MEMFD
 bool kvm_arch_supports_gmem_init_shared(struct kvm *kvm);
+bool kvm_arch_gmem_supports_provider(struct kvm *kvm);
 
 static inline u64 kvm_gmem_get_supported_flags(struct kvm *kvm)
 {
 	u64 flags = GUEST_MEMFD_FLAG_MMAP;
 
+	if (kvm_arch_gmem_supports_provider(kvm))
+		flags |= GUEST_MEMFD_FLAG_USE_PROVIDER;
+
 	if (!kvm || kvm_arch_supports_gmem_init_shared(kvm))
 		flags |= GUEST_MEMFD_FLAG_INIT_SHARED;
 
diff --git a/include/uapi/linux/kvm.h b/include/uapi/linux/kvm.h
index 419011097fa8..820448f82e22 100644
--- a/include/uapi/linux/kvm.h
+++ b/include/uapi/linux/kvm.h
@@ -1654,11 +1654,20 @@ struct kvm_memory_attributes {
 #define KVM_CREATE_GUEST_MEMFD	_IOWR(KVMIO,  0xd4, struct kvm_create_guest_memfd)
 #define GUEST_MEMFD_FLAG_MMAP		(1ULL << 0)
 #define GUEST_MEMFD_FLAG_INIT_SHARED	(1ULL << 1)
+#define GUEST_MEMFD_FLAG_USE_PROVIDER	(1ULL << 2)
 
 struct kvm_create_guest_memfd {
 	__u64 size;
 	__u64 flags;
-	__u64 reserved[6];
+	/*
+	 * With GUEST_MEMFD_FLAG_USE_PROVIDER: a memory provider file whose
+	 * pages back this guest_memfd.  The provider may take
+	 * any range back at any time; the guest and every host mapping of
+	 * this guest_memfd follow.  Must be 0 otherwise.
+	 */
+	__s32 provider_fd;
+	__u32 pad;
+	__u64 reserved[5];
 };
 
 #define KVM_PRE_FAULT_MEMORY	_IOWR(KVMIO, 0xd5, struct kvm_pre_fault_memory)
diff --git a/tools/include/uapi/linux/kvm.h b/tools/include/uapi/linux/kvm.h
index d0c0c8605976..0f4d1dd0e931 100644
--- a/tools/include/uapi/linux/kvm.h
+++ b/tools/include/uapi/linux/kvm.h
@@ -1644,11 +1644,20 @@ struct kvm_memory_attributes {
 #define KVM_CREATE_GUEST_MEMFD	_IOWR(KVMIO,  0xd4, struct kvm_create_guest_memfd)
 #define GUEST_MEMFD_FLAG_MMAP		(1ULL << 0)
 #define GUEST_MEMFD_FLAG_INIT_SHARED	(1ULL << 1)
+#define GUEST_MEMFD_FLAG_USE_PROVIDER	(1ULL << 2)
 
 struct kvm_create_guest_memfd {
 	__u64 size;
 	__u64 flags;
-	__u64 reserved[6];
+	/*
+	 * With GUEST_MEMFD_FLAG_USE_PROVIDER: a memory provider file whose
+	 * pages back this guest_memfd.  The provider may take
+	 * any range back at any time; the guest and every host mapping of
+	 * this guest_memfd follow.  Must be 0 otherwise.
+	 */
+	__s32 provider_fd;
+	__u32 pad;
+	__u64 reserved[5];
 };
 
 #define KVM_PRE_FAULT_MEMORY	_IOWR(KVMIO, 0xd5, struct kvm_pre_fault_memory)
diff --git a/tools/testing/selftests/kvm/guest_memfd_test.c b/tools/testing/selftests/kvm/guest_memfd_test.c
index 2233d871a38f..91ff10ac6274 100644
--- a/tools/testing/selftests/kvm/guest_memfd_test.c
+++ b/tools/testing/selftests/kvm/guest_memfd_test.c
@@ -405,6 +405,12 @@ static void test_guest_memfd_flags(struct kvm_vm *vm)
 
 	for (flag = BIT(0); flag; flag <<= 1) {
 		fd = __vm_create_guest_memfd(vm, page_size, flag);
+		/* USE_PROVIDER also needs MMAP and a provider file. */
+		if (flag == GUEST_MEMFD_FLAG_USE_PROVIDER && (flag & valid_flags)) {
+			TEST_ASSERT(fd < 0 && errno == EINVAL,
+				    "guest_memfd() with USE_PROVIDER alone should fail with EINVAL");
+			continue;
+		}
 		if (flag & valid_flags) {
 			TEST_ASSERT(fd >= 0,
 				    "guest_memfd() with flag '0x%lx' should succeed",
diff --git a/virt/kvm/Kconfig b/virt/kvm/Kconfig
index 794976b88c6f..cfb6c4e51128 100644
--- a/virt/kvm/Kconfig
+++ b/virt/kvm/Kconfig
@@ -105,6 +105,7 @@ config KVM_GENERIC_MEMORY_ATTRIBUTES
 
 config KVM_GUEST_MEMFD
        select XARRAY_MULTI
+       select MEM_PROVIDER
        bool
 
 config HAVE_KVM_ARCH_GMEM_PREPARE
diff --git a/virt/kvm/guest_memfd.c b/virt/kvm/guest_memfd.c
index a509f1a96c0b..aedd8630e3ea 100644
--- a/virt/kvm/guest_memfd.c
+++ b/virt/kvm/guest_memfd.c
@@ -4,6 +4,7 @@
 #include <linux/falloc.h>
 #include <linux/fs.h>
 #include <linux/kvm_host.h>
+#include <linux/mem_provider.h>
 #include <linux/mempolicy.h>
 #include <linux/pseudo_fs.h>
 #include <linux/pagemap.h>
@@ -44,6 +45,9 @@ struct gmem_inode {
 	struct list_head gmem_file_list;
 
 	u64 flags;
+
+	/* The provider of the pages, with GUEST_MEMFD_FLAG_USE_PROVIDER. */
+	struct mem_provider_attachment att;
 };
 
 static __always_inline struct gmem_inode *GMEM_I(struct inode *inode)
@@ -654,7 +658,222 @@ static const struct kvm_gmem_ops kvm_gmem_native_ops = {
 	.fallocate = kvm_gmem_native_fallocate,
 };
 
-static int __kvm_gmem_create(struct kvm *kvm, loff_t size, u64 flags)
+/*
+ * Can @kvm use a memory provider?  @kvm is NULL for the system-wide capability.
+ * An architecture opts in by overriding this.  VMs whose memory is private or
+ * encrypted are not supported: each frame would need preparing before use and
+ * reclaiming when the provider takes it back.
+ */
+bool __weak kvm_arch_gmem_supports_provider(struct kvm *kvm)
+{
+	return false;
+}
+
+/*
+ * A guest_memfd backed by a memory provider (include/linux/mem_provider.h).
+ *
+ * The provider owns the frames, and guest_memfd keeps no state for them.  A
+ * guest fault asks the provider for the frame each time.  A fault that
+ * races with a change in the provider is retried by KVM, because the
+ * provider revokes the range after the change and the revoke opens and
+ * closes KVM's invalidation window.  Host mappings are made by guest_memfd
+ * from the same frames and attributes, and removed by the revoke.
+ *
+ * Bindings and release are the same as for the native guest_memfd.
+ */
+static void kvm_gmem_provider_revoke(struct mem_provider_attachment *att,
+				     loff_t offset, loff_t len)
+{
+	struct gmem_inode *gi = container_of(att, struct gmem_inode, att);
+	struct inode *inode = &gi->vfs_inode;
+	pgoff_t start, end;
+	struct gmem_file *f;
+
+	if (offset >= att->size)
+		return;
+	len = min(len, att->size - offset);
+	start = offset >> PAGE_SHIFT;
+	end = DIV_ROUND_UP(offset + len, PAGE_SIZE);
+
+	/*
+	 * The bindings must be stable so that each start is matched by an
+	 * end.  Only VMs without private memory use a provider, but remove
+	 * both kinds of mapping so that none is missed.
+	 */
+	filemap_invalidate_lock(inode->i_mapping);
+	kvm_gmem_for_each_file(f, inode)
+		__kvm_gmem_invalidate_start(f, start, end,
+					    KVM_FILTER_SHARED | KVM_FILTER_PRIVATE);
+	unmap_mapping_range(inode->i_mapping, (loff_t)start << PAGE_SHIFT,
+			    (loff_t)(end - start) << PAGE_SHIFT, 1);
+	kvm_gmem_for_each_file(f, inode)
+		__kvm_gmem_invalidate_end(f, start, end);
+	filemap_invalidate_unlock(inode->i_mapping);
+}
+
+static int kvm_gmem_provider_get_pfn(struct file *file, struct kvm *kvm,
+				     struct kvm_memory_slot *slot, gfn_t gfn,
+				     kvm_pfn_t *pfn, struct page **page,
+				     int *max_order, bool *writable)
+{
+	struct gmem_inode *gi = GMEM_I(file_inode(file));
+	pgoff_t index = kvm_gmem_get_index(slot, gfn);
+	unsigned long frame;
+	int order, ret;
+	u32 attrs;
+
+	if (file != READ_ONCE(slot->gmem.file))
+		return -EFAULT;
+	if (xa_load(&gmem_file_of(file)->bindings, index) != slot)
+		return -EIO;
+
+	/* A large mapping needs the block aligned in both index and gfn. */
+	order = PUD_ORDER;
+	if (index != gfn)
+		order = min_t(int, order, __ffs(index ^ gfn));
+
+	ret = mem_provider_get_page(&gi->att, index, &frame, &order, &attrs);
+	if (ret)
+		return ret;
+
+	/*
+	 * KVM decides the guest's memory type itself, so accept RAM only.
+	 * -EFAULT, so that the VMM gets a memory fault exit for the access.
+	 */
+	if (mem_provider_type(attrs) != MEM_PROVIDER_TYPE_RAM)
+		return -EFAULT;
+
+	if (attrs & MEM_PROVIDER_ATTR_READONLY) {
+		if (!writable)
+			return -EPERM;
+		*writable = false;
+	}
+
+	*pfn = frame;
+	if (max_order)
+		*max_order = order;
+	return 0;
+}
+
+/* The protection for a host mapping of a frame with @attrs. */
+static pgprot_t kvm_gmem_provider_prot(struct vm_area_struct *vma, u32 attrs)
+{
+	vm_flags_t flags = vma->vm_flags;
+
+	if (attrs & MEM_PROVIDER_ATTR_READONLY)
+		flags &= ~VM_WRITE;
+	return vm_get_page_prot(flags);
+}
+
+/*
+ * Map page @vmf->pgoff of a provider-backed guest_memfd for the host.
+ *
+ * The invalidate lock is held shared across get_page() and the insert, and
+ * a revoke holds it exclusive while it unmaps the range, so a frame is never
+ * mapped after the revoke that removes it.  A writable page is mapped
+ * writable, and a read-only page without write permission.  With @mkwrite,
+ * the PTE is read only and the write must be checked: replace it under the
+ * lock, so that a change to read only cannot slip between the check and the
+ * upgrade.
+ */
+static vm_fault_t kvm_gmem_provider_map_host(struct vm_fault *vmf, bool mkwrite)
+{
+	struct vm_area_struct *vma = vmf->vma;
+	struct inode *inode = file_inode(vma->vm_file);
+	unsigned long uaddr = vmf->address & PAGE_MASK;
+	bool write = vmf->flags & FAULT_FLAG_WRITE;
+	unsigned long pfn;
+	int order = 0;
+	vm_fault_t ret;
+	u32 attrs;
+
+	if (((loff_t)vmf->pgoff << PAGE_SHIFT) >= i_size_read(inode))
+		return VM_FAULT_SIGBUS;
+
+	filemap_invalidate_lock_shared(inode->i_mapping);
+	if (mem_provider_get_page(&GMEM_I(inode)->att, vmf->pgoff, &pfn,
+				  &order, &attrs)) {
+		ret = VM_FAULT_SIGBUS;
+	} else if (mem_provider_type(attrs) != MEM_PROVIDER_TYPE_RAM ||
+		   (attrs & MEM_PROVIDER_ATTR_NO_USER_MAP) ||
+		   (write && (attrs & MEM_PROVIDER_ATTR_READONLY))) {
+		ret = VM_FAULT_SIGBUS;
+	} else {
+		if (mkwrite)
+			zap_special_vma_range(vma, uaddr, PAGE_SIZE);
+		ret = vmf_insert_pfn_prot(vma, uaddr, pfn,
+					  kvm_gmem_provider_prot(vma, attrs));
+	}
+	filemap_invalidate_unlock_shared(inode->i_mapping);
+	return ret;
+}
+
+static vm_fault_t kvm_gmem_provider_fault(struct vm_fault *vmf)
+{
+	return kvm_gmem_provider_map_host(vmf, false);
+}
+
+/* A write to a page that a read fault, or mprotect(), left read only. */
+static vm_fault_t kvm_gmem_provider_pfn_mkwrite(struct vm_fault *vmf)
+{
+	return kvm_gmem_provider_map_host(vmf, true);
+}
+
+static const struct vm_operations_struct kvm_gmem_provider_vm_ops = {
+	.fault		= kvm_gmem_provider_fault,
+	.pfn_mkwrite	= kvm_gmem_provider_pfn_mkwrite,
+};
+
+static int kvm_gmem_provider_mmap(struct file *file, struct vm_area_struct *vma)
+{
+	struct inode *inode = file_inode(file);
+	pgoff_t npages = i_size_read(inode) >> PAGE_SHIFT;
+
+	if (!kvm_gmem_supports_mmap(inode))
+		return -ENODEV;
+
+	if ((vma->vm_flags & (VM_SHARED | VM_MAYSHARE)) !=
+	    (VM_SHARED | VM_MAYSHARE))
+		return -EINVAL;
+
+	if (vma->vm_pgoff >= npages || vma_pages(vma) > npages - vma->vm_pgoff)
+		return -EINVAL;
+
+	/*
+	 * The frames may have no struct page.  They are inserted on fault, so
+	 * that a range that is revoked and comes back is reached again.
+	 */
+	vm_flags_set(vma, VM_PFNMAP | VM_IO | VM_DONTEXPAND | VM_DONTDUMP);
+	vma->vm_ops = &kvm_gmem_provider_vm_ops;
+	return 0;
+}
+
+static const struct kvm_gmem_ops kvm_gmem_provider_ops = {
+	.bind    = kvm_gmem_native_bind,
+	.unbind  = kvm_gmem_native_unbind,
+	.get_pfn = kvm_gmem_provider_get_pfn,
+	.release = kvm_gmem_native_release,
+	.mmap    = kvm_gmem_provider_mmap,
+	/* The provider decides which pages exist: no fallocate(). */
+};
+
+static int kvm_gmem_provider_attach(struct inode *inode, int provider_fd)
+{
+	struct file *file;
+	int ret;
+
+	file = fget(provider_fd);
+	if (!file)
+		return -EBADF;
+
+	ret = mem_provider_attach(&GMEM_I(inode)->att, file,
+				  i_size_read(inode), kvm_gmem_provider_revoke);
+	fput(file);
+	return ret;
+}
+
+static int __kvm_gmem_create(struct kvm *kvm, loff_t size, u64 flags,
+			     int provider_fd)
 {
 	static const char *name = "[kvm-gmem]";
 	struct gmem_file *f;
@@ -695,6 +914,12 @@ static int __kvm_gmem_create(struct kvm *kvm, loff_t size, u64 flags)
 
 	GMEM_I(inode)->flags = flags;
 
+	if (flags & GUEST_MEMFD_FLAG_USE_PROVIDER) {
+		err = kvm_gmem_provider_attach(inode, provider_fd);
+		if (err)
+			goto err_inode;
+	}
+
 	file = alloc_file_pseudo(inode, kvm_gmem_mnt, name, O_RDWR, &kvm_gmem_fops);
 	if (IS_ERR(file)) {
 		err = PTR_ERR(file);
@@ -703,12 +928,18 @@ static int __kvm_gmem_create(struct kvm *kvm, loff_t size, u64 flags)
 
 	file->f_flags |= O_LARGEFILE;
 	file->private_data = &f->backing;
-	f->backing.ops = &kvm_gmem_native_ops;
+	if (flags & GUEST_MEMFD_FLAG_USE_PROVIDER)
+		f->backing.ops = &kvm_gmem_provider_ops;
+	else
+		f->backing.ops = &kvm_gmem_native_ops;
 
 	kvm_get_kvm(kvm);
 	f->kvm = kvm;
 	xa_init(&f->bindings);
+	/* A provider can revoke, and walk the list, as soon as it is attached. */
+	filemap_invalidate_lock(inode->i_mapping);
 	list_add(&f->entry, &GMEM_I(inode)->gmem_file_list);
+	filemap_invalidate_unlock(inode->i_mapping);
 
 	fd_install(fd, file);
 	return fd;
@@ -735,7 +966,20 @@ int kvm_gmem_create(struct kvm *kvm, struct kvm_create_guest_memfd *args)
 	if (size <= 0 || !PAGE_ALIGNED(size))
 		return -EINVAL;
 
-	return __kvm_gmem_create(kvm, size, flags);
+	if (args->pad ||
+	    (!(flags & GUEST_MEMFD_FLAG_USE_PROVIDER) && args->provider_fd))
+		return -EINVAL;
+
+	/*
+	 * Without MMAP the memslot is not gmem-only, and a VM without private
+	 * memory would fault through the slot's userspace address instead of
+	 * the provider.
+	 */
+	if ((flags & GUEST_MEMFD_FLAG_USE_PROVIDER) &&
+	    !(flags & GUEST_MEMFD_FLAG_MMAP))
+		return -EINVAL;
+
+	return __kvm_gmem_create(kvm, size, flags, args->provider_fd);
 }
 
 /*
@@ -911,9 +1155,14 @@ static void kvm_gmem_native_unbind(struct file *slot_file, struct kvm *kvm,
 	 * bindings.  I.e. reaching this point means kvm_gmem_release() hasn't
 	 * yet destroyed the bindings or freed the gmem_file, and can't do so
 	 * until the caller drops slots_lock.
+	 *
+	 * A memory provider can still revoke, and walk the bindings, until
+	 * the inode is evicted, so take the invalidate lock in this case too.
 	 */
 	if (!file) {
+		filemap_invalidate_lock(slot_file->f_mapping);
 		__kvm_gmem_unbind(slot, gmem_file_of(slot_file));
+		filemap_invalidate_unlock(slot_file->f_mapping);
 		return;
 	}
 
@@ -1219,6 +1468,7 @@ static struct inode *kvm_gmem_alloc_inode(struct super_block *sb)
 	mpol_shared_policy_init(&gi->policy, NULL);
 
 	gi->flags = 0;
+	gi->att.ops = NULL;
 	INIT_LIST_HEAD(&gi->gmem_file_list);
 	return &gi->vfs_inode;
 }
@@ -1233,10 +1483,19 @@ static void kvm_gmem_free_inode(struct inode *inode)
 	kmem_cache_free(kvm_gmem_inode_cachep, GMEM_I(inode));
 }
 
+static void kvm_gmem_evict_inode(struct inode *inode)
+{
+	/* Every file is closed, so nothing maps the provider's frames. */
+	mem_provider_detach(&GMEM_I(inode)->att);
+	truncate_inode_pages_final(&inode->i_data);
+	clear_inode(inode);
+}
+
 static const struct super_operations kvm_gmem_super_operations = {
 	.statfs		= simple_statfs,
 	.alloc_inode	= kvm_gmem_alloc_inode,
 	.destroy_inode	= kvm_gmem_destroy_inode,
+	.evict_inode	= kvm_gmem_evict_inode,
 	.free_inode	= kvm_gmem_free_inode,
 };
 


^ permalink raw reply related	[flat|nested] 10+ messages in thread

* [PATCH 4/9] iommufd: Track the domains of pages that are not pinned
  2026-10-06 18:32 ` [RFC PATCH 0/9] mm: Memory providers for guest_memfd and iommufd Fred Griffoul
                     ` (2 preceding siblings ...)
  2026-10-06 18:32   ` [PATCH 3/9] KVM: guest_memfd: Add a memory provider backing Fred Griffoul
@ 2026-10-06 18:32   ` Fred Griffoul
  2026-10-06 18:32   ` [PATCH 5/9] iommufd: Map memory provider files Fred Griffoul
                     ` (4 subsequent siblings)
  8 siblings, 0 replies; 10+ messages in thread
From: Fred Griffoul @ 2026-10-06 18:32 UTC (permalink / raw)
  To: Paolo Bonzini, Sean Christopherson, Marc Zyngier, Oliver Upton,
	Andrew Morton, David Hildenbrand, Alexander Viro,
	Christian Brauner, Jan Kara, Jason Gunthorpe, Kevin Tian,
	Joerg Roedel, Will Deacon, Robin Murphy, Thomas Gleixner,
	Ingo Molnar, Borislav Petkov, Dave Hansen, x86, H . Peter Anvin,
	Jonathan Corbet, Shuah Khan
  Cc: David Woodhouse, Ackerley Tng, Lorenzo Stoakes, Liam R . Howlett,
	Vlastimil Babka, Mike Rapoport, Suren Baghdasaryan, Michal Hocko,
	Joey Gouly, Suzuki K Poulose, Zenghui Yu, Steffen Eiden,
	linux-kernel, kvm, kvmarm, iommu, linux-fsdevel, linux-mm,
	linux-kselftest

From: Fred Griffoul <fgriffo@amazon.co.uk>

Memory provider pages, added in the next patch, are not pinned and can
change under the domains that map them, as a dma-buf can. They need the
same list of (area, domain) pairs that the dma-buf revoke uses.

Move that list from struct iopt_pages_dmabuf to struct iopt_pages.

No functional change.

Signed-off-by: Fred Griffoul <fgriffo@amazon.co.uk>
---
 drivers/iommu/iommufd/io_pagetable.c | 24 +++++-----
 drivers/iommu/iommufd/io_pagetable.h | 33 +++++++++-----
 drivers/iommu/iommufd/pages.c        | 67 ++++++++++++++--------------
 3 files changed, 67 insertions(+), 57 deletions(-)

diff --git a/drivers/iommu/iommufd/io_pagetable.c b/drivers/iommu/iommufd/io_pagetable.c
index 24d4917105d9..bcd531acc9dd 100644
--- a/drivers/iommu/iommufd/io_pagetable.c
+++ b/drivers/iommu/iommufd/io_pagetable.c
@@ -1008,14 +1008,14 @@ static void iopt_unfill_domain(struct io_pagetable *iopt,
 				WARN_ON(!area->storage_domain);
 			if (area->storage_domain == domain)
 				area->storage_domain = storage_domain;
-			if (iopt_is_dmabuf(pages)) {
+			if (iopt_pages_tracked(pages)) {
 				if (!iopt_dmabuf_revoked(pages))
 					iopt_area_unmap_domain(area, domain);
-				iopt_dmabuf_untrack_domain(pages, area, domain);
+				iopt_pages_untrack_domain(pages, area, domain);
 			}
 			mutex_unlock(&pages->mutex);
 
-			if (!iopt_is_dmabuf(pages))
+			if (!iopt_pages_tracked(pages))
 				iopt_area_unmap_domain(area, domain);
 		}
 		return;
@@ -1033,8 +1033,8 @@ static void iopt_unfill_domain(struct io_pagetable *iopt,
 		WARN_ON(area->storage_domain != domain);
 		area->storage_domain = NULL;
 		iopt_area_unfill_domain(area, pages, domain);
-		if (iopt_is_dmabuf(pages))
-			iopt_dmabuf_untrack_domain(pages, area, domain);
+		if (iopt_pages_tracked(pages))
+			iopt_pages_untrack_domain(pages, area, domain);
 		mutex_unlock(&pages->mutex);
 	}
 }
@@ -1065,15 +1065,15 @@ static int iopt_fill_domain(struct io_pagetable *iopt,
 			continue;
 
 		guard(mutex)(&pages->mutex);
-		if (iopt_is_dmabuf(pages)) {
-			rc = iopt_dmabuf_track_domain(pages, area, domain);
+		if (iopt_pages_tracked(pages)) {
+			rc = iopt_pages_track_domain(pages, area, domain);
 			if (rc)
 				goto out_unfill;
 		}
 		rc = iopt_area_fill_domain(area, domain);
 		if (rc) {
-			if (iopt_is_dmabuf(pages))
-				iopt_dmabuf_untrack_domain(pages, area, domain);
+			if (iopt_pages_tracked(pages))
+				iopt_pages_untrack_domain(pages, area, domain);
 			goto out_unfill;
 		}
 		if (!area->storage_domain) {
@@ -1102,8 +1102,8 @@ static int iopt_fill_domain(struct io_pagetable *iopt,
 			area->storage_domain = NULL;
 		}
 		iopt_area_unfill_domain(area, pages, domain);
-		if (iopt_is_dmabuf(pages))
-			iopt_dmabuf_untrack_domain(pages, area, domain);
+		if (iopt_pages_tracked(pages))
+			iopt_pages_untrack_domain(pages, area, domain);
 		mutex_unlock(&pages->mutex);
 	}
 	return rc;
@@ -1315,7 +1315,7 @@ static int iopt_area_split(struct iopt_area *area, unsigned long iova)
 		return -EBUSY;
 
 	/* Maintaining the domains_itree below is a bit complicated */
-	if (iopt_is_dmabuf(pages))
+	if (iopt_pages_tracked(pages))
 		return -EOPNOTSUPP;
 
 	if (new_start & (alignment - 1) ||
diff --git a/drivers/iommu/iommufd/io_pagetable.h b/drivers/iommu/iommufd/io_pagetable.h
index 887ed94474c7..5389227eb6ff 100644
--- a/drivers/iommu/iommufd/io_pagetable.h
+++ b/drivers/iommu/iommufd/io_pagetable.h
@@ -70,15 +70,14 @@ void iopt_area_unfill_domain(struct iopt_area *area, struct iopt_pages *pages,
 void iopt_area_unmap_domain(struct iopt_area *area,
 			    struct iommu_domain *domain);
 
-int iopt_dmabuf_track_domain(struct iopt_pages *pages, struct iopt_area *area,
-			     struct iommu_domain *domain);
-void iopt_dmabuf_untrack_domain(struct iopt_pages *pages,
-				struct iopt_area *area,
-				struct iommu_domain *domain);
-int iopt_dmabuf_track_all_domains(struct iopt_area *area,
-				  struct iopt_pages *pages);
-void iopt_dmabuf_untrack_all_domains(struct iopt_area *area,
-				     struct iopt_pages *pages);
+int iopt_pages_track_domain(struct iopt_pages *pages, struct iopt_area *area,
+			    struct iommu_domain *domain);
+void iopt_pages_untrack_domain(struct iopt_pages *pages, struct iopt_area *area,
+			       struct iommu_domain *domain);
+int iopt_pages_track_all_domains(struct iopt_area *area,
+				 struct iopt_pages *pages);
+void iopt_pages_untrack_all_domains(struct iopt_area *area,
+				    struct iopt_pages *pages);
 
 static inline unsigned long iopt_area_index(struct iopt_area *area)
 {
@@ -194,7 +193,8 @@ enum iopt_address_type {
 	IOPT_ADDRESS_DMABUF,
 };
 
-struct iopt_pages_dmabuf_track {
+/* An area of the pages mapped into a domain, for pages that are not pinned. */
+struct iopt_pages_track {
 	struct iommu_domain *domain;
 	struct iopt_area *area;
 	struct list_head elm;
@@ -205,7 +205,6 @@ struct iopt_pages_dmabuf {
 	struct phys_vec phys;
 	/* Always PAGE_SIZE aligned */
 	unsigned long start;
-	struct list_head tracker;
 	/*
 	 * true if the exporter's phys is CPU RAM (map with IOMMU_CACHE, no
 	 * IOMMU_MMIO); false for MMIO/BAR memory (map with IOMMU_MMIO).  Set
@@ -253,6 +252,12 @@ struct iopt_pages {
 	struct rb_root_cached access_itree;
 	/* Of iopt_area::pages_node */
 	struct rb_root_cached domains_itree;
+	/*
+	 * Of iopt_pages_track::elm.  Pages that are not pinned can change
+	 * under the domains that map them, so every (area, domain) that maps
+	 * them is listed here.  See iopt_pages_tracked().
+	 */
+	struct list_head tracker;
 };
 
 static inline bool iopt_is_dmabuf(struct iopt_pages *pages)
@@ -262,6 +267,12 @@ static inline bool iopt_is_dmabuf(struct iopt_pages *pages)
 	return pages->type == IOPT_ADDRESS_DMABUF;
 }
 
+/* The pages are not pinned, so their domains are tracked in pages->tracker. */
+static inline bool iopt_pages_tracked(struct iopt_pages *pages)
+{
+	return iopt_is_dmabuf(pages);
+}
+
 static inline bool iopt_dmabuf_revoked(struct iopt_pages *pages)
 {
 	lockdep_assert_held(&pages->mutex);
diff --git a/drivers/iommu/iommufd/pages.c b/drivers/iommu/iommufd/pages.c
index f9b2ae6d7e96..d68f6eea836d 100644
--- a/drivers/iommu/iommufd/pages.c
+++ b/drivers/iommu/iommufd/pages.c
@@ -1392,6 +1392,7 @@ static struct iopt_pages *iopt_alloc_pages(unsigned long start_byte,
 	pages->npages = DIV_ROUND_UP(length + start_byte, PAGE_SIZE);
 	pages->access_itree = RB_ROOT_CACHED;
 	pages->domains_itree = RB_ROOT_CACHED;
+	INIT_LIST_HEAD(&pages->tracker);
 	pages->writable = writable;
 	if (capable(CAP_IPC_LOCK))
 		pages->account_mode = IOPT_PAGES_ACCOUNT_NONE;
@@ -1442,13 +1443,13 @@ struct iopt_pages *iopt_alloc_file_pages(struct file *file,
 static void iopt_revoke_notify(struct dma_buf_attachment *attach)
 {
 	struct iopt_pages *pages = attach->importer_priv;
-	struct iopt_pages_dmabuf_track *track;
+	struct iopt_pages_track *track;
 
 	guard(mutex)(&pages->mutex);
 	if (iopt_dmabuf_revoked(pages))
 		return;
 
-	list_for_each_entry(track, &pages->dmabuf.tracker, elm) {
+	list_for_each_entry(track, &pages->tracker, elm) {
 		struct iopt_area *area = track->area;
 
 		iopt_area_unmap_domain_range(area, track->domain,
@@ -1598,7 +1599,6 @@ struct iopt_pages *iopt_alloc_dmabuf_pages(struct iommufd_ctx *ictx,
 	pages->account_mode = IOPT_PAGES_ACCOUNT_NONE;
 	pages->type = IOPT_ADDRESS_DMABUF;
 	pages->dmabuf.start = start - start_byte;
-	INIT_LIST_HEAD(&pages->dmabuf.tracker);
 
 	rc = iopt_map_dmabuf(ictx, pages, dmabuf);
 	if (rc) {
@@ -1609,16 +1609,16 @@ struct iopt_pages *iopt_alloc_dmabuf_pages(struct iommufd_ctx *ictx,
 	return pages;
 }
 
-int iopt_dmabuf_track_domain(struct iopt_pages *pages, struct iopt_area *area,
-			     struct iommu_domain *domain)
+int iopt_pages_track_domain(struct iopt_pages *pages, struct iopt_area *area,
+			    struct iommu_domain *domain)
 {
-	struct iopt_pages_dmabuf_track *track;
+	struct iopt_pages_track *track;
 
 	lockdep_assert_held(&pages->mutex);
-	if (WARN_ON(!iopt_is_dmabuf(pages)))
+	if (WARN_ON(!iopt_pages_tracked(pages)))
 		return -EINVAL;
 
-	list_for_each_entry(track, &pages->dmabuf.tracker, elm)
+	list_for_each_entry(track, &pages->tracker, elm)
 		if (WARN_ON(track->domain == domain && track->area == area))
 			return -EINVAL;
 
@@ -1627,21 +1627,20 @@ int iopt_dmabuf_track_domain(struct iopt_pages *pages, struct iopt_area *area,
 		return -ENOMEM;
 	track->domain = domain;
 	track->area = area;
-	list_add_tail(&track->elm, &pages->dmabuf.tracker);
+	list_add_tail(&track->elm, &pages->tracker);
 
 	return 0;
 }
 
-void iopt_dmabuf_untrack_domain(struct iopt_pages *pages,
-				struct iopt_area *area,
-				struct iommu_domain *domain)
+void iopt_pages_untrack_domain(struct iopt_pages *pages, struct iopt_area *area,
+			       struct iommu_domain *domain)
 {
-	struct iopt_pages_dmabuf_track *track;
+	struct iopt_pages_track *track;
 
 	lockdep_assert_held(&pages->mutex);
-	WARN_ON(!iopt_is_dmabuf(pages));
+	WARN_ON(!iopt_pages_tracked(pages));
 
-	list_for_each_entry(track, &pages->dmabuf.tracker, elm) {
+	list_for_each_entry(track, &pages->tracker, elm) {
 		if (track->domain == domain && track->area == area) {
 			list_del(&track->elm);
 			kfree(track);
@@ -1651,36 +1650,36 @@ void iopt_dmabuf_untrack_domain(struct iopt_pages *pages,
 	WARN_ON(true);
 }
 
-int iopt_dmabuf_track_all_domains(struct iopt_area *area,
-				  struct iopt_pages *pages)
+int iopt_pages_track_all_domains(struct iopt_area *area,
+				 struct iopt_pages *pages)
 {
-	struct iopt_pages_dmabuf_track *track;
+	struct iopt_pages_track *track;
 	struct iommu_domain *domain;
 	unsigned long index;
 	int rc;
 
-	list_for_each_entry(track, &pages->dmabuf.tracker, elm)
+	list_for_each_entry(track, &pages->tracker, elm)
 		if (WARN_ON(track->area == area))
 			return -EINVAL;
 
 	xa_for_each(&area->iopt->domains, index, domain) {
-		rc = iopt_dmabuf_track_domain(pages, area, domain);
+		rc = iopt_pages_track_domain(pages, area, domain);
 		if (rc)
 			goto err_untrack;
 	}
 	return 0;
 err_untrack:
-	iopt_dmabuf_untrack_all_domains(area, pages);
+	iopt_pages_untrack_all_domains(area, pages);
 	return rc;
 }
 
-void iopt_dmabuf_untrack_all_domains(struct iopt_area *area,
-				     struct iopt_pages *pages)
+void iopt_pages_untrack_all_domains(struct iopt_area *area,
+				    struct iopt_pages *pages)
 {
-	struct iopt_pages_dmabuf_track *track;
-	struct iopt_pages_dmabuf_track *tmp;
+	struct iopt_pages_track *track;
+	struct iopt_pages_track *tmp;
 
-	list_for_each_entry_safe(track, tmp, &pages->dmabuf.tracker,
+	list_for_each_entry_safe(track, tmp, &pages->tracker,
 				 elm) {
 		if (track->area == area) {
 			list_del(&track->elm);
@@ -1697,6 +1696,7 @@ void iopt_release_pages(struct kref *kref)
 	WARN_ON(!RB_EMPTY_ROOT(&pages->domains_itree.rb_root));
 	WARN_ON(pages->npinned);
 	WARN_ON(!xa_empty(&pages->pinned_pfns));
+	WARN_ON(!list_empty(&pages->tracker));
 	if (iopt_is_dmabuf(pages) && pages->dmabuf.attach) {
 		struct dma_buf *dmabuf = pages->dmabuf.attach->dmabuf;
 
@@ -1705,7 +1705,6 @@ void iopt_release_pages(struct kref *kref)
 		dma_resv_unlock(dmabuf->resv);
 		dma_buf_detach(dmabuf, pages->dmabuf.attach);
 		dma_buf_put(dmabuf);
-		WARN_ON(!list_empty(&pages->dmabuf.tracker));
 	} else if (pages->type == IOPT_ADDRESS_FILE) {
 		fput(pages->file);
 	}
@@ -1790,7 +1789,7 @@ static void __iopt_area_unfill_domain(struct iopt_area *area,
 
 	lockdep_assert_held(&pages->mutex);
 
-	if (iopt_is_dmabuf(pages)) {
+	if (iopt_pages_tracked(pages)) {
 		if (WARN_ON(iopt_dmabuf_revoked(pages)))
 			return;
 		iopt_area_unmap_domain_range(area, domain, start_index,
@@ -1958,8 +1957,8 @@ int iopt_area_fill_domains(struct iopt_area *area, struct iopt_pages *pages)
 		return 0;
 
 	mutex_lock(&pages->mutex);
-	if (iopt_is_dmabuf(pages)) {
-		rc = iopt_dmabuf_track_all_domains(area, pages);
+	if (iopt_pages_tracked(pages)) {
+		rc = iopt_pages_track_all_domains(area, pages);
 		if (rc)
 			goto out_unlock;
 	}
@@ -2024,8 +2023,8 @@ int iopt_area_fill_domains(struct iopt_area *area, struct iopt_pages *pages)
 	}
 	pfn_reader_destroy(&pfns);
 out_untrack:
-	if (iopt_is_dmabuf(pages))
-		iopt_dmabuf_untrack_all_domains(area, pages);
+	if (iopt_pages_tracked(pages))
+		iopt_pages_untrack_all_domains(area, pages);
 out_unlock:
 	mutex_unlock(&pages->mutex);
 	return rc;
@@ -2065,8 +2064,8 @@ void iopt_area_unfill_domains(struct iopt_area *area, struct iopt_pages *pages)
 		WARN_ON(RB_EMPTY_NODE(&area->pages_node.rb));
 	interval_tree_remove(&area->pages_node, &pages->domains_itree);
 	iopt_area_unfill_domain(area, pages, area->storage_domain);
-	if (iopt_is_dmabuf(pages))
-		iopt_dmabuf_untrack_all_domains(area, pages);
+	if (iopt_pages_tracked(pages))
+		iopt_pages_untrack_all_domains(area, pages);
 	area->storage_domain = NULL;
 out_unlock:
 	mutex_unlock(&pages->mutex);


^ permalink raw reply related	[flat|nested] 10+ messages in thread

* [PATCH 5/9] iommufd: Map memory provider files
  2026-10-06 18:32 ` [RFC PATCH 0/9] mm: Memory providers for guest_memfd and iommufd Fred Griffoul
                     ` (3 preceding siblings ...)
  2026-10-06 18:32   ` [PATCH 4/9] iommufd: Track the domains of pages that are not pinned Fred Griffoul
@ 2026-10-06 18:32   ` Fred Griffoul
  2026-10-06 18:32   ` [PATCH 6/9] iommufd/selftest: Add mock-domain IOVA queries Fred Griffoul
                     ` (3 subsequent siblings)
  8 siblings, 0 replies; 10+ messages in thread
From: Fred Griffoul @ 2026-10-06 18:32 UTC (permalink / raw)
  To: Paolo Bonzini, Sean Christopherson, Marc Zyngier, Oliver Upton,
	Andrew Morton, David Hildenbrand, Alexander Viro,
	Christian Brauner, Jan Kara, Jason Gunthorpe, Kevin Tian,
	Joerg Roedel, Will Deacon, Robin Murphy, Thomas Gleixner,
	Ingo Molnar, Borislav Petkov, Dave Hansen, x86, H . Peter Anvin,
	Jonathan Corbet, Shuah Khan
  Cc: David Woodhouse, Ackerley Tng, Lorenzo Stoakes, Liam R . Howlett,
	Vlastimil Babka, Mike Rapoport, Suren Baghdasaryan, Michal Hocko,
	Joey Gouly, Suzuki K Poulose, Zenghui Yu, Steffen Eiden,
	linux-kernel, kvm, kvmarm, iommu, linux-fsdevel, linux-mm,
	linux-kselftest

From: Fred Griffoul <fgriffo@amazon.co.uk>

IOMMU_IOAS_MAP_FILE pins the pages it maps, which cannot work for a
memory provider file: the frames may have no struct page, and the
provider can take any of them back at any time.

Accept a provider file. iommufd neither pins nor accounts its pages.
It asks the provider for each frame when it fills a domain, and on a
revoke it unmaps the range in every domain and maps what the provider
backs now. A device that accesses the range in between faults.

Each frame gets a PAGE_SIZE entry, so a partial revoke never splits a
large IOMMU page. Holes stay unmapped, read-only pages are mapped
without IOMMU_WRITE, and MMIO pages with IOMMU_MMIO. A provider that
returns frame 0 makes the map fail with -EINVAL.

A revoke would race a dirty bitmap read and lose the dirty bits of the
range, so provider pages and a dirty tracking domain cannot share an
IOAS: whichever comes second fails with -EOPNOTSUPP. A provider
mapping cannot be split by a partial unmap, and in-kernel accesses to
it are refused.

Signed-off-by: Fred Griffoul <fgriffo@amazon.co.uk>
---
 drivers/iommu/iommufd/Kconfig        |   1 +
 drivers/iommu/iommufd/io_pagetable.c |  13 +-
 drivers/iommu/iommufd/io_pagetable.h |  23 ++-
 drivers/iommu/iommufd/pages.c        | 273 ++++++++++++++++++++++++++-
 include/uapi/linux/iommufd.h         |  11 +-
 5 files changed, 315 insertions(+), 6 deletions(-)

diff --git a/drivers/iommu/iommufd/Kconfig b/drivers/iommu/iommufd/Kconfig
index 455bac0351f2..65d71e28c12b 100644
--- a/drivers/iommu/iommufd/Kconfig
+++ b/drivers/iommu/iommufd/Kconfig
@@ -8,6 +8,7 @@ config IOMMUFD
 	select INTERVAL_TREE
 	select INTERVAL_TREE_SPAN_ITER
 	select IOMMU_API
+	select MEM_PROVIDER
 	default n
 	help
 	  Provides /dev/iommu, the user API to control the IOMMU subsystem as
diff --git a/drivers/iommu/iommufd/io_pagetable.c b/drivers/iommu/iommufd/io_pagetable.c
index bcd531acc9dd..3eca7f4f83d3 100644
--- a/drivers/iommu/iommufd/io_pagetable.c
+++ b/drivers/iommu/iommufd/io_pagetable.c
@@ -217,6 +217,7 @@ static int iopt_insert_area(struct io_pagetable *iopt, struct iopt_area *area,
 		return -EPERM;
 
 	area->iommu_prot = iommu_prot;
+	area->provider = iopt_is_provider(pages);
 	area->page_offset = start_byte % PAGE_SIZE;
 	if (area->page_offset & (iopt->iova_alignment - 1))
 		return -EINVAL;
@@ -289,6 +290,9 @@ static int iopt_alloc_area_pages(struct io_pagetable *iopt,
 		case IOPT_ADDRESS_DMABUF:
 			start = elm->start_byte + elm->pages->dmabuf.start;
 			break;
+		case IOPT_ADDRESS_PROVIDER:
+			start = elm->start_byte + elm->pages->provider.start;
+			break;
 		}
 		rc = iopt_alloc_iova(iopt, dst_iova, start, length);
 		if (rc)
@@ -515,8 +519,13 @@ int iopt_map_file_pages(struct iommufd_ctx *ictx, struct io_pagetable *iopt,
 		if (!file)
 			return -EBADF;
 
-		pages = iopt_alloc_file_pages(file, start_byte, start, length,
-					      iommu_prot & IOMMU_WRITE);
+		/* A memory provider file, or else a memfd. */
+		pages = iopt_alloc_provider_pages(file, start, length,
+						  iommu_prot & IOMMU_WRITE);
+		if (PTR_ERR(pages) == -ENODEV)
+			pages = iopt_alloc_file_pages(file, start_byte, start,
+						      length,
+						      iommu_prot & IOMMU_WRITE);
 		fput(file);
 		if (IS_ERR(pages))
 			return PTR_ERR(pages);
diff --git a/drivers/iommu/iommufd/io_pagetable.h b/drivers/iommu/iommufd/io_pagetable.h
index 5389227eb6ff..3f65380e91d2 100644
--- a/drivers/iommu/iommufd/io_pagetable.h
+++ b/drivers/iommu/iommufd/io_pagetable.h
@@ -8,6 +8,7 @@
 #include <linux/dma-buf.h>
 #include <linux/interval_tree.h>
 #include <linux/kref.h>
+#include <linux/mem_provider.h>
 #include <linux/mutex.h>
 #include <linux/xarray.h>
 
@@ -48,6 +49,8 @@ struct iopt_area {
 	/* IOMMU_READ, IOMMU_WRITE, etc */
 	int iommu_prot;
 	bool prevent_access : 1;
+	/* The pages come from a memory provider and may have holes */
+	bool provider : 1;
 	unsigned int num_accesses;
 	unsigned int num_locks;
 };
@@ -191,6 +194,7 @@ enum iopt_address_type {
 	IOPT_ADDRESS_USER = 0,
 	IOPT_ADDRESS_FILE,
 	IOPT_ADDRESS_DMABUF,
+	IOPT_ADDRESS_PROVIDER,
 };
 
 /* An area of the pages mapped into a domain, for pages that are not pinned. */
@@ -214,6 +218,12 @@ struct iopt_pages_dmabuf {
 	bool is_cpu_ram;
 };
 
+struct iopt_pages_provider {
+	struct mem_provider_attachment att;
+	/* Byte offset in the provider file, always PAGE_SIZE aligned */
+	unsigned long start;
+};
+
 /*
  * This holds a pinned page list for multiple areas of IO address space. The
  * pages always originate from a linear chunk of userspace VA. Multiple
@@ -243,6 +253,8 @@ struct iopt_pages {
 		};
 		/* IOPT_ADDRESS_DMABUF */
 		struct iopt_pages_dmabuf dmabuf;
+		/* IOPT_ADDRESS_PROVIDER */
+		struct iopt_pages_provider provider;
 	};
 	bool writable:1;
 	u8 account_mode;
@@ -267,10 +279,15 @@ static inline bool iopt_is_dmabuf(struct iopt_pages *pages)
 	return pages->type == IOPT_ADDRESS_DMABUF;
 }
 
+static inline bool iopt_is_provider(struct iopt_pages *pages)
+{
+	return pages->type == IOPT_ADDRESS_PROVIDER;
+}
+
 /* The pages are not pinned, so their domains are tracked in pages->tracker. */
 static inline bool iopt_pages_tracked(struct iopt_pages *pages)
 {
-	return iopt_is_dmabuf(pages);
+	return iopt_is_dmabuf(pages) || iopt_is_provider(pages);
 }
 
 static inline bool iopt_dmabuf_revoked(struct iopt_pages *pages)
@@ -292,6 +309,10 @@ struct iopt_pages *iopt_alloc_dmabuf_pages(struct iommufd_ctx *ictx,
 					   unsigned long start_byte,
 					   unsigned long start,
 					   unsigned long length, bool writable);
+struct iopt_pages *iopt_alloc_provider_pages(struct file *file,
+					     unsigned long start,
+					     unsigned long length,
+					     bool writable);
 void iopt_release_pages(struct kref *kref);
 static inline void iopt_put_pages(struct iopt_pages *pages)
 {
diff --git a/drivers/iommu/iommufd/pages.c b/drivers/iommu/iommufd/pages.c
index d68f6eea836d..99c95602ac4e 100644
--- a/drivers/iommu/iommufd/pages.c
+++ b/drivers/iommu/iommufd/pages.c
@@ -52,6 +52,7 @@
 #include <linux/iommu.h>
 #include <linux/iommufd.h>
 #include <linux/kthread.h>
+#include <linux/mem_provider.h>
 #include <linux/overflow.h>
 #include <linux/slab.h>
 #include <linux/sched/mm.h>
@@ -238,6 +239,35 @@ static void iommu_unmap_nofail(struct iommu_domain *domain, unsigned long iova,
 	WARN_ON(ret != size);
 }
 
+/*
+ * Provider pages are mapped with PAGE_SIZE entries, and may have holes.  Some
+ * IOMMU drivers warn when asked to unmap an IOVA that is not mapped, so find
+ * the runs that are mapped and unmap each one, without touching the holes.
+ * iopt_provider_map() refuses frame 0, which iova_to_phys() reports for
+ * an IOVA that is not mapped.
+ */
+static void iopt_provider_unmap(struct iopt_area *area,
+				struct iommu_domain *domain,
+				unsigned long start_index,
+				unsigned long last_index)
+{
+	unsigned long iova = iopt_area_index_to_iova(area, start_index);
+	size_t left = (last_index - start_index + 1) * PAGE_SIZE;
+
+	while (left) {
+		size_t len = 0;
+
+		while (len < left && iommu_iova_to_phys(domain, iova + len))
+			len += PAGE_SIZE;
+		if (len)
+			iommu_unmap_nofail(domain, iova, len);
+		/* Step over the run and the hole page that ended it. */
+		len = min(len + PAGE_SIZE, left);
+		iova += len;
+		left -= len;
+	}
+}
+
 static void iopt_area_unmap_domain_range(struct iopt_area *area,
 					 struct iommu_domain *domain,
 					 unsigned long start_index,
@@ -245,6 +275,11 @@ static void iopt_area_unmap_domain_range(struct iopt_area *area,
 {
 	unsigned long start_iova = iopt_area_index_to_iova(area, start_index);
 
+	if (area->provider) {
+		iopt_provider_unmap(area, domain, start_index, last_index);
+		return;
+	}
+
 	iommu_unmap_nofail(domain, start_iova,
 			   iopt_area_index_to_iova_last(area, last_index) -
 				   start_iova + 1);
@@ -1357,6 +1392,10 @@ static int pfn_reader_first(struct pfn_reader *pfns, struct iopt_pages *pages,
 	    WARN_ON(last_index < start_index))
 		return -EINVAL;
 
+	/* Provider pages are read from the provider, see iopt_provider_map() */
+	if (WARN_ON(iopt_is_provider(pages)))
+		return -EINVAL;
+
 	rc = pfn_reader_init(pfns, pages, start_index, last_index);
 	if (rc)
 		return rc;
@@ -1688,6 +1727,215 @@ void iopt_pages_untrack_all_domains(struct iopt_area *area,
 	}
 }
 
+static int iopt_provider_prot(struct iopt_area *area, u32 attrs)
+{
+	int prot = area->iommu_prot;
+
+	if (attrs & MEM_PROVIDER_ATTR_READONLY)
+		prot &= ~IOMMU_WRITE;
+	if (mem_provider_type(attrs) != MEM_PROVIDER_TYPE_RAM) {
+		prot &= ~IOMMU_CACHE;
+		prot |= IOMMU_MMIO;
+	}
+	return prot;
+}
+
+/*
+ * Map what the provider backs in [start_index, last_index] of the area into
+ * @domain, and leave the holes unmapped.  Each frame is mapped with a
+ * PAGE_SIZE entry, so that a revoke of part of a large block never has to
+ * split a larger IOMMU page.  On failure nothing in the range is mapped.
+ */
+static int iopt_provider_map(struct iopt_area *area, struct iopt_pages *pages,
+			     struct iommu_domain *domain,
+			     unsigned long start_index,
+			     unsigned long last_index)
+{
+	unsigned long base = pages->provider.start >> PAGE_SHIFT;
+	unsigned long index = start_index;
+	int rc;
+
+	lockdep_assert_held(&pages->mutex);
+
+	if ((1UL << __ffs(domain->pgsize_bitmap)) > PAGE_SIZE)
+		return -EOPNOTSUPP;
+
+	/*
+	 * A revoke unmaps under pages->mutex only, so it would race with a
+	 * dirty bitmap read, and lose the dirty bits of what it unmaps.
+	 * Provider pages are not mapped in a domain with dirty tracking.
+	 */
+	if (domain->dirty_ops)
+		return -EOPNOTSUPP;
+
+	while (index <= last_index) {
+		unsigned long pfn, iova, block_end, nr, i;
+		int order = PUD_ORDER;
+		u32 attrs;
+		int prot;
+
+		rc = mem_provider_get_page(&pages->provider.att, base + index,
+					   &pfn, &order, &attrs);
+		if (rc == -EFAULT) {
+			index++;
+			continue;
+		}
+		/* Frame 0 would look unmapped to iopt_provider_unmap(). */
+		if (!rc && !pfn)
+			rc = -EINVAL;
+		if (rc)
+			goto err_unmap;
+
+		block_end = ALIGN_DOWN(base + index, 1UL << order) +
+			    (1UL << order) - base;
+		nr = min(block_end, last_index + 1) - index;
+		iova = iopt_area_index_to_iova(area, index);
+		prot = iopt_provider_prot(area, attrs);
+
+		for (i = 0; i < nr; i++) {
+			rc = iommu_map_nosync(domain, iova + i * PAGE_SIZE,
+					      PFN_PHYS(pfn + i), PAGE_SIZE, prot,
+					      GFP_KERNEL_ACCOUNT);
+			if (rc)
+				break;
+		}
+		if (i) {
+			int sync_rc = iommu_sync_map(domain, iova, i * PAGE_SIZE);
+
+			if (!rc)
+				rc = sync_rc;
+		}
+		index += i;
+		if (rc)
+			goto err_unmap;
+	}
+	return 0;
+
+err_unmap:
+	if (index > start_index)
+		iopt_provider_unmap(area, domain, start_index, index - 1);
+	return rc;
+}
+
+/* Map the whole area into every domain of its io_pagetable. */
+static int iopt_provider_fill_domains(struct iopt_area *area,
+				      struct iopt_pages *pages)
+{
+	struct iommu_domain *domain, *undo;
+	unsigned long index, undo_index;
+	int rc;
+
+	xa_for_each(&area->iopt->domains, index, domain) {
+		rc = iopt_provider_map(area, pages, domain,
+				       iopt_area_index(area),
+				       iopt_area_last_index(area));
+		if (rc)
+			goto err_unmap;
+	}
+	return 0;
+
+err_unmap:
+	xa_for_each(&area->iopt->domains, undo_index, undo) {
+		if (undo_index >= index)
+			break;
+		iopt_provider_unmap(area, undo, iopt_area_index(area),
+				    iopt_area_last_index(area));
+	}
+	return rc;
+}
+
+/*
+ * The provider changed the frames behind [offset, offset + len).  In every
+ * domain, unmap the range and map what the provider backs now.  A device that
+ * accesses the range in between faults.
+ */
+static void iopt_provider_revoke(struct mem_provider_attachment *att,
+				 loff_t offset, loff_t len)
+{
+	struct iopt_pages *pages =
+		container_of(att, struct iopt_pages, provider.att);
+	u64 start = pages->provider.start;
+	u64 end = start + (u64)pages->npages * PAGE_SIZE;
+	struct iopt_pages_track *track;
+	unsigned long first, last;
+
+	if (offset < 0 || len <= 0 || offset >= end || offset + len <= start)
+		return;
+	first = (max_t(u64, offset, start) - start) >> PAGE_SHIFT;
+	last = (min_t(u64, offset + len, end) - start - 1) >> PAGE_SHIFT;
+
+	guard(mutex)(&pages->mutex);
+	list_for_each_entry(track, &pages->tracker, elm) {
+		struct iopt_area *area = track->area;
+		unsigned long s = max(first, iopt_area_index(area));
+		unsigned long l = min(last, iopt_area_last_index(area));
+
+		if (s > l)
+			continue;
+		iopt_provider_unmap(area, track->domain, s, l);
+		if (iopt_provider_map(area, pages, track->domain, s, l))
+			pr_warn_ratelimited("iommufd: cannot map provider pages after a revoke\n");
+	}
+}
+
+/**
+ * iopt_alloc_provider_pages() - Pages backed by a memory provider file
+ * @file: The provider file
+ * @start: Byte offset in the file, PAGE_SIZE aligned
+ * @length: Number of bytes
+ * @writable: The pages may be mapped writable
+ *
+ * The pages are not pinned.  The provider can change the frames behind them
+ * at any time, and every domain that maps them follows.
+ *
+ * Return: the pages, ERR_PTR(-ENODEV) if @file is not a provider file,
+ * ERR_PTR(-EINVAL) if @start is not page aligned, or another ERR_PTR().
+ */
+struct iopt_pages *iopt_alloc_provider_pages(struct file *file,
+					     unsigned long start,
+					     unsigned long length,
+					     bool writable)
+{
+	static struct lock_class_key pages_provider_mutex_key;
+	struct iopt_pages *pages;
+	int rc;
+
+	if (length / PAGE_SIZE >= MAX_NPFNS)
+		return ERR_PTR(-EINVAL);
+
+	pages = iopt_alloc_pages(0, length, writable);
+	if (IS_ERR(pages))
+		return pages;
+
+	/*
+	 * The pages mutex of provider pages is never held while taking the
+	 * mmap_lock, but is taken from the provider's revoke.  Split the lock
+	 * class from the pinned pages.
+	 */
+	lockdep_set_class(&pages->mutex, &pages_provider_mutex_key);
+
+	/* Provider pages are not pinned, so they are not accounted. */
+	pages->account_mode = IOPT_PAGES_ACCOUNT_NONE;
+	pages->type = IOPT_ADDRESS_PROVIDER;
+	pages->provider.start = start;
+
+	/*
+	 * The provider must cover the whole mapping, so the size is its end.
+	 * Check the alignment only once the file is known to be a provider
+	 * file: on -ENODEV the caller maps it as a memfd, which may start
+	 * anywhere.
+	 */
+	rc = mem_provider_attach(&pages->provider.att, file, start + length,
+				 iopt_provider_revoke);
+	if (!rc && !PAGE_ALIGNED(start))
+		rc = -EINVAL;
+	if (rc) {
+		iopt_put_pages(pages);
+		return ERR_PTR(rc);
+	}
+	return pages;
+}
+
 void iopt_release_pages(struct kref *kref)
 {
 	struct iopt_pages *pages = container_of(kref, struct iopt_pages, kref);
@@ -1705,6 +1953,8 @@ void iopt_release_pages(struct kref *kref)
 		dma_resv_unlock(dmabuf->resv);
 		dma_buf_detach(dmabuf, pages->dmabuf.attach);
 		dma_buf_put(dmabuf);
+	} else if (iopt_is_provider(pages)) {
+		mem_provider_detach(&pages->provider.att);
 	} else if (pages->type == IOPT_ADDRESS_FILE) {
 		fput(pages->file);
 	}
@@ -1855,6 +2105,12 @@ static void iopt_area_unfill_partial_domain(struct iopt_area *area,
  */
 void iopt_area_unmap_domain(struct iopt_area *area, struct iommu_domain *domain)
 {
+	if (area->provider) {
+		iopt_provider_unmap(area, domain, iopt_area_index(area),
+				    iopt_area_last_index(area));
+		return;
+	}
+
 	iommu_unmap_nofail(domain, iopt_area_iova(area),
 			   iopt_area_length(area));
 }
@@ -1898,6 +2154,11 @@ int iopt_area_fill_domain(struct iopt_area *area, struct iommu_domain *domain)
 	if (iopt_dmabuf_revoked(area->pages))
 		return 0;
 
+	if (iopt_is_provider(area->pages))
+		return iopt_provider_map(area, area->pages, domain,
+					 iopt_area_index(area),
+					 iopt_area_last_index(area));
+
 	rc = pfn_reader_first(&pfns, area->pages, iopt_area_index(area),
 			      iopt_area_last_index(area));
 	if (rc)
@@ -1963,7 +2224,11 @@ int iopt_area_fill_domains(struct iopt_area *area, struct iopt_pages *pages)
 			goto out_unlock;
 	}
 
-	if (!iopt_dmabuf_revoked(pages)) {
+	if (iopt_is_provider(pages)) {
+		rc = iopt_provider_fill_domains(area, pages);
+		if (rc)
+			goto out_untrack;
+	} else if (!iopt_dmabuf_revoked(pages)) {
 		rc = pfn_reader_first(&pfns, pages, iopt_area_index(area),
 				      iopt_area_last_index(area));
 		if (rc)
@@ -2402,7 +2667,7 @@ int iopt_pages_rw_access(struct iopt_pages *pages, unsigned long start_byte,
 	if ((flags & IOMMUFD_ACCESS_RW_WRITE) && !pages->writable)
 		return -EPERM;
 
-	if (iopt_is_dmabuf(pages))
+	if (iopt_is_dmabuf(pages) || iopt_is_provider(pages))
 		return -EINVAL;
 
 	if (pages->type != IOPT_ADDRESS_USER)
@@ -2491,6 +2756,10 @@ int iopt_area_add_access(struct iopt_area *area, unsigned long start_index,
 	if ((flags & IOMMUFD_ACCESS_RW_WRITE) && !pages->writable)
 		return -EPERM;
 
+	/* Provider frames may have no struct page, and are not pinned. */
+	if (iopt_is_provider(pages))
+		return -EOPNOTSUPP;
+
 	mutex_lock(&pages->mutex);
 	access = iopt_pages_get_exact_access(pages, start_index, last_index);
 	if (access) {
diff --git a/include/uapi/linux/iommufd.h b/include/uapi/linux/iommufd.h
index 0425d452d41e..b4500fc21368 100644
--- a/include/uapi/linux/iommufd.h
+++ b/include/uapi/linux/iommufd.h
@@ -224,7 +224,7 @@ struct iommu_ioas_map {
  * @size: sizeof(struct iommu_ioas_map_file)
  * @flags: same as for iommu_ioas_map
  * @ioas_id: same as for iommu_ioas_map
- * @fd: the memfd or supported dma-buf file to map
+ * @fd: the memfd, memory provider file or supported dma-buf file to map
  * @start: byte offset from start of the file to map from
  * @length: same as for iommu_ioas_map
  * @iova: same as for iommu_ioas_map
@@ -235,6 +235,15 @@ struct iommu_ioas_map {
  * VFIO PCI dma-bufs exported through VFIO_DEVICE_FEATURE_DMA_BUF, and
  * other dma-bufs may be rejected. All other arguments and semantics match
  * those of IOMMU_IOAS_MAP.
+ *
+ * A file from a memory provider is also accepted; @start must then be
+ * page aligned. Its pages are not pinned. The provider may change the memory
+ * behind any range at any time, and the mapping follows: pages that the
+ * provider does not back are left unmapped, read-only pages are mapped
+ * without write permission, and device memory is mapped as MMIO. A device
+ * access to a page that is being changed, or that is not backed, faults. Such
+ * a mapping cannot be split by a partial unmap, and in-kernel accesses to it
+ * are refused. It is not supported in an IOAS with a dirty tracking domain.
  */
 struct iommu_ioas_map_file {
 	__u32 size;


^ permalink raw reply related	[flat|nested] 10+ messages in thread

* [PATCH 6/9] iommufd/selftest: Add mock-domain IOVA queries
  2026-10-06 18:32 ` [RFC PATCH 0/9] mm: Memory providers for guest_memfd and iommufd Fred Griffoul
                     ` (4 preceding siblings ...)
  2026-10-06 18:32   ` [PATCH 5/9] iommufd: Map memory provider files Fred Griffoul
@ 2026-10-06 18:32   ` Fred Griffoul
  2026-10-06 18:32   ` [PATCH 7/9] iommufd/selftest: Add a mock memory provider Fred Griffoul
                     ` (2 subsequent siblings)
  8 siblings, 0 replies; 10+ messages in thread
From: Fred Griffoul @ 2026-10-06 18:32 UTC (permalink / raw)
  To: Paolo Bonzini, Sean Christopherson, Marc Zyngier, Oliver Upton,
	Andrew Morton, David Hildenbrand, Alexander Viro,
	Christian Brauner, Jan Kara, Jason Gunthorpe, Kevin Tian,
	Joerg Roedel, Will Deacon, Robin Murphy, Thomas Gleixner,
	Ingo Molnar, Borislav Petkov, Dave Hansen, x86, H . Peter Anvin,
	Jonathan Corbet, Shuah Khan
  Cc: David Woodhouse, Ackerley Tng, Lorenzo Stoakes, Liam R . Howlett,
	Vlastimil Babka, Mike Rapoport, Suren Baghdasaryan, Michal Hocko,
	Joey Gouly, Suzuki K Poulose, Zenghui Yu, Steffen Eiden,
	linux-kernel, kvm, kvmarm, iommu, linux-fsdevel, linux-mm,
	linux-kselftest

From: Fred Griffoul <fgriffo@amazon.co.uk>

A test of memory that can be taken back or replaced does not know in
advance which frame a device reaches. IOMMU_TEST_OP_MD_CHECK_MAP only
compares against a known buffer.

Add MD_CHECK_MAPPED, which checks that a range is fully mapped or fully
unmapped, and MD_IOVA_TO_PHYS, which returns the frame behind an IOVA,
or 0 if it is unmapped. Both refuse an IOVA outside the domain's
aperture.

Both hold domains_rwsem for writing so that a userspace unmap cannot
free a page-table level during the walk.

Signed-off-by: Fred Griffoul <fgriffo@amazon.co.uk>
---
 drivers/iommu/iommufd/iommufd_test.h |  16 ++++
 drivers/iommu/iommufd/selftest.c     | 109 +++++++++++++++++++++++++++
 2 files changed, 125 insertions(+)

diff --git a/drivers/iommu/iommufd/iommufd_test.h b/drivers/iommu/iommufd/iommufd_test.h
index 52b78cbcc920..28fd9c43edc4 100644
--- a/drivers/iommu/iommufd/iommufd_test.h
+++ b/drivers/iommu/iommufd/iommufd_test.h
@@ -31,6 +31,8 @@ enum {
 	IOMMU_TEST_OP_PASID_CHECK_HWPT,
 	IOMMU_TEST_OP_DMABUF_GET,
 	IOMMU_TEST_OP_DMABUF_REVOKE,
+	IOMMU_TEST_OP_MD_CHECK_MAPPED,
+	IOMMU_TEST_OP_MD_IOVA_TO_PHYS,
 };
 
 enum {
@@ -193,6 +195,20 @@ struct iommu_test_cmd {
 			__s32 dmabuf_fd;
 			__u32 revoked;
 		} dmabuf_revoke;
+		struct {
+			/*
+			 * 1: every page in [iova, iova+length) must be mapped;
+			 * 0: none of them may be. Mixed is an error.
+			 */
+			__u32 mapped;
+			__u32 __reserved;
+			__aligned_u64 iova;
+			__aligned_u64 length;
+		} check_mapped;
+		struct {
+			__aligned_u64 iova;
+			__aligned_u64 out_phys;	/* 0 if unmapped */
+		} iova_to_phys;
 	};
 	__u32 last;
 };
diff --git a/drivers/iommu/iommufd/selftest.c b/drivers/iommu/iommufd/selftest.c
index af07c642a526..1f5cd2d00fda 100644
--- a/drivers/iommu/iommufd/selftest.c
+++ b/drivers/iommu/iommufd/selftest.c
@@ -2031,6 +2031,107 @@ static int iommufd_test_dmabuf_get(struct iommufd_ucmd *ucmd,
 	return rc;
 }
 
+/*
+ * True if [iova, iova + length) lies inside the domain's aperture.  Outside
+ * it the page table returns an error code from iova_to_phys(), not 0.
+ */
+static bool mock_domain_covers(struct mock_iommu_domain *mock,
+			       unsigned long iova, size_t length)
+{
+	struct iommu_domain_geometry *geo = &mock->domain.geometry;
+	unsigned long last;
+
+	if (check_add_overflow(iova, length - 1, &last))
+		return false;
+	return iova >= geo->aperture_start && last <= geo->aperture_end;
+}
+
+/*
+ * iova_to_phys() walks the page table, so it must not run while an unmap
+ * frees a level of it.  A userspace unmap holds the IOAS domains_rwsem for
+ * reading while it unmaps the domains, so holding it for writing keeps such
+ * unmaps away.  A memory provider or dma-buf revoke unmaps under its pages
+ * mutex only, so tests must not run these queries while a revoke of the
+ * range is in progress.
+ */
+static struct rw_semaphore *
+mock_domain_unmap_lock(struct iommufd_hw_pagetable *hwpt)
+{
+	return &to_hwpt_paging(hwpt)->ioas->iopt.domains_rwsem;
+}
+
+/*
+ * Report the physical address the mock domain resolves @iova to, or 0 if
+ * it is unmapped.  Lets a test check that two IOVAs share one frame, or that
+ * an IOVA moved to another frame, without knowing the frames in advance.
+ */
+static int iommufd_test_md_iova_to_phys(struct iommufd_ucmd *ucmd,
+					unsigned int mockpt_id,
+					unsigned long iova)
+{
+	struct iommu_test_cmd *cmd = ucmd->cmd;
+	struct iommufd_hw_pagetable *hwpt;
+	struct mock_iommu_domain *mock;
+	unsigned int page_size;
+	int rc;
+
+	hwpt = get_md_pagetable(ucmd, mockpt_id, &mock);
+	if (IS_ERR(hwpt))
+		return PTR_ERR(hwpt);
+
+	page_size = 1 << __ffs(mock->domain.pgsize_bitmap);
+	if (iova % page_size || !mock_domain_covers(mock, iova, page_size)) {
+		rc = -EINVAL;
+		goto out_put;
+	}
+	down_write(mock_domain_unmap_lock(hwpt));
+	cmd->iova_to_phys.out_phys =
+		mock->domain.ops->iova_to_phys(&mock->domain, iova);
+	up_write(mock_domain_unmap_lock(hwpt));
+	rc = iommufd_ucmd_respond(ucmd, sizeof(*cmd));
+out_put:
+	iommufd_put_object(ucmd->ictx, &hwpt->obj);
+	return rc;
+}
+
+static int iommufd_test_md_check_mapped(struct iommufd_ucmd *ucmd,
+					unsigned int mockpt_id,
+					unsigned long iova, size_t length,
+					bool mapped)
+{
+	struct iommufd_hw_pagetable *hwpt;
+	struct mock_iommu_domain *mock;
+	unsigned int page_size;
+	int rc = 0;
+
+	hwpt = get_md_pagetable(ucmd, mockpt_id, &mock);
+	if (IS_ERR(hwpt))
+		return PTR_ERR(hwpt);
+
+	page_size = 1 << __ffs(mock->domain.pgsize_bitmap);
+	if (iova % page_size || length % page_size || !length ||
+	    !mock_domain_covers(mock, iova, length)) {
+		rc = -EINVAL;
+		goto out_put;
+	}
+
+	down_write(mock_domain_unmap_lock(hwpt));
+	for (; length; length -= page_size, iova += page_size) {
+		bool is_mapped =
+			mock->domain.ops->iova_to_phys(&mock->domain, iova) != 0;
+
+		if (is_mapped != mapped) {
+			rc = -ENOENT;
+			break;
+		}
+	}
+	up_write(mock_domain_unmap_lock(hwpt));
+
+out_put:
+	iommufd_put_object(ucmd->ictx, &hwpt->obj);
+	return rc;
+}
+
 static int iommufd_test_dmabuf_revoke(struct iommufd_ucmd *ucmd, int fd,
 				      bool revoked)
 {
@@ -2143,6 +2244,14 @@ int iommufd_test(struct iommufd_ucmd *ucmd)
 		return iommufd_test_dmabuf_revoke(ucmd,
 						  cmd->dmabuf_revoke.dmabuf_fd,
 						  cmd->dmabuf_revoke.revoked);
+	case IOMMU_TEST_OP_MD_CHECK_MAPPED:
+		return iommufd_test_md_check_mapped(ucmd, cmd->id,
+						    cmd->check_mapped.iova,
+						    cmd->check_mapped.length,
+						    cmd->check_mapped.mapped);
+	case IOMMU_TEST_OP_MD_IOVA_TO_PHYS:
+		return iommufd_test_md_iova_to_phys(ucmd, cmd->id,
+						    cmd->iova_to_phys.iova);
 	default:
 		return -EOPNOTSUPP;
 	}


^ permalink raw reply related	[flat|nested] 10+ messages in thread

* [PATCH 7/9] iommufd/selftest: Add a mock memory provider
  2026-10-06 18:32 ` [RFC PATCH 0/9] mm: Memory providers for guest_memfd and iommufd Fred Griffoul
                     ` (5 preceding siblings ...)
  2026-10-06 18:32   ` [PATCH 6/9] iommufd/selftest: Add mock-domain IOVA queries Fred Griffoul
@ 2026-10-06 18:32   ` Fred Griffoul
  2026-10-06 18:32   ` [PATCH 8/9] samples/kvm: Add a memory provider sample Fred Griffoul
  2026-10-06 18:32   ` [PATCH 9/9] KVM: selftests: Test a memory provider shared by KVM and iommufd Fred Griffoul
  8 siblings, 0 replies; 10+ messages in thread
From: Fred Griffoul @ 2026-10-06 18:32 UTC (permalink / raw)
  To: Paolo Bonzini, Sean Christopherson, Marc Zyngier, Oliver Upton,
	Andrew Morton, David Hildenbrand, Alexander Viro,
	Christian Brauner, Jan Kara, Jason Gunthorpe, Kevin Tian,
	Joerg Roedel, Will Deacon, Robin Murphy, Thomas Gleixner,
	Ingo Molnar, Borislav Petkov, Dave Hansen, x86, H . Peter Anvin,
	Jonathan Corbet, Shuah Khan
  Cc: David Woodhouse, Ackerley Tng, Lorenzo Stoakes, Liam R . Howlett,
	Vlastimil Babka, Mike Rapoport, Suren Baghdasaryan, Michal Hocko,
	Joey Gouly, Suzuki K Poulose, Zenghui Yu, Steffen Eiden,
	linux-kernel, kvm, kvmarm, iommu, linux-fsdevel, linux-mm,
	linux-kselftest

From: Fred Griffoul <fgriffo@amazon.co.uk>

The provider path of IOMMU_IOAS_MAP_FILE needs a provider to test
against.

Add a mock provider whose pages can be made holes, read only, MMIO or
another shared frame, with a revoke after each change. Test that holes
stay unmapped, that a revoke changes only its range, that new domains
see the current state, and that provider pages and dirty tracking
domains cannot share an IOAS.

Signed-off-by: Fred Griffoul <fgriffo@amazon.co.uk>
---
 drivers/iommu/iommufd/iommufd_test.h          |  20 ++
 drivers/iommu/iommufd/selftest.c              | 177 ++++++++++++++++++
 tools/testing/selftests/iommu/iommufd.c       | 119 ++++++++++++
 tools/testing/selftests/iommu/iommufd_utils.h |  64 +++++++
 4 files changed, 380 insertions(+)

diff --git a/drivers/iommu/iommufd/iommufd_test.h b/drivers/iommu/iommufd/iommufd_test.h
index 28fd9c43edc4..5e5023dc8070 100644
--- a/drivers/iommu/iommufd/iommufd_test.h
+++ b/drivers/iommu/iommufd/iommufd_test.h
@@ -33,6 +33,17 @@ enum {
 	IOMMU_TEST_OP_DMABUF_REVOKE,
 	IOMMU_TEST_OP_MD_CHECK_MAPPED,
 	IOMMU_TEST_OP_MD_IOVA_TO_PHYS,
+	IOMMU_TEST_OP_PROVIDER_GET,
+	IOMMU_TEST_OP_PROVIDER_SET,
+};
+
+/* Page states of the mock memory provider, IOMMU_TEST_OP_PROVIDER_SET */
+enum {
+	MOCK_PROVIDER_BACKED = 0,
+	MOCK_PROVIDER_HOLE,
+	MOCK_PROVIDER_READONLY,
+	MOCK_PROVIDER_OTHER,
+	MOCK_PROVIDER_MMIO,
 };
 
 enum {
@@ -209,6 +220,15 @@ struct iommu_test_cmd {
 			__aligned_u64 iova;
 			__aligned_u64 out_phys;	/* 0 if unmapped */
 		} iova_to_phys;
+		struct {
+			__u32 length;
+		} provider_get;
+		struct {
+			__s32 fd;
+			__u32 state;
+			__aligned_u64 offset;
+			__aligned_u64 length;
+		} provider_set;
 	};
 	__u32 last;
 };
diff --git a/drivers/iommu/iommufd/selftest.c b/drivers/iommu/iommufd/selftest.c
index 1f5cd2d00fda..270c41e6dd77 100644
--- a/drivers/iommu/iommufd/selftest.c
+++ b/drivers/iommu/iommufd/selftest.c
@@ -10,6 +10,7 @@
 #include <linux/fault-inject.h>
 #include <linux/file.h>
 #include <linux/iommu.h>
+#include <linux/mem_provider.h>
 #include <linux/platform_device.h>
 #include <linux/slab.h>
 #include <linux/xarray.h>
@@ -2132,6 +2133,174 @@ static int iommufd_test_md_check_mapped(struct iommufd_ucmd *ucmd,
 	return rc;
 }
 
+/*
+ * A mock memory provider.  Each page of the file is backed by its own page of
+ * RAM, or is a hole, read only, the one page that every page in the OTHER
+ * state shares, or reported as MMIO.
+ * IOMMU_TEST_OP_PROVIDER_SET changes the state of a range and revokes it.
+ */
+struct iommufd_test_provider {
+	struct mem_provider_file mpf;
+	size_t npages;
+	struct page **pages;
+	struct page *other;
+	u8 *state;
+};
+
+static void iommufd_test_provider_free(struct iommufd_test_provider *tp)
+{
+	size_t i;
+
+	for (i = 0; i != tp->npages; i++)
+		if (tp->pages[i])
+			__free_page(tp->pages[i]);
+	if (tp->other)
+		__free_page(tp->other);
+	kfree(tp->state);
+	kfree(tp->pages);
+	kfree(tp);
+}
+
+static int iommufd_test_provider_release(struct inode *inode,
+					 struct file *file)
+{
+	iommufd_test_provider_free(file->private_data);
+	return 0;
+}
+
+static struct mem_provider_file *
+iommufd_test_provider_attach(struct file *file, loff_t size)
+{
+	struct iommufd_test_provider *tp = file->private_data;
+
+	if (size > (loff_t)tp->npages * PAGE_SIZE)
+		return ERR_PTR(-EINVAL);
+	return &tp->mpf;
+}
+
+static void iommufd_test_provider_detach(struct mem_provider_file *mpf)
+{
+}
+
+static int iommufd_test_provider_get_page(struct mem_provider_file *mpf,
+					  pgoff_t index, unsigned long *pfn,
+					  int *max_order, u32 *attrs)
+{
+	struct iommufd_test_provider *tp =
+		container_of(mpf, struct iommufd_test_provider, mpf);
+
+	*max_order = 0;
+	switch (READ_ONCE(tp->state[index])) {
+	case MOCK_PROVIDER_HOLE:
+		return -EFAULT;
+	case MOCK_PROVIDER_READONLY:
+		*attrs = MEM_PROVIDER_ATTR_READONLY;
+		break;
+	case MOCK_PROVIDER_OTHER:
+		*pfn = page_to_pfn(tp->other);
+		return 0;
+	case MOCK_PROVIDER_MMIO:
+		*attrs = MEM_PROVIDER_TYPE_MMIO;
+		break;
+	}
+	*pfn = page_to_pfn(tp->pages[index]);
+	return 0;
+}
+
+static const struct mem_provider_ops iommufd_test_provider_ops = {
+	.attach = iommufd_test_provider_attach,
+	.detach = iommufd_test_provider_detach,
+	.get_page = iommufd_test_provider_get_page,
+};
+
+static const struct mem_provider_fops iommufd_test_provider_fops = {
+	.fops = {
+		.owner = THIS_MODULE,
+		.fop_flags = FOP_MEM_PROVIDER,
+		.release = iommufd_test_provider_release,
+	},
+	.ops = &iommufd_test_provider_ops,
+};
+
+static int iommufd_test_provider_get(struct iommufd_ucmd *ucmd, size_t len)
+{
+	struct iommufd_test_provider *tp;
+	struct file *file;
+	size_t i;
+	int fd, rc;
+
+	len = ALIGN(len, PAGE_SIZE);
+	if (len == 0 || len > PAGE_SIZE * 512)
+		return -EINVAL;
+
+	tp = kzalloc_obj(*tp);
+	if (!tp)
+		return -ENOMEM;
+	mem_provider_file_init(&tp->mpf);
+	tp->npages = len / PAGE_SIZE;
+	tp->pages = kcalloc(tp->npages, sizeof(*tp->pages), GFP_KERNEL);
+	tp->state = kcalloc(tp->npages, sizeof(*tp->state), GFP_KERNEL);
+	tp->other = alloc_page(GFP_KERNEL | __GFP_ZERO);
+	if (!tp->pages || !tp->state || !tp->other) {
+		rc = -ENOMEM;
+		goto err_free;
+	}
+	for (i = 0; i != tp->npages; i++) {
+		tp->pages[i] = alloc_page(GFP_KERNEL | __GFP_ZERO);
+		if (!tp->pages[i]) {
+			rc = -ENOMEM;
+			goto err_free;
+		}
+	}
+
+	fd = get_unused_fd_flags(O_CLOEXEC);
+	if (fd < 0) {
+		rc = fd;
+		goto err_free;
+	}
+	file = anon_inode_getfile("[iommufd-test-provider]",
+				  &iommufd_test_provider_fops.fops, tp, O_RDWR);
+	if (IS_ERR(file)) {
+		put_unused_fd(fd);
+		rc = PTR_ERR(file);
+		goto err_free;
+	}
+	fd_install(fd, file);
+	return fd;
+
+err_free:
+	iommufd_test_provider_free(tp);
+	return rc;
+}
+
+static int iommufd_test_provider_set(struct iommufd_ucmd *ucmd, int fd,
+				     unsigned int state, u64 offset, u64 length)
+{
+	struct iommufd_test_provider *tp;
+	u64 end;
+	size_t i;
+
+	if (state > MOCK_PROVIDER_MMIO)
+		return -EINVAL;
+
+	CLASS(fd, f)(fd);
+	if (fd_empty(f))
+		return -EBADF;
+	if (fd_file(f)->f_op != &iommufd_test_provider_fops.fops)
+		return -EINVAL;
+	tp = fd_file(f)->private_data;
+
+	if (!length || check_add_overflow(offset, length, &end) ||
+	    !PAGE_ALIGNED(offset) || !PAGE_ALIGNED(length) ||
+	    end > (u64)tp->npages * PAGE_SIZE)
+		return -EINVAL;
+
+	for (i = offset / PAGE_SIZE; i != end / PAGE_SIZE; i++)
+		WRITE_ONCE(tp->state[i], state);
+	mem_provider_revoke(&tp->mpf, offset, length);
+	return 0;
+}
+
 static int iommufd_test_dmabuf_revoke(struct iommufd_ucmd *ucmd, int fd,
 				      bool revoked)
 {
@@ -2252,6 +2421,14 @@ int iommufd_test(struct iommufd_ucmd *ucmd)
 	case IOMMU_TEST_OP_MD_IOVA_TO_PHYS:
 		return iommufd_test_md_iova_to_phys(ucmd, cmd->id,
 						    cmd->iova_to_phys.iova);
+	case IOMMU_TEST_OP_PROVIDER_GET:
+		return iommufd_test_provider_get(ucmd,
+						 cmd->provider_get.length);
+	case IOMMU_TEST_OP_PROVIDER_SET:
+		return iommufd_test_provider_set(ucmd, cmd->provider_set.fd,
+						 cmd->provider_set.state,
+						 cmd->provider_set.offset,
+						 cmd->provider_set.length);
 	default:
 		return -EOPNOTSUPP;
 	}
diff --git a/tools/testing/selftests/iommu/iommufd.c b/tools/testing/selftests/iommu/iommufd.c
index d44b34b05757..ebf81ca3af18 100644
--- a/tools/testing/selftests/iommu/iommufd.c
+++ b/tools/testing/selftests/iommu/iommufd.c
@@ -1627,6 +1627,125 @@ TEST_F(iommufd_ioas, dmabuf_revoke)
 	close(dfd);
 }
 
+TEST_F(iommufd_ioas, provider_simple)
+{
+	size_t buf_size = PAGE_SIZE * 4;
+	__u64 iova;
+	int pfd;
+
+	test_cmd_get_provider(buf_size, &pfd);
+	test_err_ioctl_ioas_map_file(EINVAL, pfd, 0, 0, &iova);
+	test_err_ioctl_ioas_map_file(EINVAL, pfd, 0, buf_size + PAGE_SIZE,
+				     &iova);
+	test_err_ioctl_ioas_map_file(EINVAL, pfd, 1, PAGE_SIZE, &iova);
+	test_ioctl_ioas_map_file(pfd, 0, buf_size, &iova);
+	if (variant->mock_domains)
+		test_cmd_md_check_mapped(self->hwpt_id, iova, buf_size, true);
+
+	/* The mapping keeps the provider file alive */
+	close(pfd);
+	if (variant->mock_domains)
+		test_cmd_md_check_mapped(self->hwpt_id, iova, buf_size, true);
+	test_ioctl_ioas_unmap(iova, buf_size);
+}
+
+TEST_F(iommufd_ioas, provider_revoke)
+{
+	size_t buf_size = PAGE_SIZE * 8;
+	__u64 phys0, phys1;
+	__u32 hwpt_id;
+	__u64 iova;
+	int pfd;
+
+	if (!variant->mock_domains)
+		SKIP(return, "Needs a mock domain");
+
+	test_cmd_get_provider(buf_size, &pfd);
+
+	/* A page that is a hole when the file is mapped stays unmapped */
+	test_cmd_provider_set(pfd, MOCK_PROVIDER_HOLE, PAGE_SIZE, PAGE_SIZE);
+	test_ioctl_ioas_map_file(pfd, 0, buf_size, &iova);
+	test_cmd_md_check_mapped(self->hwpt_id, iova, PAGE_SIZE, true);
+	test_cmd_md_check_mapped(self->hwpt_id, iova + PAGE_SIZE, PAGE_SIZE,
+				 false);
+	test_cmd_md_check_mapped(self->hwpt_id, iova + 2 * PAGE_SIZE,
+				 buf_size - 2 * PAGE_SIZE, true);
+
+	/* Give the page back */
+	test_cmd_provider_set(pfd, MOCK_PROVIDER_BACKED, PAGE_SIZE, PAGE_SIZE);
+	test_cmd_md_check_mapped(self->hwpt_id, iova, buf_size, true);
+
+	/* Take back the middle, and only the middle goes */
+	test_cmd_provider_set(pfd, MOCK_PROVIDER_HOLE, 2 * PAGE_SIZE,
+			      4 * PAGE_SIZE);
+	test_cmd_md_check_mapped(self->hwpt_id, iova, 2 * PAGE_SIZE, true);
+	test_cmd_md_check_mapped(self->hwpt_id, iova + 2 * PAGE_SIZE,
+				 4 * PAGE_SIZE, false);
+	test_cmd_md_check_mapped(self->hwpt_id, iova + 6 * PAGE_SIZE,
+				 2 * PAGE_SIZE, true);
+
+	/* A domain added now sees the same holes */
+	test_cmd_hwpt_alloc(self->device_id, self->ioas_id, 0, &hwpt_id);
+	test_cmd_md_check_mapped(hwpt_id, iova, 2 * PAGE_SIZE, true);
+	test_cmd_md_check_mapped(hwpt_id, iova + 2 * PAGE_SIZE, 4 * PAGE_SIZE,
+				 false);
+	test_cmd_md_check_mapped(hwpt_id, iova + 6 * PAGE_SIZE, 2 * PAGE_SIZE,
+				 true);
+	test_ioctl_destroy(hwpt_id);
+
+	/* Put the same other frame behind two pages */
+	phys0 = test_cmd_md_iova_to_phys(self->hwpt_id, iova);
+	test_cmd_provider_set(pfd, MOCK_PROVIDER_OTHER, 0, 2 * PAGE_SIZE);
+	phys1 = test_cmd_md_iova_to_phys(self->hwpt_id, iova);
+	ASSERT_NE(0, phys1);
+	ASSERT_NE(phys0, phys1);
+	ASSERT_EQ(phys1,
+		  test_cmd_md_iova_to_phys(self->hwpt_id, iova + PAGE_SIZE));
+
+	/* Read-only and MMIO pages stay mapped */
+	test_cmd_provider_set(pfd, MOCK_PROVIDER_READONLY, 0, buf_size / 2);
+	test_cmd_provider_set(pfd, MOCK_PROVIDER_MMIO, buf_size / 2,
+			      buf_size / 2);
+	test_cmd_md_check_mapped(self->hwpt_id, iova, buf_size, true);
+
+	/* A change past the end of the file is refused */
+	EXPECT_ERRNO(EINVAL, _test_cmd_provider_set(self->fd, pfd,
+						    MOCK_PROVIDER_HOLE,
+						    buf_size, PAGE_SIZE));
+
+	test_ioctl_ioas_unmap(iova, buf_size);
+	close(pfd);
+}
+
+TEST_F(iommufd_ioas, provider_dirty_tracking)
+{
+	size_t buf_size = PAGE_SIZE * 4;
+	__u32 hwpt_id;
+	__u64 iova;
+	int pfd;
+
+	if (!variant->mock_domains)
+		SKIP(return, "Needs a mock domain");
+
+	test_cmd_get_provider(buf_size, &pfd);
+
+	/* No dirty tracking domain over provider pages ... */
+	test_ioctl_ioas_map_file(pfd, 0, buf_size, &iova);
+	test_err_hwpt_alloc(EOPNOTSUPP, self->device_id, self->ioas_id,
+			    IOMMU_HWPT_ALLOC_DIRTY_TRACKING, &hwpt_id);
+	test_ioctl_ioas_unmap(iova, buf_size);
+
+	/* ... and no provider pages under a dirty tracking domain. */
+	test_cmd_hwpt_alloc(self->device_id, self->ioas_id,
+			    IOMMU_HWPT_ALLOC_DIRTY_TRACKING, &hwpt_id);
+	test_err_ioctl_ioas_map_file(EOPNOTSUPP, pfd, 0, buf_size, &iova);
+	test_ioctl_destroy(hwpt_id);
+
+	test_ioctl_ioas_map_file(pfd, 0, buf_size, &iova);
+	test_ioctl_ioas_unmap(iova, buf_size);
+	close(pfd);
+}
+
 FIXTURE(iommufd_mock_domain)
 {
 	int fd;
diff --git a/tools/testing/selftests/iommu/iommufd_utils.h b/tools/testing/selftests/iommu/iommufd_utils.h
index b4928cbd4d9c..d1bd0a283db2 100644
--- a/tools/testing/selftests/iommu/iommufd_utils.h
+++ b/tools/testing/selftests/iommu/iommufd_utils.h
@@ -593,6 +593,70 @@ static int _test_cmd_revoke_dmabuf(int fd, int dmabuf_fd, bool revoked)
 #define test_cmd_revoke_dmabuf(dmabuf_fd, revoke) \
 	ASSERT_EQ(0, _test_cmd_revoke_dmabuf(self->fd, dmabuf_fd, revoke))
 
+static int _test_cmd_get_provider(int fd, size_t len, int *out_fd)
+{
+	struct iommu_test_cmd cmd = {
+		.size = sizeof(cmd),
+		.op = IOMMU_TEST_OP_PROVIDER_GET,
+		.provider_get = { .length = len },
+	};
+
+	*out_fd = ioctl(fd, IOMMU_TEST_CMD, &cmd);
+	if (*out_fd < 0)
+		return -1;
+	return 0;
+}
+#define test_cmd_get_provider(len, out_fd) \
+	ASSERT_EQ(0, _test_cmd_get_provider(self->fd, len, out_fd))
+
+static int _test_cmd_provider_set(int fd, int pfd, unsigned int state,
+				  __u64 offset, __u64 length)
+{
+	struct iommu_test_cmd cmd = {
+		.size = sizeof(cmd),
+		.op = IOMMU_TEST_OP_PROVIDER_SET,
+		.provider_set = { .fd = pfd, .state = state,
+				  .offset = offset, .length = length },
+	};
+
+	return ioctl(fd, IOMMU_TEST_CMD, &cmd);
+}
+#define test_cmd_provider_set(pfd, state, offset, length) \
+	ASSERT_EQ(0, _test_cmd_provider_set(self->fd, pfd, state, offset, length))
+
+static int _test_cmd_md_check_mapped(int fd, __u32 hwpt_id, __u64 iova,
+				     __u64 length, bool mapped)
+{
+	struct iommu_test_cmd cmd = {
+		.size = sizeof(cmd),
+		.op = IOMMU_TEST_OP_MD_CHECK_MAPPED,
+		.id = hwpt_id,
+		.check_mapped = { .mapped = mapped, .iova = iova,
+				  .length = length },
+	};
+
+	return ioctl(fd, IOMMU_TEST_CMD, &cmd);
+}
+#define test_cmd_md_check_mapped(hwpt_id, iova, length, mapped)		\
+	ASSERT_EQ(0, _test_cmd_md_check_mapped(self->fd, hwpt_id, iova,	\
+					       length, mapped))
+
+static __u64 _test_cmd_md_iova_to_phys(int fd, __u32 hwpt_id, __u64 iova)
+{
+	struct iommu_test_cmd cmd = {
+		.size = sizeof(cmd),
+		.op = IOMMU_TEST_OP_MD_IOVA_TO_PHYS,
+		.id = hwpt_id,
+		.iova_to_phys = { .iova = iova },
+	};
+
+	if (ioctl(fd, IOMMU_TEST_CMD, &cmd))
+		return -1;
+	return cmd.iova_to_phys.out_phys;
+}
+#define test_cmd_md_iova_to_phys(hwpt_id, iova) \
+	_test_cmd_md_iova_to_phys(self->fd, hwpt_id, iova)
+
 static int _test_ioctl_destroy(int fd, unsigned int id)
 {
 	struct iommu_destroy cmd = {


^ permalink raw reply related	[flat|nested] 10+ messages in thread

* [PATCH 8/9] samples/kvm: Add a memory provider sample
  2026-10-06 18:32 ` [RFC PATCH 0/9] mm: Memory providers for guest_memfd and iommufd Fred Griffoul
                     ` (6 preceding siblings ...)
  2026-10-06 18:32   ` [PATCH 7/9] iommufd/selftest: Add a mock memory provider Fred Griffoul
@ 2026-10-06 18:32   ` Fred Griffoul
  2026-10-06 18:32   ` [PATCH 9/9] KVM: selftests: Test a memory provider shared by KVM and iommufd Fred Griffoul
  8 siblings, 0 replies; 10+ messages in thread
From: Fred Griffoul @ 2026-10-06 18:32 UTC (permalink / raw)
  To: Paolo Bonzini, Sean Christopherson, Marc Zyngier, Oliver Upton,
	Andrew Morton, David Hildenbrand, Alexander Viro,
	Christian Brauner, Jan Kara, Jason Gunthorpe, Kevin Tian,
	Joerg Roedel, Will Deacon, Robin Murphy, Thomas Gleixner,
	Ingo Molnar, Borislav Petkov, Dave Hansen, x86, H . Peter Anvin,
	Jonathan Corbet, Shuah Khan
  Cc: David Woodhouse, Ackerley Tng, Lorenzo Stoakes, Liam R . Howlett,
	Vlastimil Babka, Mike Rapoport, Suren Baghdasaryan, Michal Hocko,
	Joey Gouly, Suzuki K Poulose, Zenghui Yu, Steffen Eiden,
	linux-kernel, kvm, kvmarm, iommu, linux-fsdevel, linux-mm,
	linux-kselftest

From: Fred Griffoul <fgriffo@amazon.co.uk>

The memory provider interface has no in-tree provider that a VMM can
use.

Add mem_provider_sample. It lends a region, either a fixed range given
with addr= and len= or memory from alloc_contig_pages(), as child files
that a VMM passes to both guest_memfd and iommufd. A control device
moves pages between children, takes them back and makes them read
only, and every change revokes the range. It uses no KVM, dma-buf or
iommufd symbol.

The fixed range is not System RAM, so the sample keeps a write-back
memremap() of it while a file over it may be mapped into userspace.
Without it, PAT would make the range uncached on x86, and the VMM's
mapping would be uncached while the guest's is write-back. The sample
is x86 only, as guest_memfd accepts providers only there.

David's gmem_provider sample is unchanged.

Signed-off-by: Fred Griffoul <fgriffo@amazon.co.uk>
---
 MAINTAINERS                       |   1 +
 samples/Kconfig                   |  17 +
 samples/Makefile                  |   1 +
 samples/kvm/Makefile              |   1 +
 samples/kvm/mem_provider_sample.c | 855 ++++++++++++++++++++++++++++++
 samples/kvm/mem_provider_sample.h | 135 +++++
 6 files changed, 1010 insertions(+)
 create mode 100644 samples/kvm/mem_provider_sample.c
 create mode 100644 samples/kvm/mem_provider_sample.h

diff --git a/MAINTAINERS b/MAINTAINERS
index 6cab075a3ff7..b6b47e132e16 100644
--- a/MAINTAINERS
+++ b/MAINTAINERS
@@ -17366,6 +17366,7 @@ L:	linux-mm@kvack.org
 S:	Maintained
 F:	include/linux/mem_provider.h
 F:	mm/mem_provider.c
+F:	samples/kvm/mem_provider_sample.*
 
 MEMORY TECHNOLOGY DEVICES (MTD)
 M:	Miquel Raynal <miquel.raynal@bootlin.com>
diff --git a/samples/Kconfig b/samples/Kconfig
index d26a03dea072..e4c4aface3ab 100644
--- a/samples/Kconfig
+++ b/samples/Kconfig
@@ -344,6 +344,23 @@ config SAMPLE_KVM_GMEM_PROVIDER
 
 	  If unsure, say N.
 
+config SAMPLE_KVM_MEM_PROVIDER
+	tristate "Build sample memory provider -- loadable module only"
+	depends on MEM_PROVIDER && CONTIG_ALLOC && X86_64 && m
+	help
+	  This builds a sample memory provider (include/linux/mem_provider.h).
+	  Its files can back a guest_memfd and be mapped by iommufd.  Loaded
+	  with addr= and len=, it lends a fixed physical range that has no
+	  struct page, for example memory hidden with memmap= on the command
+	  line.  Without them, it allocates a contiguous region with
+	  alloc_contig_pages().
+
+	  It shows how a memory owner moves pages between VMs, takes them
+	  back or makes them read only, and how guest_memfd and iommufd
+	  follow each change.
+
+	  If unsure, say N.
+
 endif # SAMPLES
 
 config HAVE_SAMPLE_FTRACE_DIRECT
diff --git a/samples/Makefile b/samples/Makefile
index e85397e5e34f..0565141d1f45 100644
--- a/samples/Makefile
+++ b/samples/Makefile
@@ -38,6 +38,7 @@ subdir-$(CONFIG_SAMPLE_WATCHDOG)	+= watchdog
 subdir-$(CONFIG_SAMPLE_WATCH_QUEUE)	+= watch_queue
 obj-$(CONFIG_SAMPLE_KMEMLEAK)		+= kmemleak/
 obj-$(CONFIG_SAMPLE_KVM_GMEM_PROVIDER)	+= kvm/
+obj-$(CONFIG_SAMPLE_KVM_MEM_PROVIDER)	+= kvm/
 obj-$(CONFIG_SAMPLE_CORESIGHT_SYSCFG)	+= coresight/
 obj-$(CONFIG_SAMPLE_FPROBE)		+= fprobe/
 obj-$(CONFIG_SAMPLES_RUST)		+= rust/
diff --git a/samples/kvm/Makefile b/samples/kvm/Makefile
index dcad6e53ea78..a885b5313804 100644
--- a/samples/kvm/Makefile
+++ b/samples/kvm/Makefile
@@ -1,2 +1,3 @@
 # SPDX-License-Identifier: GPL-2.0
 obj-$(CONFIG_SAMPLE_KVM_GMEM_PROVIDER) += gmem_provider.o
+obj-$(CONFIG_SAMPLE_KVM_MEM_PROVIDER) += mem_provider_sample.o
diff --git a/samples/kvm/mem_provider_sample.c b/samples/kvm/mem_provider_sample.c
new file mode 100644
index 000000000000..8180959ab78f
--- /dev/null
+++ b/samples/kvm/mem_provider_sample.c
@@ -0,0 +1,855 @@
+// SPDX-License-Identifier: GPL-2.0
+/*
+ * mem_provider_sample - a sample memory provider.
+ *
+ * The module owns a region of physical memory and lends it through provider
+ * files (include/linux/mem_provider.h).  A VMM passes the same file to
+ * KVM_CREATE_GUEST_MEMFD, for the guest, and to IOMMU_IOAS_MAP_FILE, for the
+ * guest's devices.  The module calls neither KVM nor iommufd: it answers
+ * "what is page N" and revokes a range when the answer changes.
+ *
+ * The region is either:
+ *
+ *  - a fixed range given with addr= and len=, which has no struct page, for
+ *    example memory hidden from the kernel with memmap= on the command line;
+ *  - or, without those parameters, a range from alloc_contig_pages().
+ *
+ * The control device /dev/mem_provider_sample creates provider files.
+ * SETUP makes one file over the whole region.  NEW_CHILD carves part of the
+ * region into a child file, and MOVE, DONATE and RECLAIM change which child
+ * has which page.  A page can also be taken away or made read only.  See
+ * mem_provider_sample.h.
+ *
+ * Confidential VMs are not supported.
+ */
+
+#include <linux/anon_inodes.h>
+#include <linux/bitmap.h>
+#include <linux/file.h>
+#include <linux/fs.h>
+#include <linux/gfp.h>
+#include <linux/io.h>
+#include <linux/mem_provider.h>
+#include <linux/miscdevice.h>
+#include <linux/mm.h>
+#include <linux/module.h>
+#include <linux/mutex.h>
+#include <linux/slab.h>
+
+#include "mem_provider_sample.h"
+
+static unsigned long long addr;
+module_param(addr, ullong, 0444);
+MODULE_PARM_DESC(addr,
+		 "Physical base of a region with no struct page (optional)");
+
+static unsigned long long len;
+module_param(len, ullong, 0444);
+MODULE_PARM_DESC(len, "Size in bytes of that region (optional)");
+
+/*
+ * One provider file.
+ *
+ * Lock order: mps_root.lock, then mps_info.lock, then the consumers' locks
+ * (taken by their revoke callbacks).  get_page() takes no lock, so a
+ * consumer can call it from its fault paths and its revoke callback.
+ */
+struct mps_info {
+	struct mem_provider_file mpf;
+	struct mutex lock;		/* the bitmaps */
+
+	unsigned long base_pfn;
+	unsigned long npages;
+	struct page *cma_pages;		/* from alloc_contig_pages(), or NULL */
+	bool fixed;			/* a SETUP file over addr=/len= */
+	bool mmap_capable;
+	bool fixed_wb;			/* holds a ref on the fixed-region WB alias */
+
+	unsigned long *absent;		/* taken away by CTL_SET_PRESENT */
+	unsigned long *readonly;
+
+	/*
+	 * A child of the root: its first page in the root, and the pages it
+	 * has.  NULL @owned for a file made by SETUP, which has all its pages.
+	 */
+	unsigned long root_index;
+	unsigned long *owned;
+	struct list_head root_link;	/* mps_root.children, mps_root.lock */
+};
+
+/*
+ * The root region, owned by the control device.  Changes of ownership run
+ * under @lock, which is taken before any child's lock.
+ */
+static struct mps_root {
+	struct mutex lock;		/* everything below */
+	unsigned long base_pfn;
+	unsigned long npages;
+	struct page *cma_pages;
+	unsigned long *owned;		/* pages that a child has */
+	unsigned long *donated;		/* pages kept at the root */
+	struct list_head children;
+	/*
+	 * The addr=/len= region is used either by one SETUP file or by the
+	 * root, never by both, so that its frames have one owner.
+	 */
+	bool fixed_setup;
+	/* Files over the fixed region that may be mapped into userspace. */
+	unsigned int fixed_wb_users;
+} mps_root;
+
+/*
+ * A write-back mapping of the addr=/len= region, held while any file over it
+ * may be mapped into userspace (counted by mps_root.fixed_wb_users).  Without
+ * it, PAT makes the range uncached on x86, and the VMM's mapping would be
+ * uncached while the guest's is write-back.  PAT may still give a weaker type
+ * if the MTRRs do not mark the range write-back.
+ */
+static void *mps_fixed_wb;
+
+/* A file over the fixed region whose pages the host may map. */
+static bool mps_over_fixed_mappable(struct mps_info *info)
+{
+	return addr && len && !info->cma_pages && info->mmap_capable;
+}
+
+/*
+ * Hold a write-back mapping of the fixed region while any file over it may be
+ * mapped into userspace.  Called under mps_root.lock.
+ */
+static int mps_fixed_wb_get(void)
+{
+	if (mps_root.fixed_wb_users == 0) {
+		mps_fixed_wb = memremap(addr, len, MEMREMAP_WB);
+		if (!mps_fixed_wb)
+			return -ENOMEM;
+	}
+	mps_root.fixed_wb_users++;
+	return 0;
+}
+
+static void mps_fixed_wb_put(void)
+{
+	if (--mps_root.fixed_wb_users == 0) {
+		memunmap(mps_fixed_wb);
+		mps_fixed_wb = NULL;
+	}
+}
+
+/* Does @info have page @index now?  Called with or without info->lock. */
+static bool mps_page_present(struct mps_info *info, unsigned long index)
+{
+	if (info->owned && !test_bit(index, info->owned))
+		return false;
+	return !test_bit(index, info->absent);
+}
+
+/*
+ * The length of the run from @index of pages with the same state as @index:
+ * present, and the same read-only bit.
+ */
+static unsigned long mps_run(struct mps_info *info, unsigned long index)
+{
+	unsigned long end = info->npages;
+
+	end = min(end, find_next_bit(info->absent, info->npages, index + 1));
+	if (info->owned)
+		end = min(end, find_next_zero_bit(info->owned, info->npages,
+						  index + 1));
+	if (test_bit(index, info->readonly))
+		end = min(end, find_next_zero_bit(info->readonly, info->npages,
+						  index + 1));
+	else
+		end = min(end, find_next_bit(info->readonly, info->npages,
+					     index + 1));
+	return end - index;
+}
+
+/* ---- Provider operations ------------------------------------------------ */
+
+static struct mem_provider_file *
+mps_mp_attach(struct file *file, loff_t size)
+{
+	struct mps_info *info = file->private_data;
+
+	if (size > (loff_t)info->npages << PAGE_SHIFT)
+		return ERR_PTR(-EINVAL);
+	return &info->mpf;
+}
+
+static void mps_mp_detach(struct mem_provider_file *mpf)
+{
+}
+
+/*
+ * Reads the bitmaps without info->lock.  A change runs under the lock and
+ * revokes the range afterwards, so an answer that a change makes out of date
+ * is dropped by the consumer.
+ */
+static int mps_mp_get_page(struct mem_provider_file *mpf, pgoff_t index,
+			   unsigned long *pfn, int *max_order, u32 *attrs)
+{
+	struct mps_info *info = container_of(mpf, struct mps_info, mpf);
+	unsigned long run;
+
+	if (index >= info->npages)
+		return -EINVAL;
+
+	if (!mps_page_present(info, index))
+		return -EFAULT;
+
+	/* The aligned block around @index must be one run. */
+	run = mps_run(info, index);
+	while (*max_order) {
+		unsigned long first = ALIGN_DOWN(index, 1UL << *max_order);
+
+		if (first + (1UL << *max_order) <= index + run &&
+		    (first == index || mps_page_present(info, first)) &&
+		    mps_run(info, first) >= 1UL << *max_order)
+			break;
+		(*max_order)--;
+	}
+
+	*pfn = info->base_pfn + index;
+	if (test_bit(index, info->readonly))
+		*attrs |= MEM_PROVIDER_ATTR_READONLY;
+	/* The creator of the file decides whether the host may map it. */
+	if (!info->mmap_capable)
+		*attrs |= MEM_PROVIDER_ATTR_NO_USER_MAP;
+	return 0;
+}
+
+static const struct mem_provider_ops mps_mp_ops = {
+	.attach		= mps_mp_attach,
+	.detach		= mps_mp_detach,
+	.get_page	= mps_mp_get_page,
+};
+
+/* Tell the consumers that pages [start, end) of @info changed. */
+static void mps_revoke(struct mps_info *info, unsigned long start,
+		       unsigned long end)
+{
+	if (end > start)
+		mem_provider_revoke(&info->mpf, (loff_t)start << PAGE_SHIFT,
+				    (loff_t)(end - start) << PAGE_SHIFT);
+}
+
+/* ---- Provider file ------------------------------------------------------ */
+
+/* Byte range [offset, offset + len) of @info as page indices. */
+static int mps_range(struct mps_info *info, u64 offset, u64 len,
+		     unsigned long *start, unsigned long *end)
+{
+	if (!len || !PAGE_ALIGNED(offset) || !PAGE_ALIGNED(len))
+		return -EINVAL;
+	*start = offset >> PAGE_SHIFT;
+	*end = *start + (len >> PAGE_SHIFT);
+	if (*end > info->npages || *end < *start)
+		return -EINVAL;
+	return 0;
+}
+
+static void mps_set_readonly(struct mps_info *info, unsigned long start,
+			     unsigned long end, bool readonly)
+{
+	mutex_lock(&info->lock);
+	if (readonly)
+		bitmap_set(info->readonly, start, end - start);
+	else
+		bitmap_clear(info->readonly, start, end - start);
+	mps_revoke(info, start, end);
+	mutex_unlock(&info->lock);
+}
+
+static void mps_set_present(struct mps_info *info, unsigned long start,
+			    unsigned long end, bool present)
+{
+	mutex_lock(&info->lock);
+	if (present)
+		bitmap_clear(info->absent, start, end - start);
+	else
+		bitmap_set(info->absent, start, end - start);
+	mps_revoke(info, start, end);
+	mutex_unlock(&info->lock);
+}
+
+static long mps_get_stats(struct mps_info *info, void __user *uarg)
+{
+	struct mps_stats st = {};
+
+	mutex_lock(&info->lock);
+	st.region_offset = (u64)info->root_index << PAGE_SHIFT;
+	st.region_len = (u64)info->npages << PAGE_SHIFT;
+	st.owned_pages = info->owned ?
+		bitmap_weight(info->owned, info->npages) : info->npages;
+	st.absent_pages = bitmap_weight(info->absent, info->npages);
+	st.readonly_pages = bitmap_weight(info->readonly, info->npages);
+	mutex_unlock(&info->lock);
+
+	return copy_to_user(uarg, &st, sizeof(st)) ? -EFAULT : 0;
+}
+
+/*
+ * A child fd is read-only to its holder: only GET_STATS.  The owner changes a
+ * child's pages through the control device (MPS_CTL_SET_PRESENT and
+ * MPS_CTL_SET_READONLY), so a VMM that holds a child cannot.
+ */
+static long mps_file_ioctl(struct file *file, unsigned int cmd,
+			   unsigned long arg)
+{
+	struct mps_info *info = file->private_data;
+	void __user *uarg = (void __user *)arg;
+
+	if (cmd == MPS_GET_STATS)
+		return mps_get_stats(info, uarg);
+	return -ENOTTY;
+}
+
+/* Every consumer has detached: each one holds a reference to the file. */
+static int mps_file_release(struct inode *inode, struct file *file)
+{
+	struct mps_info *info = file->private_data;
+	unsigned long i;
+
+	/*
+	 * A child gives the pages it has back to the root.  Pages it donated
+	 * stay at the root, and pages moved out belong to another child.
+	 */
+	if (info->owned) {
+		mutex_lock(&mps_root.lock);
+		list_del(&info->root_link);
+		for_each_set_bit(i, info->owned, info->npages)
+			__clear_bit(info->root_index + i, mps_root.owned);
+		mutex_unlock(&mps_root.lock);
+		bitmap_free(info->owned);
+	}
+	if (addr && len && !info->cma_pages) {
+		mutex_lock(&mps_root.lock);
+		if (info->fixed)
+			mps_root.fixed_setup = false;
+		if (info->fixed_wb)
+			mps_fixed_wb_put();
+		mutex_unlock(&mps_root.lock);
+	}
+
+	if (info->cma_pages)
+		free_contig_range(info->base_pfn, info->npages);
+	bitmap_free(info->absent);
+	bitmap_free(info->readonly);
+	kfree(info);
+	return 0;
+}
+
+static const struct mem_provider_fops mps_file_fops = {
+	.fops = {
+		.owner		= THIS_MODULE,
+		.fop_flags	= FOP_MEM_PROVIDER,
+		.release	= mps_file_release,
+		.unlocked_ioctl	= mps_file_ioctl,
+		.compat_ioctl	= compat_ptr_ioctl,
+	},
+	.ops = &mps_mp_ops,
+};
+
+/*
+ * Make a provider file over [base_pfn, base_pfn + npages).  The fd is
+ * reserved but not installed, so that the caller can finish setting up
+ * @info before another thread can reach it.
+ */
+static int mps_new_file(unsigned long base_pfn, unsigned long npages,
+			u32 flags, struct mps_info **infop,
+			 struct file **filep)
+{
+	struct mps_info *info;
+	struct file *file;
+	int fd, ret;
+
+	info = kzalloc_obj(*info);
+	if (!info)
+		return -ENOMEM;
+	mem_provider_file_init(&info->mpf);
+	info->base_pfn = base_pfn;
+	info->npages = npages;
+	info->mmap_capable = flags & MPS_FLAG_MMAP_CAPABLE;
+	info->absent = bitmap_zalloc(npages, GFP_KERNEL);
+	info->readonly = bitmap_zalloc(npages, GFP_KERNEL);
+	if (!info->absent || !info->readonly) {
+		ret = -ENOMEM;
+		goto err_free;
+	}
+	mutex_init(&info->lock);
+	INIT_LIST_HEAD(&info->root_link);
+
+	fd = get_unused_fd_flags(O_CLOEXEC);
+	if (fd < 0) {
+		ret = fd;
+		goto err_free;
+	}
+	file = anon_inode_getfile("[mem-provider-sample]", &mps_file_fops.fops,
+				  info, O_RDWR);
+	if (IS_ERR(file)) {
+		put_unused_fd(fd);
+		ret = PTR_ERR(file);
+		goto err_free;
+	}
+	*infop = info;
+	*filep = file;
+	return fd;
+
+err_free:
+	bitmap_free(info->absent);
+	bitmap_free(info->readonly);
+	kfree(info);
+	return ret;
+}
+
+static struct page *mps_alloc_region(unsigned long npages)
+{
+	struct page *pages;
+
+	pages = alloc_contig_pages(npages, GFP_KERNEL | __GFP_ZERO,
+				   numa_node_id(), NULL);
+	return pages;
+}
+
+static long mps_ctl_setup(void __user *uarg)
+{
+	struct mps_setup setup;
+	unsigned long base_pfn, npages;
+	struct page *pages = NULL;
+	struct mps_info *info;
+	struct file *file;
+	int fd;
+
+	if (copy_from_user(&setup, uarg, sizeof(setup)))
+		return -EFAULT;
+	if ((setup.flags & ~MPS_FLAG_MMAP_CAPABLE) || setup.pad)
+		return -EINVAL;
+
+	if (addr && len) {
+		/* The fixed region has one owner: this file or the root. */
+		mutex_lock(&mps_root.lock);
+		if (mps_root.fixed_setup || mps_root.npages) {
+			mutex_unlock(&mps_root.lock);
+			return -EBUSY;
+		}
+		fd = mps_new_file(addr >> PAGE_SHIFT, len >> PAGE_SHIFT,
+				  setup.flags, &info, &file);
+		if (fd >= 0 && mps_over_fixed_mappable(info)) {
+			if (mps_fixed_wb_get()) {
+				put_unused_fd(fd);
+				fput(file);
+				fd = -ENOMEM;
+			} else {
+				info->fixed_wb = true;
+			}
+		}
+		if (fd >= 0) {
+			info->fixed = true;
+			mps_root.fixed_setup = true;
+		}
+		mutex_unlock(&mps_root.lock);
+		if (fd >= 0)
+			fd_install(fd, file);
+		return fd;
+	}
+
+	if (!setup.size || !PAGE_ALIGNED(setup.size))
+		return -EINVAL;
+	npages = setup.size >> PAGE_SHIFT;
+	pages = mps_alloc_region(npages);
+	if (!pages)
+		return -ENOMEM;
+	base_pfn = page_to_pfn(pages);
+
+	fd = mps_new_file(base_pfn, npages, setup.flags, &info, &file);
+	if (fd < 0) {
+		free_contig_range(base_pfn, npages);
+		return fd;
+	}
+	info->cma_pages = pages;
+	fd_install(fd, file);
+	return fd;
+}
+
+/* ---- Root and children -------------------------------------------------- */
+
+/* Create the root on the first NEW_CHILD, from addr=/len= or of @size. */
+static int mps_root_ensure(u64 size)
+{
+	struct page *pages = NULL;
+	unsigned long npages;
+
+	lockdep_assert_held(&mps_root.lock);
+	if (mps_root.npages)
+		return 0;
+
+	if (addr && len) {
+		if (mps_root.fixed_setup)
+			return -EBUSY;
+		mps_root.base_pfn = addr >> PAGE_SHIFT;
+		npages = len >> PAGE_SHIFT;
+	} else {
+		if (!size || !PAGE_ALIGNED(size))
+			return -EINVAL;
+		npages = size >> PAGE_SHIFT;
+		pages = mps_alloc_region(npages);
+		if (!pages)
+			return -ENOMEM;
+		mps_root.base_pfn = page_to_pfn(pages);
+	}
+	mps_root.owned = bitmap_zalloc(npages, GFP_KERNEL);
+	mps_root.donated = bitmap_zalloc(npages, GFP_KERNEL);
+	if (!mps_root.owned || !mps_root.donated) {
+		bitmap_free(mps_root.owned);
+		bitmap_free(mps_root.donated);
+		mps_root.owned = NULL;
+		mps_root.donated = NULL;
+		if (pages)
+			free_contig_range(page_to_pfn(pages), npages);
+		return -ENOMEM;
+	}
+	mps_root.cma_pages = pages;
+	mps_root.npages = npages;
+	return 0;
+}
+
+static void mps_root_teardown(void)
+{
+	if (!mps_root.npages)
+		return;
+	WARN_ON(!list_empty(&mps_root.children));
+	if (mps_root.cma_pages)
+		free_contig_range(mps_root.base_pfn, mps_root.npages);
+	bitmap_free(mps_root.owned);
+	bitmap_free(mps_root.donated);
+}
+
+/* A page-aligned byte range of the root as page indices [first, last). */
+static int mps_root_range(u64 offset, u64 length, unsigned long *first,
+			  unsigned long *last)
+{
+	if (!length || !PAGE_ALIGNED(offset) || !PAGE_ALIGNED(length))
+		return -EINVAL;
+	*first = offset >> PAGE_SHIFT;
+	*last = *first + (length >> PAGE_SHIFT);
+	if (*last <= *first || *last > mps_root.npages)
+		return -EINVAL;
+	return 0;
+}
+
+/* Return a referenced child file made by NEW_CHILD. */
+static struct file *mps_get_child(int fd, struct mps_info **infop)
+{
+	struct file *file = fget(fd);
+
+	if (!file)
+		return ERR_PTR(-EBADF);
+	if (file->f_op != &mps_file_fops.fops ||
+	    !((struct mps_info *)file->private_data)->owned) {
+		fput(file);
+		return ERR_PTR(-EINVAL);
+	}
+	*infop = file->private_data;
+	return file;
+}
+
+static long mps_ctl_new_child(void __user *uarg)
+{
+	struct mps_new_child nc;
+	unsigned long first, last, i, grant = 0;
+	struct mps_info *info;
+	unsigned long *owned;
+	struct file *file;
+	int fd, ret;
+
+	if (copy_from_user(&nc, uarg, sizeof(nc)))
+		return -EFAULT;
+	if ((nc.flags & ~MPS_FLAG_MMAP_CAPABLE) || nc.pad)
+		return -EINVAL;
+
+	mutex_lock(&mps_root.lock);
+	ret = mps_root_ensure(nc.offset + nc.len);
+	if (ret)
+		goto out_unlock;
+	ret = mps_root_range(nc.offset, nc.len, &first, &last);
+	if (ret)
+		goto out_unlock;
+
+	/* A child that would get no page is likely a mistake. */
+	for (i = first; i < last; i++)
+		if (!test_bit(i, mps_root.owned) &&
+		    !test_bit(i, mps_root.donated))
+			grant++;
+	if (!grant) {
+		ret = -EBUSY;
+		goto out_unlock;
+	}
+
+	owned = bitmap_zalloc(last - first, GFP_KERNEL);
+	if (!owned) {
+		ret = -ENOMEM;
+		goto out_unlock;
+	}
+	fd = mps_new_file(mps_root.base_pfn + first, last - first, nc.flags,
+			  &info, &file);
+	if (fd < 0) {
+		bitmap_free(owned);
+		ret = fd;
+		goto out_unlock;
+	}
+	if (mps_over_fixed_mappable(info)) {
+		if (mps_fixed_wb_get()) {
+			put_unused_fd(fd);
+			fput(file);
+			bitmap_free(owned);
+			ret = -ENOMEM;
+			goto out_unlock;
+		}
+		info->fixed_wb = true;
+	}
+
+	/* No other thread can reach the child until fd_install(). */
+	for (i = first; i < last; i++) {
+		if (test_bit(i, mps_root.owned) ||
+		    test_bit(i, mps_root.donated))
+			continue;
+		__set_bit(i - first, owned);
+		__set_bit(i, mps_root.owned);
+	}
+	info->owned = owned;
+	info->root_index = first;
+	list_add(&info->root_link, &mps_root.children);
+	mutex_unlock(&mps_root.lock);
+	fd_install(fd, file);
+	return fd;
+
+out_unlock:
+	mutex_unlock(&mps_root.lock);
+	return ret;
+}
+
+/* Take [first, last) of the root from @info.  Called under mps_root.lock. */
+static void mps_child_lose(struct mps_info *info, unsigned long first,
+			   unsigned long last)
+{
+	unsigned long s = first - info->root_index;
+	unsigned long e = last - info->root_index;
+
+	mutex_lock(&info->lock);
+	bitmap_clear(info->owned, s, e - s);
+	mps_revoke(info, s, e);
+	mutex_unlock(&info->lock);
+	bitmap_clear(mps_root.owned, first, last - first);
+}
+
+/* Give [first, last) of the root to @info.  Called under mps_root.lock. */
+static void mps_child_gain(struct mps_info *info, unsigned long first,
+			   unsigned long last)
+{
+	unsigned long s = first - info->root_index;
+	unsigned long e = last - info->root_index;
+
+	mutex_lock(&info->lock);
+	bitmap_set(info->owned, s, e - s);
+	mps_revoke(info, s, e);
+	mutex_unlock(&info->lock);
+	bitmap_set(mps_root.owned, first, last - first);
+}
+
+static bool mps_child_covers(struct mps_info *info, unsigned long first,
+			     unsigned long last)
+{
+	return first >= info->root_index &&
+	       last <= info->root_index + info->npages;
+}
+
+static bool mps_child_owns(struct mps_info *info, unsigned long first,
+			   unsigned long last)
+{
+	unsigned long s = first - info->root_index;
+	unsigned long e = last - info->root_index;
+
+	return mps_child_covers(info, first, last) &&
+	       find_next_zero_bit(info->owned, e, s) >= e;
+}
+
+static long mps_ctl_move(void __user *uarg)
+{
+	struct mps_move mv;
+	struct mps_info *src, *dst;
+	unsigned long first, last;
+	struct file *sf, *df;
+	long ret;
+
+	if (copy_from_user(&mv, uarg, sizeof(mv)))
+		return -EFAULT;
+	sf = mps_get_child(mv.src_fd, &src);
+	if (IS_ERR(sf))
+		return PTR_ERR(sf);
+	df = mps_get_child(mv.dst_fd, &dst);
+	if (IS_ERR(df)) {
+		fput(sf);
+		return PTR_ERR(df);
+	}
+
+	mutex_lock(&mps_root.lock);
+	ret = mps_root_range(mv.offset, mv.len, &first, &last);
+	if (ret)
+		goto out;
+	if (src == dst || !mps_child_owns(src, first, last) ||
+	    !mps_child_covers(dst, first, last)) {
+		ret = -EINVAL;
+		goto out;
+	}
+	/* Revoke from the source before the destination gets the pages. */
+	mps_child_lose(src, first, last);
+	mps_child_gain(dst, first, last);
+out:
+	mutex_unlock(&mps_root.lock);
+	fput(df);
+	fput(sf);
+	return ret;
+}
+
+static long mps_ctl_donate(void __user *uarg, bool reclaim)
+{
+	struct mps_donate d;
+	unsigned long first, last;
+	struct mps_info *info;
+	struct file *file;
+	long ret;
+
+	if (copy_from_user(&d, uarg, sizeof(d)))
+		return -EFAULT;
+	if (d.pad)
+		return -EINVAL;
+	file = mps_get_child(d.fd, &info);
+	if (IS_ERR(file))
+		return PTR_ERR(file);
+
+	mutex_lock(&mps_root.lock);
+	ret = mps_root_range(d.offset, d.len, &first, &last);
+	if (ret)
+		goto out;
+	if (!reclaim) {
+		if (!mps_child_owns(info, first, last)) {
+			ret = -EINVAL;
+			goto out;
+		}
+		mps_child_lose(info, first, last);
+		bitmap_set(mps_root.donated, first, last - first);
+	} else {
+		if (!mps_child_covers(info, first, last) ||
+		    find_next_zero_bit(mps_root.donated, last, first) < last) {
+			ret = -EINVAL;
+			goto out;
+		}
+		bitmap_clear(mps_root.donated, first, last - first);
+		mps_child_gain(info, first, last);
+	}
+out:
+	mutex_unlock(&mps_root.lock);
+	fput(file);
+	return ret;
+}
+
+/*
+ * CTL_SET_READONLY and CTL_SET_PRESENT change a child's pages, by the owner on
+ * the control device.  A VMM that holds the child fd cannot, so it cannot undo
+ * them.
+ */
+static long mps_ctl_set_range(void __user *uarg, bool readonly)
+{
+	struct mps_ctl_range cr;
+	unsigned long start, end;
+	struct mps_info *info;
+	struct file *file;
+	int ret;
+
+	if (copy_from_user(&cr, uarg, sizeof(cr)))
+		return -EFAULT;
+	if (cr.value > 1)
+		return -EINVAL;
+	file = mps_get_child(cr.fd, &info);
+	if (IS_ERR(file))
+		return PTR_ERR(file);
+	ret = mps_range(info, cr.offset, cr.len, &start, &end);
+	if (!ret) {
+		if (readonly)
+			mps_set_readonly(info, start, end, cr.value);
+		else
+			mps_set_present(info, start, end, cr.value);
+	}
+	fput(file);
+	return ret;
+}
+
+static long mps_ctl_ioctl(struct file *file, unsigned int cmd,
+			  unsigned long arg)
+{
+	void __user *uarg = (void __user *)arg;
+
+	switch (cmd) {
+	case MPS_SETUP:
+		return mps_ctl_setup(uarg);
+	case MPS_NEW_CHILD:
+		return mps_ctl_new_child(uarg);
+	case MPS_MOVE:
+		return mps_ctl_move(uarg);
+	case MPS_DONATE:
+		return mps_ctl_donate(uarg, false);
+	case MPS_RECLAIM:
+		return mps_ctl_donate(uarg, true);
+	case MPS_CTL_SET_READONLY:
+		return mps_ctl_set_range(uarg, true);
+	case MPS_CTL_SET_PRESENT:
+		return mps_ctl_set_range(uarg, false);
+	default:
+		return -ENOTTY;
+	}
+}
+
+static const struct file_operations mps_ctl_fops = {
+	.owner		= THIS_MODULE,
+	.unlocked_ioctl	= mps_ctl_ioctl,
+	.compat_ioctl	= compat_ptr_ioctl,
+};
+
+static struct miscdevice mps_dev = {
+	.minor	= MISC_DYNAMIC_MINOR,
+	.name	= "mem_provider_sample",
+	.fops	= &mps_ctl_fops,
+};
+
+static int __init mps_init(void)
+{
+	int ret;
+
+	if ((addr || len) &&
+	    (!addr || !len || !PAGE_ALIGNED(addr) || !PAGE_ALIGNED(len))) {
+		pr_err("mem_provider_sample: addr= and len= must both be set and page aligned\n");
+		return -EINVAL;
+	}
+
+	mutex_init(&mps_root.lock);
+	INIT_LIST_HEAD(&mps_root.children);
+
+	ret = misc_register(&mps_dev);
+	if (ret)
+		return ret;
+	return 0;
+}
+module_init(mps_init);
+
+static void __exit mps_exit(void)
+{
+	misc_deregister(&mps_dev);
+	mps_root_teardown();
+	if (mps_fixed_wb)
+		memunmap(mps_fixed_wb);
+}
+module_exit(mps_exit);
+
+MODULE_LICENSE("GPL");
+MODULE_DESCRIPTION("Sample memory provider for guest_memfd and iommufd");
diff --git a/samples/kvm/mem_provider_sample.h b/samples/kvm/mem_provider_sample.h
new file mode 100644
index 000000000000..3be453756d84
--- /dev/null
+++ b/samples/kvm/mem_provider_sample.h
@@ -0,0 +1,135 @@
+/* SPDX-License-Identifier: GPL-2.0 */
+#ifndef _SAMPLES_KVM_MEM_PROVIDER_SAMPLE_H
+#define _SAMPLES_KVM_MEM_PROVIDER_SAMPLE_H
+
+#include <linux/ioctl.h>
+#include <linux/types.h>
+
+/*
+ * A sample memory provider (include/linux/mem_provider.h).
+ *
+ * Every fd that this module returns is a provider file.  Pass it as
+ * provider_fd to KVM_CREATE_GUEST_MEMFD with GUEST_MEMFD_FLAG_USE_PROVIDER,
+ * and as fd to IOMMU_IOAS_MAP_FILE.  Both consumers follow every change that
+ * the ioctls below make.
+ */
+
+/* Flags for struct mps_setup and struct mps_new_child */
+/* The host may map the pages; otherwise they are NO_USER_MAP */
+#define MPS_FLAG_MMAP_CAPABLE	(1u << 0)
+
+/*
+ * ioctl on /dev/mem_provider_sample: create a provider fd.
+ *
+ * If the module was loaded with addr= and len=, the fd covers that fixed
+ * range, which has no struct page, and @size is ignored.  The fixed range
+ * then has one owner.  SETUP fails with -EBUSY while a SETUP file uses it,
+ * or once the first MPS_NEW_CHILD has made it the root of the children, which
+ * lasts until the module is unloaded.  MPS_NEW_CHILD fails with -EBUSY while
+ * a SETUP file uses it.  Otherwise the fd covers @size bytes from
+ * alloc_contig_pages().
+ */
+struct mps_setup {
+	__u32 flags;	/* MPS_FLAG_* */
+	__u32 pad;
+	__u64 size;	/* bytes, page aligned */
+};
+
+#define MPS_IOCTL_BASE		'P'
+#define MPS_SETUP \
+	_IOW(MPS_IOCTL_BASE, 1, struct mps_setup)
+
+/*
+ * ---- A tree of owners --------------------------------------------------
+ *
+ * The control device owns one region, the root.  NEW_CHILD carves part of
+ * it into a new provider fd for one VM.  The owner can MOVE pages between
+ * children, DONATE pages to the root, so that no child has them, and
+ * RECLAIM them.  Each change revokes the range, so KVM and iommufd follow.
+ *
+ * The owner changes a child's pages through the control device
+ * (CTL_SET_READONLY, CTL_SET_PRESENT).  A child fd is read-only to its
+ * holder: a VMM can only GET_STATS, never change the memory.
+ */
+
+/*
+ * ioctl on /dev/mem_provider_sample: carve [@offset, @offset + @len) of the
+ * root into a new child fd.  The child gets every page of the range that no
+ * other child has and that is not donated.  Ranges of children may overlap,
+ * which is how a page can MOVE between them.  Returns the child fd.
+ */
+struct mps_new_child {
+	__u32 flags;	/* MPS_FLAG_* */
+	__u32 pad;
+	__u64 offset;	/* bytes into the root, page aligned */
+	__u64 len;	/* bytes, page aligned */
+};
+
+#define MPS_NEW_CHILD \
+	_IOW(MPS_IOCTL_BASE, 5, struct mps_new_child)
+
+/*
+ * ioctl on /dev/mem_provider_sample: move [@offset, @offset + @len) of the root
+ * from child @src_fd to child @dst_fd.  @src_fd must have every page of the
+ * range, and the range must be inside @dst_fd's.  The pages are revoked from
+ * the source before the destination gets them, so the consumers of the two
+ * children never map them at the same time.
+ */
+struct mps_move {
+	__s32 src_fd;
+	__s32 dst_fd;
+	__u64 offset;	/* bytes into the root, page aligned */
+	__u64 len;	/* bytes, page aligned */
+};
+
+#define MPS_MOVE \
+	_IOW(MPS_IOCTL_BASE, 6, struct mps_move)
+
+/*
+ * ioctls on /dev/mem_provider_sample: DONATE takes [@offset, @offset + @len) of
+ * the root from child @fd, which must have all of it, and keeps it at the
+ * root.  RECLAIM gives a donated range to child @fd, whose range must
+ * contain it.  The sample does not scrub the pages.
+ */
+struct mps_donate {
+	__s32 fd;
+	__u32 pad;
+	__u64 offset;	/* bytes into the root, page aligned */
+	__u64 len;	/* bytes, page aligned */
+};
+
+#define MPS_DONATE \
+	_IOW(MPS_IOCTL_BASE, 7, struct mps_donate)
+#define MPS_RECLAIM \
+	_IOW(MPS_IOCTL_BASE, 8, struct mps_donate)
+
+/* ioctl on a child fd: read the child's state, for tests. */
+struct mps_stats {
+	__u64 region_offset;	/* the child's range in the root */
+	__u64 region_len;
+	__u64 owned_pages;	/* pages the child has */
+	__u64 absent_pages;	/* pages taken away by CTL_SET_PRESENT */
+	__u64 readonly_pages;
+};
+
+#define MPS_GET_STATS \
+	_IOR(MPS_IOCTL_BASE, 10, struct mps_stats)
+
+/*
+ * ioctls on /dev/mem_provider_sample: CTL_SET_READONLY or CTL_SET_PRESENT on
+ * child @fd, by the owner.  @value is the readonly or present value, and
+ * @offset and @len are a byte range within the child, page aligned.
+ */
+struct mps_ctl_range {
+	__s32 fd;
+	__u32 value;
+	__u64 offset;	/* bytes into the child, page aligned */
+	__u64 len;	/* bytes, page aligned */
+};
+
+#define MPS_CTL_SET_READONLY \
+	_IOW(MPS_IOCTL_BASE, 11, struct mps_ctl_range)
+#define MPS_CTL_SET_PRESENT \
+	_IOW(MPS_IOCTL_BASE, 12, struct mps_ctl_range)
+
+#endif /* _SAMPLES_KVM_MEM_PROVIDER_SAMPLE_H */


^ permalink raw reply related	[flat|nested] 10+ messages in thread

* [PATCH 9/9] KVM: selftests: Test a memory provider shared by KVM and iommufd
  2026-10-06 18:32 ` [RFC PATCH 0/9] mm: Memory providers for guest_memfd and iommufd Fred Griffoul
                     ` (7 preceding siblings ...)
  2026-10-06 18:32   ` [PATCH 8/9] samples/kvm: Add a memory provider sample Fred Griffoul
@ 2026-10-06 18:32   ` Fred Griffoul
  8 siblings, 0 replies; 10+ messages in thread
From: Fred Griffoul @ 2026-10-06 18:32 UTC (permalink / raw)
  To: Paolo Bonzini, Sean Christopherson, Marc Zyngier, Oliver Upton,
	Andrew Morton, David Hildenbrand, Alexander Viro,
	Christian Brauner, Jan Kara, Jason Gunthorpe, Kevin Tian,
	Joerg Roedel, Will Deacon, Robin Murphy, Thomas Gleixner,
	Ingo Molnar, Borislav Petkov, Dave Hansen, x86, H . Peter Anvin,
	Jonathan Corbet, Shuah Khan
  Cc: David Woodhouse, Ackerley Tng, Lorenzo Stoakes, Liam R . Howlett,
	Vlastimil Babka, Mike Rapoport, Suren Baghdasaryan, Michal Hocko,
	Joey Gouly, Suzuki K Poulose, Zenghui Yu, Steffen Eiden,
	linux-kernel, kvm, kvmarm, iommu, linux-fsdevel, linux-mm,
	linux-kselftest

From: Fred Griffoul <fgriffo@amazon.co.uk>

Test that guest, host and device mappings of one provider file follow
the provider's changes.

Two VMs get child files of the sample provider, each passed to
guest_memfd and to an iommufd mock domain. The test moves a range
between them, donates and reclaims pages, and makes a page read only.
After each change it checks the guest, a host mapping and the device
mapping. Pages are mapped on the host before a change, so the test
checks that the revoke zaps them.

Signed-off-by: Fred Griffoul <fgriffo@amazon.co.uk>
---
 tools/testing/selftests/kvm/Makefile.kvm      |   1 +
 .../selftests/kvm/x86/mem_provider_test.c     | 582 ++++++++++++++++++
 2 files changed, 583 insertions(+)
 create mode 100644 tools/testing/selftests/kvm/x86/mem_provider_test.c

diff --git a/tools/testing/selftests/kvm/Makefile.kvm b/tools/testing/selftests/kvm/Makefile.kvm
index 4d7082448cea..ccee373b4c13 100644
--- a/tools/testing/selftests/kvm/Makefile.kvm
+++ b/tools/testing/selftests/kvm/Makefile.kvm
@@ -84,6 +84,7 @@ TEST_GEN_PROGS_x86 += x86/gmem_provider_revoke_test
 TEST_GEN_PROGS_x86 += x86/gmem_provider_readonly_test
 TEST_GEN_PROGS_x86 += x86/gmem_provider_iommufd_test
 TEST_GEN_PROGS_x86 += x86/gmem_provider_vfio_test
+TEST_GEN_PROGS_x86 += x86/mem_provider_test
 TEST_GEN_PROGS_x86 += x86/hwcr_msr_test
 TEST_GEN_PROGS_x86 += x86/hyperv_clock
 TEST_GEN_PROGS_x86 += x86/hyperv_cpuid
diff --git a/tools/testing/selftests/kvm/x86/mem_provider_test.c b/tools/testing/selftests/kvm/x86/mem_provider_test.c
new file mode 100644
index 000000000000..bd7ae44f05d7
--- /dev/null
+++ b/tools/testing/selftests/kvm/x86/mem_provider_test.c
@@ -0,0 +1,582 @@
+// SPDX-License-Identifier: GPL-2.0
+/*
+ * mem_provider_test - one memory provider shared by KVM and iommufd.
+ *
+ * A control process and two VMMs, in one process:
+ *
+ *  ctl:  holds /dev/mem_provider_sample, the owner of the memory.  Creates a
+ *        child provider file per VM, moves pages between children, donates
+ *        and reclaims them, and makes them read only.
+ *        Never touches a VM.
+ *  vmm:  holds one child file, one VM and one iommufd IOAS on a mock domain.
+ *        Passes the child file to KVM_CREATE_GUEST_MEMFD and to
+ *        IOMMU_IOAS_MAP_FILE.  Never touches the control device.
+ *
+ * Scenarios:
+ *
+ *  1. Launch:    each guest writes to its memory, the host sees the write,
+ *                and the device mapping covers the pages the child has.
+ *  2. Move:      a range leaves A for B.  A's guest and A's host mapping
+ *                fault on it and A's device loses only that range.  B's
+ *                guest and host read what A wrote, and B's device reaches
+ *                the frame A had.
+ *  3. Donate:    a range leaves A for the root and comes back on reclaim,
+ *                for the guest, the host mapping and the device.
+ *  4. Read only: a guest write to a read-only page exits to the VMM and a
+ *                host write raises SIGBUS; both land once the page is
+ *                writable again.
+ *  5. Window:    a host mapping of the guest_memfd faults after the range it
+ *                covers is donated.
+ *
+ * Requires samples/kvm/mem_provider_sample.ko and CONFIG_IOMMUFD_TEST.
+ */
+#include <fcntl.h>
+#include <errno.h>
+#include <setjmp.h>
+#include <signal.h>
+#include <stdint.h>
+#include <stdio.h>
+#include <string.h>
+#include <unistd.h>
+#include <sys/ioctl.h>
+#include <sys/mman.h>
+
+#include <linux/iommufd.h>
+
+#include "test_util.h"
+#include "kvm_util.h"
+#include "processor.h"
+
+/*
+ * The iommufd mock domain test interface.  The iommufd selftest helpers
+ * bring their own harness, which does not mix with the KVM selftest library.
+ */
+#include "../../../../../drivers/iommu/iommufd/iommufd_test.h"
+
+/* Mirrors samples/kvm/mem_provider_sample.h */
+#define MPS_FLAG_MMAP_CAPABLE	(1u << 0)
+
+struct mps_new_child { __u32 flags, pad; __u64 offset, len; };
+struct mps_move { __s32 src_fd, dst_fd; __u64 offset, len; };
+struct mps_donate { __s32 fd; __u32 pad; __u64 offset, len; };
+struct mps_ctl_range { __s32 fd; __u32 value; __u64 offset, len; };
+struct mps_stats {
+	__u64 region_offset, region_len, owned_pages, absent_pages, readonly_pages;
+};
+
+#define MPS_NEW_CHILD		_IOW('P', 5, struct mps_new_child)
+#define MPS_MOVE		_IOW('P', 6, struct mps_move)
+#define MPS_DONATE		_IOW('P', 7, struct mps_donate)
+#define MPS_RECLAIM		_IOW('P', 8, struct mps_donate)
+#define MPS_GET_STATS		_IOR('P', 10, struct mps_stats)
+#define MPS_CTL_SET_READONLY	_IOW('P', 11, struct mps_ctl_range)
+
+#define PAGE		0x1000ULL
+#define CHILD_SIZE	0x400000ULL		/* 4 MiB per child */
+#define GPA		(1ULL << 32)		/* each VM maps its child here */
+#define IOVA		(1ULL << 28)		/* inside the mock aperture */
+#define MAGIC_A		0xa11ce000a11ce000ULL
+#define MAGIC_B		0xb0bb0bb0b0bb0bb0ULL
+
+/*
+ * The ranges of A and B in the root overlap on [SHARED_OFF, +SHARED_LEN).
+ * A is created first and has its whole range, B is created second and has
+ * its range without the overlap.  The overlap can then move from A to B.
+ */
+#define A_OFF		0ULL
+#define SHARED_LEN	(4 * PAGE)
+#define B_OFF		(CHILD_SIZE - SHARED_LEN)
+#define SHARED_OFF	B_OFF
+#define ROOT_SIZE	(B_OFF + CHILD_SIZE)
+
+#define A_SHARED_GPA	(GPA + (SHARED_OFF - A_OFF))
+#define B_SHARED_GPA	(GPA + (SHARED_OFF - B_OFF))
+#define A_SHARED_IOVA	(IOVA + (SHARED_OFF - A_OFF))
+#define B_SHARED_IOVA	(IOVA + (SHARED_OFF - B_OFF))
+
+/* A page of A that never moves. */
+#define RO_OFF		(64 * PAGE)
+#define RO_GPA		(GPA + RO_OFF)
+
+/* Pages of A that are donated and reclaimed. */
+#define DON_OFF		(8 * PAGE)
+#define DON_GPA		(GPA + DON_OFF)
+#define DON_IOVA	(IOVA + DON_OFF)
+
+/* A page that B has from the start. */
+#define B_HOME_OFF	(CHILD_SIZE / 2)
+#define B_HOME_GPA	(GPA + B_HOME_OFF)
+
+struct guest_args {
+	uint64_t write_gpa;	/* 0 to skip */
+	uint64_t write_val;
+	uint64_t read_gpa;	/* 0 to skip; the value goes to GUEST_SYNC */
+};
+
+static void guest_code(struct guest_args *a)
+{
+	if (a->write_gpa)
+		*(volatile uint64_t *)a->write_gpa = a->write_val;
+	if (a->read_gpa)
+		GUEST_SYNC(*(volatile uint64_t *)a->read_gpa);
+	GUEST_DONE();
+}
+
+/* ---- Control process ---------------------------------------------------- */
+
+static int ctl;
+
+static int ctl_new_child(uint64_t off, uint64_t len)
+{
+	struct mps_new_child nc = {
+		.flags = MPS_FLAG_MMAP_CAPABLE,
+		.offset = off, .len = len,
+	};
+	int fd = ioctl(ctl, MPS_NEW_CHILD, &nc);
+
+	TEST_ASSERT(fd >= 0, "NEW_CHILD(%#llx, %#llx) errno=%d",
+		    (unsigned long long)off, (unsigned long long)len, errno);
+	return fd;
+}
+
+static int ctl_move(int src, int dst, uint64_t off, uint64_t len)
+{
+	struct mps_move mv = { .src_fd = src, .dst_fd = dst,
+			       .offset = off, .len = len };
+
+	return ioctl(ctl, MPS_MOVE, &mv) ? -errno : 0;
+}
+
+static void ctl_donate(int fd, uint64_t off, uint64_t len, bool reclaim)
+{
+	struct mps_donate d = { .fd = fd, .offset = off, .len = len };
+
+	TEST_ASSERT(!ioctl(ctl, reclaim ? MPS_RECLAIM : MPS_DONATE, &d),
+		    "%s errno=%d", reclaim ? "RECLAIM" : "DONATE", errno);
+}
+
+static void ctl_readonly(int fd, uint64_t off, bool ro)
+{
+	struct mps_ctl_range cr = { .fd = fd, .value = ro, .offset = off,
+				    .len = PAGE };
+
+	TEST_ASSERT(!ioctl(ctl, MPS_CTL_SET_READONLY, &cr),
+		    "CTL_SET_READONLY errno=%d", errno);
+}
+
+/* ---- VMM ---------------------------------------------------------------- */
+
+struct vmm {
+	const char *name;
+	int child;		/* the provider file */
+	int gmem;		/* the guest_memfd backed by it */
+	int iommufd;
+	uint32_t ioas, stdev, hwpt;
+	struct kvm_vm *vm;
+	struct kvm_vcpu *vcpu;
+	void *hva;		/* host mapping of the whole guest_memfd */
+	gva_t args_gva;
+};
+
+static void vmm_create(struct vmm *v)
+{
+	v->vm = vm_create_with_one_vcpu(&v->vcpu, guest_code);
+	v->args_gva = vm_alloc_page(v->vm);
+}
+
+static bool vmm_iova_mapped(struct vmm *v, uint64_t iova, uint64_t len,
+			    bool mapped)
+{
+	struct iommu_test_cmd cmd = {
+		.size = sizeof(cmd), .op = IOMMU_TEST_OP_MD_CHECK_MAPPED,
+		.id = v->hwpt,
+		.check_mapped = { .mapped = mapped, .iova = iova,
+				  .length = len },
+	};
+
+	return !ioctl(v->iommufd, IOMMU_TEST_CMD, &cmd);
+}
+
+static uint64_t vmm_iova_phys(struct vmm *v, uint64_t iova)
+{
+	struct iommu_test_cmd cmd = {
+		.size = sizeof(cmd), .op = IOMMU_TEST_OP_MD_IOVA_TO_PHYS,
+		.id = v->hwpt,
+		.iova_to_phys = { .iova = iova },
+	};
+
+	if (ioctl(v->iommufd, IOMMU_TEST_CMD, &cmd))
+		return 0;
+	return cmd.iova_to_phys.out_phys;
+}
+
+/* Give the child file to KVM and to iommufd. */
+static void vmm_attach(struct vmm *v, int child_fd)
+{
+	struct iommu_ioas_alloc alloc = { .size = sizeof(alloc) };
+	struct kvm_create_guest_memfd gm = {
+		.size = CHILD_SIZE,
+		.flags = GUEST_MEMFD_FLAG_MMAP | GUEST_MEMFD_FLAG_USE_PROVIDER,
+		.provider_fd = child_fd,
+	};
+	struct iommu_test_cmd mock = {
+		.size = sizeof(mock), .op = IOMMU_TEST_OP_MOCK_DOMAIN,
+	};
+	struct iommu_ioas_map_file map = {
+		.size = sizeof(map),
+		.flags = IOMMU_IOAS_MAP_FIXED_IOVA | IOMMU_IOAS_MAP_READABLE |
+			 IOMMU_IOAS_MAP_WRITEABLE,
+		.fd = child_fd, .start = 0, .length = CHILD_SIZE, .iova = IOVA,
+	};
+	int r;
+
+	v->child = child_fd;
+
+	/* KVM: a guest_memfd backed by the child, behind one memslot. */
+	v->gmem = __vm_ioctl(v->vm, KVM_CREATE_GUEST_MEMFD, &gm);
+	TEST_ASSERT(v->gmem >= 0, "%s: KVM_CREATE_GUEST_MEMFD errno=%d",
+		    v->name, errno);
+	v->hva = mmap(NULL, CHILD_SIZE, PROT_READ | PROT_WRITE, MAP_SHARED,
+		      v->gmem, 0);
+	TEST_ASSERT(v->hva != MAP_FAILED, "%s: mmap(guest_memfd) errno=%d",
+		    v->name, errno);
+	r = __vm_set_user_memory_region2(v->vm, 10, KVM_MEM_GUEST_MEMFD, GPA,
+					 CHILD_SIZE, v->hva, v->gmem, 0);
+	TEST_ASSERT(!r, "%s: SET_USER_MEMORY_REGION2 errno=%d", v->name,
+		    errno);
+	virt_map(v->vm, GPA, GPA, CHILD_SIZE / PAGE);
+
+	/* iommufd: the same child file, in an IOAS on a mock domain. */
+	v->iommufd = open("/dev/iommu", O_RDWR);
+	__TEST_REQUIRE(v->iommufd >= 0, "iommufd unavailable");
+	TEST_ASSERT(!ioctl(v->iommufd, IOMMU_IOAS_ALLOC, &alloc),
+		    "IOAS_ALLOC errno=%d", errno);
+	v->ioas = alloc.out_ioas_id;
+
+	mock.id = v->ioas;
+	__TEST_REQUIRE(!ioctl(v->iommufd, IOMMU_TEST_CMD, &mock),
+		       "no iommufd mock domain (CONFIG_IOMMUFD_TEST?)");
+	v->stdev = mock.mock_domain.out_stdev_id;
+	v->hwpt = mock.mock_domain.out_hwpt_id;
+
+	map.ioas_id = v->ioas;
+	TEST_ASSERT(!ioctl(v->iommufd, IOMMU_IOAS_MAP_FILE, &map),
+		    "%s: IOAS_MAP_FILE(provider) errno=%d", v->name, errno);
+}
+
+/*
+ * Continue the guest.  KVM_RUN fails with EFAULT for KVM_EXIT_MEMORY_FAULT,
+ * which several scenarios expect; the expect_*() helpers check the exit.
+ */
+static void vmm_resume(struct vmm *v)
+{
+	int r = _vcpu_run(v->vcpu);
+
+	TEST_ASSERT(!r || (errno == EFAULT &&
+			   v->vcpu->run->exit_reason == KVM_EXIT_MEMORY_FAULT),
+		    "%s: KVM_RUN r=%d errno=%d exit=%s", v->name, r, errno,
+		    exit_reason_str(v->vcpu->run->exit_reason));
+}
+
+static void vmm_run(struct vmm *v, uint64_t write_gpa, uint64_t write_val,
+		    uint64_t read_gpa)
+{
+	struct guest_args *a = addr_gva2hva(v->vm, v->args_gva);
+
+	a->write_gpa = write_gpa;
+	a->write_val = write_val;
+	a->read_gpa = read_gpa;
+	vcpu_arch_set_entry_point(v->vcpu, guest_code);
+	vcpu_args_set(v->vcpu, 1, v->args_gva);
+	vmm_resume(v);
+}
+
+static void expect_done(struct vmm *v)
+{
+	struct kvm_run *run = v->vcpu->run;
+	struct ucall uc;
+
+	TEST_ASSERT(run->exit_reason != KVM_EXIT_MEMORY_FAULT,
+		    "%s: unexpected memory fault at gpa %#llx", v->name,
+		    (unsigned long long)run->memory_fault.gpa);
+	TEST_ASSERT(get_ucall(v->vcpu, &uc) == UCALL_DONE, "%s: guest exit %s",
+		    v->name, exit_reason_str(run->exit_reason));
+}
+
+static uint64_t expect_sync_then_done(struct vmm *v)
+{
+	struct ucall uc;
+	uint64_t val;
+
+	TEST_ASSERT(get_ucall(v->vcpu, &uc) == UCALL_SYNC, "%s: guest exit %s",
+		    v->name, exit_reason_str(v->vcpu->run->exit_reason));
+	val = uc.args[1];
+	vmm_resume(v);
+	expect_done(v);
+	return val;
+}
+
+static void expect_memory_fault(struct vmm *v, uint64_t gpa)
+{
+	struct kvm_run *run = v->vcpu->run;
+
+	TEST_ASSERT(run->exit_reason == KVM_EXIT_MEMORY_FAULT,
+		    "%s: want KVM_EXIT_MEMORY_FAULT, got %s", v->name,
+		    exit_reason_str(run->exit_reason));
+	TEST_ASSERT(run->memory_fault.gpa == gpa,
+		    "%s: fault at gpa %#llx, want %#llx", v->name,
+		    (unsigned long long)run->memory_fault.gpa,
+		    (unsigned long long)gpa);
+}
+
+static void vmm_stats(struct vmm *v, struct mps_stats *st)
+{
+	TEST_ASSERT(!ioctl(v->child, MPS_GET_STATS, st),
+		    "%s: GET_STATS errno=%d", v->name, errno);
+}
+
+static void vmm_destroy(struct vmm *v)
+{
+	close(v->iommufd);
+	kvm_vm_free(v->vm);
+	munmap(v->hva, CHILD_SIZE);
+	close(v->gmem);
+	close(v->child);
+}
+
+static sigjmp_buf host_jmp;
+
+static void host_sig(int sig)
+{
+	siglongjmp(host_jmp, sig);
+}
+
+/*
+ * Read @addr, or write @val to it, through a host mapping.  Return the
+ * signal that the access raised, or 0.
+ */
+static int host_access(volatile uint64_t *addr, bool write, uint64_t val)
+{
+	struct sigaction sa = { .sa_handler = host_sig }, old_bus, old_segv;
+	int sig;
+
+	sigaction(SIGBUS, &sa, &old_bus);
+	sigaction(SIGSEGV, &sa, &old_segv);
+	sig = sigsetjmp(host_jmp, 1);
+	if (!sig) {
+		if (write)
+			*addr = val;
+		else
+			(void)*addr;
+	}
+	sigaction(SIGBUS, &old_bus, NULL);
+	sigaction(SIGSEGV, &old_segv, NULL);
+	return sig;
+}
+
+#define host_read_faults(p)	(host_access(p, false, 0) == SIGBUS)
+#define host_write_faults(p, v)	(host_access(p, true, v) == SIGBUS)
+
+/* ---- Scenarios ---------------------------------------------------------- */
+
+static void scenario_launch(struct vmm *a, struct vmm *b)
+{
+	pr_info("1. launch\n");
+	vmm_run(a, GPA, MAGIC_A, 0);
+	expect_done(a);
+	vmm_run(b, B_HOME_GPA, MAGIC_B, 0);
+	expect_done(b);
+	TEST_ASSERT(*(volatile uint64_t *)a->hva == MAGIC_A,
+		    "A: the guest write is not visible to the host");
+	TEST_ASSERT(*(volatile uint64_t *)(b->hva + B_HOME_OFF) == MAGIC_B,
+		    "B: the guest write is not visible to the host");
+
+	/* A has its whole range; B lacks the shared window. */
+	TEST_ASSERT(vmm_iova_mapped(a, IOVA, CHILD_SIZE, true),
+		    "A: device mapping incomplete");
+	TEST_ASSERT(vmm_iova_mapped(b, B_SHARED_IOVA, SHARED_LEN, false),
+		    "B: device maps a window that B does not have");
+	TEST_ASSERT(vmm_iova_mapped(b, B_SHARED_IOVA + SHARED_LEN,
+				    CHILD_SIZE - SHARED_LEN, true),
+		    "B: device mapping incomplete");
+}
+
+static void scenario_move(struct vmm *a, struct vmm *b)
+{
+	struct mps_stats st;
+	uint64_t frame;
+	int r;
+
+	pr_info("2. move A -> B\n");
+	vmm_run(a, A_SHARED_GPA, MAGIC_A, 0);
+	expect_done(a);
+	frame = vmm_iova_phys(a, A_SHARED_IOVA);
+	TEST_ASSERT(frame, "A: window not mapped before the move");
+	/* Map the window on the host too, so that the move must zap it. */
+	TEST_ASSERT(*(volatile uint64_t *)(a->hva + SHARED_OFF - A_OFF) ==
+		    MAGIC_A, "A: host mapping does not see the guest write");
+
+	/* B does not have the window yet. */
+	vmm_run(b, 0, 0, B_SHARED_GPA);
+	expect_memory_fault(b, B_SHARED_GPA);
+
+	/* A move outside the destination's range is refused. */
+	r = ctl_move(a->child, b->child, RO_OFF, PAGE);
+	TEST_ASSERT(r == -EINVAL, "MOVE outside B's range returned %d", r);
+
+	r = ctl_move(a->child, b->child, SHARED_OFF, SHARED_LEN);
+	TEST_ASSERT(!r, "MOVE returned %d", r);
+
+	vmm_stats(a, &st);
+	TEST_ASSERT(st.owned_pages == (CHILD_SIZE - SHARED_LEN) / PAGE,
+		    "A has %llu pages after the move",
+		    (unsigned long long)st.owned_pages);
+	vmm_stats(b, &st);
+	TEST_ASSERT(st.owned_pages == CHILD_SIZE / PAGE,
+		    "B has %llu pages after the move",
+		    (unsigned long long)st.owned_pages);
+
+	/* A's device lost the window and only the window. */
+	TEST_ASSERT(vmm_iova_mapped(a, A_SHARED_IOVA, SHARED_LEN, false),
+		    "A: moved window still mapped");
+	TEST_ASSERT(vmm_iova_mapped(a, IOVA, CHILD_SIZE - SHARED_LEN, true),
+		    "A: the rest of the mapping went too");
+
+	/* A's guest and A's host mapping fault on the window. */
+	vmm_run(a, 0, 0, A_SHARED_GPA);
+	expect_memory_fault(a, A_SHARED_GPA);
+	TEST_ASSERT(host_read_faults(a->hva + SHARED_OFF - A_OFF),
+		    "A: host mapping of the moved window still readable");
+
+	/* B reads what A wrote, and B's device reaches A's frame. */
+	vmm_run(b, 0, 0, B_SHARED_GPA);
+	TEST_ASSERT(expect_sync_then_done(b) == MAGIC_A,
+		    "B does not see A's write");
+	TEST_ASSERT(*(volatile uint64_t *)(b->hva + SHARED_OFF - B_OFF) ==
+		    MAGIC_A, "B: host mapping does not see A's write");
+	TEST_ASSERT(vmm_iova_phys(b, B_SHARED_IOVA) == frame,
+		    "B: window is not on the frame A had");
+}
+
+static void scenario_donate(struct vmm *a)
+{
+	struct mps_stats st;
+
+	pr_info("3. donate and reclaim\n");
+	/* Map the page on the host, so that the donation must zap it. */
+	TEST_ASSERT(!host_access(a->hva + DON_OFF, false, 0),
+		    "A: host mapping faults before the donation");
+	ctl_donate(a->child, DON_OFF, 2 * PAGE, false);
+	vmm_stats(a, &st);
+	TEST_ASSERT(st.owned_pages == (CHILD_SIZE - SHARED_LEN) / PAGE - 2,
+		    "A has %llu pages after the donation",
+		    (unsigned long long)st.owned_pages);
+	TEST_ASSERT(vmm_iova_mapped(a, DON_IOVA, 2 * PAGE, false),
+		    "A: donated pages still mapped for the device");
+	TEST_ASSERT(host_read_faults(a->hva + DON_OFF),
+		    "A: host mapping of a donated page still readable");
+	vmm_run(a, 0, 0, DON_GPA);
+	expect_memory_fault(a, DON_GPA);
+
+	ctl_donate(a->child, DON_OFF, 2 * PAGE, true);
+	vmm_stats(a, &st);
+	TEST_ASSERT(st.owned_pages == (CHILD_SIZE - SHARED_LEN) / PAGE,
+		    "A has %llu pages after the reclaim",
+		    (unsigned long long)st.owned_pages);
+	TEST_ASSERT(vmm_iova_mapped(a, DON_IOVA, 2 * PAGE, true),
+		    "A: reclaimed pages not mapped again for the device");
+	TEST_ASSERT(!host_access(a->hva + DON_OFF, false, 0),
+		    "A: host mapping of a reclaimed page faults");
+	vmm_run(a, 0, 0, DON_GPA);
+	expect_sync_then_done(a);
+}
+
+static void scenario_readonly(struct vmm *a)
+{
+	volatile uint64_t *page = (volatile uint64_t *)(a->hva + RO_OFF);
+
+	pr_info("4. read only\n");
+	*page = MAGIC_B;
+
+	ctl_readonly(a->child, RO_OFF, true);
+	TEST_ASSERT(vmm_iova_mapped(a, IOVA + RO_OFF, PAGE, true),
+		    "A: read-only page not mapped for the device");
+
+	vmm_run(a, RO_GPA, MAGIC_A, 0);
+	expect_memory_fault(a, RO_GPA);
+	TEST_ASSERT(host_write_faults(page, MAGIC_A),
+		    "A: host write to a read-only page did not fault");
+	TEST_ASSERT(*page == MAGIC_B, "write to a read-only page landed");
+
+	/* Clear the bit; the guest retries the same write. */
+	ctl_readonly(a->child, RO_OFF, false);
+	vmm_resume(a);
+	expect_done(a);
+	TEST_ASSERT(*page == MAGIC_A, "write after clearing read only lost");
+	TEST_ASSERT(!host_write_faults(page, MAGIC_B),
+		    "A: host write after clearing read only faulted");
+	TEST_ASSERT(*page == MAGIC_B, "host write after clearing read only lost");
+}
+
+static void scenario_window(struct vmm *a)
+{
+	volatile uint64_t *win;
+
+	pr_info("5. host mapping\n");
+	win = mmap(NULL, PAGE, PROT_READ | PROT_WRITE, MAP_SHARED, a->gmem,
+		   DON_OFF);
+	TEST_ASSERT(win != MAP_FAILED, "mmap(guest_memfd, page) errno=%d",
+		    errno);
+	*win = MAGIC_A;
+	TEST_ASSERT(*(volatile uint64_t *)(a->hva + DON_OFF) == MAGIC_A,
+		    "the two host mappings disagree");
+
+	ctl_donate(a->child, DON_OFF, PAGE, false);
+
+	TEST_ASSERT(host_read_faults(win),
+		    "host mapping still readable after the donation");
+
+	ctl_donate(a->child, DON_OFF, PAGE, true);
+	TEST_ASSERT(*win == MAGIC_A, "host mapping did not come back");
+	munmap((void *)win, PAGE);
+}
+
+int main(void)
+{
+	struct vmm a = { .name = "A" }, b = { .name = "B" };
+	struct mps_new_child probe = { .offset = ROOT_SIZE - PAGE, .len = PAGE };
+	int fd;
+
+	TEST_REQUIRE(kvm_check_cap(KVM_CAP_GUEST_MEMFD_FLAGS) &
+		     GUEST_MEMFD_FLAG_USE_PROVIDER);
+
+	ctl = open("/dev/mem_provider_sample", O_RDWR);
+	__TEST_REQUIRE(ctl >= 0, "mem_provider_sample not loaded");
+
+	/*
+	 * Without addr= and len=, the first NEW_CHILD sizes the root, so ask
+	 * for its last page first and give it back.  A fixed root that is too
+	 * small makes the test skip.
+	 */
+	fd = ioctl(ctl, MPS_NEW_CHILD, &probe);
+	__TEST_REQUIRE(fd >= 0, "provider root smaller than %#llx bytes",
+		       (unsigned long long)ROOT_SIZE);
+	close(fd);
+
+	vmm_create(&a);
+	vmm_create(&b);
+	vmm_attach(&a, ctl_new_child(A_OFF, CHILD_SIZE));
+	vmm_attach(&b, ctl_new_child(B_OFF, CHILD_SIZE));
+
+	scenario_launch(&a, &b);
+	scenario_move(&a, &b);
+	scenario_donate(&a);
+	scenario_readonly(&a);
+	scenario_window(&a);
+
+	vmm_destroy(&a);
+	vmm_destroy(&b);
+	close(ctl);
+	pr_info("mem_provider: all scenarios passed\n");
+	return 0;
+}


^ permalink raw reply related	[flat|nested] 10+ messages in thread

end of thread, other threads:[~2026-10-06 18:33 UTC | newest]

Thread overview: 10+ messages (download: mbox.gz follow: Atom feed
-- links below jump to the message on this page --
     [not found] <20260720111259.122911-1-dwmw2@infradead.org>
2026-10-06 18:32 ` [RFC PATCH 0/9] mm: Memory providers for guest_memfd and iommufd Fred Griffoul
2026-10-06 18:32   ` [PATCH 1/9] KVM: guest_memfd: Add a writable result to get_pfn() Fred Griffoul
2026-10-06 18:32   ` [PATCH 2/9] mm: Add memory providers Fred Griffoul
2026-10-06 18:32   ` [PATCH 3/9] KVM: guest_memfd: Add a memory provider backing Fred Griffoul
2026-10-06 18:32   ` [PATCH 4/9] iommufd: Track the domains of pages that are not pinned Fred Griffoul
2026-10-06 18:32   ` [PATCH 5/9] iommufd: Map memory provider files Fred Griffoul
2026-10-06 18:32   ` [PATCH 6/9] iommufd/selftest: Add mock-domain IOVA queries Fred Griffoul
2026-10-06 18:32   ` [PATCH 7/9] iommufd/selftest: Add a mock memory provider Fred Griffoul
2026-10-06 18:32   ` [PATCH 8/9] samples/kvm: Add a memory provider sample Fred Griffoul
2026-10-06 18:32   ` [PATCH 9/9] KVM: selftests: Test a memory provider shared by KVM and iommufd Fred Griffoul

This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox