All of lore.kernel.org
 help / color / mirror / Atom feed
* [PATCH 0/7] x86: Address Space Isolation, part 1: per-domain area mapping rework
@ 2026-08-20 17:43 George Dunlap
  2026-08-20 17:43 ` [PATCH 1/7] x86/mm: allocate the per-domain page-tables from the xenheap George Dunlap
                   ` (7 more replies)
  0 siblings, 8 replies; 17+ messages in thread
From: George Dunlap @ 2026-08-20 17:43 UTC (permalink / raw)
  To: xen-devel
  Cc: Jan Beulich, Andrew Cooper, Roger Pau Monné, Teddy Astie,
	Anthony PERARD, Michal Orzel, Julien Grall, Stefano Stabellini

This is the first batch of patches continuing the x86 Address Space
Isolation (ASI) work that Roger posted as "x86: adventures in Address
Space Isolation" (v1 [1], v2 [2]).  I've taken over finishing it up
and getting it upstream.

Rather than re-posting the whole stack (nearly 60 patches) each time,
I'd like to run this as a rolling series: post a reviewable slice from
the front, drop patches as they are committed, and append the next
ones as they mature.  Each batch should stand on its own; the cover
letter of each will say where it sits in the larger picture.

A "map" of the entire series -- grouped into logical chunks, with the
dependencies between patches -- is maintained here:

https://xenbits.xenproject.org/people/gdunlap/asi-series-deps.html

Note that the graph above is a work in progress; dependency lines may
change as more of the series is vetted.  Note also that the full
series includes a design doc as patch 1; that's not ready for
publication yet, so patches 1-7 of this series correspond to nodes 2-8
of the graph.

The problem this slice addresses: the per-domain area already has
central machinery for building its page-tables and for managing the
backing pages it owns itself (create_perdomain_mapping() and friends);
what it lacks is a way to install a caller's own pages at a chosen
address.  The PV GDT and LDT code fills that gap privately, by having
create_perdomain_mapping() hand back aliases of the L1 tables it
builds and stashing them in d->arch.pv.gdt_ldt_l1tab; mapping updates
are then written through the stash, bypassing the interface.  That
arrangement assumes a single, domain-wide set of per-domain
page-tables, which stops holding once the per-domain area becomes
per-vCPU (the next slices).

This slice closes the gap centrally:

 - Patch 1 moves the per-domain page-table allocations from the
   domheap to the xenheap.  The page-tables (not the data pages they
   map) are then reachable through their always-mapped alias from any
   context, so walking them needs no mapcache -- including from the
   context switch.
 - Patch 2 introduces populate_perdomain_mapping() on top: a single
   writer for the per-domain area, installing caller-owned pages at a
   chosen address by walking the always-mapped page-tables.
 - Patches 3-5 convert the Xen-GDT slot, the guest GDT, and the guest
   LDT paths to it.
 - Patch 6 removes the stash.  Patch 7 simplifies
   create_perdomain_mapping(), whose L1-capture mode existed only to
   build the stash.

One point reviewers may want to look at specifically: patch 1 changes
where the per-domain page-tables are allocated from, and its commit
message discusses the (minor) NUMA-placement consequence.

Relative to v2: patch 1 is new -- v2 kept the page-tables in the
domheap and walked them through map_domain_page(), with a linear-map
fast path for the currently-running vCPU; making the page-tables
always-mapped lets one plain walk serve every caller and context, and
the context switch keeps its existing structure.  The populate patch
is split from its first user; the LDT demand-map now goes through
populate_perdomain_mapping() rather than writing linear entries;
pv_destroy_gdt() keeps mapping torn-down slots read-only to the zero
page (in v2 they became empty -- a guest-visible partial revert of
cf6d39f819); and the domain -> vCPU parameter switches move to the
next slice.  Per-patch changes are noted below each patch's "---".

Testing:
 - Each patch builds (x86_64, CONFIG_DEBUG=y); tier-1 qemu boot at
   the tip.
 - The series passes the Xen GitLab CI pipeline, including the
   hardware runner:
    https://gitlab.com/xen-project/hardware/xen-staging/-/pipelines/2776674713
 - On an Intel NUC, debug build: XTF pv64 and pv32pae suites (the
   latter with cet=no-shstk,no-ibt pv=32, since CET disables PV32);
   plus an LDT exerciser in a PV Linux guest (modify_ldt() with 1-16
   page LDTs, demand-faulting every page, LAR beyond the limit,
   shrinking, teardown; also with the guest's vCPUs bounced across
   pCPUs) -- thousands of rounds, no assertions or "unable to map"
   reports.

[1] https://lore.kernel.org/xen-devel/20240726152206.28411-1-roger.pau@citrix.com/
[2] https://lore.kernel.org/xen-devel/20250108142659.99490-1-roger.pau@citrix.com/

George Dunlap (1):
  x86/mm: allocate the per-domain page-tables from the xenheap

Roger Pau Monné (6):
  x86/mm: introduce populate_perdomain_mapping()
  x86/pv: use populate_perdomain_mapping() to map the Xen GDT
  x86/pv: set/clear guest GDT mappings using
    populate_perdomain_mapping()
  x86/pv: update guest LDT mappings using
    {populate,destroy}_perdomain_mapping()
  x86/pv: remove stashing of GDT/LDT L1 page-tables
  x86/mm: simplify create_perdomain_mapping() interface

 xen/arch/x86/domain.c               |  17 ++-
 xen/arch/x86/domain_page.c          |  10 +-
 xen/arch/x86/hvm/hvm.c              |   2 +-
 xen/arch/x86/include/asm/desc.h     |   6 +-
 xen/arch/x86/include/asm/domain.h   |  14 +-
 xen/arch/x86/include/asm/mm.h       |   9 +-
 xen/arch/x86/mm.c                   | 225 +++++++++++++++++-----------
 xen/arch/x86/pv/descriptor-tables.c |  57 ++++---
 xen/arch/x86/pv/domain.c            |  16 +-
 xen/arch/x86/pv/mm.c                |  16 +-
 xen/arch/x86/smpboot.c              |  14 +-
 xen/arch/x86/traps.c                |   4 +-
 xen/arch/x86/x86_64/mm.c            |   3 +-
 13 files changed, 221 insertions(+), 172 deletions(-)


base-commit: 669f8c502aeaa538f8407013ba04461e71361ed4
-- 
2.55.0



^ permalink raw reply	[flat|nested] 17+ messages in thread

* [PATCH 1/7] x86/mm: allocate the per-domain page-tables from the xenheap
  2026-08-20 17:43 [PATCH 0/7] x86: Address Space Isolation, part 1: per-domain area mapping rework George Dunlap
@ 2026-08-20 17:43 ` George Dunlap
  2026-08-20 17:43 ` [PATCH 2/7] x86/mm: introduce populate_perdomain_mapping() George Dunlap
                   ` (6 subsequent siblings)
  7 siblings, 0 replies; 17+ messages in thread
From: George Dunlap @ 2026-08-20 17:43 UTC (permalink / raw)
  To: xen-devel
  Cc: Jan Beulich, Andrew Cooper, Roger Pau Monné, Teddy Astie,
	Anthony PERARD, Michal Orzel, Julien Grall, Stefano Stabellini

The per-domain area's L3, L2 and L1 page-tables are currently
domain-heap pages with no owner, and accessed through
map_domain_page() whenever they need editing.  That has two costs.
First, every edit not going through the linear mappings of the loaded
page-tables goes through map_domain_page() per level.  Even now that is
sometimes a mapcache round trip; once the direct map becomes on-demand
it always will be.  Second, the mapcache is not usable everywhere;
specifically, it cannot be called in the context switch before the
incoming vCPU's page-tables are loaded.

Currently the Xen slot of a PV vcpu's full GDT is written during context
switch; that works by special-casing the GDT/LDT L1 tables into the
xenheap and stashing their addresses in d->arch.pv.gdt_ldt_l1tab.
Later in the series the mappings of the guest's root page-table (for
the XPTI root sync) and of the CPU's own stack (for per-CPU stack
isolation) need writing at context switch too, and per-vCPU roots add
writes on the PV kernel/user switch and new-CR3 paths.

Instead of adding more ad-hoc pointers, or arranging to be able to
call map_domain_page() from within a context switch, allocate *all* of
the per-domain page-tables from the xenheap.  Keep a pointer directly
to the L3 page within the xenheap, and when walking the page-tables,
use the MFN in the entry to reconstruct the virtual address of the
xenheap page directly.  Xenheap pages are mapped in every context, so
the tables can be edited from anywhere through their always-mapped
alias: no mapcache, no special case for the context switch.  Under the
planned on-demand direct map, xenheap pages remain mapped at that
alias for their lifetime, so this stays true.

It may seem strange for a series whose end goal is to move things out
of global mappings to start by requiring a further class of pages to
stay globally mapped.  The series is about protecting *guest* data.
The pages in question here hold page-table entries only -- MFNs and
flags, reachable through the linear mappings whenever the tables are
loaded anyway -- not guest data; XPTI's per-CPU root page-tables are
xenheap pages by the same reasoning.  The pages mapped *by* these
tables (GDT/LDT frames, the mapcache's targets, the compat
argument-translation area) are unaffected and stay domain-heap pages.

The "capture" mode of create_perdomain_mapping() no longer needs to take
its L1s from a different heap than the others; it now only records the
pointers.  The is_xen_heap_page() check in free_perdomain_mappings() goes
away with it.  No change for callers.

Two side effects to note.  First, the tables are subject to the
xenheap allocation limit (within the PV-visible direct map), as the
stashed L1s, the GDTs and XPTI's root page-tables already are -- a few
pages per vcpu at most.

Second, the xenheap allocator takes no domain, so there is no
round-robin over the domain's node affinity: so we place the tables
with MEMF_node(domain_to_node(d)), on the node of the domain's first
vcpu (or the allocating CPU's node before it exists), as the stashed
L1s already were.

Assisted-by: Claude Code:claude-fable-5
Signed-off-by: George Dunlap <gwd@xenproject.org>
---
 xen/arch/x86/domain.c             |   4 +-
 xen/arch/x86/include/asm/domain.h |   3 +-
 xen/arch/x86/mm.c                 | 121 ++++++++++--------------------
 3 files changed, 45 insertions(+), 83 deletions(-)

diff --git a/xen/arch/x86/domain.c b/xen/arch/x86/domain.c
index 996b50af7a..163a2c97ae 100644
--- a/xen/arch/x86/domain.c
+++ b/xen/arch/x86/domain.c
@@ -2015,8 +2015,8 @@ void cf_check paravirt_ctxt_switch_to(struct vcpu *v)
 
     if ( root_pgt )
         root_pgt[root_table_offset(PERDOMAIN_VIRT_START)] =
-            l4e_from_page(v->domain->arch.perdomain_l3_pg,
-                          __PAGE_HYPERVISOR_RW);
+            l4e_from_paddr(__pa(v->domain->arch.perdomain_l3),
+                           __PAGE_HYPERVISOR_RW);
 
     if ( unlikely(v->arch.dr7 & DR7_ACTIVE_MASK) )
         activate_debugregs(v);
diff --git a/xen/arch/x86/include/asm/domain.h b/xen/arch/x86/include/asm/domain.h
index 2d0a915410..5275bb10ea 100644
--- a/xen/arch/x86/include/asm/domain.h
+++ b/xen/arch/x86/include/asm/domain.h
@@ -332,7 +332,8 @@ struct monitor_write_data {
 
 struct arch_domain
 {
-    struct page_info *perdomain_l3_pg;
+    /* Xenheap page: the per-domain page-tables are always mapped. */
+    l3_pgentry_t *perdomain_l3;
 
     /* I/O-port admin-specified access capabilities. */
     struct rangeset *ioport_caps;
diff --git a/xen/arch/x86/mm.c b/xen/arch/x86/mm.c
index b158742408..38a0f984fc 100644
--- a/xen/arch/x86/mm.c
+++ b/xen/arch/x86/mm.c
@@ -1679,7 +1679,7 @@ void init_xen_l4_slots(l4_pgentry_t *l4t, mfn_t l4mfn,
 
     /* Slot 260: Per-domain mappings. */
     l4t[l4_table_offset(PERDOMAIN_VIRT_START)] =
-        l4e_from_page(d->arch.perdomain_l3_pg, __PAGE_HYPERVISOR_RW);
+        l4e_from_mfn(virt_to_mfn(d->arch.perdomain_l3), __PAGE_HYPERVISOR_RW);
 
     /* Slot 4: Per-domain mappings mirror. */
     BUILD_BUG_ON(IS_ENABLED(CONFIG_PV32) &&
@@ -6219,54 +6219,47 @@ int create_perdomain_mapping(struct domain *d, unsigned long va,
     l3_pgentry_t *l3tab;
     l2_pgentry_t *l2tab;
     l1_pgentry_t *l1tab;
+    unsigned int memflags = MEMF_node(domain_to_node(d));
     int rc = 0;
 
     ASSERT(va >= PERDOMAIN_VIRT_START &&
            va < PERDOMAIN_VIRT_SLOT(PERDOMAIN_SLOTS));
 
-    if ( !d->arch.perdomain_l3_pg )
+    /*
+     * The per-domain page-tables come from the xenheap, so that they can be
+     * edited from any context through their always-mapped alias, without
+     * going through the mapcache -- in particular from the context switch,
+     * which writes the incoming vcpu's tables before loading them.
+     */
+    l3tab = d->arch.perdomain_l3;
+    if ( !l3tab )
     {
-        pg = alloc_domheap_page(d, MEMF_no_owner);
-        if ( !pg )
+        l3tab = alloc_xenheap_pages(0, memflags);
+        if ( !l3tab )
             return -ENOMEM;
-        l3tab = __map_domain_page(pg);
         clear_page(l3tab);
-        d->arch.perdomain_l3_pg = pg;
-        if ( !nr )
-        {
-            unmap_domain_page(l3tab);
-            return 0;
-        }
+        d->arch.perdomain_l3 = l3tab;
     }
-    else if ( !nr )
+
+    if ( !nr )
         return 0;
-    else
-        l3tab = __map_domain_page(d->arch.perdomain_l3_pg);
 
     ASSERT(!l3_table_offset(va ^ (va + nr * PAGE_SIZE - 1)));
 
     if ( !(l3e_get_flags(l3tab[l3_table_offset(va)]) & _PAGE_PRESENT) )
     {
-        pg = alloc_domheap_page(d, MEMF_no_owner);
-        if ( !pg )
-        {
-            unmap_domain_page(l3tab);
+        l2tab = alloc_xenheap_pages(0, memflags);
+        if ( !l2tab )
             return -ENOMEM;
-        }
-        l2tab = __map_domain_page(pg);
         clear_page(l2tab);
-        l3tab[l3_table_offset(va)] = l3e_from_page(pg, __PAGE_HYPERVISOR_RW);
+        l3tab[l3_table_offset(va)] = l3e_from_mfn(virt_to_mfn(l2tab),
+                                                  __PAGE_HYPERVISOR_RW);
     }
     else
-        l2tab = map_l2t_from_l3e(l3tab[l3_table_offset(va)]);
-
-    unmap_domain_page(l3tab);
+        l2tab = maddr_to_virt(l3e_get_paddr(l3tab[l3_table_offset(va)]));
 
     if ( !pl1tab && !ppg )
-    {
-        unmap_domain_page(l2tab);
         return 0;
-    }
 
     for ( l1tab = NULL; !rc && nr--; )
     {
@@ -6274,33 +6267,22 @@ int create_perdomain_mapping(struct domain *d, unsigned long va,
 
         if ( !(l2e_get_flags(*pl2e) & _PAGE_PRESENT) )
         {
+            l1tab = alloc_xenheap_pages(0, memflags);
+            if ( !l1tab )
+            {
+                rc = -ENOMEM;
+                break;
+            }
             if ( pl1tab && !IS_NIL(pl1tab) )
             {
-                l1tab = alloc_xenheap_pages(0, MEMF_node(domain_to_node(d)));
-                if ( !l1tab )
-                {
-                    rc = -ENOMEM;
-                    break;
-                }
                 ASSERT(!pl1tab[l2_table_offset(va)]);
                 pl1tab[l2_table_offset(va)] = l1tab;
-                pg = virt_to_page(l1tab);
-            }
-            else
-            {
-                pg = alloc_domheap_page(d, MEMF_no_owner);
-                if ( !pg )
-                {
-                    rc = -ENOMEM;
-                    break;
-                }
-                l1tab = __map_domain_page(pg);
             }
             clear_page(l1tab);
-            *pl2e = l2e_from_page(pg, __PAGE_HYPERVISOR_RW);
+            *pl2e = l2e_from_mfn(virt_to_mfn(l1tab), __PAGE_HYPERVISOR_RW);
         }
         else if ( !l1tab )
-            l1tab = map_l1t_from_l2e(*pl2e);
+            l1tab = maddr_to_virt(l2e_get_paddr(*pl2e));
 
         if ( ppg &&
              !(l1e_get_flags(l1tab[l1_table_offset(va)]) & _PAGE_PRESENT) )
@@ -6321,15 +6303,10 @@ int create_perdomain_mapping(struct domain *d, unsigned long va,
 
         va += PAGE_SIZE;
         if ( rc || !nr || !l1_table_offset(va) )
-        {
-            /* Note that this is a no-op for the alloc_xenheap_page() case. */
-            unmap_domain_page(l1tab);
             l1tab = NULL;
-        }
     }
 
     ASSERT(!l1tab);
-    unmap_domain_page(l2tab);
 
     return rc;
 }
@@ -6343,15 +6320,15 @@ void destroy_perdomain_mapping(struct domain *d, unsigned long va,
            va < PERDOMAIN_VIRT_SLOT(PERDOMAIN_SLOTS));
     ASSERT(!nr || !l3_table_offset(va ^ (va + nr * PAGE_SIZE - 1)));
 
-    if ( !d->arch.perdomain_l3_pg )
+    l3tab = d->arch.perdomain_l3;
+    if ( !l3tab )
         return;
 
-    l3tab = __map_domain_page(d->arch.perdomain_l3_pg);
     pl3e = l3tab + l3_table_offset(va);
 
     if ( l3e_get_flags(*pl3e) & _PAGE_PRESENT )
     {
-        const l2_pgentry_t *l2tab = map_l2t_from_l3e(*pl3e);
+        const l2_pgentry_t *l2tab = maddr_to_virt(l3e_get_paddr(*pl3e));
         const l2_pgentry_t *pl2e = l2tab + l2_table_offset(va);
         unsigned int i = l1_table_offset(va);
 
@@ -6359,7 +6336,7 @@ void destroy_perdomain_mapping(struct domain *d, unsigned long va,
         {
             if ( l2e_get_flags(*pl2e) & _PAGE_PRESENT )
             {
-                l1_pgentry_t *l1tab = map_l1t_from_l2e(*pl2e);
+                l1_pgentry_t *l1tab = maddr_to_virt(l2e_get_paddr(*pl2e));
 
                 for ( ; nr && i < L1_PAGETABLE_ENTRIES; --nr, ++i )
                 {
@@ -6367,8 +6344,6 @@ void destroy_perdomain_mapping(struct domain *d, unsigned long va,
                         free_domheap_page(l1e_get_page(l1tab[i]));
                     l1tab[i] = l1e_empty();
                 }
-
-                unmap_domain_page(l1tab);
             }
             else if ( nr + i < L1_PAGETABLE_ENTRIES )
                 break;
@@ -6378,60 +6353,46 @@ void destroy_perdomain_mapping(struct domain *d, unsigned long va,
             ++pl2e;
             i = 0;
         }
-
-        unmap_domain_page(l2tab);
     }
-
-    unmap_domain_page(l3tab);
 }
 
 void free_perdomain_mappings(struct domain *d)
 {
-    l3_pgentry_t *l3tab;
+    l3_pgentry_t *l3tab = d->arch.perdomain_l3;
     unsigned int i;
 
-    if ( !d->arch.perdomain_l3_pg )
+    if ( !l3tab )
         return;
 
-    l3tab = __map_domain_page(d->arch.perdomain_l3_pg);
-
     for ( i = 0; i < PERDOMAIN_SLOTS; ++i)
         if ( l3e_get_flags(l3tab[i]) & _PAGE_PRESENT )
         {
-            struct page_info *l2pg = l3e_get_page(l3tab[i]);
-            l2_pgentry_t *l2tab = __map_domain_page(l2pg);
+            l2_pgentry_t *l2tab = maddr_to_virt(l3e_get_paddr(l3tab[i]));
             unsigned int j;
 
             for ( j = 0; j < L2_PAGETABLE_ENTRIES; ++j )
                 if ( l2e_get_flags(l2tab[j]) & _PAGE_PRESENT )
                 {
-                    struct page_info *l1pg = l2e_get_page(l2tab[j]);
+                    l1_pgentry_t *l1tab =
+                        maddr_to_virt(l2e_get_paddr(l2tab[j]));
 
                     if ( l2e_get_flags(l2tab[j]) & _PAGE_AVAIL0 )
                     {
-                        l1_pgentry_t *l1tab = __map_domain_page(l1pg);
                         unsigned int k;
 
                         for ( k = 0; k < L1_PAGETABLE_ENTRIES; ++k )
                             if ( perdomain_l1e_needs_freeing(l1tab[k]) )
                                 free_domheap_page(l1e_get_page(l1tab[k]));
-
-                        unmap_domain_page(l1tab);
                     }
 
-                    if ( is_xen_heap_page(l1pg) )
-                        free_xenheap_page(page_to_virt(l1pg));
-                    else
-                        free_domheap_page(l1pg);
+                    free_xenheap_page(l1tab);
                 }
 
-            unmap_domain_page(l2tab);
-            free_domheap_page(l2pg);
+            free_xenheap_page(l2tab);
         }
 
-    unmap_domain_page(l3tab);
-    free_domheap_page(d->arch.perdomain_l3_pg);
-    d->arch.perdomain_l3_pg = NULL;
+    free_xenheap_page(l3tab);
+    d->arch.perdomain_l3 = NULL;
 }
 
 static void write_sss_token(unsigned long *ptr)
-- 
2.55.0



^ permalink raw reply related	[flat|nested] 17+ messages in thread

* [PATCH 2/7] x86/mm: introduce populate_perdomain_mapping()
  2026-08-20 17:43 [PATCH 0/7] x86: Address Space Isolation, part 1: per-domain area mapping rework George Dunlap
  2026-08-20 17:43 ` [PATCH 1/7] x86/mm: allocate the per-domain page-tables from the xenheap George Dunlap
@ 2026-08-20 17:43 ` George Dunlap
  2026-08-20 17:43 ` [PATCH 3/7] x86/pv: use populate_perdomain_mapping() to map the Xen GDT George Dunlap
                   ` (5 subsequent siblings)
  7 siblings, 0 replies; 17+ messages in thread
From: George Dunlap @ 2026-08-20 17:43 UTC (permalink / raw)
  To: xen-devel
  Cc: Roger Pau Monné, Jan Beulich, Andrew Cooper,
	Roger Pau Monné, Teddy Astie, Anthony PERARD, Michal Orzel,
	Julien Grall, Stefano Stabellini

From: Roger Pau Monné <roger.pau@citrix.com>

The per-domain area already has central machinery for building its
page-tables and for managing the backing pages it owns itself:
create_perdomain_mapping() / destroy_perdomain_mapping(), used by the
mapcache bitmaps, the compat argument-translation area, and the
GDT/LDT slots alike.  What the interface lacks is a way to install a
caller's own pages at a chosen address.  The PV GDT and LDT code
open-codes its own modifications to the per-domain area, capturing
aliases of its L1 tables at creation time
(create_perdomain_mapping()'s pl1tab argument) and stashing them in
d->arch.pv.gdt_ldt_l1tab.

Introduce populate_perdomain_mapping(v, va, mfn, nr, flags) to close
this gap, giving the perdomain area's rules a single place to live.
populate_perdomain_mapping writes the given MFNs, with the given
page-table flags, into v's view of the per-domain area by walking its
per-domain page-tables.  Those are xenheap pages, reached through
their always-mapped alias, so the walk involves no mapping and is
usable from any context -- including the context switch, before the
incoming vcpu's page-tables are loaded.  Callers don't need to know
where the page-tables live, how the area is structured, or whether it
is per-domain or per-vcpu.

We require the range to already have been populated down to the L1
tables by create_perdomain_mapping().  TLB flushing is left to the
caller.  A present entry not owned by the area (!_PAGE_AVAIL0) is
replaced.  A present entry owned by the area (_PAGE_AVAIL0, installed
by create_perdomain_mapping() itself) is freed and replaced: such a
page is referenced only by the mapping, so displacing it without
freeing it would leak it.  Nothing in this series replaces area-owned
backing, so the free is marked ASSERT_UNREACHABLE(); note that freeing
requires a context where the allocator may be entered -- IRQs enabled,
not in interrupt context (see ASSERT_ALLOC_CONTEXT()) -- so any future
caller replacing area-owned backing must not do so from the context
switch path.  Missing page-table structure is a hypervisor bug and
BUG_ON(): there is no safe continuation, least of all from the context
switch, where the next descriptor fetch through an unmapped GDT slot
would be fatal.

Subsequent patches convert the users of the stashed L1 tables to this
interface, starting with the Xen slots of the full GDT; the stash --
which could in any case not represent per-vcpu mappings without being
replicated for every vcpu and slot -- is then removed, leaving
create_perdomain_mapping() to manage only the page-table structure and
the pages the area owns itself.  Later parts of the series use the new
interface for their own mappings rather than adding further
mechanisms.

Signed-off-by: Roger Pau Monné <roger.pau@citrix.com>
Assisted-by: Claude Code:claude-fable-5, Claude Code:claude-opus-4-8
Signed-off-by: George Dunlap <gwd@xenproject.org>
---
Changes since the previously posted version:
- Split the introduction of populate_perdomain_mapping() from its first
   user (previously one patch: "x86/pv: introduce function to populate
   perdomain area and use it to map Xen GDT").
- Drop the linear-map fast path and the sync_local_execstate() call:
   with the per-domain page-tables in the xenheap (previous patch) the
   walk needs no mapping, so a single path serves all callers and
   contexts.
- Keep the ASSERT_UNREACHABLE() + free_domheap_page() handling of a
   replaced area-owned entry, and document the allocation-context
   requirement it places on callers replacing such entries.  BUG_ON()
   missing page-table structure, instead of domain_crash().
- Take the page-table flags as a parameter (the Xen GDT and guest GDT
   slots want RW mappings; the zero page backing torn-down GDT slots is
   mapped read-only, as today).
- Document the contract in a header comment.
- Make the mfn parameter const and nr unsigned int, matching
   {create,destroy}_perdomain_mapping().
- Drop the unused cr3_mfn() helper.

Considered, but not done to limit churn against the previously posted
version: splitting the interface into a "populate" variant (any present
entry is a bug) and an "update" variant (replacement expected), so that
call sites declare their intent and unexpected collisions become
detectable.  Of the eventual call sites in the wider series, roughly
half are of each kind.  Could be done as a follow-up if there is
interest.
---
 xen/arch/x86/include/asm/mm.h |  3 ++
 xen/arch/x86/mm.c             | 68 +++++++++++++++++++++++++++++++++++
 2 files changed, 71 insertions(+)

diff --git a/xen/arch/x86/include/asm/mm.h b/xen/arch/x86/include/asm/mm.h
index 2254a7e3fe..1888807394 100644
--- a/xen/arch/x86/include/asm/mm.h
+++ b/xen/arch/x86/include/asm/mm.h
@@ -606,6 +606,9 @@ int compat_arch_memory_op(unsigned long cmd, XEN_GUEST_HANDLE_PARAM(void) arg);
 int create_perdomain_mapping(struct domain *d, unsigned long va,
                              unsigned int nr, l1_pgentry_t **pl1tab,
                              struct page_info **ppg);
+void populate_perdomain_mapping(const struct vcpu *v, unsigned long va,
+                                const mfn_t *mfn, unsigned int nr,
+                                unsigned int flags);
 void destroy_perdomain_mapping(struct domain *d, unsigned long va,
                                unsigned int nr);
 void free_perdomain_mappings(struct domain *d);
diff --git a/xen/arch/x86/mm.c b/xen/arch/x86/mm.c
index 38a0f984fc..1810971677 100644
--- a/xen/arch/x86/mm.c
+++ b/xen/arch/x86/mm.c
@@ -6311,6 +6311,74 @@ int create_perdomain_mapping(struct domain *d, unsigned long va,
     return rc;
 }
 
+/*
+ * Map @nr pages, @mfn[0..nr-1], at consecutive pages from @va in v's view of
+ * the per-domain area, with page-table @flags.  The range must lie within a
+ * single per-domain slot, and must already have been plumbed down to the L1
+ * tables by create_perdomain_mapping(): missing structure is a bug.  A
+ * present entry not owned by the area (no _PAGE_AVAIL0) is silently
+ * replaced, as that is how callers update their mappings; a present
+ * area-owned entry is freed and replaced, which constrains the calling
+ * context (see the comment in the body).  No TLB flushing is done: the
+ * caller decides whether the old translations can still be cached
+ * anywhere.
+ *
+ * The walk goes through the always-mapped xenheap alias of the per-domain
+ * page-tables, so it needs nothing from the current address space and is
+ * usable from any context -- including the context switch, before the
+ * incoming vcpu's page-tables are loaded.
+ */
+void populate_perdomain_mapping(const struct vcpu *v, unsigned long va,
+                                const mfn_t *mfn, unsigned int nr,
+                                unsigned int flags)
+{
+    l1_pgentry_t *l1tab = NULL, *pl1e;
+    const l3_pgentry_t *l3tab;
+    const l2_pgentry_t *l2tab;
+
+    ASSERT(va >= PERDOMAIN_VIRT_START &&
+           va < PERDOMAIN_VIRT_SLOT(PERDOMAIN_SLOTS));
+    ASSERT(!nr || !l3_table_offset(va ^ (va + nr * PAGE_SIZE - 1)));
+    /* Area-owned pages are installed by create_perdomain_mapping() only. */
+    ASSERT(!(flags & _PAGE_AVAIL0));
+
+    l3tab = v->domain->arch.perdomain_l3;
+    BUG_ON(!l3tab);
+    BUG_ON(!(l3e_get_flags(l3tab[l3_table_offset(va)]) & _PAGE_PRESENT));
+
+    l2tab = maddr_to_virt(l3e_get_paddr(l3tab[l3_table_offset(va)]));
+
+    for ( ; nr--; va += PAGE_SIZE, mfn++ )
+    {
+        if ( !l1tab || !l1_table_offset(va) )
+        {
+            const l2_pgentry_t *pl2e = l2tab + l2_table_offset(va);
+
+            BUG_ON(!(l2e_get_flags(*pl2e) & _PAGE_PRESENT));
+            l1tab = maddr_to_virt(l2e_get_paddr(*pl2e));
+        }
+
+        pl1e = &l1tab[l1_table_offset(va)];
+
+        /*
+         * An area-owned entry (installed by create_perdomain_mapping(),
+         * marked _PAGE_AVAIL0) holds the only reference to its page, so
+         * displacing it means freeing it.  Nothing in this series replaces
+         * area-owned backing, hence the ASSERT_UNREACHABLE(); any future
+         * caller doing so must run where freeing is permitted -- IRQs
+         * enabled, not in interrupt context (see ASSERT_ALLOC_CONTEXT())
+         * -- which the context switch path is not.
+         */
+        if ( unlikely(perdomain_l1e_needs_freeing(*pl1e)) )
+        {
+            ASSERT_UNREACHABLE();
+            free_domheap_page(l1e_get_page(*pl1e));
+        }
+
+        l1e_write(pl1e, l1e_from_mfn(*mfn, flags));
+    }
+}
+
 void destroy_perdomain_mapping(struct domain *d, unsigned long va,
                                unsigned int nr)
 {
-- 
2.55.0



^ permalink raw reply related	[flat|nested] 17+ messages in thread

* [PATCH 3/7] x86/pv: use populate_perdomain_mapping() to map the Xen GDT
  2026-08-20 17:43 [PATCH 0/7] x86: Address Space Isolation, part 1: per-domain area mapping rework George Dunlap
  2026-08-20 17:43 ` [PATCH 1/7] x86/mm: allocate the per-domain page-tables from the xenheap George Dunlap
  2026-08-20 17:43 ` [PATCH 2/7] x86/mm: introduce populate_perdomain_mapping() George Dunlap
@ 2026-08-20 17:43 ` George Dunlap
  2026-08-21 20:59   ` Andrew Cooper
  2026-08-20 17:43 ` [PATCH 4/7] x86/pv: set/clear guest GDT mappings using populate_perdomain_mapping() George Dunlap
                   ` (4 subsequent siblings)
  7 siblings, 1 reply; 17+ messages in thread
From: George Dunlap @ 2026-08-20 17:43 UTC (permalink / raw)
  To: xen-devel
  Cc: Roger Pau Monné, Jan Beulich, Andrew Cooper,
	Roger Pau Monné, Teddy Astie, Anthony PERARD, Michal Orzel,
	Julien Grall, Stefano Stabellini

From: Roger Pau Monné <roger.pau@citrix.com>

Currently, update_xen_slot_in_full_gdt() uses the stashed pointer in
d->arch.pv.gdt_ldt_l1tab to update the incoming vcpu's page-tables with
Xen's GDT, by writing a stashed per-cpu copy of a pre-baked L1 entry
(either 64-bit or compat version).

Now that all perdomain pagetables are allocated from the xenheap and
their root L3 stashed in d->arch.perdomain_l3,
d->arch.pv.gdt_ldt_l1tab is redundant.  Remove one user by switching
update_xen_slot_in_full_gdt() to using populate_perdomain_mapping().

Since populate_perdomain_mapping() takes an mfn rather than an l1e,
cache the mfn of the per-cpu page instead.  As a side effect, this
consolidates the setting of the flags into a single place.  Continue
to check that the per-cpu value we're using has been initialized:
per-CPU data starts out zeroed, and no GDT can live at MFN 0.

What was one store through a cached pointer is now an out-of-line
walk (L3 pointer, L3e, L2e, then the L1e) on every PV context switch:
three dependent loads, negligible next to the CR3 write that follows.

Signed-off-by: Roger Pau Monné <roger.pau@citrix.com>
Assisted-by: Claude Code:claude-fable-5, Claude Code:claude-opus-4-8
Signed-off-by: George Dunlap <gwd@xenproject.org>
---
Changes since the previously posted version:
- populate_perdomain_mapping() introduction split into the previous
   patch; this patch is now just the Xen GDT conversion.
- Retain the "GDT MFN cached" check as ASSERT(mfn_x(mfn)).
---
 xen/arch/x86/domain.c           | 13 +++++++++----
 xen/arch/x86/include/asm/desc.h |  6 ++++--
 xen/arch/x86/smpboot.c          | 14 +++++---------
 xen/arch/x86/traps.c            |  4 ++--
 4 files changed, 20 insertions(+), 17 deletions(-)

diff --git a/xen/arch/x86/domain.c b/xen/arch/x86/domain.c
index 163a2c97ae..f5b2eef95a 100644
--- a/xen/arch/x86/domain.c
+++ b/xen/arch/x86/domain.c
@@ -2062,11 +2062,16 @@ static always_inline bool need_full_gdt(const struct domain *d)
 
 static void update_xen_slot_in_full_gdt(const struct vcpu *v, unsigned int cpu)
 {
-    ASSERT(per_cpu(gdt_l1e, cpu).l1); /* Confirm these have been cached. */
+    mfn_t mfn = !is_pv_32bit_vcpu(v) ? per_cpu(gdt_mfn, cpu)
+                                     : per_cpu(compat_gdt_mfn, cpu);
 
-    l1e_write(pv_gdt_ptes(v) + FIRST_RESERVED_GDT_PAGE,
-              !is_pv_32bit_vcpu(v) ? per_cpu(gdt_l1e, cpu)
-                                   : per_cpu(compat_gdt_l1e, cpu));
+    /* Confirm the GDT MFNs have been cached (MFN 0 is never a GDT). */
+    ASSERT(mfn_x(mfn));
+
+    populate_perdomain_mapping(v,
+                               GDT_VIRT_START(v) +
+                               (FIRST_RESERVED_GDT_PAGE << PAGE_SHIFT),
+                               &mfn, 1, __PAGE_HYPERVISOR_RW);
 }
 
 static void load_full_gdt(const struct vcpu *v, unsigned int cpu)
diff --git a/xen/arch/x86/include/asm/desc.h b/xen/arch/x86/include/asm/desc.h
index dcbdac3ff7..a26df9fbe9 100644
--- a/xen/arch/x86/include/asm/desc.h
+++ b/xen/arch/x86/include/asm/desc.h
@@ -44,6 +44,8 @@
 
 #ifndef __ASSEMBLER__
 
+#include <xen/mm-frame.h>
+
 #define GUEST_KERNEL_RPL(d) (is_pv_32bit_domain(d) ? 1 : 3)
 
 /* Fix up the RPL of a guest segment selector. */
@@ -136,10 +138,10 @@ struct __packed desc_ptr {
 
 extern seg_desc_t boot_gdt[];
 DECLARE_PER_CPU(seg_desc_t *, gdt);
-DECLARE_PER_CPU(l1_pgentry_t, gdt_l1e);
+DECLARE_PER_CPU(mfn_t, gdt_mfn);
 extern seg_desc_t boot_compat_gdt[];
 DECLARE_PER_CPU(seg_desc_t *, compat_gdt);
-DECLARE_PER_CPU(l1_pgentry_t, compat_gdt_l1e);
+DECLARE_PER_CPU(mfn_t, compat_gdt_mfn);
 DECLARE_PER_CPU(bool, full_gdt_loaded);
 
 static inline void lgdt(const struct desc_ptr *gdtr)
diff --git a/xen/arch/x86/smpboot.c b/xen/arch/x86/smpboot.c
index 84e9e4beed..9246945506 100644
--- a/xen/arch/x86/smpboot.c
+++ b/xen/arch/x86/smpboot.c
@@ -1085,8 +1085,7 @@ static int cpu_smpboot_alloc(unsigned int cpu)
     if ( gdt == NULL )
         goto out;
     per_cpu(gdt, cpu) = gdt;
-    per_cpu(gdt_l1e, cpu) =
-        l1e_from_pfn(virt_to_mfn(gdt), __PAGE_HYPERVISOR_RW);
+    per_cpu(gdt_mfn, cpu) = _mfn(virt_to_mfn(gdt));
     memcpy(gdt, boot_gdt, NR_RESERVED_GDT_PAGES * PAGE_SIZE);
     BUILD_BUG_ON(NR_CPUS > 0x10000);
     gdt[PER_CPU_GDT_ENTRY - FIRST_RESERVED_GDT_ENTRY].a = cpu;
@@ -1095,8 +1094,7 @@ static int cpu_smpboot_alloc(unsigned int cpu)
     per_cpu(compat_gdt, cpu) = gdt = alloc_xenheap_pages(0, memflags);
     if ( gdt == NULL )
         goto out;
-    per_cpu(compat_gdt_l1e, cpu) =
-        l1e_from_pfn(virt_to_mfn(gdt), __PAGE_HYPERVISOR_RW);
+    per_cpu(compat_gdt_mfn, cpu) = _mfn(virt_to_mfn(gdt));
     memcpy(gdt, boot_compat_gdt, NR_RESERVED_GDT_PAGES * PAGE_SIZE);
     gdt[PER_CPU_GDT_ENTRY - FIRST_RESERVED_GDT_ENTRY].a = cpu;
 #endif
@@ -1174,15 +1172,13 @@ void __init smp_prepare_cpus(void)
     print_cpu_info(0);
 
     /*
-     * Cache {,compat_}gdt_l1e for the BSP now that physically relocation is
+     * Cache {,compat_}gdt_mfn for the BSP now that physically relocation is
      * done.  It must be after physical relocation of Xen, and before the
      * first context_switch().
      */
-    this_cpu(gdt_l1e) =
-        l1e_from_pfn(virt_to_mfn(boot_gdt), __PAGE_HYPERVISOR_RW);
+    this_cpu(gdt_mfn) = _mfn(virt_to_mfn(boot_gdt));
     if ( IS_ENABLED(CONFIG_PV32) )
-        this_cpu(compat_gdt_l1e) =
-            l1e_from_pfn(virt_to_mfn(boot_compat_gdt), __PAGE_HYPERVISOR_RW);
+        this_cpu(compat_gdt_mfn) = _mfn(virt_to_mfn(boot_compat_gdt));
 
     boot_cpu_physical_apicid = get_apic_id();
     x86_cpu_to_apicid[0] = boot_cpu_physical_apicid;
diff --git a/xen/arch/x86/traps.c b/xen/arch/x86/traps.c
index 1774966305..2fcaa413f7 100644
--- a/xen/arch/x86/traps.c
+++ b/xen/arch/x86/traps.c
@@ -71,10 +71,10 @@ DEFINE_PER_CPU(uint64_t, efer);
 static DEFINE_PER_CPU(unsigned long, last_extable_addr);
 
 DEFINE_PER_CPU_READ_MOSTLY(seg_desc_t *, gdt);
-DEFINE_PER_CPU_READ_MOSTLY(l1_pgentry_t, gdt_l1e);
+DEFINE_PER_CPU_READ_MOSTLY(mfn_t, gdt_mfn);
 #ifdef CONFIG_PV32
 DEFINE_PER_CPU_READ_MOSTLY(seg_desc_t *, compat_gdt);
-DEFINE_PER_CPU_READ_MOSTLY(l1_pgentry_t, compat_gdt_l1e);
+DEFINE_PER_CPU_READ_MOSTLY(mfn_t, compat_gdt_mfn);
 #endif
 
 /*
-- 
2.55.0



^ permalink raw reply related	[flat|nested] 17+ messages in thread

* [PATCH 4/7] x86/pv: set/clear guest GDT mappings using populate_perdomain_mapping()
  2026-08-20 17:43 [PATCH 0/7] x86: Address Space Isolation, part 1: per-domain area mapping rework George Dunlap
                   ` (2 preceding siblings ...)
  2026-08-20 17:43 ` [PATCH 3/7] x86/pv: use populate_perdomain_mapping() to map the Xen GDT George Dunlap
@ 2026-08-20 17:43 ` George Dunlap
  2026-08-20 17:43 ` [PATCH 5/7] x86/pv: update guest LDT mappings using {populate,destroy}_perdomain_mapping() George Dunlap
                   ` (3 subsequent siblings)
  7 siblings, 0 replies; 17+ messages in thread
From: George Dunlap @ 2026-08-20 17:43 UTC (permalink / raw)
  To: xen-devel
  Cc: Roger Pau Monné, Jan Beulich, Andrew Cooper,
	Roger Pau Monné, Teddy Astie, Anthony PERARD, Michal Orzel,
	Julien Grall, Stefano Stabellini

From: Roger Pau Monné <roger.pau@citrix.com>

Until the previous patch, update_xen_slot_in_full_gdt() used the
stashed pointer in d->arch.pv.gdt_ldt_l1tab to update the incoming
vCPU's page tables with Xen's GDT; this was necessary because
perdomain pagetables were mapped from the domheap by default, and
map_domain_page() couldn't be called in a context switch.  Having a
handy pointer to an always-mapped version of the GDT/LDT L1 table,
other sites which modify the table started using it for convenience,
even if they weren't called from within a context switch.  One example
is pv_{set,destroy}_gdt.

Now that all perdomain pagetables are allocated from the xenheap and
their root L3 stashed in d->arch.perdomain_l3,
d->arch.pv.gdt_ldt_l1tab is redundant.  The previous patch removed one
user by modifying update_xen_slot_in_full_gdt() to call
populate_perdomain_mapping().  Continue that process by switching both
pv_{set,destroy}_gdt() to it as well.

pv_destroy_gdt() currently loops over the L1 entries directly,
extracting the MFN from each, dropping the type and reference unless
it was the zero page, and replacing the entry with a read-only mapping
of the zero page.  Since we no longer have the L1 to hand, drop the
references using v->arch.pv.gdt_frames[] instead, and install the
zero-page mappings with a single populate_perdomain_mapping() call.
This makes gdt_frames[] consistently the source of truth for MFNs.

Behaviour is unchanged: torn-down slots map the zero page read-only,
as they have since cf6d39f819 ("x86/PV: properly populate descriptor
tables"), so that LAR/LSL/VERR/VERW on a selector beyond the guest's
limit clear ZF as on native rather than taking a #PF-converted #GP --
and as every PV vCPU's unused slots do from the start, pv_set_gdt()
tearing down the old GDT (zero page included) before installing the
new one.

In the case of pv_set_gdt, we have a slightly awkward situation with
types.  The ABI with the guest uses unsigned long[], but
populate_perdomain_mapping wants an array of mfn_t.
v->arch.pv.gdt_frames being unsigned long means we can just copy from
it across the guest ABI with no conversions.  We could in theory
convert it to mfn_t[] instead, and then pass v->arch.pv.gdt_frames
into populate_perdomain_mapping; but then we'd need to add a
conversion on all the places where frames are copied out.  We choose
instead to copy frames into a temporary mfn_t array on the stack to
pass into populate_perdomain_mapping.

Signed-off-by: Roger Pau Monné <roger.pau@citrix.com>
Assisted-by: Claude Code:claude-fable-5, Claude Code:claude-opus-4-8
Signed-off-by: George Dunlap <gwd@xenproject.org>
---
Changes since the previously posted version:
- Retain the gdt_ents zeroing when tearing down the GDT (its removal
   was queried by Jan).
- Map torn-down slots read-only to the zero page (via the
   populate_perdomain_mapping() flags parameter) rather than removing
   the mappings with destroy_perdomain_mapping(): empty slots would be
   a guest-visible partial revert of cf6d39f819 (see the commit
   message).  With the destroy call gone, its v->arch.cr3 guard --
   also queried by Jan -- goes too: the zero-page rewrite runs
   unconditionally.
- Keep gdt_frames[] as unsigned long[] rather than switching it to
   mfn_t[] as Jan suggested; the commit message explains the
   trade-off.
- Retitle: destroy_perdomain_mapping() is no longer used here.
---
 xen/arch/x86/pv/descriptor-tables.c | 37 ++++++++++++++++++-----------
 1 file changed, 23 insertions(+), 14 deletions(-)

diff --git a/xen/arch/x86/pv/descriptor-tables.c b/xen/arch/x86/pv/descriptor-tables.c
index 8a32b9ae5c..5dda5bffe3 100644
--- a/xen/arch/x86/pv/descriptor-tables.c
+++ b/xen/arch/x86/pv/descriptor-tables.c
@@ -49,33 +49,42 @@ bool pv_destroy_ldt(struct vcpu *v)
 
 void pv_destroy_gdt(struct vcpu *v)
 {
-    l1_pgentry_t *pl1e = pv_gdt_ptes(v);
-    mfn_t zero_mfn = _mfn(virt_to_mfn(zero_page));
-    l1_pgentry_t zero_l1e = l1e_from_mfn(zero_mfn, __PAGE_HYPERVISOR_RO);
+    const mfn_t zero_mfn = _mfn(virt_to_mfn(zero_page));
+    mfn_t zero_mfns[ARRAY_SIZE(v->arch.pv.gdt_frames)];
     unsigned int i;
 
     ASSERT(v == current || !vcpu_cpu_dirty(v));
 
     v->arch.pv.gdt_ents = 0;
-    for ( i = 0; i < FIRST_RESERVED_GDT_PAGE; i++ )
+
+    for ( i = 0; i < ARRAY_SIZE(zero_mfns); i++ )
     {
-        mfn_t mfn = l1e_get_mfn(pl1e[i]);
+        zero_mfns[i] = zero_mfn;
 
-        if ( (l1e_get_flags(pl1e[i]) & _PAGE_PRESENT) &&
-             !mfn_eq(mfn, zero_mfn) )
-            put_page_and_type(mfn_to_page(mfn));
+        /* MFN 0 can never pass get_page_and_type(), so 0 marks unused slots. */
+        if ( !v->arch.pv.gdt_frames[i] )
+            continue;
 
-        l1e_write(&pl1e[i], zero_l1e);
+        put_page_and_type(mfn_to_page(_mfn(v->arch.pv.gdt_frames[i])));
         v->arch.pv.gdt_frames[i] = 0;
     }
+
+    /*
+     * Point every slot at the zero page, read-only: a descriptor fetch from
+     * the unused part of the GDT then finds a not-present descriptor rather
+     * than a missing mapping, so LAR/LSL/VERR/VERW on a selector beyond the
+     * guest's limit clear ZF as they do on native, instead of faulting.
+     */
+    populate_perdomain_mapping(v, GDT_VIRT_START(v), zero_mfns,
+                               ARRAY_SIZE(zero_mfns), __PAGE_HYPERVISOR_RO);
 }
 
 int pv_set_gdt(struct vcpu *v, const unsigned long frames[],
                unsigned int entries)
 {
     struct domain *d = v->domain;
-    l1_pgentry_t *pl1e;
     unsigned int i, nr_frames = DIV_ROUND_UP(entries, 512);
+    mfn_t mfns[ARRAY_SIZE(v->arch.pv.gdt_frames)];
 
     ASSERT(v == current || !vcpu_cpu_dirty(v));
 
@@ -90,6 +99,8 @@ int pv_set_gdt(struct vcpu *v, const unsigned long frames[],
         if ( !mfn_valid(mfn) ||
              !get_page_and_type(mfn_to_page(mfn), d, PGT_seg_desc_page) )
             goto fail;
+
+        mfns[i] = mfn;
     }
 
     /* Tear down the old GDT. */
@@ -97,12 +108,10 @@ int pv_set_gdt(struct vcpu *v, const unsigned long frames[],
 
     /* Install the new GDT. */
     v->arch.pv.gdt_ents = entries;
-    pl1e = pv_gdt_ptes(v);
     for ( i = 0; i < nr_frames; i++ )
-    {
         v->arch.pv.gdt_frames[i] = frames[i];
-        l1e_write(&pl1e[i], l1e_from_pfn(frames[i], __PAGE_HYPERVISOR_RW));
-    }
+    populate_perdomain_mapping(v, GDT_VIRT_START(v), mfns, nr_frames,
+                               __PAGE_HYPERVISOR_RW);
 
     return 0;
 
-- 
2.55.0



^ permalink raw reply related	[flat|nested] 17+ messages in thread

* [PATCH 5/7] x86/pv: update guest LDT mappings using {populate,destroy}_perdomain_mapping()
  2026-08-20 17:43 [PATCH 0/7] x86: Address Space Isolation, part 1: per-domain area mapping rework George Dunlap
                   ` (3 preceding siblings ...)
  2026-08-20 17:43 ` [PATCH 4/7] x86/pv: set/clear guest GDT mappings using populate_perdomain_mapping() George Dunlap
@ 2026-08-20 17:43 ` George Dunlap
  2026-08-20 17:43 ` [PATCH 6/7] x86/pv: remove stashing of GDT/LDT L1 page-tables George Dunlap
                   ` (2 subsequent siblings)
  7 siblings, 0 replies; 17+ messages in thread
From: George Dunlap @ 2026-08-20 17:43 UTC (permalink / raw)
  To: xen-devel
  Cc: Roger Pau Monné, Jan Beulich, Andrew Cooper,
	Roger Pau Monné, Teddy Astie, Anthony PERARD, Michal Orzel,
	Julien Grall, Stefano Stabellini

From: Roger Pau Monné <roger.pau@citrix.com>

Until two patches ago, update_xen_slot_in_full_gdt() used the stashed
pointer in d->arch.pv.gdt_ldt_l1tab to update the incoming vCPU's page
tables with Xen's GDT; this was necessary because perdomain pagetables
were mapped from the domheap by default, and map_domain_page() couldn't
be called in a context switch.  Having a handy pointer to an
always-mapped version of the GDT/LDT L1 table, other sites which modify
the table started using it for convenience, even if they weren't called
from within a context switch.  These include pv_map_ldt_shadow_page()
and pv_destroy_ldt().

Now that all perdomain pagetables are allocated from the xenheap and
their root L3 stashed in d->arch.perdomain_l3,
d->arch.pv.gdt_ldt_l1tab is redundant.  The previous two patches
removed the GDT users; continue that process by refactoring the LDT
sites as well.

pv_map_ldt_shadow_page() is, by definition, always modifying the
currently-running vCPU: it runs from the #PF handler for a guest-mode
descriptor fetch, and running the guest implies its page tables are
loaded.  So it could simply write the linear recursive mappings
directly.  Go through populate_perdomain_mapping() anyway, to keep a
single writer for the per-domain area.

For pv_destroy_ldt(), use destroy_perdomain_mapping().

Previously, pv_destroy_ldt() used the L1 LDT entries themselves to
determine which MFNs to drop type and count references to.  Since we
don't have the L1 handy, we must now keep the MFNs corresponding to L1
slots in an array in the vCPU structure, as we do in the GDT case.

Note that mappings_dropped (the return value of pv_destroy_ldt()) now
reflects the *number of valid MFNs in this array*, not *the number of
non-empty L1 entries*.  This introduces an invariant we must maintain:
pv_map_ldt_shadow_page() writes both the array entry and the mapping,
and pv_destroy_ldt() clears both, so the two stay in lockstep.

Also note that, unlike pv_destroy_gdt() from the previous patch,
pv_destroy_ldt() doesn't fill in the values with zero_l1e (see
61031e64d3), so there's no change here.

Signed-off-by: Roger Pau Monné <roger.pau@citrix.com>
Assisted-by: Claude Code:claude-fable-5, Claude Code:claude-opus-4-8
Signed-off-by: George Dunlap <gwd@xenproject.org>
---
Changes since the previously posted version:
- Initialise ldt_frames[] ahead of the first fail-able initialisation step.
- Call destroy_perdomain_mapping() with its existing domain parameter;
   the switch to a vCPU parameter moves to a future patch.
- Use populate_perdomain_mapping() in pv_map_ldt_shadow_page() rather
   than open-coding the linear-map write; retitle accordingly.
- Comment the INVALID_MFN skip in pv_destroy_ldt(): the LDT is
   demand-faulted, so its pages may be sparsely mapped (Alejandro's
   question on v2).
- Rewrite commit message (including describing the ldt_frames[]
   array's role directly, as Jan asked).
---
 xen/arch/x86/include/asm/domain.h   |  2 ++
 xen/arch/x86/pv/descriptor-tables.c | 20 +++++++++++---------
 xen/arch/x86/pv/domain.c            |  4 ++++
 xen/arch/x86/pv/mm.c                | 16 ++++++++++++----
 4 files changed, 29 insertions(+), 13 deletions(-)

diff --git a/xen/arch/x86/include/asm/domain.h b/xen/arch/x86/include/asm/domain.h
index 5275bb10ea..2c9efd59bf 100644
--- a/xen/arch/x86/include/asm/domain.h
+++ b/xen/arch/x86/include/asm/domain.h
@@ -542,6 +542,8 @@ struct pv_vcpu
     struct trap_info *trap_ctxt;
 
     unsigned long gdt_frames[FIRST_RESERVED_GDT_PAGE];
+    /* Max LDT entries is 8192, so 8192 * 8 = 64KiB (16 pages). */
+    mfn_t ldt_frames[16];
     unsigned long ldt_base;
     unsigned int gdt_ents, ldt_ents;
 
diff --git a/xen/arch/x86/pv/descriptor-tables.c b/xen/arch/x86/pv/descriptor-tables.c
index 5dda5bffe3..261bf29c90 100644
--- a/xen/arch/x86/pv/descriptor-tables.c
+++ b/xen/arch/x86/pv/descriptor-tables.c
@@ -20,28 +20,30 @@
  */
 bool pv_destroy_ldt(struct vcpu *v)
 {
-    l1_pgentry_t *pl1e;
+    const unsigned int nr_frames = ARRAY_SIZE(v->arch.pv.ldt_frames);
     unsigned int i, mappings_dropped = 0;
-    struct page_info *page;
 
     ASSERT(!in_irq());
 
     ASSERT(v == current || !vcpu_cpu_dirty(v));
 
-    pl1e = pv_ldt_ptes(v);
+    destroy_perdomain_mapping(v->domain, LDT_VIRT_START(v), nr_frames);
 
-    for ( i = 0; i < 16; i++ )
+    for ( i = 0; i < nr_frames; i++ )
     {
-        if ( !(l1e_get_flags(pl1e[i]) & _PAGE_PRESENT) )
-            continue;
+        mfn_t mfn = v->arch.pv.ldt_frames[i];
+        struct page_info *page;
 
-        page = l1e_get_page(pl1e[i]);
-        l1e_write(&pl1e[i], l1e_empty());
-        mappings_dropped++;
+        /* The LDT is demand-faulted, so its pages may be sparsely mapped. */
+        if ( mfn_eq(mfn, INVALID_MFN) )
+            continue;
 
+        v->arch.pv.ldt_frames[i] = INVALID_MFN;
+        page = mfn_to_page(mfn);
         ASSERT_PAGE_IS_TYPE(page, PGT_seg_desc_page);
         ASSERT_PAGE_IS_DOMAIN(page, v->domain);
         put_page_and_type(page);
+        mappings_dropped++;
     }
 
     return mappings_dropped;
diff --git a/xen/arch/x86/pv/domain.c b/xen/arch/x86/pv/domain.c
index 0c42ae58aa..7ddab1949f 100644
--- a/xen/arch/x86/pv/domain.c
+++ b/xen/arch/x86/pv/domain.c
@@ -340,10 +340,14 @@ void pv_vcpu_destroy(struct vcpu *v)
 int pv_vcpu_initialise(struct vcpu *v)
 {
     struct domain *d = v->domain;
+    unsigned int i;
     int rc;
 
     ASSERT(!is_idle_domain(d));
 
+    for ( i = 0; i < ARRAY_SIZE(v->arch.pv.ldt_frames); i++ )
+        v->arch.pv.ldt_frames[i] = INVALID_MFN;
+
     rc = pv_create_gdt_ldt_l1tab(v);
     if ( rc )
         return rc;
diff --git a/xen/arch/x86/pv/mm.c b/xen/arch/x86/pv/mm.c
index 5378299b8c..da280d7757 100644
--- a/xen/arch/x86/pv/mm.c
+++ b/xen/arch/x86/pv/mm.c
@@ -53,7 +53,8 @@ bool pv_map_ldt_shadow_page(unsigned int offset)
     struct vcpu *curr = current;
     struct domain *currd = curr->domain;
     struct page_info *page;
-    l1_pgentry_t gl1e, *pl1e, nl1e;
+    l1_pgentry_t gl1e;
+    mfn_t mfn;
     unsigned long linear = curr->arch.pv.ldt_base + offset;
 
     BUG_ON(in_irq());
@@ -87,10 +88,17 @@ bool pv_map_ldt_shadow_page(unsigned int offset)
         return false;
     }
 
-    pl1e = &pv_ldt_ptes(curr)[offset >> PAGE_SHIFT];
-    nl1e = l1e_from_pfn(l1e_get_pfn(gl1e), __PAGE_HYPERVISOR_RW);
+    mfn = page_to_mfn(page);
+    curr->arch.pv.ldt_frames[offset >> PAGE_SHIFT] = mfn;
 
-    l1e_write(pl1e, nl1e);
+    /*
+     * Running the guest implies its page-tables are loaded, so the linear
+     * mappings would do; go through the interface anyway to keep a single
+     * writer for the per-domain area.
+     */
+    populate_perdomain_mapping(curr,
+                               LDT_VIRT_START(curr) + (offset & PAGE_MASK),
+                               &mfn, 1, __PAGE_HYPERVISOR_RW);
 
     return true;
 }
-- 
2.55.0



^ permalink raw reply related	[flat|nested] 17+ messages in thread

* [PATCH 6/7] x86/pv: remove stashing of GDT/LDT L1 page-tables
  2026-08-20 17:43 [PATCH 0/7] x86: Address Space Isolation, part 1: per-domain area mapping rework George Dunlap
                   ` (4 preceding siblings ...)
  2026-08-20 17:43 ` [PATCH 5/7] x86/pv: update guest LDT mappings using {populate,destroy}_perdomain_mapping() George Dunlap
@ 2026-08-20 17:43 ` George Dunlap
  2026-08-20 17:43 ` [PATCH 7/7] x86/mm: simplify create_perdomain_mapping() interface George Dunlap
  2026-08-21  8:45 ` [PATCH 0/7] x86: Address Space Isolation, part 1: per-domain area mapping rework Jan Beulich
  7 siblings, 0 replies; 17+ messages in thread
From: George Dunlap @ 2026-08-20 17:43 UTC (permalink / raw)
  To: xen-devel
  Cc: Roger Pau Monné, Jan Beulich, Andrew Cooper,
	Roger Pau Monné, Teddy Astie, Anthony PERARD, Michal Orzel,
	Julien Grall, Stefano Stabellini

From: Roger Pau Monné <roger.pau@citrix.com>

There are no remaining users of the stashed L1 page-tables in
d->arch.pv.gdt_ldt_l1tab.  Remove it, and all helpers.

pv_create_gdt_ldt_l1tab() now passes NIL() rather than the stash
array.  create_perdomain_mapping() still eagerly allocates the L1
tables covering the GDT/LDT range, but their addresses are no longer
handed back.  Doing this is necessary because
populate_perdomain_mapping() only fills existing tables, and treats
missing structure as a bug.

Signed-off-by: Roger Pau Monné <roger.pau@citrix.com>
Assisted-by: Claude Code:claude-fable-5
Signed-off-by: George Dunlap <gwd@xenproject.org>
---
Changes since the previously posted version:
- Note the implications of changing from pointer to NIL() in
   pv_create_gdt_ldt_l1tab().  In v2 this also changed where new
   GDT/LDT L1 tables were allocated from: upstream's capture mode
   takes them from the xenheap (the stashed pointer has to stay
   usable), the NIL() mode from the domheap.  Here they come from the
   xenheap in all modes ("x86/mm: allocate the per-domain page-tables
   from the xenheap"), so the switch only stops the addresses being
   handed back.
---
 xen/arch/x86/include/asm/domain.h |  9 ---------
 xen/arch/x86/pv/domain.c          | 10 +---------
 2 files changed, 1 insertion(+), 18 deletions(-)

diff --git a/xen/arch/x86/include/asm/domain.h b/xen/arch/x86/include/asm/domain.h
index 2c9efd59bf..7eab2ff597 100644
--- a/xen/arch/x86/include/asm/domain.h
+++ b/xen/arch/x86/include/asm/domain.h
@@ -287,8 +287,6 @@ struct time_scale {
 
 struct pv_domain
 {
-    l1_pgentry_t **gdt_ldt_l1tab;
-
     atomic_t nr_l4_pages;
 
     /* Is a 32-bit PV guest? */
@@ -525,13 +523,6 @@ struct arch_domain
 #define has_pirq(d)        (!!((d)->arch.emulation_flags & X86_EMU_USE_PIRQ))
 #define has_vpci(d)        (!!((d)->arch.emulation_flags & X86_EMU_VPCI))
 
-#define gdt_ldt_pt_idx(v) \
-      ((v)->vcpu_id >> (PAGETABLE_ORDER - GDT_LDT_VCPU_SHIFT))
-#define pv_gdt_ptes(v) \
-    ((v)->domain->arch.pv.gdt_ldt_l1tab[gdt_ldt_pt_idx(v)] + \
-     (((v)->vcpu_id << GDT_LDT_VCPU_SHIFT) & (L1_PAGETABLE_ENTRIES - 1)))
-#define pv_ldt_ptes(v) (pv_gdt_ptes(v) + 16)
-
 struct pv_vcpu
 {
     /* map_domain_page() mapping cache. */
diff --git a/xen/arch/x86/pv/domain.c b/xen/arch/x86/pv/domain.c
index 7ddab1949f..35d1761c9c 100644
--- a/xen/arch/x86/pv/domain.c
+++ b/xen/arch/x86/pv/domain.c
@@ -315,7 +315,7 @@ static int pv_create_gdt_ldt_l1tab(struct vcpu *v)
 {
     return create_perdomain_mapping(v->domain, GDT_VIRT_START(v),
                                     1U << GDT_LDT_VCPU_SHIFT,
-                                    v->domain->arch.pv.gdt_ldt_l1tab,
+                                    NIL(l1_pgentry_t *),
                                     NULL);
 }
 
@@ -389,8 +389,6 @@ void pv_domain_destroy(struct domain *d)
                               GDT_LDT_MBYTES << (20 - PAGE_SHIFT));
 
     XFREE(d->arch.pv.cpuidmasks);
-
-    FREE_XENHEAP_PAGE(d->arch.pv.gdt_ldt_l1tab);
 }
 
 void noreturn cf_check continue_pv_domain(void);
@@ -406,12 +404,6 @@ int pv_domain_initialise(struct domain *d)
 
     pv_l1tf_domain_init(d);
 
-    d->arch.pv.gdt_ldt_l1tab =
-        alloc_xenheap_pages(0, MEMF_node(domain_to_node(d)));
-    if ( !d->arch.pv.gdt_ldt_l1tab )
-        goto fail;
-    clear_page(d->arch.pv.gdt_ldt_l1tab);
-
     if ( levelling_caps & ~LCAP_faulting &&
          (d->arch.pv.cpuidmasks = xmemdup(&cpuidmask_defaults)) == NULL )
         goto fail;
-- 
2.55.0



^ permalink raw reply related	[flat|nested] 17+ messages in thread

* [PATCH 7/7] x86/mm: simplify create_perdomain_mapping() interface
  2026-08-20 17:43 [PATCH 0/7] x86: Address Space Isolation, part 1: per-domain area mapping rework George Dunlap
                   ` (5 preceding siblings ...)
  2026-08-20 17:43 ` [PATCH 6/7] x86/pv: remove stashing of GDT/LDT L1 page-tables George Dunlap
@ 2026-08-20 17:43 ` George Dunlap
  2026-08-21  8:45 ` [PATCH 0/7] x86: Address Space Isolation, part 1: per-domain area mapping rework Jan Beulich
  7 siblings, 0 replies; 17+ messages in thread
From: George Dunlap @ 2026-08-20 17:43 UTC (permalink / raw)
  To: xen-devel
  Cc: Roger Pau Monné, Jan Beulich, Andrew Cooper,
	Roger Pau Monné, Teddy Astie, Anthony PERARD, Michal Orzel,
	Julien Grall, Stefano Stabellini

From: Roger Pau Monné <roger.pau@citrix.com>

create_perdomain_mapping()'s interface is richer than any caller needs.
The paging structure for the requested range is built to a depth
selected by the pl1tab and ppg arguments, each of which distinguishes
NULL from NIL() from a real pointer:

 - nr == 0: only ensure the per-domain L3 exists; nothing else is
   allocated, and the other arguments are ignored.
 - pl1tab == a pointer: allocate the L1 tables covering the range, and
   return their addresses in the array -- the mode that existed to
   build the GDT/LDT stash.
 - pl1tab == NIL(): allocate the L1 tables, and return nothing.
 - pl1tab == NULL: do not plumb L1 tables for their own sake (they are
   still allocated on demand if data-page population requires them).
 - ppg == a pointer: allocate and install zeroed data pages across the
   range, and return their struct page_info pointers in the array.
 - ppg == NIL(): allocate and install the zeroed data pages, but hand
   nothing back; the pages are reachable only through the mapping.
 - ppg == NULL: do not allocate data pages.
 - both NULL, nr > 0: stop after the slot's L2; do not plumb L1 tables
   at all.

Very few of these modes have users now.  The last user of the
pl1tab capture mode was removed when we removed the GDT/LDT stash.
The ppg capture mode never had any users.  Nothing uses the both-NULL
L2-only mode with nr != 0.  What remains is exactly one bit of
information: whether the caller wants the range populated with zeroed,
area-owned data pages, or merely plumbed down to the L1 tables, ready
for populate_perdomain_mapping() to install caller-owned pages.

Replace the two arguments with a boolean expressing that bit.  With the
stashing mode gone the NIL()/IS_NIL() macros lose their last user, so
drop them as well; and document the resulting interface.

No caller changes behaviour: every existing call maps onto the boolean
exactly.

Signed-off-by: Roger Pau Monné <roger.pau@citrix.com>
Assisted-by: Claude Code:claude-fable-5
Signed-off-by: George Dunlap <gwd@xenproject.org>
---
Changes since the previously posted version:
- Drop the now-unused NIL()/IS_NIL() macros as requested during review
- Describe the prior interface in the commit message and add a doc
   comment for the simplified one.
---
 xen/arch/x86/domain_page.c    | 10 ++++-----
 xen/arch/x86/hvm/hvm.c        |  2 +-
 xen/arch/x86/include/asm/mm.h |  6 +-----
 xen/arch/x86/mm.c             | 40 +++++++++++++++++++++++------------
 xen/arch/x86/pv/domain.c      |  4 +---
 xen/arch/x86/x86_64/mm.c      |  3 +--
 6 files changed, 35 insertions(+), 30 deletions(-)

diff --git a/xen/arch/x86/domain_page.c b/xen/arch/x86/domain_page.c
index 72c00194f3..1c1deeeebc 100644
--- a/xen/arch/x86/domain_page.c
+++ b/xen/arch/x86/domain_page.c
@@ -246,8 +246,7 @@ int mapcache_domain_init(struct domain *d)
     spin_lock_init(&dcache->lock);
 
     return create_perdomain_mapping(d, (unsigned long)dcache->inuse,
-                                    2 * bitmap_pages + 1,
-                                    NIL(l1_pgentry_t *), NULL);
+                                    2 * bitmap_pages + 1, false);
 }
 
 int mapcache_vcpu_init(struct vcpu *v)
@@ -264,16 +263,15 @@ int mapcache_vcpu_init(struct vcpu *v)
     if ( ents > dcache->entries )
     {
         /* Populate page tables. */
-        int rc = create_perdomain_mapping(d, MAPCACHE_VIRT_START, ents,
-                                          NIL(l1_pgentry_t *), NULL);
+        int rc = create_perdomain_mapping(d, MAPCACHE_VIRT_START, ents, false);
 
         /* Populate bit maps. */
         if ( !rc )
             rc = create_perdomain_mapping(d, (unsigned long)dcache->inuse,
-                                          nr, NULL, NIL(struct page_info *));
+                                          nr, true);
         if ( !rc )
             rc = create_perdomain_mapping(d, (unsigned long)dcache->garbage,
-                                          nr, NULL, NIL(struct page_info *));
+                                          nr, true);
 
         if ( rc )
             return rc;
diff --git a/xen/arch/x86/hvm/hvm.c b/xen/arch/x86/hvm/hvm.c
index 955fc062a5..e8fbe7ba27 100644
--- a/xen/arch/x86/hvm/hvm.c
+++ b/xen/arch/x86/hvm/hvm.c
@@ -620,7 +620,7 @@ int hvm_domain_initialise(struct domain *d,
     INIT_LIST_HEAD(&d->arch.hvm.mmcfg_regions);
     INIT_LIST_HEAD(&d->arch.hvm.msix_tables);
 
-    rc = create_perdomain_mapping(d, PERDOMAIN_VIRT_START, 0, NULL, NULL);
+    rc = create_perdomain_mapping(d, PERDOMAIN_VIRT_START, 0, false);
     if ( rc )
         goto fail;
 
diff --git a/xen/arch/x86/include/asm/mm.h b/xen/arch/x86/include/asm/mm.h
index 1888807394..30eaec9179 100644
--- a/xen/arch/x86/include/asm/mm.h
+++ b/xen/arch/x86/include/asm/mm.h
@@ -600,12 +600,8 @@ long arch_memory_op(unsigned long cmd, XEN_GUEST_HANDLE_PARAM(void) arg);
 long subarch_memory_op(unsigned long cmd, XEN_GUEST_HANDLE_PARAM(void) arg);
 int compat_arch_memory_op(unsigned long cmd, XEN_GUEST_HANDLE_PARAM(void) arg);
 
-#define NIL(type) ((type *)-sizeof(type))
-#define IS_NIL(ptr) (!((uintptr_t)(ptr) + sizeof(*(ptr))))
-
 int create_perdomain_mapping(struct domain *d, unsigned long va,
-                             unsigned int nr, l1_pgentry_t **pl1tab,
-                             struct page_info **ppg);
+                             unsigned int nr, bool populate);
 void populate_perdomain_mapping(const struct vcpu *v, unsigned long va,
                                 const mfn_t *mfn, unsigned int nr,
                                 unsigned int flags);
diff --git a/xen/arch/x86/mm.c b/xen/arch/x86/mm.c
index 1810971677..aa26d12ac2 100644
--- a/xen/arch/x86/mm.c
+++ b/xen/arch/x86/mm.c
@@ -6211,9 +6211,33 @@ static bool perdomain_l1e_needs_freeing(l1_pgentry_t l1e)
            (_PAGE_PRESENT | _PAGE_AVAIL0);
 }
 
+/*
+ * Ensure the paging structure for [va, va + nr * PAGE_SIZE) of d's
+ * per-domain area is in place, allocating whichever levels are missing:
+ * the (domain-wide) L3 root, the slot's L2, and all L1 tables covering
+ * the range.  The page-tables come from the xenheap, so that they stay
+ * reachable through their always-mapped alias; populated pages come from
+ * the domain heap.  The range must lie within a single per-domain slot
+ * (one L3 entry), and already-present levels and entries are left
+ * untouched, so calls are idempotent over existing ranges.
+ *
+ * nr == 0: only ensure the per-domain L3 itself exists; populate is
+ * ignored.  Used to set the area up before any sub-range is known.
+ *
+ * populate == false: stop once the L1 tables are in place.  The range is
+ * then ready for caller-owned pages to be mapped and unmapped via
+ * populate_perdomain_mapping() / destroy_perdomain_mapping(), which only
+ * fill (or clear) existing tables; populate treats missing structure as a
+ * bug, destroy skips it.
+ *
+ * populate == true: additionally install a freshly allocated, zeroed page
+ * at every not-yet-present entry in the range.  Such pages are marked
+ * _PAGE_AVAIL0, "owned by the per-domain area": teardown frees them (see
+ * perdomain_l1e_needs_freeing()), whereas caller-owned mappings are only
+ * ever unmapped.
+ */
 int create_perdomain_mapping(struct domain *d, unsigned long va,
-                             unsigned int nr, l1_pgentry_t **pl1tab,
-                             struct page_info **ppg)
+                             unsigned int nr, bool populate)
 {
     struct page_info *pg;
     l3_pgentry_t *l3tab;
@@ -6258,9 +6282,6 @@ int create_perdomain_mapping(struct domain *d, unsigned long va,
     else
         l2tab = maddr_to_virt(l3e_get_paddr(l3tab[l3_table_offset(va)]));
 
-    if ( !pl1tab && !ppg )
-        return 0;
-
     for ( l1tab = NULL; !rc && nr--; )
     {
         l2_pgentry_t *pl2e = l2tab + l2_table_offset(va);
@@ -6273,26 +6294,19 @@ int create_perdomain_mapping(struct domain *d, unsigned long va,
                 rc = -ENOMEM;
                 break;
             }
-            if ( pl1tab && !IS_NIL(pl1tab) )
-            {
-                ASSERT(!pl1tab[l2_table_offset(va)]);
-                pl1tab[l2_table_offset(va)] = l1tab;
-            }
             clear_page(l1tab);
             *pl2e = l2e_from_mfn(virt_to_mfn(l1tab), __PAGE_HYPERVISOR_RW);
         }
         else if ( !l1tab )
             l1tab = maddr_to_virt(l2e_get_paddr(*pl2e));
 
-        if ( ppg &&
+        if ( populate &&
              !(l1e_get_flags(l1tab[l1_table_offset(va)]) & _PAGE_PRESENT) )
         {
             pg = alloc_domheap_page(d, MEMF_no_owner);
             if ( pg )
             {
                 clear_domain_page(page_to_mfn(pg));
-                if ( !IS_NIL(ppg) )
-                    *ppg++ = pg;
                 l1tab[l1_table_offset(va)] =
                     l1e_from_page(pg, __PAGE_HYPERVISOR_RW | _PAGE_AVAIL0);
                 l2e_add_flags(*pl2e, _PAGE_AVAIL0);
diff --git a/xen/arch/x86/pv/domain.c b/xen/arch/x86/pv/domain.c
index 35d1761c9c..15a8238aff 100644
--- a/xen/arch/x86/pv/domain.c
+++ b/xen/arch/x86/pv/domain.c
@@ -314,9 +314,7 @@ int switch_compat(struct domain *d)
 static int pv_create_gdt_ldt_l1tab(struct vcpu *v)
 {
     return create_perdomain_mapping(v->domain, GDT_VIRT_START(v),
-                                    1U << GDT_LDT_VCPU_SHIFT,
-                                    NIL(l1_pgentry_t *),
-                                    NULL);
+                                    1U << GDT_LDT_VCPU_SHIFT, false);
 }
 
 static void pv_destroy_gdt_ldt_l1tab(struct vcpu *v)
diff --git a/xen/arch/x86/x86_64/mm.c b/xen/arch/x86/x86_64/mm.c
index 8eadab7933..ffeda06e08 100644
--- a/xen/arch/x86/x86_64/mm.c
+++ b/xen/arch/x86/x86_64/mm.c
@@ -733,8 +733,7 @@ void __init zap_low_mappings(void)
 int setup_compat_arg_xlat(struct vcpu *v)
 {
     return create_perdomain_mapping(v->domain, ARG_XLAT_START(v),
-                                    PFN_UP(COMPAT_ARG_XLAT_SIZE),
-                                    NULL, NIL(struct page_info *));
+                                    PFN_UP(COMPAT_ARG_XLAT_SIZE), true);
 }
 
 void free_compat_arg_xlat(struct vcpu *v)
-- 
2.55.0



^ permalink raw reply related	[flat|nested] 17+ messages in thread

* Re: [PATCH 0/7] x86: Address Space Isolation, part 1: per-domain area mapping rework
  2026-08-20 17:43 [PATCH 0/7] x86: Address Space Isolation, part 1: per-domain area mapping rework George Dunlap
                   ` (6 preceding siblings ...)
  2026-08-20 17:43 ` [PATCH 7/7] x86/mm: simplify create_perdomain_mapping() interface George Dunlap
@ 2026-08-21  8:45 ` Jan Beulich
  2026-08-21 15:17   ` George Dunlap
  7 siblings, 1 reply; 17+ messages in thread
From: Jan Beulich @ 2026-08-21  8:45 UTC (permalink / raw)
  To: George Dunlap
  Cc: Andrew Cooper, Roger Pau Monné, Teddy Astie, Anthony PERARD,
	Michal Orzel, Julien Grall, Stefano Stabellini, xen-devel

On 20.08.2026 19:43, George Dunlap wrote:
> One point reviewers may want to look at specifically: patch 1 changes
> where the per-domain page-tables are allocated from, and its commit
> message discusses the (minor) NUMA-placement consequence.

While I don't recall which recent patch (series) it was, I can't very well
say "no new xenheap allocations please" there without also saying so here.
I've read over patch 1's description, and while it tries to justify this
accordingly, I still remain concerned. I think we simply have to accept
the mapping overhead, to avoid allocating from a pool which - over time -
is representing a decreasing portion of total memory systems have (on
average, and not even considering systems with extremely sparse memory
layouts, and with perhaps PDX compression not doing good enough to
compensate).

Jan


^ permalink raw reply	[flat|nested] 17+ messages in thread

* Re: [PATCH 0/7] x86: Address Space Isolation, part 1: per-domain area mapping rework
  2026-08-21  8:45 ` [PATCH 0/7] x86: Address Space Isolation, part 1: per-domain area mapping rework Jan Beulich
@ 2026-08-21 15:17   ` George Dunlap
  2026-08-21 15:36     ` George Dunlap
  2026-08-24  9:02     ` Jan Beulich
  0 siblings, 2 replies; 17+ messages in thread
From: George Dunlap @ 2026-08-21 15:17 UTC (permalink / raw)
  To: Jan Beulich
  Cc: Andrew Cooper, Roger Pau Monné, Teddy Astie, Anthony PERARD,
	Michal Orzel, Julien Grall, Stefano Stabellini, xen-devel

On Fri, Aug 21, 2026 at 9:45 AM Jan Beulich <jbeulich@suse.com> wrote:
>
> On 20.08.2026 19:43, George Dunlap wrote:
> > One point reviewers may want to look at specifically: patch 1 changes
> > where the per-domain page-tables are allocated from, and its commit
> > message discusses the (minor) NUMA-placement consequence.
>
> While I don't recall which recent patch (series) it was, I can't very well
> say "no new xenheap allocations please" there without also saying so here.
> I've read over patch 1's description, and while it tries to justify this
> accordingly, I still remain concerned. I think we simply have to accept
> the mapping overhead, to avoid allocating from a pool which - over time -
> is representing a decreasing portion of total memory systems have (on
> average, and not even considering systems with extremely sparse memory
> layouts, and with perhaps PDX compression not doing good enough to
> compensate).

You should certainly have the same resistance to adding new xenheap
allocations.  But looking at the numbers, I don't see that we're
anywhere near the point where we say, "Absolutely no new xenheap
allocations, regardless of the cost."  domheap+vmap looks like it was
cheap and easy alternative for Teddy, but the alternatives here
aren't, compared to the cost of extra xenheap allocations.

So let's lay everything out.

My understanding is that we have the following two issues allocating
things in the xenheap on systems larger than 4T:

- The total amount of xenheap space is limited to 4 TiB of virtual
  address space.  On some systems, this may correspond to 4 TiB of
  actual RAM; but on a machine whose RAM layout is sparser, the
  actual RAM addressable in this window may be far less

- It's not symmetric NUMA-wise; so the larger the system, the more of
  the xenheap will end up being from the same NUMA node.  This will
  limit Xen's ability to have NUMA-local data structures, and its
  ability to give NUMA-local data to guests running on node 0.

Looking at this series as a whole, although the first patch adds pages
to the xenheap, the end goal of the rest of the work is to remove
pages from the xenheap. Things added in:

- Making the perdomain area per-vCPU, with its pagetables allocated
  from the xenheap, adds a per-vCPU L3 plus an L2+L1 pair for each
  slot in use.  This totals 5 pages/vCPU for HVM guests and 8 pages/vCPU
  for PV guests.

  (Note that the GDT/LDT L1s are already allocated from the xenheap
  today, but per-domain rather than per-vCPU.)

Things removed:

- Per-pCPU stacks -- 8 xenheap pages / pCPU

- AMD VMCB - one xenheap page / vCPU

- VMX guest MSR area: 1 page per vCPU

- sub-page XSAVE areas (~2.7 KiB/vCPU of xmalloc pool today; planned
  follow-on work aggregating other miscellaneous xmalloc'd guest state
  should take this to about a page per vCPU)

To do some math: current security-supported limits for x86 are 4096
pCPUs on a 12TiB system.  Suppose we have an 8:1 vCPU:pCPU ratio, and
an average of 8 vcpus per domain.  So 32768 total vCPUs and 4096
domains.  On a Full ASI system, vcpu-pt on all domains, per-CPU stacks
on, all Intel HVM domains, we get numbers like the following:

Added to xenheap:

- Per-vCPU tables, 5/vCPU (L3; mapcache L2+L1; state-window L2+L1):
  5 × 32,768 = 163,840 pages = 640 MiB
- Per-pCPU stack tables, 2/pCPU: 2 × 4,096 = 8,192 pages = 32 MiB
  (→ 0: these are only written at CPU bring-up and tear-down, so we
  have already moved them to the domheap in the working branch --
  which also makes them NUMA-local unconditionally)
- Per-domain tables: replaced by the per-vCPU sets in vcpu-pt mode → 0
- Total added: 172,032 pages = 672 MiB

Removed from xenheap:

- Stacks, 8/pCPU: 8 × 4,096 = 32,768 pages = 128 MiB
- XSAVE, ~2.7 KiB/vCPU from the xmalloc pools: 32,768 × 2.7 KiB ≈
  21,600 pages ≈ 86 MiB (0 if guests get AMX — those areas are domheap
  today)
- VMX guest MSR page: lazily allocated, typically absent → 0 (upper
  bound 128 MiB if every vCPU used one)
- Total removed: ≈ 54,400 pages ≈ 214 MiB

Net: +117,600 pages ~ +458 MiB — against a 4 TiB window (0.011%), on
a 12 TiB host (0.0036%).

An AMD HVM fleet would include VMCB removal (32,768 pages = 128 MiB) →
net +330 MiB. A PV fleet is the worst case — 8 tables/vCPU (add
GDT/LDT L2+L1 and the per-vCPU root) → 1,056 MiB added, 214 removed,
net +842 MiB.

I have explored a number of other options, to various levels of depth.

One is map_domain_page_irqoff(): If the caller promises to keep
interrupts disabled until unmap_domain_page_irqoff(), we can safely
perform maps in a context switch without having to worry about
sync_lazy_execstate.  (This was actually implemented and almost sent
on Tuesday evening, when I noticed your review of Roger's v2 saying,
"Question is whether it's a good idea in the first place to start
using map_domain_page() from the context switch path.  Surely there
are possible alternatives.")  This maps all vcpu pages from the
domheap, adding nothing to the xenheap *or* the vmap area.  But it
costs 9 map/unmap pairs *per context switch*.

I absolutely reject the idea that because on a 12TiB system with 32k
PV vCPUs, we take up an extra 0.02% of the xenheap area, that a laptop
running QubesOS has to do 9 maps and unmaps per context switch.  That
is not a valid cost/benefits tradeoff.  In the worst case we could
just add a switch to such a system, allowing people who find their
xenheap too full to use the mapcache version instead.  (We could even
turn this on automatically at boot based on projected xenheap
utilization.)

There are other options I've explored:

- domheap + vmap; basically, allocate from domheap, map in the vmap
  area.  On paper this sounds like the same thing; the problem is that
  we don't have a simple MFN -> VA mapping, as we do in the xenheap
  case, so the walk is a lot harder; we start to have to do lookups,
  significantly increasing the cost over simple memory reads and math.
  (This is the difference from the intremap table on the VT-d thread:
  that's a leaf structure reached from a single pointer, so a
  permanent vmap costs nothing there.  Pagetable hierarchies are
  exactly the case where the MFN -> VA step is critical: each entry
  read yields an MFN, which the walk has to turn into the next VA.)
  And if we're concerned about "xenheap creep", when we have a 4 TiB
  ceiling, shouldn't we also be worried about "vmap creep", when we
  have a 64 GiB ceiling?

- Stash everything we need; basically, an extension of the current
  gdt_ldt_l1tab functionality.  Allocate everything from the domheap,
  map it in the vmap area (moving gdt_ldt_l1tab there as well), keep
  pointers to all the things we need to modify on context switch, so
  we don't need to walk the tables.  This would basically be, three
  pointers per vCPU: a pointer to its GDT/LDT L1, a pointer to its
  per-vCPU L3, and a pointer to the per-vCPU root_pgt.  (This would
  put ~384 MiB of mappings into the 64 GiB vmap region -- 0.6%, shared
  with ioremap and the fixmap -- to avoid 0.02% of the xenheap
  window.)

Both the vmap options have two complications, compared to the posted
option.  One thing to worry about here would be the additional stress
on the vmap allocator: It's a linear bitmap scan under one global
lock, designed for dozens-to-hundreds of ioremaps, not ~100k
long-lived single-page mappings (32k vCPUs x 3 pages per vCPU in the
"stash everything" case).

The second is that we begin to run into bootstrapping issues.  With
the xenheap approach, we can begin building and walking pagetables
very early in boot in the same manner in which they'll be walked
throughout Xen's lifecycle.  With the vmap approach, we need to deal
with the fact that the vmap area itself isn't up until later.

The final option I looked at was mapping the incoming vcpu's linear
map to edit it ("altlinmap").  That still adds a map/unmap per context
switch, and requires some additional complication to handle
ASI/non-ASI systems.

Xen already consistently allocates its page tables from the xenheap
whenever it needs to access them during a context switch:
alloc_xen_pagetable() has allocated from the domheap since Hongyan's
directmap-removal preparation (those tables are only ever walked in
contexts where map_domain_page() works), but XPTI's per-CPU root_pgt
is alloc_xenheap_page(), precisely because it has to be written on the
context-switch path.  The same for the PV GDT / LDT L1 tables.  The
series follows the same rule for the same reason.

Ultimately, I think there's a lot of wisdom in the saying, "Premature
optimization is the root of all evil."  As I said, it's certainly
right to be on our guard against adding things to xenheap, and look at
alternatives; but we're nowhere near the point where we need to say,
"Absolutely nothing added, regardless of the cost."  The design here
is not locking us into the pages long-term; alternate designs have a
significant cost in terms of authoring, reviewing, code complexity and
maintenance, and code performance.  At such time as we find systems
where the xenheap allocations introduced in this series become a
problem, we have a number of potential ways to mitigate the problem,
including switching to mapcache *on systems with the problem*, or
switching to a number of the other more complicated approaches.

 -George


^ permalink raw reply	[flat|nested] 17+ messages in thread

* Re: [PATCH 0/7] x86: Address Space Isolation, part 1: per-domain area mapping rework
  2026-08-21 15:17   ` George Dunlap
@ 2026-08-21 15:36     ` George Dunlap
  2026-08-24  9:02     ` Jan Beulich
  1 sibling, 0 replies; 17+ messages in thread
From: George Dunlap @ 2026-08-21 15:36 UTC (permalink / raw)
  To: Jan Beulich
  Cc: Andrew Cooper, Roger Pau Monné, Teddy Astie, Anthony PERARD,
	Michal Orzel, Julien Grall, Stefano Stabellini, xen-devel

On Fri, Aug 21, 2026 at 4:17 PM George Dunlap <gwd@xenproject.org> wrote:

> Ultimately, I think there's a lot of wisdom in the saying, "Premature
> optimization is the root of all evil."  As I said, it's certainly
> right to be on our guard against adding things to xenheap, and look at
> alternatives; but we're nowhere near the point where we need to say,
> "Absolutely nothing added, regardless of the cost."  The design here
> is not locking us into the pages long-term; alternate designs have a
> significant cost in terms of authoring, reviewing, code complexity and
> maintenance, and code performance.

To get a sense of how much this approach simplifies things, look at
how patch 1 of this series simplified create_perdomain_mapping; and
look at how much simpler patch 2 is compared to Roger's implementation
[1].  The xenheap option is way easier to review and maintain.

 -George

[1] https://marc.info/?l=xen-devel&m=173634640624259


^ permalink raw reply	[flat|nested] 17+ messages in thread

* Re: [PATCH 3/7] x86/pv: use populate_perdomain_mapping() to map the Xen GDT
  2026-08-20 17:43 ` [PATCH 3/7] x86/pv: use populate_perdomain_mapping() to map the Xen GDT George Dunlap
@ 2026-08-21 20:59   ` Andrew Cooper
  0 siblings, 0 replies; 17+ messages in thread
From: Andrew Cooper @ 2026-08-21 20:59 UTC (permalink / raw)
  To: George Dunlap, xen-devel
  Cc: Andrew Cooper, Roger Pau Monné, Jan Beulich,
	Roger Pau Monné, Teddy Astie, Anthony PERARD, Michal Orzel,
	Julien Grall, Stefano Stabellini

On 20/08/2026 6:43 pm, George Dunlap wrote:
> diff --git a/xen/arch/x86/traps.c b/xen/arch/x86/traps.c
> index 1774966305..2fcaa413f7 100644
> --- a/xen/arch/x86/traps.c
> +++ b/xen/arch/x86/traps.c
> @@ -71,10 +71,10 @@ DEFINE_PER_CPU(uint64_t, efer);
>  static DEFINE_PER_CPU(unsigned long, last_extable_addr);
>  
>  DEFINE_PER_CPU_READ_MOSTLY(seg_desc_t *, gdt);
> -DEFINE_PER_CPU_READ_MOSTLY(l1_pgentry_t, gdt_l1e);
> +DEFINE_PER_CPU_READ_MOSTLY(mfn_t, gdt_mfn);
>  #ifdef CONFIG_PV32
>  DEFINE_PER_CPU_READ_MOSTLY(seg_desc_t *, compat_gdt);
> -DEFINE_PER_CPU_READ_MOSTLY(l1_pgentry_t, compat_gdt_l1e);
> +DEFINE_PER_CPU_READ_MOSTLY(mfn_t, compat_gdt_mfn);
>  #endif

I am not convinced by this change.

The reason we had the L1e stashed in the first place is because mfn <->
maddr conversions are expensive.  Stashing the L1e in this way proved to
be a win in the context switch path, despite the fragility it added by
needing to maintain a second form of the same data.

mfn <-> maddr conversions have changed expense since the optimisation
was first put in, but one form is now even more expensive than when the
optimisation was first put in.

Stashing the L1e means doing the conversion once at boot.  Anything else
means doing it on every context switch path.

So what this patch is doing is still keeping the double copy (the
fragility) but reintroducing the expensive part of the operation into
the context switch path.  If you can't keep it being L1e, there's
probably no point keeping the optimisation at all.

~Andrew


^ permalink raw reply	[flat|nested] 17+ messages in thread

* Re: [PATCH 0/7] x86: Address Space Isolation, part 1: per-domain area mapping rework
  2026-08-21 15:17   ` George Dunlap
  2026-08-21 15:36     ` George Dunlap
@ 2026-08-24  9:02     ` Jan Beulich
  2026-08-25 11:42       ` George Dunlap
  1 sibling, 1 reply; 17+ messages in thread
From: Jan Beulich @ 2026-08-24  9:02 UTC (permalink / raw)
  To: George Dunlap
  Cc: Andrew Cooper, Roger Pau Monné, Teddy Astie, Anthony PERARD,
	Michal Orzel, Julien Grall, Stefano Stabellini, xen-devel

On 21.08.2026 17:17, George Dunlap wrote:
> On Fri, Aug 21, 2026 at 9:45 AM Jan Beulich <jbeulich@suse.com> wrote:
>> On 20.08.2026 19:43, George Dunlap wrote:
>>> One point reviewers may want to look at specifically: patch 1 changes
>>> where the per-domain page-tables are allocated from, and its commit
>>> message discusses the (minor) NUMA-placement consequence.
>>
>> While I don't recall which recent patch (series) it was, I can't very well
>> say "no new xenheap allocations please" there without also saying so here.
>> I've read over patch 1's description, and while it tries to justify this
>> accordingly, I still remain concerned. I think we simply have to accept
>> the mapping overhead, to avoid allocating from a pool which - over time -
>> is representing a decreasing portion of total memory systems have (on
>> average, and not even considering systems with extremely sparse memory
>> layouts, and with perhaps PDX compression not doing good enough to
>> compensate).
> 
> You should certainly have the same resistance to adding new xenheap
> allocations.  But looking at the numbers, I don't see that we're
> anywhere near the point where we say, "Absolutely no new xenheap
> allocations, regardless of the cost."

Well, that depends, and in part on the longer term plans with ASI. It has
been my (silent) assumption that eventually the directmap would go away
altogether when ASI is in use, with the VA space freed (almost?) all
becoming available for vmap(). With the disappearance of directmap, the
xenheap would naturally disappear as well. Hence putting stuff there in
new work actually adds to our technical debt.

>  domheap+vmap looks like it was
> cheap and easy alternative for Teddy, but the alternatives here
> aren't, compared to the cost of extra xenheap allocations.
> 
> So let's lay everything out.
> 
> My understanding is that we have the following two issues allocating
> things in the xenheap on systems larger than 4T:
> 
> - The total amount of xenheap space is limited to 4 TiB of virtual
>   address space.  On some systems, this may correspond to 4 TiB of
>   actual RAM; but on a machine whose RAM layout is sparser, the
>   actual RAM addressable in this window may be far less
> 
> - It's not symmetric NUMA-wise; so the larger the system, the more of
>   the xenheap will end up being from the same NUMA node.  This will
>   limit Xen's ability to have NUMA-local data structures, and its
>   ability to give NUMA-local data to guests running on node 0.
> 
> Looking at this series as a whole, although the first patch adds pages
> to the xenheap, the end goal of the rest of the work is to remove
> pages from the xenheap. Things added in:
> 
> - Making the perdomain area per-vCPU, with its pagetables allocated
>   from the xenheap, adds a per-vCPU L3 plus an L2+L1 pair for each
>   slot in use.  This totals 5 pages/vCPU for HVM guests and 8 pages/vCPU
>   for PV guests.
> 
>   (Note that the GDT/LDT L1s are already allocated from the xenheap
>   today, but per-domain rather than per-vCPU.)
> 
> Things removed:
> 
> - Per-pCPU stacks -- 8 xenheap pages / pCPU
> 
> - AMD VMCB - one xenheap page / vCPU
> 
> - VMX guest MSR area: 1 page per vCPU
> 
> - sub-page XSAVE areas (~2.7 KiB/vCPU of xmalloc pool today; planned
>   follow-on work aggregating other miscellaneous xmalloc'd guest state
>   should take this to about a page per vCPU)
> 
> To do some math: current security-supported limits for x86 are 4096
> pCPUs on a 12TiB system.  Suppose we have an 8:1 vCPU:pCPU ratio, and
> an average of 8 vcpus per domain.  So 32768 total vCPUs and 4096
> domains.  On a Full ASI system, vcpu-pt on all domains, per-CPU stacks
> on, all Intel HVM domains, we get numbers like the following:
> 
> Added to xenheap:
> 
> - Per-vCPU tables, 5/vCPU (L3; mapcache L2+L1; state-window L2+L1):
>   5 × 32,768 = 163,840 pages = 640 MiB
> - Per-pCPU stack tables, 2/pCPU: 2 × 4,096 = 8,192 pages = 32 MiB
>   (→ 0: these are only written at CPU bring-up and tear-down, so we
>   have already moved them to the domheap in the working branch --
>   which also makes them NUMA-local unconditionally)
> - Per-domain tables: replaced by the per-vCPU sets in vcpu-pt mode → 0
> - Total added: 172,032 pages = 672 MiB
> 
> Removed from xenheap:
> 
> - Stacks, 8/pCPU: 8 × 4,096 = 32,768 pages = 128 MiB
> - XSAVE, ~2.7 KiB/vCPU from the xmalloc pools: 32,768 × 2.7 KiB ≈
>   21,600 pages ≈ 86 MiB (0 if guests get AMX — those areas are domheap
>   today)
> - VMX guest MSR page: lazily allocated, typically absent → 0 (upper
>   bound 128 MiB if every vCPU used one)
> - Total removed: ≈ 54,400 pages ≈ 214 MiB
> 
> Net: +117,600 pages ~ +458 MiB — against a 4 TiB window (0.011%), on
> a 12 TiB host (0.0036%).

While these percentiles in particular of course look very tiny, they are
applicable only on systems having no meaningful gaps in the physical
address map. And even more generally I find all of these calculations
only partly convincing, not the least because you start out from numbers
which look pretty contrived when comparing to actual systems which would
run the new code. (Using more realistic real-system values may end up
going in favor of what you want to convey, or it may not.)

> I have explored a number of other options, to various levels of depth.
> 
> One is map_domain_page_irqoff(): If the caller promises to keep
> interrupts disabled until unmap_domain_page_irqoff(), we can safely
> perform maps in a context switch without having to worry about
> sync_lazy_execstate.  (This was actually implemented and almost sent
> on Tuesday evening, when I noticed your review of Roger's v2 saying,
> "Question is whether it's a good idea in the first place to start
> using map_domain_page() from the context switch path.  Surely there
> are possible alternatives.")  This maps all vcpu pages from the
> domheap, adding nothing to the xenheap *or* the vmap area.  But it
> costs 9 map/unmap pairs *per context switch*.

But why would not using vmap() be a necessary conclusion of my initial
comment? All I'm objecting to are new uses of the xenheap.

> I absolutely reject the idea that because on a 12TiB system with 32k
> PV vCPUs, we take up an extra 0.02% of the xenheap area, that a laptop
> running QubesOS has to do 9 maps and unmaps per context switch.  That
> is not a valid cost/benefits tradeoff.  In the worst case we could
> just add a switch to such a system, allowing people who find their
> xenheap too full to use the mapcache version instead.  (We could even
> turn this on automatically at boot based on projected xenheap
> utilization.)

Maybe, yet extra overhead may be a necessary (but hopefully only
transient) price to pay in the course of the transformation.

> There are other options I've explored:
> 
> - domheap + vmap; basically, allocate from domheap, map in the vmap
>   area.  On paper this sounds like the same thing; the problem is that
>   we don't have a simple MFN -> VA mapping, as we do in the xenheap
>   case, so the walk is a lot harder; we start to have to do lookups,
>   significantly increasing the cost over simple memory reads and math.
>   (This is the difference from the intremap table on the VT-d thread:
>   that's a leaf structure reached from a single pointer, so a
>   permanent vmap costs nothing there.  Pagetable hierarchies are
>   exactly the case where the MFN -> VA step is critical: each entry
>   read yields an MFN, which the walk has to turn into the next VA.)

The pages used here are entirely private to logic handling those page
tables. Hence a struct page_info field can very likely be used to stash
the VA of a permanent mapping. (Feels like similarly I must have
suggested this somewhere else recently, yet I don't recall the context.)

>   And if we're concerned about "xenheap creep", when we have a 4 TiB
>   ceiling, shouldn't we also be worried about "vmap creep", when we
>   have a 64 GiB ceiling?

Absolutely, and I have been mentioning the need to consider growing this
area in a number of situations (one iirc again pretty recently).

> - Stash everything we need; basically, an extension of the current
>   gdt_ldt_l1tab functionality.  Allocate everything from the domheap,
>   map it in the vmap area (moving gdt_ldt_l1tab there as well), keep
>   pointers to all the things we need to modify on context switch, so
>   we don't need to walk the tables.  This would basically be, three
>   pointers per vCPU: a pointer to its GDT/LDT L1, a pointer to its
>   per-vCPU L3, and a pointer to the per-vCPU root_pgt.  (This would
>   put ~384 MiB of mappings into the 64 GiB vmap region -- 0.6%, shared
>   with ioremap and the fixmap -- to avoid 0.02% of the xenheap
>   window.)
> 
> Both the vmap options have two complications, compared to the posted
> option.  One thing to worry about here would be the additional stress
> on the vmap allocator: It's a linear bitmap scan under one global
> lock, designed for dozens-to-hundreds of ioremaps, not ~100k
> long-lived single-page mappings (32k vCPUs x 3 pages per vCPU in the
> "stash everything" case).

Indeed, heavier use of that allocator may require work to be done there.

> The second is that we begin to run into bootstrapping issues.  With
> the xenheap approach, we can begin building and walking pagetables
> very early in boot in the same manner in which they'll be walked
> throughout Xen's lifecycle.  With the vmap approach, we need to deal
> with the fact that the vmap area itself isn't up until later.

Valid concern, yet surely possible to deal with.

> The final option I looked at was mapping the incoming vcpu's linear
> map to edit it ("altlinmap").  That still adds a map/unmap per context
> switch, and requires some additional complication to handle
> ASI/non-ASI systems.
> 
> Xen already consistently allocates its page tables from the xenheap
> whenever it needs to access them during a context switch:
> alloc_xen_pagetable() has allocated from the domheap since Hongyan's
> directmap-removal preparation (those tables are only ever walked in
> contexts where map_domain_page() works), but XPTI's per-CPU root_pgt
> is alloc_xenheap_page(), precisely because it has to be written on the
> context-switch path.  The same for the PV GDT / LDT L1 tables.  The
> series follows the same rule for the same reason.

"Rule" is a strong word. XPTI at the time needed to be done quickly.
The inability to map_domain_page() from the context switch path left
xenheap as the only viable option. Whereas with ASI, as said at the
top, phasing out directmap (and hence xenheap) as a concept is (imo) a
mid- to long-term goal.

> Ultimately, I think there's a lot of wisdom in the saying, "Premature
> optimization is the root of all evil."  As I said, it's certainly
> right to be on our guard against adding things to xenheap, and look at
> alternatives; but we're nowhere near the point where we need to say,
> "Absolutely nothing added, regardless of the cost."  The design here
> is not locking us into the pages long-term; alternate designs have a
> significant cost in terms of authoring, reviewing, code complexity and
> maintenance, and code performance.  At such time as we find systems
> where the xenheap allocations introduced in this series become a
> problem, we have a number of potential ways to mitigate the problem,
> including switching to mapcache *on systems with the problem*, or
> switching to a number of the other more complicated approaches.

I'm a little puzzled by you talking of "optimization" (premature or
not) here. In my initial reply I did point out a functional aspect, and
I made clear that I'm aware that this is going to have a performance
impact. I.e. quite the opposite of "optimization".

Jan


^ permalink raw reply	[flat|nested] 17+ messages in thread

* Re: [PATCH 0/7] x86: Address Space Isolation, part 1: per-domain area mapping rework
  2026-08-24  9:02     ` Jan Beulich
@ 2026-08-25 11:42       ` George Dunlap
  2026-08-25 12:00         ` Juergen Gross
  2026-08-25 13:28         ` Jan Beulich
  0 siblings, 2 replies; 17+ messages in thread
From: George Dunlap @ 2026-08-25 11:42 UTC (permalink / raw)
  To: Jan Beulich
  Cc: Andrew Cooper, Roger Pau Monné, Teddy Astie, Anthony PERARD,
	Michal Orzel, Julien Grall, Stefano Stabellini, xen-devel

On Mon, Aug 24, 2026 at 10:02 AM Jan Beulich <jbeulich@suse.com> wrote:
> On 21.08.2026 17:17, George Dunlap wrote:
> > On Fri, Aug 21, 2026 at 9:45 AM Jan Beulich <jbeulich@suse.com> wrote:
> >> On 20.08.2026 19:43, George Dunlap wrote:
> >>> One point reviewers may want to look at specifically: patch 1 changes
> >>> where the per-domain page-tables are allocated from, and its commit
> >>> message discusses the (minor) NUMA-placement consequence.
> >>
> >> While I don't recall which recent patch (series) it was, I can't very well
> >> say "no new xenheap allocations please" there without also saying so here.
> >> I've read over patch 1's description, and while it tries to justify this
> >> accordingly, I still remain concerned. I think we simply have to accept
> >> the mapping overhead, to avoid allocating from a pool which - over time -
> >> is representing a decreasing portion of total memory systems have (on
> >> average, and not even considering systems with extremely sparse memory
> >> layouts, and with perhaps PDX compression not doing good enough to
> >> compensate).
> >
> > You should certainly have the same resistance to adding new xenheap
> > allocations.  But looking at the numbers, I don't see that we're
> > anywhere near the point where we say, "Absolutely no new xenheap
> > allocations, regardless of the cost."
>
> Well, that depends, and in part on the longer term plans with ASI. It has
> been my (silent) assumption that eventually the directmap would go away
> altogether when ASI is in use, with the VA space freed (almost?) all
> becoming available for vmap(). With the disappearance of directmap, the
> xenheap would naturally disappear as well. Hence putting stuff there in
> new work actually adds to our technical debt.

I do think that makes sense as a long-term goal.  Actually, I asked
Fable to do an audit of xenheap allocations, asking it to classify
them as to whether they needed xenheap's key properties (easy mfn <->
va conversion, available in early boot), and it reckoned only about 4%
of the allocations (by volume) in the example system I did numbers for
needed xenheap-specific properties.  Lots of things just need *some*
global mapping somewhere, and would probably actually overall benefit
from being moved to the vmap, so they wouldn't be restricted to a
single NUMA node.  Grant frames were the biggest chunk.

The xenheap / vmap split is certainly something I would now consider a
large piece of technical debt; moving towards paying that off is
certainly something I think worth achieving, provided the rest of the
maintainers are on board.

(Imagine how differently this conversation would have gone, if in the
first email you had said, "Actually, I thought one of the main goals
of this series was to get rid of the xenheap altogether, so that we
could switch to having a single large vmap area instead?")

> While these percentiles in particular of course look very tiny, they are
> applicable only on systems having no meaningful gaps in the physical
> address map. And even more generally I find all of these calculations
> only partly convincing, not the least because you start out from numbers
> which look pretty contrived when comparing to actual systems which would
> run the new code. (Using more realistic real-system values may end up
> going in favor of what you want to convey, or it may not.)

To be honest, I'm inclined to think that they're not very convincing
because you don't actually have an idea what the problem is.  You
didn't specify what you were worried about, so I tried to guess a
scenario that I considered 95th-percentile worse case.  I don't know
what kinds of sparse memory layout machines you have in mind -- are
they written down anywhere, so that contributors can read and
understand what they need to consider *before* implementing?  Even now
you haven't even said what about my scenario you consider unrealistic,
much less told me parameters you think are more realistic.

I don't even know exactly what failure mode you're worried about.  Two
kinds of potential failures I know about:
 - Performance impacted because pages can't be NUMA-local
 - Toolstack operations (including domain creation) fail because
xenheap has been exhausted.

If you're willing to accept 9 map/unmap operations on *all* systems,
then NUMA-non-local accesses can't be that big of an issue for you.

On the fairly largish system / load that I tried to estimate, the
total xenheap usage was less than 3GiB in the worst case.  Let's
double that just for safety sake: Do there exist systems whose memory
is so sparse that even with PDX compression, they can't even scrounge
together 6GiB below the 4TiB limit?  If so, I think a much better
solution would be to document that such systems may be able to support
a lower degree of oversubscribing than most systems, and leave it at
that.

In short: I can't imagine a scenario where a larger xenheap is an
issue we should be concerned about.

It's not up to me to guess what sorts of numbers would allay your
concern.  If you want me to consider a large xenheap to be a problem
on its own, it is now your job to articulate, first, at least one
target system (hardware and configuration) you think would be
problematic;  and secondly, exactly what bad thing you're worried
about happening.  Only then do I have any hope of addressing your
concerns.  Until that time, I don't consider "the xenheap is getting
too large" objection to be valid.

Objections I will consider:
- xenheap has poorer NUMA locality
- The xenheap/vmap split is a big ugly unnecessary bit of technical
debt; Xen would be far better if we could get rid of the xenheap
altogether.  Every additional user of xenheap is another patch in a
series converting xenheap to vmap.

> > One is map_domain_page_irqoff(): If the caller promises to keep
> > interrupts disabled until unmap_domain_page_irqoff(), we can safely
> > perform maps in a context switch without having to worry about
> > sync_lazy_execstate.  (This was actually implemented and almost sent
> > on Tuesday evening, when I noticed your review of Roger's v2 saying,
> > "Question is whether it's a good idea in the first place to start
> > using map_domain_page() from the context switch path.  Surely there
> > are possible alternatives.")  This maps all vcpu pages from the
> > domheap, adding nothing to the xenheap *or* the vmap area.  But it
> > costs 9 map/unmap pairs *per context switch*.
>
> But why would not using vmap() be a necessary conclusion of my initial
> comment? All I'm objecting to are new uses of the xenheap.

I'm trying to list all the advantages and disadvantages of the various
options I've explored.  You've agreed that growing the vmap region is
*also* something we need to worry about; and that the vmap allocator
may not be ready to become a performance-critical part of the system.
Furthermore, as Roger pointed out privately, regardless of where the
global mapping lives (vmap or sparsely-mapped xenheap), having a
global mapping at all means global TLB flushes whenever we destroy a
vCPU; and in any case, in principle we'd like to avoid exposing any
data whatsoever.  Using the mapcache avoids all those problems, for a
different cost.

> > Ultimately, I think there's a lot of wisdom in the saying, "Premature
> > optimization is the root of all evil."
> > ...
> I'm a little puzzled by you talking of "optimization" (premature or
> not) here. In my initial reply I did point out a functional aspect, and
> I made clear that I'm aware that this is going to have a performance
> impact. I.e. quite the opposite of "optimization".

"Optimize" in the terms of "improve", not necessarily in terms of cycle count.

The point of the principle is to say this:  First, build it correctly,
in a way that is simple, clear, robust, and easy to write, review, and
maintain.  *Then*, after you've measured that there is a problem,
where there is a problem, and so on, should you put in extra effort
and add extra complication, only to areas where you know there will be
some material benefit.

You've argued that we should avoid using xenheap because it will have
some negative impacts on large systems with sparse memory layouts; in
other words, you're asking me to *optimize* for those use cases, at
the expense of more typical systems.  In isolation, this principle
would say: take the xenheap option first, as it's clean and fast in
the common case, and measure it on a target systems (or at least,
estimate what the impact would be based on modeling).  Once you have
reason to believe there will be a problem, then introduce code
complications based on the actual issue you find.

> > There are other options I've explored:
> >
> > - domheap + vmap; basically, allocate from domheap, map in the vmap
> >   area.  On paper this sounds like the same thing; the problem is that
> >   we don't have a simple MFN -> VA mapping, as we do in the xenheap
> >   case, so the walk is a lot harder; we start to have to do lookups,
> >   significantly increasing the cost over simple memory reads and math.
> >   (This is the difference from the intremap table on the VT-d thread:
> >   that's a leaf structure reached from a single pointer, so a
> >   permanent vmap costs nothing there.  Pagetable hierarchies are
> >   exactly the case where the MFN -> VA step is critical: each entry
> >   read yields an MFN, which the walk has to turn into the next VA.)
>
> The pages used here are entirely private to logic handling those page
> tables. Hence a struct page_info field can very likely be used to stash
> the VA of a permanent mapping. (Feels like similarly I must have
> suggested this somewhere else recently, yet I don't recall the context.)

This is an interesting idea, particularly for a full xenheap -> vmap
change.  Probably too complicated for this series (see below).

> Absolutely, and I have been mentioning the need to consider growing this
> area in a number of situations (one iirc again pretty recently).
...
> Indeed, heavier use of that allocator may require work to be done there.
...
> Valid concern, yet surely possible to deal with.

One thing you do need to consider:  There are at least 45 patches to
get to the most basic form of extra security (no direct-map, FPU/XSAVE
moved to domheap); and another 13 after that to move to per-cpu
stacks.  I'm engaged until November to work on this.  If we don't have
significant progress by then, there may be no ASI at all (at least for
a long time), and thus no hope of getting rid of the xenheap.  If
every batch of 7 patches takes a month to get through, we're not going
to be anywhere close by November.  So you need to be strategic about
what kinds of additional work you ask me to do: what does a solution
look like that is both technically acceptable, and achieves measurable
progress by November?

If we had all the time in the world, we could consider trying to
convert the entire xenheap to vmap as the first step.  (Even on the
fairly large system I tried to describe, the xenheap was only around
3GiB; still plenty of room in the vmap area to get us by until the
direct map is gone.)  I don't think that's really viable, as there's
quite a long tail of allocations that would probably end up being
haggled over before we even began the ASI series itself.

So let's try to take stock.  We can't safely remove the direct-map
unless we have per-vcpu mapcaches.  We can't really say we've isolated
the system while all pCPU stacks, with random bits of guest state, are
visible to all other pCPUs. We can't have per-vcpu mapcaches or
per-CPU stack maps unless we have per-cpu root pagetables for PV
guests.  Both require modifying per-pCPU bits of pagetables of the
incoming vCPU on a context switch.

We have four ways of mapping in general: xenheap, vmap, mapcache, or
(for the pagetables) the linear map.

For the first three, we have several different ways of arriving at the
entries.  Both Roger's v2 and my v1 start at the top and walk down the
pagetables.  For the mapcache, this seems relatively heavy.  I thought
xenheap would be just simple math, but with PDX on all the time,
that's more expensive than it looks.  vmap would require looking into
stashing a pointer into an unused (by xenheap pages) portion of the
struct page_info.

But the other approach is to stash references to just the page we need
to modify -- basically, rather than get rid of gdt_ldt_l1tab, add two
more instances.  For xenheap or vmap, this would be pointers to the
virtual addresses; but it's also possible to do in the mapcache
version, by stashing the mfn of the exact table we need to map.  In
all cases, for the context switch, it's just three references.

Both global-mapping options expose Xen pagetables.  We've agreed these
are not sensitive, but also in general our posture is that we
shouldn't reveal anything unless it buys us something.  They also both
require host-wide TLB shootdowns on vCPU tear-down.

In the spirit of "measure before optimizing", I did some tests of the
mapcache-walk variant.  On my NUC, at the end of the series, I get:

- baseline: 1480 cycles / context switch
- xenheap-walk: 2460 cycles / context switch
- mapcache-walk: 6440 cycles / context switch

By default Xen has a context switch rate limit of 1ms, so the
difference isn't measurable.  If you disable the ratelimit and do a
"ping flood" microbenchmark, the mapcache-walk reduces performance by
a whopping 70% (41k pings per second -> 12k pings per second).

If it weren't for the general intent to move away form xenheap, I'd
argue more strenuously that we should take the series as I've posted
it.  As it stands, I think the performance of mapcache-walk is
acceptable enough for a first cut, particularly given that we have two
potential optimizations already (caching MFNs rather than walking
pagetables as an easy option, switching to vmap as a slightly more
complicated one).

It's annoying that Tuesday evening I didn't know that you hated
xenheap allocations, and had only a mild distaste for mapping in a
context switch, or I might have implemented vmap instead.  At any
rate, I'll move forward with mapcache-walk, which we can later look at
optimizing by stashing the relevant MFNs so we can avoid the walk.

If anyone doesn't like that, let me know sooner rather than later, so
we can avoid wasting more time.

 -George

 -George


^ permalink raw reply	[flat|nested] 17+ messages in thread

* Re: [PATCH 0/7] x86: Address Space Isolation, part 1: per-domain area mapping rework
  2026-08-25 11:42       ` George Dunlap
@ 2026-08-25 12:00         ` Juergen Gross
  2026-08-25 13:28         ` Jan Beulich
  1 sibling, 0 replies; 17+ messages in thread
From: Juergen Gross @ 2026-08-25 12:00 UTC (permalink / raw)
  To: George Dunlap, Jan Beulich
  Cc: Andrew Cooper, Roger Pau Monné, Teddy Astie, Anthony PERARD,
	Michal Orzel, Julien Grall, Stefano Stabellini, xen-devel


[-- Attachment #1.1.1: Type: text/plain, Size: 2805 bytes --]

On 25.08.26 13:42, George Dunlap wrote:
> On Mon, Aug 24, 2026 at 10:02 AM Jan Beulich <jbeulich@suse.com> wrote:
>> On 21.08.2026 17:17, George Dunlap wrote:
>>> On Fri, Aug 21, 2026 at 9:45 AM Jan Beulich <jbeulich@suse.com> wrote:
>>>> On 20.08.2026 19:43, George Dunlap wrote:
>>>>> One point reviewers may want to look at specifically: patch 1 changes
>>>>> where the per-domain page-tables are allocated from, and its commit
>>>>> message discusses the (minor) NUMA-placement consequence.
>>>>
>>>> While I don't recall which recent patch (series) it was, I can't very well
>>>> say "no new xenheap allocations please" there without also saying so here.
>>>> I've read over patch 1's description, and while it tries to justify this
>>>> accordingly, I still remain concerned. I think we simply have to accept
>>>> the mapping overhead, to avoid allocating from a pool which - over time -
>>>> is representing a decreasing portion of total memory systems have (on
>>>> average, and not even considering systems with extremely sparse memory
>>>> layouts, and with perhaps PDX compression not doing good enough to
>>>> compensate).
>>>
>>> You should certainly have the same resistance to adding new xenheap
>>> allocations.  But looking at the numbers, I don't see that we're
>>> anywhere near the point where we say, "Absolutely no new xenheap
>>> allocations, regardless of the cost."
>>
>> Well, that depends, and in part on the longer term plans with ASI. It has
>> been my (silent) assumption that eventually the directmap would go away
>> altogether when ASI is in use, with the VA space freed (almost?) all
>> becoming available for vmap(). With the disappearance of directmap, the
>> xenheap would naturally disappear as well. Hence putting stuff there in
>> new work actually adds to our technical debt.
> 
> I do think that makes sense as a long-term goal.  Actually, I asked
> Fable to do an audit of xenheap allocations, asking it to classify
> them as to whether they needed xenheap's key properties (easy mfn <->
> va conversion, available in early boot), and it reckoned only about 4%
> of the allocations (by volume) in the example system I did numbers for
> needed xenheap-specific properties.  Lots of things just need *some*
> global mapping somewhere, and would probably actually overall benefit
> from being moved to the vmap, so they wouldn't be restricted to a
> single NUMA node.  Grant frames were the biggest chunk.

Please note that I'm currently working on a patch series which will need
to convert grant frames and some other xenheap allocations to use domheap
and vmap().

Currently this is just a proof of concept for Xen summit, but I expect
this to become mature after some feedback I hope to get there.


Juergen

[-- Attachment #1.1.2: OpenPGP public key --]
[-- Type: application/pgp-keys, Size: 3743 bytes --]

[-- Attachment #2: OpenPGP digital signature --]
[-- Type: application/pgp-signature, Size: 495 bytes --]

^ permalink raw reply	[flat|nested] 17+ messages in thread

* Re: [PATCH 0/7] x86: Address Space Isolation, part 1: per-domain area mapping rework
  2026-08-25 11:42       ` George Dunlap
  2026-08-25 12:00         ` Juergen Gross
@ 2026-08-25 13:28         ` Jan Beulich
  2026-08-25 14:40           ` George Dunlap
  1 sibling, 1 reply; 17+ messages in thread
From: Jan Beulich @ 2026-08-25 13:28 UTC (permalink / raw)
  To: George Dunlap
  Cc: Andrew Cooper, Roger Pau Monné, Teddy Astie, Anthony PERARD,
	Michal Orzel, Julien Grall, Stefano Stabellini, xen-devel

On 25.08.2026 13:42, George Dunlap wrote:
> On Mon, Aug 24, 2026 at 10:02 AM Jan Beulich <jbeulich@suse.com> wrote:
>> While these percentiles in particular of course look very tiny, they are
>> applicable only on systems having no meaningful gaps in the physical
>> address map. And even more generally I find all of these calculations
>> only partly convincing, not the least because you start out from numbers
>> which look pretty contrived when comparing to actual systems which would
>> run the new code. (Using more realistic real-system values may end up
>> going in favor of what you want to convey, or it may not.)
> 
> To be honest, I'm inclined to think that they're not very convincing
> because you don't actually have an idea what the problem is.  You
> didn't specify what you were worried about, so I tried to guess a
> scenario that I considered 95th-percentile worse case.  I don't know
> what kinds of sparse memory layout machines you have in mind -- are
> they written down anywhere, so that contributors can read and
> understand what they need to consider *before* implementing?  Even now
> you haven't even said what about my scenario you consider unrealistic,
> much less told me parameters you think are more realistic.

What I specifically considered unrealistic is that you use huge pCPU and
vCPU counts. Yes, you're trying to do a worst case estimate, yet at the
same time you're assuming huge amounts of memory to be available (which
doesn't represent a "worst case").

As to sparse layouts - ones which have led to the two forms of PDX
compression are well known (I think). The need for more recent (offset)
form is a good example of what could go wrong here: New machines can
always come with new layouts, potentially requiring new compressions
approaches. So what I'm concerned about is effectively _any_ sparse
layout that we may encounter without having a suitable PDX compression
method readily available.

> I don't even know exactly what failure mode you're worried about.  Two
> kinds of potential failures I know about:
>  - Performance impacted because pages can't be NUMA-local
>  - Toolstack operations (including domain creation) fail because
> xenheap has been exhausted.

One thing I can't help thinking you keep overlooking throughout your
reply: xenheap and domheap aren't separate. There being only a
relatively small part of it needed for the worst case estimate you did
means nothing as to exhausting the xenheap in practice: Almost the
entirety of it (with the DMA reserve being somewhat protected) can be
used to build domains. Once in that state, allocations would fail no
matter that large swathes of domheap might (have become) available
(again).

That said, with what you indicated at the very bottom of your reply,
it looks like this part of the discussion has become largely moot.

Jan


^ permalink raw reply	[flat|nested] 17+ messages in thread

* Re: [PATCH 0/7] x86: Address Space Isolation, part 1: per-domain area mapping rework
  2026-08-25 13:28         ` Jan Beulich
@ 2026-08-25 14:40           ` George Dunlap
  0 siblings, 0 replies; 17+ messages in thread
From: George Dunlap @ 2026-08-25 14:40 UTC (permalink / raw)
  To: Jan Beulich
  Cc: Andrew Cooper, Roger Pau Monné, Teddy Astie, Anthony PERARD,
	Michal Orzel, Julien Grall, Stefano Stabellini, xen-devel

On Tue, Aug 25, 2026 at 2:28 PM Jan Beulich <jbeulich@suse.com> wrote:
>
> On 25.08.2026 13:42, George Dunlap wrote:
> > On Mon, Aug 24, 2026 at 10:02 AM Jan Beulich <jbeulich@suse.com> wrote:
> >> While these percentiles in particular of course look very tiny, they are
> >> applicable only on systems having no meaningful gaps in the physical
> >> address map. And even more generally I find all of these calculations
> >> only partly convincing, not the least because you start out from numbers
> >> which look pretty contrived when comparing to actual systems which would
> >> run the new code. (Using more realistic real-system values may end up
> >> going in favor of what you want to convey, or it may not.)
> >
> > To be honest, I'm inclined to think that they're not very convincing
> > because you don't actually have an idea what the problem is.  You
> > didn't specify what you were worried about, so I tried to guess a
> > scenario that I considered 95th-percentile worse case.  I don't know
> > what kinds of sparse memory layout machines you have in mind -- are
> > they written down anywhere, so that contributors can read and
> > understand what they need to consider *before* implementing?  Even now
> > you haven't even said what about my scenario you consider unrealistic,
> > much less told me parameters you think are more realistic.
>
> What I specifically considered unrealistic is that you use huge pCPU and
> vCPU counts. Yes, you're trying to do a worst case estimate, yet at the
> same time you're assuming huge amounts of memory to be available (which
> doesn't represent a "worst case").

Let me point out that you still haven't named exact numbers -- you're
still offloading that to me to try to guess or imagine.

The v1 series I posted adds a few pages per vCPU and a few pages per
pCPU into the xenheap. The problem is using up too much of the xenheap
address space.  So obviously to make a reasonable worst-case that
you're not going to dismiss as contrived, I need to maximize my pCPU
count and vCPU count.  pCPUs is easy -- we're documented as supporting
4096.  How many is a reasonable number of domains and vcpus?  Well, in
general, pCPUs are an effective limit to how many vCPUs you have total
on the system; an 8:1 vCPU overcommit is high, but not preposterously
high.

I don't understand your point about huge amounts of memory.  If you're
talking about *total RAM used*, it doesn't matter whether it comes
from the domheap or the xenheap.  The only possible reason to say
domheap is OK but xenheap is not is if you're concerned about RAM
above the 4TiB boundary.  Which can only happen on system with large
amounts of RAM, or systems with really sparse memory layouts.  Does
the analysis really change at all whether you're using 12TiB or 6TiB?

> As to sparse layouts - ones which have led to the two forms of PDX
> compression are well known (I think). The need for more recent (offset)
> form is a good example of what could go wrong here: New machines can
> always come with new layouts, potentially requiring new compressions
> approaches. So what I'm concerned about is effectively _any_ sparse
> layout that we may encounter without having a suitable PDX compression
> method readily available.

"There may be some new layout that doesn't compress well" -- it's not
uncommon for random bits of new hardware not to work well until we
supply a patch to fix it.  The position you're supporting is
effectively: "We must absolutely avoid a situation where some unknown
system is temporarily restricted in how many vCPUs it can create due
to a sparse address space, even if it means making the context switch
4x as expensive for every single current user."  I just don't think
that's a reasonable position in any shape or form.

> > I don't even know exactly what failure mode you're worried about.  Two
> > kinds of potential failures I know about:
> >  - Performance impacted because pages can't be NUMA-local
> >  - Toolstack operations (including domain creation) fail because
> > xenheap has been exhausted.
>
> One thing I can't help thinking you keep overlooking throughout your
> reply: xenheap and domheap aren't separate. There being only a
> relatively small part of it needed for the worst case estimate you did
> means nothing as to exhausting the xenheap in practice: Almost the
> entirety of it (with the DMA reserve being somewhat protected) can be
> used to build domains. Once in that state, allocations would fail no
> matter that large swathes of domheap might (have become) available
> (again).

Right, so if Xen allocates too much domheap from the directmap region,
the xenheap may be not be able to allocate any more, even if there's
plenty of memory.

So something like the following:

We have 8TiB of RAM, 4 nodes, 2TiB per node.  The user wants to start
4 2TiB guests, one pinned to each node; so she starts d1 on node 1, d2
on node 2, then tries d3 and can't start it because although there's
still 4TiB of RAM left, all the memory below 4TiB was handed out to
guests already.

Is that what you had in mind?

If I didn't agree that the xenheap represents technical debt that
needs to be removed anyway, I'd say a simpler solution would be to do
do some simple xenheap reservation, based on various factors
(including number of pCPUs, and the total amount of RAM).  Reserving
6GiB on an 8TiB system would have very little impact (the first guest
would either need to be a bit smaller, or have, and would make the
whole problem go away essentially.  The first guest would either need
to be less than 0.1% smaller, or have 0.1% of its pages on a different
node.

> That said, with what you indicated at the very bottom of your reply,
> it looks like this part of the discussion has become largely moot.

Yes, but I also want to challenge your operating principles -- to get
you to state more clearly what you're concerned about.  Also, in order
to either get you to relax a bit about the xenheap growing, or to help
you articulate more clearly what problems which contributors need to
address.

 -George


^ permalink raw reply	[flat|nested] 17+ messages in thread

end of thread, other threads:[~2026-08-25 14:41 UTC | newest]

Thread overview: 17+ messages (download: mbox.gz follow: Atom feed
-- links below jump to the message on this page --
2026-08-20 17:43 [PATCH 0/7] x86: Address Space Isolation, part 1: per-domain area mapping rework George Dunlap
2026-08-20 17:43 ` [PATCH 1/7] x86/mm: allocate the per-domain page-tables from the xenheap George Dunlap
2026-08-20 17:43 ` [PATCH 2/7] x86/mm: introduce populate_perdomain_mapping() George Dunlap
2026-08-20 17:43 ` [PATCH 3/7] x86/pv: use populate_perdomain_mapping() to map the Xen GDT George Dunlap
2026-08-21 20:59   ` Andrew Cooper
2026-08-20 17:43 ` [PATCH 4/7] x86/pv: set/clear guest GDT mappings using populate_perdomain_mapping() George Dunlap
2026-08-20 17:43 ` [PATCH 5/7] x86/pv: update guest LDT mappings using {populate,destroy}_perdomain_mapping() George Dunlap
2026-08-20 17:43 ` [PATCH 6/7] x86/pv: remove stashing of GDT/LDT L1 page-tables George Dunlap
2026-08-20 17:43 ` [PATCH 7/7] x86/mm: simplify create_perdomain_mapping() interface George Dunlap
2026-08-21  8:45 ` [PATCH 0/7] x86: Address Space Isolation, part 1: per-domain area mapping rework Jan Beulich
2026-08-21 15:17   ` George Dunlap
2026-08-21 15:36     ` George Dunlap
2026-08-24  9:02     ` Jan Beulich
2026-08-25 11:42       ` George Dunlap
2026-08-25 12:00         ` Juergen Gross
2026-08-25 13:28         ` Jan Beulich
2026-08-25 14:40           ` George Dunlap

This is an external index of several public inboxes,
see mirroring instructions on how to clone and mirror
all data and code used by this external index.