* [PATCH v2 01/14] x86/domain_page: introduce IRQs-off variants of {,un}map_domain_page()
2026-09-02 9:43 [PATCH v2 00/14] x86: Address Space Isolation, part 2: asi= option and per-vCPU page tables George Dunlap
@ 2026-09-02 9:43 ` George Dunlap
2026-09-03 14:07 ` Jan Beulich
2026-09-02 9:43 ` [PATCH v2 02/14] x86/mm: introduce populate_perdomain_mapping() George Dunlap
` (12 subsequent siblings)
13 siblings, 1 reply; 44+ messages in thread
From: George Dunlap @ 2026-09-02 9:43 UTC (permalink / raw)
To: xen-devel
Cc: George Dunlap, Jan Beulich, Andrew Cooper, Roger Pau Monné,
Alejandro Vallejo, Teddy Astie, Anthony PERARD, Michal Orzel,
Julien Grall, Stefano Stabellini
From: George Dunlap <gwd@xenproject.org>
Currently, map_domain_page() cannot be called in the context switch
path. However, Xen already needs to update the slot of an incoming PV
vcpu's GDT during context switch; and when we soon switch to per-vCPU
root pagetables, we'll have to modify two more places.
Xen currently solves the problem by special-casing the GDT/LDT L1
tables to be allocated from the xenheap, and stashing a pointer to its
address in the xenheap in the domain struct. Rather than add more Xen
pagetable pages to the xenheap, introduce a version of map_domain_page
which can be called from the context switch path.
The reason map_domain_page() cannot be called from the context switch
path is x86's lazy context-switch state. Mapcache mappings are
created in the page-tables that are loaded on the pCPU. When Xen is
in a lazy context-switch state, current is the idle vCPU while the
previously-running vCPU's page-tables remain loaded. If in this
state, another pcpu wants access to the lazily-swapped-out vcpu's
state, it will send a FLUSH_VCPU_STATE IPI to the processor, which
will call sync_local_execstate().
sync_local_execstate() is implemented internally by calling a full
__context_switch(). In addition to copying the processor state into
the vcpu structure, this also switches the loaded pagetables to the
idle vcpu's, which would in turn cause mappings created before the IPI
to disappear mid-use. Therefore, mappings cannot be held in the
mapcache when a FLUSH_VCPU_STATE IPI may execute. To this end,
map_domain_page() calls sync_local_execstate() itself proactively when
it detects a lazy context-switch state. This guarantees that the
pagetables will remain consistent at least until the next context
switch.
But of course, that synchronization must not be triggered from the
context switch path itself: sync_local_execstate() ends up in
__context_switch(), so a call made while a context switch is in
progress would recurse, and the assertions along that path (current
being the idle vCPU) don't hold there either.
A full synchronization is sufficient to prevent a FLUSH_VCPU_STATE IPI
from switching the pagetables; however, it is not necessary. It
suffices to maintain interrupts disabled from before the page is
mapped until after it is unmapped. This condition is satisfied for
the mappings used on the context switch path.
Introduce {,un}map_domain_page_irqoff() variants for callers which
guarantee that interrupts remain disabled from the map until the
matching unmap. Under that guarantee the synchronization is
unnecessary rather than merely inconvenient: no IPI can be delivered
while the mapping is in use, so the lazy state cannot change under the
caller's feet, and this_cpu(pgtable_vcpu) accurately identifies the
mapcache to use (see 622c9a5ba95d "x86/mm: accurately track which vCPU
page-tables are loaded"). The variants assert that interrupts are
disabled on entry; the rest of the contract remains the caller's
responsibility.
This will be used by the next patch, which introduces a function which
will be used to modify the incoming vCPU's per-domain mappings from
within __context_switch(); it will also be used in future ASI
patches (tearing down and establishing per-CPU stack mappings during
context switch).
No functional change for existing callers.
Assisted-by: Claude Code:claude-fable-5
Signed-off-by: George Dunlap <gwd@xenproject.org>
---
Changes in v2:
- New in this version. Replaces "x86/mm: allocate the per-domain
page-tables from the xenheap".
NB an alternate approach would be to disable the lazy context switch
entirely. This simplifies the Xen code in general, and makes a patch
like this completely unnecessary, as then map_domain_page would itself
be safe to call in a context switch. Tests show, however, that simply
removing the lazy context switch measurably hurts wake-heavy
workloads: on a wake-paced ping flood, throughput drops 22%
(round-trip latency 13→17 µs) on my NUC.
Another approach is to take up the GDT/LDT L1 technique instead. v1
of the series made all perdomain pagetables allocated out of the
xenheap; but this was objected to due to the additional xenheap
allocations. An alternate version would allocate from the domheap,
and then make permanent mappings in the vmap instead. Another
potential performance improvement would be stashing the exact pages we
want to map, rather than needing to walk from the L3 each time. We
leave both of these for future work.
---
xen/arch/x86/domain_page.c | 53 ++++++++++++++++++++++++++++-------
xen/include/xen/domain_page.h | 14 +++++++++
2 files changed, 57 insertions(+), 10 deletions(-)
diff --git a/xen/arch/x86/domain_page.c b/xen/arch/x86/domain_page.c
index 72c00194f3..1fc1580e62 100644
--- a/xen/arch/x86/domain_page.c
+++ b/xen/arch/x86/domain_page.c
@@ -18,7 +18,7 @@
#include <asm/hardirq.h>
#include <asm/setup.h>
-static inline struct vcpu *mapcache_current_vcpu(void)
+static inline struct vcpu *mapcache_current_vcpu(bool irqs_off)
{
struct vcpu *v = this_cpu(pgtable_vcpu);
struct vcpu *curr = current;
@@ -36,8 +36,15 @@ static inline struct vcpu *mapcache_current_vcpu(void)
* to the idle vCPU now, otherwise an incoming FLUSH_VCPU_STATE IPI would
* change the page tables under our feet an invalidate any in-use mapcache
* entries.
+ *
+ * Callers of the irqs_off variants instead guarantee that interrupts stay
+ * disabled until the matching unmap: no IPI can be delivered while the
+ * mapping is in use, so the lazy state cannot change under our feet and
+ * pgtable_vcpu identifies the right mapcache directly. This also makes
+ * those variants usable from the context switch path itself, where
+ * calling sync_local_execstate() would recurse into __context_switch().
*/
- if ( unlikely(this_cpu(curr_vcpu) != curr) )
+ if ( !irqs_off && unlikely(this_cpu(curr_vcpu) != curr) )
{
ASSERT(curr == idle_vcpu[smp_processor_id()]);
sync_local_execstate();
@@ -46,10 +53,12 @@ static inline struct vcpu *mapcache_current_vcpu(void)
}
/*
- * At this point we can guarantee Xen is not in lazy context switch: either
- * the code above will have synced the state, or an incoming
- * FLUSH_VCPU_STATE IPI has done so behind our back. Use ACCESS_ONCE to
- * ensure the compiler never returns the locally cached pgtable_vcpu value.
+ * At this point either Xen is not in a lazy context switch (the code
+ * above will have synced the state, or an incoming FLUSH_VCPU_STATE IPI
+ * has done so behind our back), or the caller holds interrupts disabled
+ * and the state cannot change until it re-enables them. Use ACCESS_ONCE
+ * to ensure the compiler never returns the locally cached pgtable_vcpu
+ * value.
*/
return ACCESS_ONCE(this_cpu(pgtable_vcpu));
}
@@ -59,7 +68,7 @@ static inline struct vcpu *mapcache_current_vcpu(void)
#define MAPCACHE_L1ENT(idx) \
__linear_l1_table[l1_linear_offset(MAPCACHE_VIRT_START + pfn_to_paddr(idx))]
-void *map_domain_page(mfn_t mfn)
+static void *do_map_domain_page(mfn_t mfn, bool irqs_off)
{
unsigned long flags;
unsigned int idx, i;
@@ -73,7 +82,7 @@ void *map_domain_page(mfn_t mfn)
return mfn_to_virt(mfn_x(mfn));
#endif
- v = mapcache_current_vcpu();
+ v = mapcache_current_vcpu(irqs_off);
if ( !v || !is_pv_vcpu(v) )
return mfn_to_virt(mfn_x(mfn));
@@ -165,7 +174,19 @@ void *map_domain_page(mfn_t mfn)
return (void *)MAPCACHE_VIRT_START + pfn_to_paddr(idx);
}
-void unmap_domain_page(const void *ptr)
+void *map_domain_page(mfn_t mfn)
+{
+ return do_map_domain_page(mfn, false);
+}
+
+void *map_domain_page_irqoff(mfn_t mfn)
+{
+ ASSERT(!local_irq_is_enabled());
+
+ return do_map_domain_page(mfn, true);
+}
+
+static void do_unmap_domain_page(const void *ptr, bool irqs_off)
{
unsigned int idx;
struct vcpu *v;
@@ -178,7 +199,7 @@ void unmap_domain_page(const void *ptr)
ASSERT(va >= MAPCACHE_VIRT_START && va < MAPCACHE_VIRT_END);
- v = mapcache_current_vcpu();
+ v = mapcache_current_vcpu(irqs_off);
ASSERT(v && is_pv_vcpu(v));
dcache = &v->domain->arch.pv.mapcache;
@@ -223,6 +244,18 @@ void unmap_domain_page(const void *ptr)
local_irq_restore(flags);
}
+void unmap_domain_page(const void *ptr)
+{
+ do_unmap_domain_page(ptr, false);
+}
+
+void unmap_domain_page_irqoff(const void *ptr)
+{
+ ASSERT(!local_irq_is_enabled());
+
+ do_unmap_domain_page(ptr, true);
+}
+
int mapcache_domain_init(struct domain *d)
{
struct mapcache_domain *dcache = &d->arch.pv.mapcache;
diff --git a/xen/include/xen/domain_page.h b/xen/include/xen/domain_page.h
index c89b149e54..b72dffb4c7 100644
--- a/xen/include/xen/domain_page.h
+++ b/xen/include/xen/domain_page.h
@@ -31,6 +31,16 @@ void *map_domain_page(mfn_t mfn);
*/
void unmap_domain_page(const void *ptr);
+/*
+ * Variants of the above for callers which guarantee that interrupts are
+ * kept disabled from map until the matching unmap. Under that guarantee
+ * no state synchronization is required to keep the mapping valid, so these
+ * are safe to use in contexts where such a synchronization must not be
+ * triggered, in particular from the context switch path itself.
+ */
+void *map_domain_page_irqoff(mfn_t mfn);
+void unmap_domain_page_irqoff(const void *ptr);
+
/*
* Given a VA from map_domain_page(), return its underlying MFN.
*/
@@ -45,6 +55,7 @@ void *map_domain_page_global(mfn_t mfn);
void unmap_domain_page_global(const void *ptr);
#define __map_domain_page(pg) map_domain_page(page_to_mfn(pg))
+#define __map_domain_page_irqoff(pg) map_domain_page_irqoff(page_to_mfn(pg))
static inline void *__map_domain_page_global(const struct page_info *pg)
{
@@ -54,8 +65,11 @@ static inline void *__map_domain_page_global(const struct page_info *pg)
#else /* !CONFIG_ARCH_MAP_DOMAIN_PAGE */
#define map_domain_page(mfn) __mfn_to_virt(mfn_x(mfn))
+#define map_domain_page_irqoff(mfn) map_domain_page(mfn)
#define __map_domain_page(pg) page_to_virt(pg)
+#define __map_domain_page_irqoff(pg) __map_domain_page(pg)
#define unmap_domain_page(ptr) ((void)(ptr))
+#define unmap_domain_page_irqoff(ptr) unmap_domain_page(ptr)
#define domain_page_map_to_mfn(ptr) _mfn(__virt_to_mfn((unsigned long)(ptr)))
static inline void *map_domain_page_global(mfn_t mfn)
--
2.55.0
^ permalink raw reply related [flat|nested] 44+ messages in thread* Re: [PATCH v2 01/14] x86/domain_page: introduce IRQs-off variants of {,un}map_domain_page()
2026-09-02 9:43 ` [PATCH v2 01/14] x86/domain_page: introduce IRQs-off variants of {,un}map_domain_page() George Dunlap
@ 2026-09-03 14:07 ` Jan Beulich
2026-09-03 19:56 ` George Dunlap
0 siblings, 1 reply; 44+ messages in thread
From: Jan Beulich @ 2026-09-03 14:07 UTC (permalink / raw)
To: George Dunlap
Cc: George Dunlap, Andrew Cooper, Roger Pau Monné,
Alejandro Vallejo, Teddy Astie, Anthony PERARD, Michal Orzel,
Julien Grall, Stefano Stabellini, xen-devel
On 02.09.2026 11:43, George Dunlap wrote:
> From: George Dunlap <gwd@xenproject.org>
>
> Currently, map_domain_page() cannot be called in the context switch
> path. However, Xen already needs to update the slot of an incoming PV
> vcpu's GDT during context switch; and when we soon switch to per-vCPU
> root pagetables, we'll have to modify two more places.
>
> Xen currently solves the problem by special-casing the GDT/LDT L1
> tables to be allocated from the xenheap, and stashing a pointer to its
> address in the xenheap in the domain struct. Rather than add more Xen
> pagetable pages to the xenheap, introduce a version of map_domain_page
> which can be called from the context switch path.
>
> The reason map_domain_page() cannot be called from the context switch
> path is x86's lazy context-switch state. Mapcache mappings are
> created in the page-tables that are loaded on the pCPU. When Xen is
> in a lazy context-switch state, current is the idle vCPU while the
> previously-running vCPU's page-tables remain loaded. If in this
> state, another pcpu wants access to the lazily-swapped-out vcpu's
> state, it will send a FLUSH_VCPU_STATE IPI to the processor, which
> will call sync_local_execstate().
>
> sync_local_execstate() is implemented internally by calling a full
> __context_switch(). In addition to copying the processor state into
> the vcpu structure, this also switches the loaded pagetables to the
> idle vcpu's, which would in turn cause mappings created before the IPI
> to disappear mid-use. Therefore, mappings cannot be held in the
> mapcache when a FLUSH_VCPU_STATE IPI may execute. To this end,
> map_domain_page() calls sync_local_execstate() itself proactively when
> it detects a lazy context-switch state. This guarantees that the
> pagetables will remain consistent at least until the next context
> switch.
>
> But of course, that synchronization must not be triggered from the
> context switch path itself: sync_local_execstate() ends up in
> __context_switch(), so a call made while a context switch is in
> progress would recurse, and the assertions along that path (current
> being the idle vCPU) don't hold there either.
>
> A full synchronization is sufficient to prevent a FLUSH_VCPU_STATE IPI
> from switching the pagetables; however, it is not necessary. It
> suffices to maintain interrupts disabled from before the page is
> mapped until after it is unmapped. This condition is satisfied for
> the mappings used on the context switch path.
>
> Introduce {,un}map_domain_page_irqoff() variants for callers which
> guarantee that interrupts remain disabled from the map until the
> matching unmap. Under that guarantee the synchronization is
> unnecessary rather than merely inconvenient: no IPI can be delivered
> while the mapping is in use, so the lazy state cannot change under the
> caller's feet, and this_cpu(pgtable_vcpu) accurately identifies the
> mapcache to use (see 622c9a5ba95d "x86/mm: accurately track which vCPU
> page-tables are loaded"). The variants assert that interrupts are
> disabled on entry; the rest of the contract remains the caller's
> responsibility.
>
> This will be used by the next patch, which introduces a function which
> will be used to modify the incoming vCPU's per-domain mappings from
> within __context_switch(); it will also be used in future ASI
> patches (tearing down and establishing per-CPU stack mappings during
> context switch).
>
> No functional change for existing callers.
>
> Assisted-by: Claude Code:claude-fable-5
> Signed-off-by: George Dunlap <gwd@xenproject.org>
Reviewed-by: Jan Beulich <jbeulich@suse.com>
with one aspect for further consideration:
> @@ -59,7 +68,7 @@ static inline struct vcpu *mapcache_current_vcpu(void)
> #define MAPCACHE_L1ENT(idx) \
> __linear_l1_table[l1_linear_offset(MAPCACHE_VIRT_START + pfn_to_paddr(idx))]
>
> -void *map_domain_page(mfn_t mfn)
> +static void *do_map_domain_page(mfn_t mfn, bool irqs_off)
> {
do_...() commonly (but sadly not consistently) mark top-level hypercall
handlers. Personally I'd prefer if the "do" (but not the underscore) were
dropped here and ...
> @@ -165,7 +174,19 @@ void *map_domain_page(mfn_t mfn)
> return (void *)MAPCACHE_VIRT_START + pfn_to_paddr(idx);
> }
>
> -void unmap_domain_page(const void *ptr)
> +void *map_domain_page(mfn_t mfn)
> +{
> + return do_map_domain_page(mfn, false);
> +}
> +
> +void *map_domain_page_irqoff(mfn_t mfn)
> +{
> + ASSERT(!local_irq_is_enabled());
> +
> + return do_map_domain_page(mfn, true);
> +}
> +
> +static void do_unmap_domain_page(const void *ptr, bool irqs_off)
... here. Identifiers with a single leading underscore (and no following
upper-case letter) are designated for use by static functions, after all.
Jan
^ permalink raw reply [flat|nested] 44+ messages in thread* Re: [PATCH v2 01/14] x86/domain_page: introduce IRQs-off variants of {,un}map_domain_page()
2026-09-03 14:07 ` Jan Beulich
@ 2026-09-03 19:56 ` George Dunlap
0 siblings, 0 replies; 44+ messages in thread
From: George Dunlap @ 2026-09-03 19:56 UTC (permalink / raw)
To: Jan Beulich
Cc: Andrew Cooper, Roger Pau Monné, Alejandro Vallejo,
Teddy Astie, Anthony PERARD, Michal Orzel, Julien Grall,
Stefano Stabellini, xen-devel
On Thu, Sep 3, 2026 at 3:07 PM Jan Beulich <jbeulich@suse.com> wrote:
>
> On 02.09.2026 11:43, George Dunlap wrote:
> > From: George Dunlap <gwd@xenproject.org>
> >
> > Currently, map_domain_page() cannot be called in the context switch
> > path. However, Xen already needs to update the slot of an incoming PV
> > vcpu's GDT during context switch; and when we soon switch to per-vCPU
> > root pagetables, we'll have to modify two more places.
> >
> > Xen currently solves the problem by special-casing the GDT/LDT L1
> > tables to be allocated from the xenheap, and stashing a pointer to its
> > address in the xenheap in the domain struct. Rather than add more Xen
> > pagetable pages to the xenheap, introduce a version of map_domain_page
> > which can be called from the context switch path.
> >
> > The reason map_domain_page() cannot be called from the context switch
> > path is x86's lazy context-switch state. Mapcache mappings are
> > created in the page-tables that are loaded on the pCPU. When Xen is
> > in a lazy context-switch state, current is the idle vCPU while the
> > previously-running vCPU's page-tables remain loaded. If in this
> > state, another pcpu wants access to the lazily-swapped-out vcpu's
> > state, it will send a FLUSH_VCPU_STATE IPI to the processor, which
> > will call sync_local_execstate().
> >
> > sync_local_execstate() is implemented internally by calling a full
> > __context_switch(). In addition to copying the processor state into
> > the vcpu structure, this also switches the loaded pagetables to the
> > idle vcpu's, which would in turn cause mappings created before the IPI
> > to disappear mid-use. Therefore, mappings cannot be held in the
> > mapcache when a FLUSH_VCPU_STATE IPI may execute. To this end,
> > map_domain_page() calls sync_local_execstate() itself proactively when
> > it detects a lazy context-switch state. This guarantees that the
> > pagetables will remain consistent at least until the next context
> > switch.
> >
> > But of course, that synchronization must not be triggered from the
> > context switch path itself: sync_local_execstate() ends up in
> > __context_switch(), so a call made while a context switch is in
> > progress would recurse, and the assertions along that path (current
> > being the idle vCPU) don't hold there either.
> >
> > A full synchronization is sufficient to prevent a FLUSH_VCPU_STATE IPI
> > from switching the pagetables; however, it is not necessary. It
> > suffices to maintain interrupts disabled from before the page is
> > mapped until after it is unmapped. This condition is satisfied for
> > the mappings used on the context switch path.
> >
> > Introduce {,un}map_domain_page_irqoff() variants for callers which
> > guarantee that interrupts remain disabled from the map until the
> > matching unmap. Under that guarantee the synchronization is
> > unnecessary rather than merely inconvenient: no IPI can be delivered
> > while the mapping is in use, so the lazy state cannot change under the
> > caller's feet, and this_cpu(pgtable_vcpu) accurately identifies the
> > mapcache to use (see 622c9a5ba95d "x86/mm: accurately track which vCPU
> > page-tables are loaded"). The variants assert that interrupts are
> > disabled on entry; the rest of the contract remains the caller's
> > responsibility.
> >
> > This will be used by the next patch, which introduces a function which
> > will be used to modify the incoming vCPU's per-domain mappings from
> > within __context_switch(); it will also be used in future ASI
> > patches (tearing down and establishing per-CPU stack mappings during
> > context switch).
> >
> > No functional change for existing callers.
> >
> > Assisted-by: Claude Code:claude-fable-5
> > Signed-off-by: George Dunlap <gwd@xenproject.org>
>
> Reviewed-by: Jan Beulich <jbeulich@suse.com>
> with one aspect for further consideration:
>
> > @@ -59,7 +68,7 @@ static inline struct vcpu *mapcache_current_vcpu(void)
> > #define MAPCACHE_L1ENT(idx) \
> > __linear_l1_table[l1_linear_offset(MAPCACHE_VIRT_START + pfn_to_paddr(idx))]
> >
> > -void *map_domain_page(mfn_t mfn)
> > +static void *do_map_domain_page(mfn_t mfn, bool irqs_off)
> > {
>
> do_...() commonly (but sadly not consistently) mark top-level hypercall
> handlers. Personally I'd prefer if the "do" (but not the underscore) were
> dropped here and ...
>
> > @@ -165,7 +174,19 @@ void *map_domain_page(mfn_t mfn)
> > return (void *)MAPCACHE_VIRT_START + pfn_to_paddr(idx);
> > }
> >
> > -void unmap_domain_page(const void *ptr)
> > +void *map_domain_page(mfn_t mfn)
> > +{
> > + return do_map_domain_page(mfn, false);
> > +}
> > +
> > +void *map_domain_page_irqoff(mfn_t mfn)
> > +{
> > + ASSERT(!local_irq_is_enabled());
> > +
> > + return do_map_domain_page(mfn, true);
> > +}
> > +
> > +static void do_unmap_domain_page(const void *ptr, bool irqs_off)
>
> ... here. Identifiers with a single leading underscore (and no following
> upper-case letter) are designated for use by static functions, after all.
Thanks -- I'll make that change for v3 if nobody objects.
-George
^ permalink raw reply [flat|nested] 44+ messages in thread
* [PATCH v2 02/14] x86/mm: introduce populate_perdomain_mapping()
2026-09-02 9:43 [PATCH v2 00/14] x86: Address Space Isolation, part 2: asi= option and per-vCPU page tables George Dunlap
2026-09-02 9:43 ` [PATCH v2 01/14] x86/domain_page: introduce IRQs-off variants of {,un}map_domain_page() George Dunlap
@ 2026-09-02 9:43 ` George Dunlap
2026-09-03 15:57 ` Jan Beulich
2026-09-04 5:47 ` Jan Beulich
2026-09-02 9:43 ` [PATCH v2 03/14] x86/pv: use populate_perdomain_mapping() to map the Xen GDT George Dunlap
` (11 subsequent siblings)
13 siblings, 2 replies; 44+ messages in thread
From: George Dunlap @ 2026-09-02 9:43 UTC (permalink / raw)
To: xen-devel
Cc: Roger Pau Monné, Jan Beulich, Andrew Cooper,
Roger Pau Monné, Alejandro Vallejo, Teddy Astie,
Anthony PERARD, Michal Orzel, Julien Grall, Stefano Stabellini,
George Dunlap
From: Roger Pau Monné <roger.pau@citrix.com>
The per-domain area already has central machinery for building its
page-tables and for managing the backing pages it owns itself:
create_perdomain_mapping() / destroy_perdomain_mapping(), used by the
mapcache bitmaps, the compat argument-translation area, and the
GDT/LDT slots alike. What the interface lacks is a way to install a
caller's own pages at a chosen address. The PV GDT and LDT code
open-codes its own modifications to the per-domain area, capturing
aliases of its L1 tables at creation time
(create_perdomain_mapping()'s pl1tab argument) and stashing them in
d->arch.pv.gdt_ldt_l1tab.
Introduce populate_perdomain_mapping(v, va, mfn, nr, flags) to close
this gap, giving the perdomain area's rules a single place to live.
populate_perdomain_mapping writes the given MFNs, with the given
page-table flags, into v's view of the per-domain area: through the
recursive linear mappings when v's page-tables are loaded on the
current pCPU, or by walking the per-domain page-table structures
otherwise. Callers don't need to know where the page-tables live,
how the area is structured, or whether it is per-domain or per-vcpu.
The fast path is keyed off this_cpu(pgtable_vcpu) rather than current:
following 622c9a5ba95d ("x86/mm: accurately track which vCPU
page-tables are loaded") that's the accurate way to tell whether the
linear mappings reach v's per-domain area, and it copes with the
transient states where current doesn't match the loaded page-tables
(e.g. the _toggle_guest_pt() error window, or mid context switch). It
also removes any need to call sync_local_execstate(): when the vCPU's
page-tables aren't loaded, the walk instead maps the per-domain
page-table pages with the map_domain_page_irqoff() variants, holding
interrupts off for its duration, and so is usable from any context --
including the context switch, before the incoming vcpu's page-tables
are loaded.
We require the range to already have been populated down to the L1
tables by create_perdomain_mapping(). TLB flushing is left to the
caller. A present entry not owned by the area (!_PAGE_AVAIL0) is
replaced. A present entry owned by the area (_PAGE_AVAIL0, installed
by create_perdomain_mapping() itself) is freed and replaced: such a
page is referenced only by the mapping, so displacing it without
freeing it would leak it. Nothing in this series replaces area-owned
backing, so the free is marked ASSERT_UNREACHABLE(); note that freeing
requires a context where the allocator may be entered -- IRQs enabled,
not in interrupt context (see ASSERT_ALLOC_CONTEXT()) -- so any future
caller replacing area-owned backing must not do so from the context
switch path, nor anywhere the slow-path walk (which holds IRQs off)
can be taken. Missing page-table structure is a hypervisor bug and
BUG_ON(): there is no safe continuation, least of all from the context
switch, where the next descriptor fetch through an unmapped GDT slot
would be fatal.
Subsequent patches convert the users of the stashed L1 tables to this
interface, starting with the Xen slots of the full GDT; the stash --
which could in any case not represent per-vcpu mappings without being
replicated for every vcpu and slot -- is then removed, leaving
create_perdomain_mapping() to manage only the page-table structure and
the pages the area owns itself. Later parts of the series use the new
interface for their own mappings rather than adding further
mechanisms.
Signed-off-by: Roger Pau Monné <roger.pau@citrix.com>
Assisted-by: Claude Code:claude-fable-5, Claude Code:claude-opus-4-8
Signed-off-by: George Dunlap <gwd@xenproject.org>
---
Changes in v2:
- Drop the xenheap allocation of the per-domain page-tables (v1's
patch 1); the walk instead maps the page-table pages with the new
map_domain_page_irqoff() variants, holding interrupts off for the
duration.
- Re-introduce the linear-map fast path for when the target vCPU's
page-tables are loaded on the current pCPU, now keyed off
this_cpu(pgtable_vcpu).
Changes since the previously posted version:
- Split the introduction of populate_perdomain_mapping() from its first
user (previously one patch: "x86/pv: introduce function to populate
perdomain area and use it to map Xen GDT").
- Drop the linear-map fast path and the sync_local_execstate() call:
with the per-domain page-tables in the xenheap (previous patch) the
walk needs no mapping, so a single path serves all callers and
contexts.
- Keep the ASSERT_UNREACHABLE() + free_domheap_page() handling of a
replaced area-owned entry, and document the allocation-context
requirement it places on callers replacing such entries. BUG_ON()
missing page-table structure, instead of domain_crash().
- Take the page-table flags as a parameter (the Xen GDT and guest GDT
slots want RW mappings; the zero page backing torn-down GDT slots is
mapped read-only, as today).
- Document the contract in a header comment.
- Make the mfn parameter const and nr unsigned int, matching
{create,destroy}_perdomain_mapping().
- Drop the unused cr3_mfn() helper.
Considered, but not done to limit churn against the previously posted
version: splitting the interface into a "populate" variant (any present
entry is a bug) and an "update" variant (replacement expected), so that
call sites declare their intent and unexpected collisions become
detectable. Of the eventual call sites in the wider series, roughly
half are of each kind. Could be done as a follow-up if there is
interest.
---
xen/arch/x86/include/asm/mm.h | 3 +
xen/arch/x86/mm.c | 124 ++++++++++++++++++++++++++++++++++
2 files changed, 127 insertions(+)
diff --git a/xen/arch/x86/include/asm/mm.h b/xen/arch/x86/include/asm/mm.h
index 2254a7e3fe..1888807394 100644
--- a/xen/arch/x86/include/asm/mm.h
+++ b/xen/arch/x86/include/asm/mm.h
@@ -606,6 +606,9 @@ int compat_arch_memory_op(unsigned long cmd, XEN_GUEST_HANDLE_PARAM(void) arg);
int create_perdomain_mapping(struct domain *d, unsigned long va,
unsigned int nr, l1_pgentry_t **pl1tab,
struct page_info **ppg);
+void populate_perdomain_mapping(const struct vcpu *v, unsigned long va,
+ const mfn_t *mfn, unsigned int nr,
+ unsigned int flags);
void destroy_perdomain_mapping(struct domain *d, unsigned long va,
unsigned int nr);
void free_perdomain_mappings(struct domain *d);
diff --git a/xen/arch/x86/mm.c b/xen/arch/x86/mm.c
index b158742408..552559ecf1 100644
--- a/xen/arch/x86/mm.c
+++ b/xen/arch/x86/mm.c
@@ -6334,6 +6334,130 @@ int create_perdomain_mapping(struct domain *d, unsigned long va,
return rc;
}
+/*
+ * Map @nr pages, @mfn[0..nr-1], at consecutive pages from @va in v's view of
+ * the per-domain area, with page-table @flags. The range must lie within a
+ * single per-domain slot, and must already have been plumbed down to the L1
+ * tables by create_perdomain_mapping(): missing structure is a bug. A
+ * present entry not owned by the area (no _PAGE_AVAIL0) is silently
+ * replaced, as that is how callers update their mappings; a present
+ * area-owned entry is freed and replaced, which constrains the calling
+ * context (see the comment in the body). No TLB flushing is done: the
+ * caller decides whether the old translations can still be cached
+ * anywhere.
+ *
+ * When v's page-tables are loaded on this pCPU the L1 entries are reached
+ * through the recursive linear mappings; otherwise the walk maps the
+ * per-domain page-table pages transiently with IRQs off, so it needs
+ * nothing from the current address space and is usable from any context --
+ * including the context switch, before the incoming vcpu's page-tables are
+ * loaded.
+ */
+void populate_perdomain_mapping(const struct vcpu *v, unsigned long va,
+ const mfn_t *mfn, unsigned int nr,
+ unsigned int flags)
+{
+ l1_pgentry_t *l1tab = NULL, *pl1e;
+ const l3_pgentry_t *l3tab;
+ const l2_pgentry_t *l2tab;
+ struct domain *d = v->domain;
+ unsigned long irq_flags;
+
+ ASSERT(va >= PERDOMAIN_VIRT_START &&
+ va < PERDOMAIN_VIRT_SLOT(PERDOMAIN_SLOTS));
+ ASSERT(!nr || !l3_table_offset(va ^ (va + nr * PAGE_SIZE - 1)));
+ /* Area-owned pages are installed by create_perdomain_mapping() only. */
+ ASSERT(!(flags & _PAGE_AVAIL0));
+
+ if ( likely(this_cpu(pgtable_vcpu) == v) )
+ {
+ unsigned int i;
+
+ /*
+ * Fast path: v's page-tables are loaded on this pCPU, so the L1
+ * entries can be reached using the recursive linear mappings.
+ */
+ pl1e = &__linear_l1_table[l1_linear_offset(va)];
+
+ for ( i = 0; i < nr; i++, pl1e++ )
+ {
+ /*
+ * An area-owned entry (installed by create_perdomain_mapping(),
+ * marked _PAGE_AVAIL0) holds the only reference to its page, so
+ * displacing it means freeing it. Nothing in this series
+ * replaces area-owned backing, hence the ASSERT_UNREACHABLE();
+ * any future caller doing so must run where freeing is
+ * permitted -- IRQs enabled, not in interrupt context (see
+ * ASSERT_ALLOC_CONTEXT()) -- which the context switch path is
+ * not.
+ */
+ if ( unlikely(perdomain_l1e_needs_freeing(*pl1e)) )
+ {
+ ASSERT_UNREACHABLE();
+ free_domheap_page(l1e_get_page(*pl1e));
+ }
+ l1e_write(pl1e, l1e_from_mfn(mfn[i], flags));
+ }
+
+ return;
+ }
+
+ BUG_ON(!d->arch.perdomain_l3_pg);
+
+ /*
+ * Slow path: walk v's per-domain page-table pages. All mappings are
+ * local to this function, so disabling interrupts for the duration of
+ * the walk satisfies the map_domain_page_irqoff() contract. This in
+ * turn makes this function usable from the context switch path, where
+ * a plain map_domain_page() could recurse into __context_switch() via
+ * sync_local_execstate().
+ */
+ local_irq_save(irq_flags);
+
+ l3tab = __map_domain_page_irqoff(d->arch.perdomain_l3_pg);
+
+ /*
+ * Missing page-table structure is a hypervisor bug: there is no safe
+ * continuation, least of all from the context switch, where the next
+ * descriptor fetch through an unmapped GDT slot would be fatal.
+ */
+ BUG_ON(!(l3e_get_flags(l3tab[l3_table_offset(va)]) & _PAGE_PRESENT));
+
+ l2tab = map_domain_page_irqoff(l3e_get_mfn(l3tab[l3_table_offset(va)]));
+
+ for ( ; nr--; va += PAGE_SIZE, mfn++ )
+ {
+ if ( !l1tab || !l1_table_offset(va) )
+ {
+ const l2_pgentry_t *pl2e = l2tab + l2_table_offset(va);
+
+ BUG_ON(!(l2e_get_flags(*pl2e) & _PAGE_PRESENT));
+
+ unmap_domain_page_irqoff(l1tab);
+ l1tab = map_domain_page_irqoff(l2e_get_mfn(*pl2e));
+ }
+
+ pl1e = &l1tab[l1_table_offset(va)];
+
+ /*
+ * As the fast path -- and the slow path holds IRQs off throughout,
+ * so replacing area-owned backing here is never permitted.
+ */
+ if ( unlikely(perdomain_l1e_needs_freeing(*pl1e)) )
+ {
+ ASSERT_UNREACHABLE();
+ free_domheap_page(l1e_get_page(*pl1e));
+ }
+ l1e_write(pl1e, l1e_from_mfn(*mfn, flags));
+ }
+
+ unmap_domain_page_irqoff(l1tab);
+ unmap_domain_page_irqoff(l2tab);
+ unmap_domain_page_irqoff(l3tab);
+
+ local_irq_restore(irq_flags);
+}
+
void destroy_perdomain_mapping(struct domain *d, unsigned long va,
unsigned int nr)
{
--
2.55.0
^ permalink raw reply related [flat|nested] 44+ messages in thread* Re: [PATCH v2 02/14] x86/mm: introduce populate_perdomain_mapping()
2026-09-02 9:43 ` [PATCH v2 02/14] x86/mm: introduce populate_perdomain_mapping() George Dunlap
@ 2026-09-03 15:57 ` Jan Beulich
2026-09-03 21:27 ` George Dunlap
2026-09-04 5:47 ` Jan Beulich
1 sibling, 1 reply; 44+ messages in thread
From: Jan Beulich @ 2026-09-03 15:57 UTC (permalink / raw)
To: George Dunlap
Cc: Roger Pau Monné, Andrew Cooper, Alejandro Vallejo,
Teddy Astie, Anthony PERARD, Michal Orzel, Julien Grall,
Stefano Stabellini, George Dunlap, xen-devel
On 02.09.2026 11:43, George Dunlap wrote:
> --- a/xen/arch/x86/mm.c
> +++ b/xen/arch/x86/mm.c
> @@ -6334,6 +6334,130 @@ int create_perdomain_mapping(struct domain *d, unsigned long va,
> return rc;
> }
>
> +/*
> + * Map @nr pages, @mfn[0..nr-1], at consecutive pages from @va in v's view of
> + * the per-domain area, with page-table @flags. The range must lie within a
> + * single per-domain slot, and must already have been plumbed down to the L1
> + * tables by create_perdomain_mapping(): missing structure is a bug. A
> + * present entry not owned by the area (no _PAGE_AVAIL0) is silently
> + * replaced, as that is how callers update their mappings; a present
> + * area-owned entry is freed and replaced, which constrains the calling
> + * context (see the comment in the body). No TLB flushing is done: the
> + * caller decides whether the old translations can still be cached
> + * anywhere.
> + *
> + * When v's page-tables are loaded on this pCPU the L1 entries are reached
> + * through the recursive linear mappings; otherwise the walk maps the
> + * per-domain page-table pages transiently with IRQs off, so it needs
> + * nothing from the current address space and is usable from any context --
> + * including the context switch, before the incoming vcpu's page-tables are
> + * loaded.
> + */
> +void populate_perdomain_mapping(const struct vcpu *v, unsigned long va,
> + const mfn_t *mfn, unsigned int nr,
> + unsigned int flags)
> +{
> + l1_pgentry_t *l1tab = NULL, *pl1e;
> + const l3_pgentry_t *l3tab;
> + const l2_pgentry_t *l2tab;
> + struct domain *d = v->domain;
> + unsigned long irq_flags;
> +
> + ASSERT(va >= PERDOMAIN_VIRT_START &&
> + va < PERDOMAIN_VIRT_SLOT(PERDOMAIN_SLOTS));
> + ASSERT(!nr || !l3_table_offset(va ^ (va + nr * PAGE_SIZE - 1)));
> + /* Area-owned pages are installed by create_perdomain_mapping() only. */
> + ASSERT(!(flags & _PAGE_AVAIL0));
> +
> + if ( likely(this_cpu(pgtable_vcpu) == v) )
> + {
> + unsigned int i;
> +
> + /*
> + * Fast path: v's page-tables are loaded on this pCPU, so the L1
> + * entries can be reached using the recursive linear mappings.
> + */
> + pl1e = &__linear_l1_table[l1_linear_offset(va)];
As mentioned elsewhere, I'm concerned of this (or really any) new use of
the linear page tables. (Which, ftaod, isn't an objection.)
> + for ( i = 0; i < nr; i++, pl1e++ )
> + {
> + /*
> + * An area-owned entry (installed by create_perdomain_mapping(),
> + * marked _PAGE_AVAIL0) holds the only reference to its page, so
> + * displacing it means freeing it. Nothing in this series
> + * replaces area-owned backing, hence the ASSERT_UNREACHABLE();
> + * any future caller doing so must run where freeing is
> + * permitted -- IRQs enabled, not in interrupt context (see
> + * ASSERT_ALLOC_CONTEXT()) -- which the context switch path is
> + * not.
> + */
> + if ( unlikely(perdomain_l1e_needs_freeing(*pl1e)) )
> + {
> + ASSERT_UNREACHABLE();
> + free_domheap_page(l1e_get_page(*pl1e));
> + }
> + l1e_write(pl1e, l1e_from_mfn(mfn[i], flags));
> + }
> +
> + return;
> + }
> +
> + BUG_ON(!d->arch.perdomain_l3_pg);
> +
> + /*
> + * Slow path: walk v's per-domain page-table pages. All mappings are
> + * local to this function, so disabling interrupts for the duration of
> + * the walk satisfies the map_domain_page_irqoff() contract. This in
> + * turn makes this function usable from the context switch path, where
> + * a plain map_domain_page() could recurse into __context_switch() via
> + * sync_local_execstate().
> + */
> + local_irq_save(irq_flags);
> +
> + l3tab = __map_domain_page_irqoff(d->arch.perdomain_l3_pg);
> +
> + /*
> + * Missing page-table structure is a hypervisor bug: there is no safe
> + * continuation, least of all from the context switch, where the next
> + * descriptor fetch through an unmapped GDT slot would be fatal.
> + */
> + BUG_ON(!(l3e_get_flags(l3tab[l3_table_offset(va)]) & _PAGE_PRESENT));
> +
> + l2tab = map_domain_page_irqoff(l3e_get_mfn(l3tab[l3_table_offset(va)]));
l3tab[] isn't used any further, so I think it wants unmapping right away. No
need to have undue pressure on the number of active mappings.
> + for ( ; nr--; va += PAGE_SIZE, mfn++ )
> + {
> + if ( !l1tab || !l1_table_offset(va) )
> + {
> + const l2_pgentry_t *pl2e = l2tab + l2_table_offset(va);
> +
> + BUG_ON(!(l2e_get_flags(*pl2e) & _PAGE_PRESENT));
> +
> + unmap_domain_page_irqoff(l1tab);
> + l1tab = map_domain_page_irqoff(l2e_get_mfn(*pl2e));
> + }
> +
> + pl1e = &l1tab[l1_table_offset(va)];
> +
> + /*
> + * As the fast path -- and the slow path holds IRQs off throughout,
> + * so replacing area-owned backing here is never permitted.
> + */
With this comment I think ...
> + if ( unlikely(perdomain_l1e_needs_freeing(*pl1e)) )
> + {
> + ASSERT_UNREACHABLE();
> + free_domheap_page(l1e_get_page(*pl1e));
... this call should be removed from here (I would have suggested to comment
it out, but Misra dislikes that iirc). Otherwise it would in principle be
reachable in release builds.
Maybe instead of ASSERT_UNREACHABLE() it should really be BUG() here.
Jan
^ permalink raw reply [flat|nested] 44+ messages in thread* Re: [PATCH v2 02/14] x86/mm: introduce populate_perdomain_mapping()
2026-09-03 15:57 ` Jan Beulich
@ 2026-09-03 21:27 ` George Dunlap
2026-09-04 5:58 ` Jan Beulich
0 siblings, 1 reply; 44+ messages in thread
From: George Dunlap @ 2026-09-03 21:27 UTC (permalink / raw)
To: Jan Beulich
Cc: Roger Pau Monné, Andrew Cooper, Alejandro Vallejo,
Teddy Astie, Anthony PERARD, Michal Orzel, Julien Grall,
Stefano Stabellini, xen-devel
On Thu, Sep 3, 2026 at 4:57 PM Jan Beulich <jbeulich@suse.com> wrote:
> > + if ( likely(this_cpu(pgtable_vcpu) == v) )
> > + {
> > + unsigned int i;
> > +
> > + /*
> > + * Fast path: v's page-tables are loaded on this pCPU, so the L1
> > + * entries can be reached using the recursive linear mappings.
> > + */
> > + pl1e = &__linear_l1_table[l1_linear_offset(va)];
>
> As mentioned elsewhere, I'm concerned of this (or really any) new use of
> the linear page tables. (Which, ftaod, isn't an objection.)
FWIW here we can drop this fast path at any time, and we get exactly
the same result as if we drop it from the patch now.
> > + l2tab = map_domain_page_irqoff(l3e_get_mfn(l3tab[l3_table_offset(va)]));
>
> l3tab[] isn't used any further, so I think it wants unmapping right away. No
> need to have undue pressure on the number of active mappings.
Ack
> > + /*
> > + * As the fast path -- and the slow path holds IRQs off throughout,
> > + * so replacing area-owned backing here is never permitted.
> > + */
>
> With this comment I think ...
>
> > + if ( unlikely(perdomain_l1e_needs_freeing(*pl1e)) )
> > + {
> > + ASSERT_UNREACHABLE();
> > + free_domheap_page(l1e_get_page(*pl1e));
>
> ... this call should be removed from here (I would have suggested to comment
> it out, but Misra dislikes that iirc). Otherwise it would in principle be
> reachable in release builds.
Ah, right -- sorry, this was safe in v1's "xenheap-walk" approach
which doesn't need to disable interrupts; with the "mapcache-walk"
approach we can't do this any more.
> Maybe instead of ASSERT_UNREACHABLE() it should really be BUG() here.
Hrm, docs/misc/xen-error-handling.txt used to have guidelines about
*which* error handling to use. Basically, BUG() is an immediate
hypervisor crash DoS in production, and should only be used when
there's no way to continue without doing something worse. Leaking
memory is a lower-grade DoS, and so would be preferable in production
(though in this case probably with a printk, so there's some hope of
figuring out what's wrong).
In theory we could allow callers who knew they'd take the fast path to
do the free; but then we couldn't just rip out the fast path without
doing more surgery.
So I'd propose: In both fast and slow paths, replace the free with a
printk (keeping the ASSERT_UNREACHABLE), with a comment explaining why
leaking is preferable to BUG.
-George
^ permalink raw reply [flat|nested] 44+ messages in thread* Re: [PATCH v2 02/14] x86/mm: introduce populate_perdomain_mapping()
2026-09-03 21:27 ` George Dunlap
@ 2026-09-04 5:58 ` Jan Beulich
0 siblings, 0 replies; 44+ messages in thread
From: Jan Beulich @ 2026-09-04 5:58 UTC (permalink / raw)
To: George Dunlap
Cc: Roger Pau Monné, Andrew Cooper, Alejandro Vallejo,
Teddy Astie, Anthony PERARD, Michal Orzel, Julien Grall,
Stefano Stabellini, xen-devel
On 03.09.2026 23:27, George Dunlap wrote:
> So I'd propose: In both fast and slow paths, replace the free with a
> printk (keeping the ASSERT_UNREACHABLE), with a comment explaining why
> leaking is preferable to BUG.
Fine with me. With this and the l3tab[] unmapping moved earlier:
Reviewed-by: Jan Beulich <jbeulich@suse.com>
Jan
^ permalink raw reply [flat|nested] 44+ messages in thread
* Re: [PATCH v2 02/14] x86/mm: introduce populate_perdomain_mapping()
2026-09-02 9:43 ` [PATCH v2 02/14] x86/mm: introduce populate_perdomain_mapping() George Dunlap
2026-09-03 15:57 ` Jan Beulich
@ 2026-09-04 5:47 ` Jan Beulich
1 sibling, 0 replies; 44+ messages in thread
From: Jan Beulich @ 2026-09-04 5:47 UTC (permalink / raw)
To: George Dunlap
Cc: Andrew Cooper, Roger Pau Monné, Alejandro Vallejo,
Teddy Astie, Anthony PERARD, Michal Orzel, Julien Grall,
Stefano Stabellini, George Dunlap, xen-devel
On 02.09.2026 11:43, George Dunlap wrote:
> --- a/xen/arch/x86/mm.c
> +++ b/xen/arch/x86/mm.c
> @@ -6334,6 +6334,130 @@ int create_perdomain_mapping(struct domain *d, unsigned long va,
> return rc;
> }
>
> +/*
> + * Map @nr pages, @mfn[0..nr-1], at consecutive pages from @va in v's view of
> + * the per-domain area, with page-table @flags. The range must lie within a
> + * single per-domain slot, and must already have been plumbed down to the L1
> + * tables by create_perdomain_mapping(): missing structure is a bug. A
> + * present entry not owned by the area (no _PAGE_AVAIL0) is silently
> + * replaced, as that is how callers update their mappings; a present
> + * area-owned entry is freed and replaced, which constrains the calling
> + * context (see the comment in the body). No TLB flushing is done: the
> + * caller decides whether the old translations can still be cached
> + * anywhere.
> + *
> + * When v's page-tables are loaded on this pCPU the L1 entries are reached
> + * through the recursive linear mappings; otherwise the walk maps the
> + * per-domain page-table pages transiently with IRQs off, so it needs
> + * nothing from the current address space and is usable from any context --
> + * including the context switch, before the incoming vcpu's page-tables are
> + * loaded.
> + */
> +void populate_perdomain_mapping(const struct vcpu *v, unsigned long va,
> + const mfn_t *mfn, unsigned int nr,
> + unsigned int flags)
> +{
> + l1_pgentry_t *l1tab = NULL, *pl1e;
> + const l3_pgentry_t *l3tab;
> + const l2_pgentry_t *l2tab;
> + struct domain *d = v->domain;
> + unsigned long irq_flags;
> +
> + ASSERT(va >= PERDOMAIN_VIRT_START &&
> + va < PERDOMAIN_VIRT_SLOT(PERDOMAIN_SLOTS));
> + ASSERT(!nr || !l3_table_offset(va ^ (va + nr * PAGE_SIZE - 1)));
> + /* Area-owned pages are installed by create_perdomain_mapping() only. */
> + ASSERT(!(flags & _PAGE_AVAIL0));
> +
> + if ( likely(this_cpu(pgtable_vcpu) == v) )
> + {
> + unsigned int i;
> +
> + /*
> + * Fast path: v's page-tables are loaded on this pCPU, so the L1
> + * entries can be reached using the recursive linear mappings.
> + */
> + pl1e = &__linear_l1_table[l1_linear_offset(va)];
> +
> + for ( i = 0; i < nr; i++, pl1e++ )
> + {
> + /*
> + * An area-owned entry (installed by create_perdomain_mapping(),
> + * marked _PAGE_AVAIL0) holds the only reference to its page, so
> + * displacing it means freeing it. Nothing in this series
> + * replaces area-owned backing, hence the ASSERT_UNREACHABLE();
> + * any future caller doing so must run where freeing is
> + * permitted -- IRQs enabled, not in interrupt context (see
> + * ASSERT_ALLOC_CONTEXT()) -- which the context switch path is
> + * not.
> + */
> + if ( unlikely(perdomain_l1e_needs_freeing(*pl1e)) )
> + {
> + ASSERT_UNREACHABLE();
> + free_domheap_page(l1e_get_page(*pl1e));
> + }
> + l1e_write(pl1e, l1e_from_mfn(mfn[i], flags));
> + }
> +
> + return;
> + }
> +
> + BUG_ON(!d->arch.perdomain_l3_pg);
> +
> + /*
> + * Slow path: walk v's per-domain page-table pages. All mappings are
> + * local to this function, so disabling interrupts for the duration of
> + * the walk satisfies the map_domain_page_irqoff() contract. This in
> + * turn makes this function usable from the context switch path, where
> + * a plain map_domain_page() could recurse into __context_switch() via
> + * sync_local_execstate().
> + */
> + local_irq_save(irq_flags);
> +
> + l3tab = __map_domain_page_irqoff(d->arch.perdomain_l3_pg);
> +
> + /*
> + * Missing page-table structure is a hypervisor bug: there is no safe
> + * continuation, least of all from the context switch, where the next
> + * descriptor fetch through an unmapped GDT slot would be fatal.
> + */
> + BUG_ON(!(l3e_get_flags(l3tab[l3_table_offset(va)]) & _PAGE_PRESENT));
> +
> + l2tab = map_domain_page_irqoff(l3e_get_mfn(l3tab[l3_table_offset(va)]));
> +
> + for ( ; nr--; va += PAGE_SIZE, mfn++ )
> + {
> + if ( !l1tab || !l1_table_offset(va) )
> + {
> + const l2_pgentry_t *pl2e = l2tab + l2_table_offset(va);
> +
> + BUG_ON(!(l2e_get_flags(*pl2e) & _PAGE_PRESENT));
> +
> + unmap_domain_page_irqoff(l1tab);
> + l1tab = map_domain_page_irqoff(l2e_get_mfn(*pl2e));
> + }
> +
> + pl1e = &l1tab[l1_table_offset(va)];
> +
> + /*
> + * As the fast path -- and the slow path holds IRQs off throughout,
> + * so replacing area-owned backing here is never permitted.
> + */
> + if ( unlikely(perdomain_l1e_needs_freeing(*pl1e)) )
> + {
> + ASSERT_UNREACHABLE();
> + free_domheap_page(l1e_get_page(*pl1e));
> + }
Just for the possible case of freeing here really becoming necessary: This
could be deferred until ...
> + l1e_write(pl1e, l1e_from_mfn(*mfn, flags));
> + }
> +
> + unmap_domain_page_irqoff(l1tab);
> + unmap_domain_page_irqoff(l2tab);
> + unmap_domain_page_irqoff(l3tab);
> +
> + local_irq_restore(irq_flags);
... here. Easily for the nr == 1 case (just requires a local variable to
hold MFN or struct page_info *), and with a slight change to the contract
with the caller (allowing mfn[] to be altered) also in the general case.
Jan
^ permalink raw reply [flat|nested] 44+ messages in thread
* [PATCH v2 03/14] x86/pv: use populate_perdomain_mapping() to map the Xen GDT
2026-09-02 9:43 [PATCH v2 00/14] x86: Address Space Isolation, part 2: asi= option and per-vCPU page tables George Dunlap
2026-09-02 9:43 ` [PATCH v2 01/14] x86/domain_page: introduce IRQs-off variants of {,un}map_domain_page() George Dunlap
2026-09-02 9:43 ` [PATCH v2 02/14] x86/mm: introduce populate_perdomain_mapping() George Dunlap
@ 2026-09-02 9:43 ` George Dunlap
2026-09-03 16:11 ` Jan Beulich
2026-09-02 9:43 ` [PATCH v2 04/14] x86/pv: set/clear guest GDT mappings using populate_perdomain_mapping() George Dunlap
` (10 subsequent siblings)
13 siblings, 1 reply; 44+ messages in thread
From: George Dunlap @ 2026-09-02 9:43 UTC (permalink / raw)
To: xen-devel
Cc: Roger Pau Monné, Jan Beulich, Andrew Cooper,
Roger Pau Monné, Alejandro Vallejo, Teddy Astie,
Anthony PERARD, Michal Orzel, Julien Grall, Stefano Stabellini,
George Dunlap
From: Roger Pau Monné <roger.pau@citrix.com>
Currently, update_xen_slot_in_full_gdt() uses the stashed direct-map
pointer in d->arch.pv.gdt_ldt_l1tab to update the incoming vcpu's
page tables with Xen's GDT, by writing a stashed per-cpu copy of a
pre-baked L1 entry (either 64-bit or compat version).
Switch this to using populate_perdomain_mapping(), which doesn't rely
on the stashed address of the l1 page in the direct map. Rather than
also stashing a pre-baked value for the payload, compute the mfn from
the per-cpu GDT pointer at use: the conversion is a handful of cycles
on a path costing thousands, and computing at use removes the
parallel {,compat_}gdt_l1e bookkeeping along with its boot-ordering
constraint (the cached value could only be generated after Xen's
physical relocation, and had to be in place before the first context
switch; a use-time lookup is correct by construction). The flags on
the final mapping are identical.
Signed-off-by: Roger Pau Monné <roger.pau@citrix.com>
Assisted-by: Claude Code:claude-fable-5, Claude Code:claude-opus-4-8
Signed-off-by: George Dunlap <gwd@xenproject.org>
---
Changes in v2:
- Drop the {,compat_}gdt_mfn caching entirely (suggested by Andrew
Cooper): compute virt_to_mfn() from the per-cpu GDT pointer at use.
The PDX lookup behind it measures ~5-10 cycles warm against a
~1,500-cycle context switch, and this removes the double
bookkeeping and the after-relocation caching constraint. The
cached-MFN assertion goes with the cache: a use-time computation
from a live pointer needs no staleness check.
Changes since the previously posted version:
- populate_perdomain_mapping() introduction split into the previous
patch; this patch is now just the Xen GDT conversion.
- Retain the "GDT MFN cached" check as ASSERT(mfn_x(mfn)).
---
xen/arch/x86/domain.c | 13 ++++++++-----
xen/arch/x86/include/asm/desc.h | 2 --
xen/arch/x86/smpboot.c | 15 ---------------
xen/arch/x86/traps.c | 2 --
4 files changed, 8 insertions(+), 24 deletions(-)
diff --git a/xen/arch/x86/domain.c b/xen/arch/x86/domain.c
index 996b50af7a..d8af06e533 100644
--- a/xen/arch/x86/domain.c
+++ b/xen/arch/x86/domain.c
@@ -2062,11 +2062,14 @@ static always_inline bool need_full_gdt(const struct domain *d)
static void update_xen_slot_in_full_gdt(const struct vcpu *v, unsigned int cpu)
{
- ASSERT(per_cpu(gdt_l1e, cpu).l1); /* Confirm these have been cached. */
-
- l1e_write(pv_gdt_ptes(v) + FIRST_RESERVED_GDT_PAGE,
- !is_pv_32bit_vcpu(v) ? per_cpu(gdt_l1e, cpu)
- : per_cpu(compat_gdt_l1e, cpu));
+ mfn_t mfn = _mfn(virt_to_mfn(!is_pv_32bit_vcpu(v)
+ ? per_cpu(gdt, cpu)
+ : per_cpu(compat_gdt, cpu)));
+
+ populate_perdomain_mapping(v,
+ GDT_VIRT_START(v) +
+ (FIRST_RESERVED_GDT_PAGE << PAGE_SHIFT),
+ &mfn, 1, __PAGE_HYPERVISOR_RW);
}
static void load_full_gdt(const struct vcpu *v, unsigned int cpu)
diff --git a/xen/arch/x86/include/asm/desc.h b/xen/arch/x86/include/asm/desc.h
index dcbdac3ff7..a860134211 100644
--- a/xen/arch/x86/include/asm/desc.h
+++ b/xen/arch/x86/include/asm/desc.h
@@ -136,10 +136,8 @@ struct __packed desc_ptr {
extern seg_desc_t boot_gdt[];
DECLARE_PER_CPU(seg_desc_t *, gdt);
-DECLARE_PER_CPU(l1_pgentry_t, gdt_l1e);
extern seg_desc_t boot_compat_gdt[];
DECLARE_PER_CPU(seg_desc_t *, compat_gdt);
-DECLARE_PER_CPU(l1_pgentry_t, compat_gdt_l1e);
DECLARE_PER_CPU(bool, full_gdt_loaded);
static inline void lgdt(const struct desc_ptr *gdtr)
diff --git a/xen/arch/x86/smpboot.c b/xen/arch/x86/smpboot.c
index 84e9e4beed..9b837a1769 100644
--- a/xen/arch/x86/smpboot.c
+++ b/xen/arch/x86/smpboot.c
@@ -1085,8 +1085,6 @@ static int cpu_smpboot_alloc(unsigned int cpu)
if ( gdt == NULL )
goto out;
per_cpu(gdt, cpu) = gdt;
- per_cpu(gdt_l1e, cpu) =
- l1e_from_pfn(virt_to_mfn(gdt), __PAGE_HYPERVISOR_RW);
memcpy(gdt, boot_gdt, NR_RESERVED_GDT_PAGES * PAGE_SIZE);
BUILD_BUG_ON(NR_CPUS > 0x10000);
gdt[PER_CPU_GDT_ENTRY - FIRST_RESERVED_GDT_ENTRY].a = cpu;
@@ -1095,8 +1093,6 @@ static int cpu_smpboot_alloc(unsigned int cpu)
per_cpu(compat_gdt, cpu) = gdt = alloc_xenheap_pages(0, memflags);
if ( gdt == NULL )
goto out;
- per_cpu(compat_gdt_l1e, cpu) =
- l1e_from_pfn(virt_to_mfn(gdt), __PAGE_HYPERVISOR_RW);
memcpy(gdt, boot_compat_gdt, NR_RESERVED_GDT_PAGES * PAGE_SIZE);
gdt[PER_CPU_GDT_ENTRY - FIRST_RESERVED_GDT_ENTRY].a = cpu;
#endif
@@ -1173,17 +1169,6 @@ void __init smp_prepare_cpus(void)
initialize_cpu_data(0); /* Final full version of the data */
print_cpu_info(0);
- /*
- * Cache {,compat_}gdt_l1e for the BSP now that physically relocation is
- * done. It must be after physical relocation of Xen, and before the
- * first context_switch().
- */
- this_cpu(gdt_l1e) =
- l1e_from_pfn(virt_to_mfn(boot_gdt), __PAGE_HYPERVISOR_RW);
- if ( IS_ENABLED(CONFIG_PV32) )
- this_cpu(compat_gdt_l1e) =
- l1e_from_pfn(virt_to_mfn(boot_compat_gdt), __PAGE_HYPERVISOR_RW);
-
boot_cpu_physical_apicid = get_apic_id();
x86_cpu_to_apicid[0] = boot_cpu_physical_apicid;
diff --git a/xen/arch/x86/traps.c b/xen/arch/x86/traps.c
index 1774966305..2ab61db167 100644
--- a/xen/arch/x86/traps.c
+++ b/xen/arch/x86/traps.c
@@ -71,10 +71,8 @@ DEFINE_PER_CPU(uint64_t, efer);
static DEFINE_PER_CPU(unsigned long, last_extable_addr);
DEFINE_PER_CPU_READ_MOSTLY(seg_desc_t *, gdt);
-DEFINE_PER_CPU_READ_MOSTLY(l1_pgentry_t, gdt_l1e);
#ifdef CONFIG_PV32
DEFINE_PER_CPU_READ_MOSTLY(seg_desc_t *, compat_gdt);
-DEFINE_PER_CPU_READ_MOSTLY(l1_pgentry_t, compat_gdt_l1e);
#endif
/*
--
2.55.0
^ permalink raw reply related [flat|nested] 44+ messages in thread* Re: [PATCH v2 03/14] x86/pv: use populate_perdomain_mapping() to map the Xen GDT
2026-09-02 9:43 ` [PATCH v2 03/14] x86/pv: use populate_perdomain_mapping() to map the Xen GDT George Dunlap
@ 2026-09-03 16:11 ` Jan Beulich
2026-09-03 22:35 ` George Dunlap
0 siblings, 1 reply; 44+ messages in thread
From: Jan Beulich @ 2026-09-03 16:11 UTC (permalink / raw)
To: George Dunlap
Cc: Andrew Cooper, Roger Pau Monné, Alejandro Vallejo,
Teddy Astie, Anthony PERARD, Michal Orzel, Julien Grall,
Stefano Stabellini, George Dunlap, xen-devel, Juergen Gross
On 02.09.2026 11:43, George Dunlap wrote:
> From: Roger Pau Monné <roger.pau@citrix.com>
>
> Currently, update_xen_slot_in_full_gdt() uses the stashed direct-map
> pointer in d->arch.pv.gdt_ldt_l1tab to update the incoming vcpu's
> page tables with Xen's GDT, by writing a stashed per-cpu copy of a
> pre-baked L1 entry (either 64-bit or compat version).
>
> Switch this to using populate_perdomain_mapping(), which doesn't rely
> on the stashed address of the l1 page in the direct map. Rather than
> also stashing a pre-baked value for the payload, compute the mfn from
> the per-cpu GDT pointer at use: the conversion is a handful of cycles
> on a path costing thousands, and computing at use removes the
> parallel {,compat_}gdt_l1e bookkeeping along with its boot-ordering
> constraint (the cached value could only be generated after Xen's
> physical relocation, and had to be in place before the first context
> switch; a use-time lookup is correct by construction). The flags on
> the final mapping are identical.
>
> Signed-off-by: Roger Pau Monné <roger.pau@citrix.com>
> Assisted-by: Claude Code:claude-fable-5, Claude Code:claude-opus-4-8
> Signed-off-by: George Dunlap <gwd@xenproject.org>
> ---
> Changes in v2:
> - Drop the {,compat_}gdt_mfn caching entirely (suggested by Andrew
> Cooper): compute virt_to_mfn() from the per-cpu GDT pointer at use.
> The PDX lookup behind it measures ~5-10 cycles warm against a
> ~1,500-cycle context switch, and this removes the double
> bookkeeping and the after-relocation caching constraint. The
> cached-MFN assertion goes with the cache: a use-time computation
> from a live pointer needs no staleness check.
This looks to contradict what 564d261687c0 ("x86/ctxt-switch: Document
and improve GDT handling") used as justification to put in place the
caching. Also Cc-ing Jürgen, who also was involved there, for possible
further insight.
Functionally the change looks okay to me, but the above will need
sorting, at the very least by specifically discussing why effectively
undoing that earlier change is okay.
Jan
^ permalink raw reply [flat|nested] 44+ messages in thread* Re: [PATCH v2 03/14] x86/pv: use populate_perdomain_mapping() to map the Xen GDT
2026-09-03 16:11 ` Jan Beulich
@ 2026-09-03 22:35 ` George Dunlap
2026-09-04 6:00 ` Jan Beulich
0 siblings, 1 reply; 44+ messages in thread
From: George Dunlap @ 2026-09-03 22:35 UTC (permalink / raw)
To: Jan Beulich
Cc: Andrew Cooper, Roger Pau Monné, Alejandro Vallejo,
Teddy Astie, Anthony PERARD, Michal Orzel, Julien Grall,
Stefano Stabellini, xen-devel, Juergen Gross
On Thu, Sep 3, 2026 at 5:11 PM Jan Beulich <jbeulich@suse.com> wrote:
>
> On 02.09.2026 11:43, George Dunlap wrote:
> > From: Roger Pau Monné <roger.pau@citrix.com>
> >
> > Currently, update_xen_slot_in_full_gdt() uses the stashed direct-map
> > pointer in d->arch.pv.gdt_ldt_l1tab to update the incoming vcpu's
> > page tables with Xen's GDT, by writing a stashed per-cpu copy of a
> > pre-baked L1 entry (either 64-bit or compat version).
> >
> > Switch this to using populate_perdomain_mapping(), which doesn't rely
> > on the stashed address of the l1 page in the direct map. Rather than
> > also stashing a pre-baked value for the payload, compute the mfn from
> > the per-cpu GDT pointer at use: the conversion is a handful of cycles
> > on a path costing thousands, and computing at use removes the
> > parallel {,compat_}gdt_l1e bookkeeping along with its boot-ordering
> > constraint (the cached value could only be generated after Xen's
> > physical relocation, and had to be in place before the first context
> > switch; a use-time lookup is correct by construction). The flags on
> > the final mapping are identical.
> >
> > Signed-off-by: Roger Pau Monné <roger.pau@citrix.com>
> > Assisted-by: Claude Code:claude-fable-5, Claude Code:claude-opus-4-8
> > Signed-off-by: George Dunlap <gwd@xenproject.org>
> > ---
> > Changes in v2:
> > - Drop the {,compat_}gdt_mfn caching entirely (suggested by Andrew
> > Cooper): compute virt_to_mfn() from the per-cpu GDT pointer at use.
> > The PDX lookup behind it measures ~5-10 cycles warm against a
> > ~1,500-cycle context switch, and this removes the double
> > bookkeeping and the after-relocation caching constraint. The
> > cached-MFN assertion goes with the cache: a use-time computation
> > from a live pointer needs no staleness check.
>
> This looks to contradict what 564d261687c0 ("x86/ctxt-switch: Document
> and improve GDT handling") used as justification to put in place the
> caching. Also Cc-ing Jürgen, who also was involved there, for possible
> further insight.
>
> Functionally the change looks okay to me, but the above will need
> sorting, at the very least by specifically discussing why effectively
> undoing that earlier change is okay.
So looking back at the thread, Jürgen measured a 14% improvement for
something that might be described as a microbenchmark before and after
the patch (a benchmark purposely trying to set up an unusual scenario
to maximize the effect of context switch overhead, not one to
represent a typical workflow). But are the numbers really plausible?
Even at an implausible 100k switches/s across the box, saving 100
cycles per switch is about 0.04% of eight 3 GHz cores.
At any rate, we're already adding several map/unmap operations, and
about to add several more. Keeping the PTE caching would require
adding a separate path that can write just PTEs, which then will
potentially further complication future paths where we need to make
sure we handle both domain-wide perdomain areas and per-vcpu areas.
If it were easy I would already have been keeping it.
I'd be inclined to say: Since we're going to be adding more
populate_perdomain_mapping() calls anyway, let's do it the simple
correct way first; and then explore the idea of stashing mfns of
frequently-mapped L1s (rather than having to walk L3 -> L2 -> L1); and
at that time look into stashing baked l1es to avoid conversions.
-George
^ permalink raw reply [flat|nested] 44+ messages in thread* Re: [PATCH v2 03/14] x86/pv: use populate_perdomain_mapping() to map the Xen GDT
2026-09-03 22:35 ` George Dunlap
@ 2026-09-04 6:00 ` Jan Beulich
2026-09-04 6:54 ` Jürgen Groß
0 siblings, 1 reply; 44+ messages in thread
From: Jan Beulich @ 2026-09-04 6:00 UTC (permalink / raw)
To: George Dunlap
Cc: Andrew Cooper, Roger Pau Monné, Alejandro Vallejo,
Teddy Astie, Anthony PERARD, Michal Orzel, Julien Grall,
Stefano Stabellini, xen-devel, Juergen Gross
On 04.09.2026 00:35, George Dunlap wrote:
> On Thu, Sep 3, 2026 at 5:11 PM Jan Beulich <jbeulich@suse.com> wrote:
>>
>> On 02.09.2026 11:43, George Dunlap wrote:
>>> From: Roger Pau Monné <roger.pau@citrix.com>
>>>
>>> Currently, update_xen_slot_in_full_gdt() uses the stashed direct-map
>>> pointer in d->arch.pv.gdt_ldt_l1tab to update the incoming vcpu's
>>> page tables with Xen's GDT, by writing a stashed per-cpu copy of a
>>> pre-baked L1 entry (either 64-bit or compat version).
>>>
>>> Switch this to using populate_perdomain_mapping(), which doesn't rely
>>> on the stashed address of the l1 page in the direct map. Rather than
>>> also stashing a pre-baked value for the payload, compute the mfn from
>>> the per-cpu GDT pointer at use: the conversion is a handful of cycles
>>> on a path costing thousands, and computing at use removes the
>>> parallel {,compat_}gdt_l1e bookkeeping along with its boot-ordering
>>> constraint (the cached value could only be generated after Xen's
>>> physical relocation, and had to be in place before the first context
>>> switch; a use-time lookup is correct by construction). The flags on
>>> the final mapping are identical.
>>>
>>> Signed-off-by: Roger Pau Monné <roger.pau@citrix.com>
>>> Assisted-by: Claude Code:claude-fable-5, Claude Code:claude-opus-4-8
>>> Signed-off-by: George Dunlap <gwd@xenproject.org>
>>> ---
>>> Changes in v2:
>>> - Drop the {,compat_}gdt_mfn caching entirely (suggested by Andrew
>>> Cooper): compute virt_to_mfn() from the per-cpu GDT pointer at use.
>>> The PDX lookup behind it measures ~5-10 cycles warm against a
>>> ~1,500-cycle context switch, and this removes the double
>>> bookkeeping and the after-relocation caching constraint. The
>>> cached-MFN assertion goes with the cache: a use-time computation
>>> from a live pointer needs no staleness check.
>>
>> This looks to contradict what 564d261687c0 ("x86/ctxt-switch: Document
>> and improve GDT handling") used as justification to put in place the
>> caching. Also Cc-ing Jürgen, who also was involved there, for possible
>> further insight.
>>
>> Functionally the change looks okay to me, but the above will need
>> sorting, at the very least by specifically discussing why effectively
>> undoing that earlier change is okay.
>
> So looking back at the thread, Jürgen measured a 14% improvement for
> something that might be described as a microbenchmark before and after
> the patch (a benchmark purposely trying to set up an unusual scenario
> to maximize the effect of context switch overhead, not one to
> represent a typical workflow). But are the numbers really plausible?
> Even at an implausible 100k switches/s across the box, saving 100
> cycles per switch is about 0.04% of eight 3 GHz cores.
>
> At any rate, we're already adding several map/unmap operations, and
> about to add several more. Keeping the PTE caching would require
> adding a separate path that can write just PTEs, which then will
> potentially further complication future paths where we need to make
> sure we handle both domain-wide perdomain areas and per-vcpu areas.
> If it were easy I would already have been keeping it.
>
> I'd be inclined to say: Since we're going to be adding more
> populate_perdomain_mapping() calls anyway, let's do it the simple
> correct way first; and then explore the idea of stashing mfns of
> frequently-mapped L1s (rather than having to walk L3 -> L2 -> L1); and
> at that time look into stashing baked l1es to avoid conversions.
Perhaps; I'd like to have Jürgen's and/or Andrew's input here, though.
Jan
^ permalink raw reply [flat|nested] 44+ messages in thread* Re: [PATCH v2 03/14] x86/pv: use populate_perdomain_mapping() to map the Xen GDT
2026-09-04 6:00 ` Jan Beulich
@ 2026-09-04 6:54 ` Jürgen Groß
2026-09-04 8:06 ` George Dunlap
0 siblings, 1 reply; 44+ messages in thread
From: Jürgen Groß @ 2026-09-04 6:54 UTC (permalink / raw)
To: Jan Beulich, George Dunlap
Cc: Andrew Cooper, Roger Pau Monné, Alejandro Vallejo,
Teddy Astie, Anthony PERARD, Michal Orzel, Julien Grall,
Stefano Stabellini, xen-devel
[-- Attachment #1.1.1: Type: text/plain, Size: 4300 bytes --]
On 04.09.26 08:00, Jan Beulich wrote:
> On 04.09.2026 00:35, George Dunlap wrote:
>> On Thu, Sep 3, 2026 at 5:11 PM Jan Beulich <jbeulich@suse.com> wrote:
>>>
>>> On 02.09.2026 11:43, George Dunlap wrote:
>>>> From: Roger Pau Monné <roger.pau@citrix.com>
>>>>
>>>> Currently, update_xen_slot_in_full_gdt() uses the stashed direct-map
>>>> pointer in d->arch.pv.gdt_ldt_l1tab to update the incoming vcpu's
>>>> page tables with Xen's GDT, by writing a stashed per-cpu copy of a
>>>> pre-baked L1 entry (either 64-bit or compat version).
>>>>
>>>> Switch this to using populate_perdomain_mapping(), which doesn't rely
>>>> on the stashed address of the l1 page in the direct map. Rather than
>>>> also stashing a pre-baked value for the payload, compute the mfn from
>>>> the per-cpu GDT pointer at use: the conversion is a handful of cycles
>>>> on a path costing thousands, and computing at use removes the
>>>> parallel {,compat_}gdt_l1e bookkeeping along with its boot-ordering
>>>> constraint (the cached value could only be generated after Xen's
>>>> physical relocation, and had to be in place before the first context
>>>> switch; a use-time lookup is correct by construction). The flags on
>>>> the final mapping are identical.
>>>>
>>>> Signed-off-by: Roger Pau Monné <roger.pau@citrix.com>
>>>> Assisted-by: Claude Code:claude-fable-5, Claude Code:claude-opus-4-8
>>>> Signed-off-by: George Dunlap <gwd@xenproject.org>
>>>> ---
>>>> Changes in v2:
>>>> - Drop the {,compat_}gdt_mfn caching entirely (suggested by Andrew
>>>> Cooper): compute virt_to_mfn() from the per-cpu GDT pointer at use.
>>>> The PDX lookup behind it measures ~5-10 cycles warm against a
>>>> ~1,500-cycle context switch, and this removes the double
>>>> bookkeeping and the after-relocation caching constraint. The
>>>> cached-MFN assertion goes with the cache: a use-time computation
>>>> from a live pointer needs no staleness check.
>>>
>>> This looks to contradict what 564d261687c0 ("x86/ctxt-switch: Document
>>> and improve GDT handling") used as justification to put in place the
>>> caching. Also Cc-ing Jürgen, who also was involved there, for possible
>>> further insight.
>>>
>>> Functionally the change looks okay to me, but the above will need
>>> sorting, at the very least by specifically discussing why effectively
>>> undoing that earlier change is okay.
>>
>> So looking back at the thread, Jürgen measured a 14% improvement for
>> something that might be described as a microbenchmark before and after
>> the patch (a benchmark purposely trying to set up an unusual scenario
>> to maximize the effect of context switch overhead, not one to
>> represent a typical workflow). But are the numbers really plausible?
>> Even at an implausible 100k switches/s across the box, saving 100
>> cycles per switch is about 0.04% of eight 3 GHz cores.
>>
>> At any rate, we're already adding several map/unmap operations, and
>> about to add several more. Keeping the PTE caching would require
>> adding a separate path that can write just PTEs, which then will
>> potentially further complication future paths where we need to make
>> sure we handle both domain-wide perdomain areas and per-vcpu areas.
>> If it were easy I would already have been keeping it.
>>
>> I'd be inclined to say: Since we're going to be adding more
>> populate_perdomain_mapping() calls anyway, let's do it the simple
>> correct way first; and then explore the idea of stashing mfns of
>> frequently-mapped L1s (rather than having to walk L3 -> L2 -> L1); and
>> at that time look into stashing baked l1es to avoid conversions.
>
> Perhaps; I'd like to have Jürgen's and/or Andrew's input here, though.
At that time I implemented core scheduling in Xen. I noticed that very
subtle changes in the context switch path could result in unexpected large
performance differences. As I had the performance test for my purpose
already set up, I used it for Andrew's patch (which was a result of my
context switch path performance findings) and really did measure the
impressive effect of it.
Note that you can't only count instructions, often cache effects and
branch predictions are dominating the performance.
Juergen
[-- Attachment #1.1.2: OpenPGP public key --]
[-- Type: application/pgp-keys, Size: 3743 bytes --]
[-- Attachment #2: OpenPGP digital signature --]
[-- Type: application/pgp-signature, Size: 495 bytes --]
^ permalink raw reply [flat|nested] 44+ messages in thread* Re: [PATCH v2 03/14] x86/pv: use populate_perdomain_mapping() to map the Xen GDT
2026-09-04 6:54 ` Jürgen Groß
@ 2026-09-04 8:06 ` George Dunlap
2026-09-04 8:29 ` Jan Beulich
0 siblings, 1 reply; 44+ messages in thread
From: George Dunlap @ 2026-09-04 8:06 UTC (permalink / raw)
To: Jürgen Groß
Cc: Jan Beulich, Andrew Cooper, Roger Pau Monné,
Alejandro Vallejo, Teddy Astie, Anthony PERARD, Michal Orzel,
Julien Grall, Stefano Stabellini, xen-devel
On Fri, Sep 4, 2026 at 7:54 AM Jürgen Groß <jgross@suse.com> wrote:
>
> On 04.09.26 08:00, Jan Beulich wrote:
> > On 04.09.2026 00:35, George Dunlap wrote:
> >> On Thu, Sep 3, 2026 at 5:11 PM Jan Beulich <jbeulich@suse.com> wrote:
> >>>
> >>> On 02.09.2026 11:43, George Dunlap wrote:
> >>>> From: Roger Pau Monné <roger.pau@citrix.com>
> >>>>
> >>>> Currently, update_xen_slot_in_full_gdt() uses the stashed direct-map
> >>>> pointer in d->arch.pv.gdt_ldt_l1tab to update the incoming vcpu's
> >>>> page tables with Xen's GDT, by writing a stashed per-cpu copy of a
> >>>> pre-baked L1 entry (either 64-bit or compat version).
> >>>>
> >>>> Switch this to using populate_perdomain_mapping(), which doesn't rely
> >>>> on the stashed address of the l1 page in the direct map. Rather than
> >>>> also stashing a pre-baked value for the payload, compute the mfn from
> >>>> the per-cpu GDT pointer at use: the conversion is a handful of cycles
> >>>> on a path costing thousands, and computing at use removes the
> >>>> parallel {,compat_}gdt_l1e bookkeeping along with its boot-ordering
> >>>> constraint (the cached value could only be generated after Xen's
> >>>> physical relocation, and had to be in place before the first context
> >>>> switch; a use-time lookup is correct by construction). The flags on
> >>>> the final mapping are identical.
> >>>>
> >>>> Signed-off-by: Roger Pau Monné <roger.pau@citrix.com>
> >>>> Assisted-by: Claude Code:claude-fable-5, Claude Code:claude-opus-4-8
> >>>> Signed-off-by: George Dunlap <gwd@xenproject.org>
> >>>> ---
> >>>> Changes in v2:
> >>>> - Drop the {,compat_}gdt_mfn caching entirely (suggested by Andrew
> >>>> Cooper): compute virt_to_mfn() from the per-cpu GDT pointer at use.
> >>>> The PDX lookup behind it measures ~5-10 cycles warm against a
> >>>> ~1,500-cycle context switch, and this removes the double
> >>>> bookkeeping and the after-relocation caching constraint. The
> >>>> cached-MFN assertion goes with the cache: a use-time computation
> >>>> from a live pointer needs no staleness check.
> >>>
> >>> This looks to contradict what 564d261687c0 ("x86/ctxt-switch: Document
> >>> and improve GDT handling") used as justification to put in place the
> >>> caching. Also Cc-ing Jürgen, who also was involved there, for possible
> >>> further insight.
> >>>
> >>> Functionally the change looks okay to me, but the above will need
> >>> sorting, at the very least by specifically discussing why effectively
> >>> undoing that earlier change is okay.
> >>
> >> So looking back at the thread, Jürgen measured a 14% improvement for
> >> something that might be described as a microbenchmark before and after
> >> the patch (a benchmark purposely trying to set up an unusual scenario
> >> to maximize the effect of context switch overhead, not one to
> >> represent a typical workflow). But are the numbers really plausible?
> >> Even at an implausible 100k switches/s across the box, saving 100
> >> cycles per switch is about 0.04% of eight 3 GHz cores.
> >>
> >> At any rate, we're already adding several map/unmap operations, and
> >> about to add several more. Keeping the PTE caching would require
> >> adding a separate path that can write just PTEs, which then will
> >> potentially further complication future paths where we need to make
> >> sure we handle both domain-wide perdomain areas and per-vcpu areas.
> >> If it were easy I would already have been keeping it.
> >>
> >> I'd be inclined to say: Since we're going to be adding more
> >> populate_perdomain_mapping() calls anyway, let's do it the simple
> >> correct way first; and then explore the idea of stashing mfns of
> >> frequently-mapped L1s (rather than having to walk L3 -> L2 -> L1); and
> >> at that time look into stashing baked l1es to avoid conversions.
> >
> > Perhaps; I'd like to have Jürgen's and/or Andrew's input here, though.
>
> At that time I implemented core scheduling in Xen. I noticed that very
> subtle changes in the context switch path could result in unexpected large
> performance differences. As I had the performance test for my purpose
> already set up, I used it for Andrew's patch (which was a result of my
> context switch path performance findings) and really did measure the
> impressive effect of it.
>
> Note that you can't only count instructions, often cache effects and
> branch predictions are dominating the performance.
Right, but:
1. That's going to be very much hardware- and workload- dependent.
Even on the same hardware, if you'd made a slight change in the
workload, you might have seen a very different result; and on
different hardware you're going to see something different again
2. As I said, we're now adding two extra map / unmaps, which is going
to perturb everything again.
If anything, your argument says we should wait until we've stopped
modifying the context switch path (which won't happen until patch 49
at least, guessing from the patch titles), and then measure things
again to see what's actually slow.
I'm sorry Jan, your position here is really inconsistent: You wave
away a partial pagetable walk with three map/unmap operations on the
context switch path as something we'll have to do in the interim, and
can optimize later, but are now threatening to make me add in
special-case codepaths and run tests to save a few memory reads and
shifts.
I could put back the mfn caching that was present in v1 of the series
(which Andy said was probably not sufficient, on balance, to make the
duplication involved worth it). Even that I think isn't really
sensible, but it's not too difficult to do. To isolate the PTE
caching effect I'd have to write an entire duplicate codepath anyway,
and then try to duplicate Jürgen's test. I don't think that's really
a reasonable ask at this point in the series.
-George
^ permalink raw reply [flat|nested] 44+ messages in thread* Re: [PATCH v2 03/14] x86/pv: use populate_perdomain_mapping() to map the Xen GDT
2026-09-04 8:06 ` George Dunlap
@ 2026-09-04 8:29 ` Jan Beulich
2026-09-04 8:50 ` George Dunlap
0 siblings, 1 reply; 44+ messages in thread
From: Jan Beulich @ 2026-09-04 8:29 UTC (permalink / raw)
To: George Dunlap
Cc: Andrew Cooper, Roger Pau Monné, Alejandro Vallejo,
Teddy Astie, Anthony PERARD, Michal Orzel, Julien Grall,
Stefano Stabellini, xen-devel, Jürgen Groß
On 04.09.2026 10:06, George Dunlap wrote:
> On Fri, Sep 4, 2026 at 7:54 AM Jürgen Groß <jgross@suse.com> wrote:
>>
>> On 04.09.26 08:00, Jan Beulich wrote:
>>> On 04.09.2026 00:35, George Dunlap wrote:
>>>> On Thu, Sep 3, 2026 at 5:11 PM Jan Beulich <jbeulich@suse.com> wrote:
>>>>>
>>>>> On 02.09.2026 11:43, George Dunlap wrote:
>>>>>> From: Roger Pau Monné <roger.pau@citrix.com>
>>>>>>
>>>>>> Currently, update_xen_slot_in_full_gdt() uses the stashed direct-map
>>>>>> pointer in d->arch.pv.gdt_ldt_l1tab to update the incoming vcpu's
>>>>>> page tables with Xen's GDT, by writing a stashed per-cpu copy of a
>>>>>> pre-baked L1 entry (either 64-bit or compat version).
>>>>>>
>>>>>> Switch this to using populate_perdomain_mapping(), which doesn't rely
>>>>>> on the stashed address of the l1 page in the direct map. Rather than
>>>>>> also stashing a pre-baked value for the payload, compute the mfn from
>>>>>> the per-cpu GDT pointer at use: the conversion is a handful of cycles
>>>>>> on a path costing thousands, and computing at use removes the
>>>>>> parallel {,compat_}gdt_l1e bookkeeping along with its boot-ordering
>>>>>> constraint (the cached value could only be generated after Xen's
>>>>>> physical relocation, and had to be in place before the first context
>>>>>> switch; a use-time lookup is correct by construction). The flags on
>>>>>> the final mapping are identical.
>>>>>>
>>>>>> Signed-off-by: Roger Pau Monné <roger.pau@citrix.com>
>>>>>> Assisted-by: Claude Code:claude-fable-5, Claude Code:claude-opus-4-8
>>>>>> Signed-off-by: George Dunlap <gwd@xenproject.org>
>>>>>> ---
>>>>>> Changes in v2:
>>>>>> - Drop the {,compat_}gdt_mfn caching entirely (suggested by Andrew
>>>>>> Cooper): compute virt_to_mfn() from the per-cpu GDT pointer at use.
>>>>>> The PDX lookup behind it measures ~5-10 cycles warm against a
>>>>>> ~1,500-cycle context switch, and this removes the double
>>>>>> bookkeeping and the after-relocation caching constraint. The
>>>>>> cached-MFN assertion goes with the cache: a use-time computation
>>>>>> from a live pointer needs no staleness check.
>>>>>
>>>>> This looks to contradict what 564d261687c0 ("x86/ctxt-switch: Document
>>>>> and improve GDT handling") used as justification to put in place the
>>>>> caching. Also Cc-ing Jürgen, who also was involved there, for possible
>>>>> further insight.
>>>>>
>>>>> Functionally the change looks okay to me, but the above will need
>>>>> sorting, at the very least by specifically discussing why effectively
>>>>> undoing that earlier change is okay.
>>>>
>>>> So looking back at the thread, Jürgen measured a 14% improvement for
>>>> something that might be described as a microbenchmark before and after
>>>> the patch (a benchmark purposely trying to set up an unusual scenario
>>>> to maximize the effect of context switch overhead, not one to
>>>> represent a typical workflow). But are the numbers really plausible?
>>>> Even at an implausible 100k switches/s across the box, saving 100
>>>> cycles per switch is about 0.04% of eight 3 GHz cores.
>>>>
>>>> At any rate, we're already adding several map/unmap operations, and
>>>> about to add several more. Keeping the PTE caching would require
>>>> adding a separate path that can write just PTEs, which then will
>>>> potentially further complication future paths where we need to make
>>>> sure we handle both domain-wide perdomain areas and per-vcpu areas.
>>>> If it were easy I would already have been keeping it.
>>>>
>>>> I'd be inclined to say: Since we're going to be adding more
>>>> populate_perdomain_mapping() calls anyway, let's do it the simple
>>>> correct way first; and then explore the idea of stashing mfns of
>>>> frequently-mapped L1s (rather than having to walk L3 -> L2 -> L1); and
>>>> at that time look into stashing baked l1es to avoid conversions.
>>>
>>> Perhaps; I'd like to have Jürgen's and/or Andrew's input here, though.
>>
>> At that time I implemented core scheduling in Xen. I noticed that very
>> subtle changes in the context switch path could result in unexpected large
>> performance differences. As I had the performance test for my purpose
>> already set up, I used it for Andrew's patch (which was a result of my
>> context switch path performance findings) and really did measure the
>> impressive effect of it.
>>
>> Note that you can't only count instructions, often cache effects and
>> branch predictions are dominating the performance.
>
> Right, but:
>
> 1. That's going to be very much hardware- and workload- dependent.
> Even on the same hardware, if you'd made a slight change in the
> workload, you might have seen a very different result; and on
> different hardware you're going to see something different again
>
> 2. As I said, we're now adding two extra map / unmaps, which is going
> to perturb everything again.
>
> If anything, your argument says we should wait until we've stopped
> modifying the context switch path (which won't happen until patch 49
> at least, guessing from the patch titles), and then measure things
> again to see what's actually slow.
>
> I'm sorry Jan,
You were replying to Jürgen, though.
> your position here is really inconsistent: You wave
> away a partial pagetable walk with three map/unmap operations on the
> context switch path as something we'll have to do in the interim, and
> can optimize later, but are now threatening to make me add in
> special-case codepaths and run tests to save a few memory reads and
> shifts.
I think you misunderstood. There was a concern raised already on v1,
and that concern wasn't covered by the patch description. In my initial
reply I said "Functionally the change looks okay to me" for a reason,
after all.
Jan
> I could put back the mfn caching that was present in v1 of the series
> (which Andy said was probably not sufficient, on balance, to make the
> duplication involved worth it). Even that I think isn't really
> sensible, but it's not too difficult to do. To isolate the PTE
> caching effect I'd have to write an entire duplicate codepath anyway,
> and then try to duplicate Jürgen's test. I don't think that's really
> a reasonable ask at this point in the series.
>
> -George
^ permalink raw reply [flat|nested] 44+ messages in thread* Re: [PATCH v2 03/14] x86/pv: use populate_perdomain_mapping() to map the Xen GDT
2026-09-04 8:29 ` Jan Beulich
@ 2026-09-04 8:50 ` George Dunlap
2026-09-04 10:11 ` Jan Beulich
2026-09-04 10:34 ` Roger Pau Monné
0 siblings, 2 replies; 44+ messages in thread
From: George Dunlap @ 2026-09-04 8:50 UTC (permalink / raw)
To: Jan Beulich
Cc: Andrew Cooper, Roger Pau Monné, Alejandro Vallejo,
Teddy Astie, Anthony PERARD, Michal Orzel, Julien Grall,
Stefano Stabellini, xen-devel, Jürgen Groß
On Fri, Sep 4, 2026 at 9:29 AM Jan Beulich <jbeulich@suse.com> wrote:
> > your position here is really inconsistent: You wave
> > away a partial pagetable walk with three map/unmap operations on the
> > context switch path as something we'll have to do in the interim, and
> > can optimize later, but are now threatening to make me add in
> > special-case codepaths and run tests to save a few memory reads and
> > shifts.
>
> I think you misunderstood. There was a concern raised already on v1,
> and that concern wasn't covered by the patch description. In my initial
> reply I said "Functionally the change looks okay to me" for a reason,
> after all.
To quote Andy's mail:
<<<
So what this patch is doing is still keeping the double copy (the
fragility) but reintroducing the expensive part of the operation into
the context switch path. If you can't keep it being L1e, there's
probably no point keeping the optimisation at all.
>>>
Basically what I took from this is;
- Andy thinks stashing any intermediate form (whether L1E or MFN) has
a technical cost (two copies that could potentially go out of sync,
thus "fragility")
- Andy thinks that the expensive part of the conversion is the MFN ->
L1E conversion, not the vaddr -> MFN conversion
- So, stashing the L1E might be a win, but stashing the MFN is unlikely to be.
- If we're not going to special-case this path, we have to pass an
MFN; and if we're going to pass an MFN, it's probably better to just
to get rid of the stashing; the extra fragility introduced doesn't pay
for itself in terms of potential performance improvement.
Note also that by the end of the series, we add two more
populate_perdomain_mapping() calls to the context switch path, at
least for ASI domains, which means another two of the "expensive" MFN
-> L1E conversions.
So v2 is doing what I understood Andy to have suggested. I agree the
meaning isn't 100% clear, though, so I may have misunderstood him.
As I've said, I'm not opposed to optimizing this path once we have the
final form functional and have measured it. Mapping the three tables
we need to modify in vmap, and stashing both the addresses and
pre-baked l1es, sounds like a perfectly reasonable thing to do,
*after* we get things functional and have had a chance to measure the
new context switch in its entirety.
-George
^ permalink raw reply [flat|nested] 44+ messages in thread
* Re: [PATCH v2 03/14] x86/pv: use populate_perdomain_mapping() to map the Xen GDT
2026-09-04 8:50 ` George Dunlap
@ 2026-09-04 10:11 ` Jan Beulich
2026-09-04 10:34 ` Roger Pau Monné
1 sibling, 0 replies; 44+ messages in thread
From: Jan Beulich @ 2026-09-04 10:11 UTC (permalink / raw)
To: George Dunlap
Cc: Andrew Cooper, Roger Pau Monné, Alejandro Vallejo,
Teddy Astie, Anthony PERARD, Michal Orzel, Julien Grall,
Stefano Stabellini, xen-devel, Jürgen Groß
On 04.09.2026 10:50, George Dunlap wrote:
> On Fri, Sep 4, 2026 at 9:29 AM Jan Beulich <jbeulich@suse.com> wrote:
>>> your position here is really inconsistent: You wave
>>> away a partial pagetable walk with three map/unmap operations on the
>>> context switch path as something we'll have to do in the interim, and
>>> can optimize later, but are now threatening to make me add in
>>> special-case codepaths and run tests to save a few memory reads and
>>> shifts.
>>
>> I think you misunderstood. There was a concern raised already on v1,
>> and that concern wasn't covered by the patch description. In my initial
>> reply I said "Functionally the change looks okay to me" for a reason,
>> after all.
>
> To quote Andy's mail:
>
> <<<
>
> So what this patch is doing is still keeping the double copy (the
> fragility) but reintroducing the expensive part of the operation into
> the context switch path. If you can't keep it being L1e, there's
> probably no point keeping the optimisation at all.
>
>>>>
>
> Basically what I took from this is;
>
> - Andy thinks stashing any intermediate form (whether L1E or MFN) has
> a technical cost (two copies that could potentially go out of sync,
> thus "fragility")
>
> - Andy thinks that the expensive part of the conversion is the MFN ->
> L1E conversion, not the vaddr -> MFN conversion
Iirc later, when discussing with me and Roger, this was somewhat adjusted.
Unfortunately the outcome of that discussion wasn't put in a reply there.
> - So, stashing the L1E might be a win, but stashing the MFN is unlikely to be.
>
> - If we're not going to special-case this path, we have to pass an
> MFN; and if we're going to pass an MFN, it's probably better to just
> to get rid of the stashing; the extra fragility introduced doesn't pay
> for itself in terms of potential performance improvement.
>
> Note also that by the end of the series, we add two more
> populate_perdomain_mapping() calls to the context switch path, at
> least for ASI domains, which means another two of the "expensive" MFN
> -> L1E conversions.
>
> So v2 is doing what I understood Andy to have suggested. I agree the
> meaning isn't 100% clear, though, so I may have misunderstood him.
>
> As I've said, I'm not opposed to optimizing this path once we have the
> final form functional and have measured it. Mapping the three tables
> we need to modify in vmap, and stashing both the addresses and
> pre-baked l1es, sounds like a perfectly reasonable thing to do,
> *after* we get things functional and have had a chance to measure the
> new context switch in its entirety.
And I (largely) agree. What I'm asking for (beyond feedback from those
who were involved in putting in the optimization) is that the removal
of that optimization be justified against the original commit's
reasoning.
Jan
^ permalink raw reply [flat|nested] 44+ messages in thread
* Re: [PATCH v2 03/14] x86/pv: use populate_perdomain_mapping() to map the Xen GDT
2026-09-04 8:50 ` George Dunlap
2026-09-04 10:11 ` Jan Beulich
@ 2026-09-04 10:34 ` Roger Pau Monné
2026-09-07 13:58 ` George Dunlap
1 sibling, 1 reply; 44+ messages in thread
From: Roger Pau Monné @ 2026-09-04 10:34 UTC (permalink / raw)
To: George Dunlap
Cc: Jan Beulich, Andrew Cooper, Alejandro Vallejo, Teddy Astie,
Anthony PERARD, Michal Orzel, Julien Grall, Stefano Stabellini,
xen-devel, Jürgen Groß
On Fri, Sep 04, 2026 at 09:50:02AM +0100, George Dunlap wrote:
> On Fri, Sep 4, 2026 at 9:29 AM Jan Beulich <jbeulich@suse.com> wrote:
> > > your position here is really inconsistent: You wave
> > > away a partial pagetable walk with three map/unmap operations on the
> > > context switch path as something we'll have to do in the interim, and
> > > can optimize later, but are now threatening to make me add in
> > > special-case codepaths and run tests to save a few memory reads and
> > > shifts.
> >
> > I think you misunderstood. There was a concern raised already on v1,
> > and that concern wasn't covered by the patch description. In my initial
> > reply I said "Functionally the change looks okay to me" for a reason,
> > after all.
>
> To quote Andy's mail:
>
> <<<
>
> So what this patch is doing is still keeping the double copy (the
> fragility) but reintroducing the expensive part of the operation into
> the context switch path. If you can't keep it being L1e, there's
> probably no point keeping the optimisation at all.
>
> >>>
>
> Basically what I took from this is;
>
> - Andy thinks stashing any intermediate form (whether L1E or MFN) has
> a technical cost (two copies that could potentially go out of sync,
> thus "fragility")
>
> - Andy thinks that the expensive part of the conversion is the MFN ->
> L1E conversion, not the vaddr -> MFN conversion
We later discussed this, and the assumption was that the expensive
part was the vaddr -> MFN translation, as that's where PDX is
involved. I expect crafting a PTE shouldn't be expensive at all, but
maybe there's something I'm missing here.
The original patch cached the MFN in an attempt to not remove the
optimization, because my understanding was that the possible expensive
part was the PDX translation. Then again I don't have real
measurements to back up any of the claims above, and hence it might
all be plain wrong.
Regards, Roger.
^ permalink raw reply [flat|nested] 44+ messages in thread
* Re: [PATCH v2 03/14] x86/pv: use populate_perdomain_mapping() to map the Xen GDT
2026-09-04 10:34 ` Roger Pau Monné
@ 2026-09-07 13:58 ` George Dunlap
0 siblings, 0 replies; 44+ messages in thread
From: George Dunlap @ 2026-09-07 13:58 UTC (permalink / raw)
To: Roger Pau Monné
Cc: Jan Beulich, Andrew Cooper, Alejandro Vallejo, Teddy Astie,
Anthony PERARD, Michal Orzel, Julien Grall, Stefano Stabellini,
xen-devel, Jürgen Groß
On Fri, Sep 4, 2026 at 11:34 AM Roger Pau Monné <roger@xenproject.org> wrote:
>
> On Fri, Sep 04, 2026 at 09:50:02AM +0100, George Dunlap wrote:
> > On Fri, Sep 4, 2026 at 9:29 AM Jan Beulich <jbeulich@suse.com> wrote:
> > > > your position here is really inconsistent: You wave
> > > > away a partial pagetable walk with three map/unmap operations on the
> > > > context switch path as something we'll have to do in the interim, and
> > > > can optimize later, but are now threatening to make me add in
> > > > special-case codepaths and run tests to save a few memory reads and
> > > > shifts.
> > >
> > > I think you misunderstood. There was a concern raised already on v1,
> > > and that concern wasn't covered by the patch description. In my initial
> > > reply I said "Functionally the change looks okay to me" for a reason,
> > > after all.
> >
> > To quote Andy's mail:
> >
> > <<<
> >
> > So what this patch is doing is still keeping the double copy (the
> > fragility) but reintroducing the expensive part of the operation into
> > the context switch path. If you can't keep it being L1e, there's
> > probably no point keeping the optimisation at all.
> >
> > >>>
> >
> > Basically what I took from this is;
> >
> > - Andy thinks stashing any intermediate form (whether L1E or MFN) has
> > a technical cost (two copies that could potentially go out of sync,
> > thus "fragility")
> >
> > - Andy thinks that the expensive part of the conversion is the MFN ->
> > L1E conversion, not the vaddr -> MFN conversion
>
> We later discussed this, and the assumption was that the expensive
> part was the vaddr -> MFN translation, as that's where PDX is
> involved. I expect crafting a PTE shouldn't be expensive at all, but
> maybe there's something I'm missing here.
>
> The original patch cached the MFN in an attempt to not remove the
> optimization, because my understanding was that the possible expensive
> part was the PDX translation. Then again I don't have real
> measurements to back up any of the claims above, and hence it might
> all be plain wrong.
I ran some tests on my NUC. Measuring cycles for *just*
update_xen_slot_in_full_gdt(), using the previous version (stashed
xenheap pointer + stashed l1e), and then mapcache walk with {stashed
mfn, full conversion} x {default, PDX forced on}.
xenheap+l1e: 38 cycles
mapcache walk, stashed mfn (v1), no PDX: 98 cycles
mapcache walk, full convervion, no PDX: 97 cycles
mapcache walk, stashed mfn (v1), PDX forced on: 128 cycles
mapcache walk, full conversion: 122 cycles
The variance is pretty high, so basically the mfn stashing didn't have
any statistically significant effect.
I'll leave it doing the full conversion for now.
-George
^ permalink raw reply [flat|nested] 44+ messages in thread
* [PATCH v2 04/14] x86/pv: set/clear guest GDT mappings using populate_perdomain_mapping()
2026-09-02 9:43 [PATCH v2 00/14] x86: Address Space Isolation, part 2: asi= option and per-vCPU page tables George Dunlap
` (2 preceding siblings ...)
2026-09-02 9:43 ` [PATCH v2 03/14] x86/pv: use populate_perdomain_mapping() to map the Xen GDT George Dunlap
@ 2026-09-02 9:43 ` George Dunlap
2026-09-07 12:50 ` Jan Beulich
2026-09-02 9:43 ` [PATCH v2 05/14] x86/pv: update guest LDT mappings using {populate,destroy}_perdomain_mapping() George Dunlap
` (9 subsequent siblings)
13 siblings, 1 reply; 44+ messages in thread
From: George Dunlap @ 2026-09-02 9:43 UTC (permalink / raw)
To: xen-devel
Cc: Roger Pau Monné, Jan Beulich, Andrew Cooper,
Roger Pau Monné, Alejandro Vallejo, Teddy Astie,
Anthony PERARD, Michal Orzel, Julien Grall, Stefano Stabellini,
George Dunlap
From: Roger Pau Monné <roger.pau@citrix.com>
Until the previous patch, update_xen_slot_in_full_gdt() used the
stashed pointer in d->arch.pv.gdt_ldt_l1tab to update the incoming
vCPU's page tables with Xen's GDT; this was previously necessary
because map_domain_page() couldn't be called in a context switch.
Having a handy pointer to an always-mapped version of the GDT/LDT L1
table, other sites which modify the table started using it for
convenience, even if they weren't called from within a context switch.
One example is pv_{set,destroy}_gdt().
The previous patch switched the main user of the stashed reference to
use populate_perdomain_mapping() instead. Continue that process by
switching both pv_{set,destroy}_gdt() to it as well.
pv_destroy_gdt() currently loops over the L1 entries directly,
extracting the MFN from each, dropping the type and reference unless
it was the zero page, and replacing the entry with a read-only mapping
of the zero page. Rather than reading from the stashed L1, drop the
references using v->arch.pv.gdt_frames[] instead, and install the
zero-page mappings with a single populate_perdomain_mapping() call.
This makes gdt_frames[] consistently the source of truth for MFNs.
Note that we must maintain the invariant introduced in cf6d39f819
("x86/PV: properly populate descriptor tables"): pv_destroy_gdt() maps
the zero page read-only in torn-down slots rather than unmapping them,
so that LAR/LSL/VERR/VERW on a selector beyond the guest's limit clear
ZF as on native rather than taking a #PF-converted #GP. (And since
pv_set_gdt() tears down the old GDT before installing the new one,
guests never see unmapped entries, only zero-page entries.)
In the case of pv_set_gdt(), we have a slightly awkward situation with
types. The ABI with the guest uses unsigned long[], but
populate_perdomain_mapping() wants an array of mfn_t.
v->arch.pv.gdt_frames being unsigned long means we can just copy from
it across the guest ABI with no conversions. We could in theory
convert it to mfn_t[] instead, and then pass v->arch.pv.gdt_frames
into populate_perdomain_mapping(); but then we'd need to add a
conversion on all the places where frames are copied out. We choose
instead to copy frames into a temporary mfn_t array on the stack to
pass into populate_perdomain_mapping().
Signed-off-by: Roger Pau Monné <roger.pau@citrix.com>
Assisted-by: Claude Code:claude-fable-5, Claude Code:claude-opus-4-8
Signed-off-by: George Dunlap <gwd@xenproject.org>
---
Changes in v2:
- Reword commit message
Changes since the previously posted version:
- Retain the gdt_ents zeroing when tearing down the GDT (its removal
was queried by Jan).
- Map torn-down slots read-only to the zero page (via the
populate_perdomain_mapping() flags parameter) rather than removing
the mappings with destroy_perdomain_mapping(): empty slots would be
a guest-visible partial revert of cf6d39f819 (see the commit
message). With the destroy call gone, its v->arch.cr3 guard --
also queried by Jan -- goes too: the zero-page rewrite runs
unconditionally.
- Keep gdt_frames[] as unsigned long[] rather than switching it to
mfn_t[] as Jan suggested; the commit message explains the
trade-off.
- Retitle: destroy_perdomain_mapping() is no longer used here.
---
xen/arch/x86/pv/descriptor-tables.c | 37 ++++++++++++++++++-----------
1 file changed, 23 insertions(+), 14 deletions(-)
diff --git a/xen/arch/x86/pv/descriptor-tables.c b/xen/arch/x86/pv/descriptor-tables.c
index 8a32b9ae5c..5dda5bffe3 100644
--- a/xen/arch/x86/pv/descriptor-tables.c
+++ b/xen/arch/x86/pv/descriptor-tables.c
@@ -49,33 +49,42 @@ bool pv_destroy_ldt(struct vcpu *v)
void pv_destroy_gdt(struct vcpu *v)
{
- l1_pgentry_t *pl1e = pv_gdt_ptes(v);
- mfn_t zero_mfn = _mfn(virt_to_mfn(zero_page));
- l1_pgentry_t zero_l1e = l1e_from_mfn(zero_mfn, __PAGE_HYPERVISOR_RO);
+ const mfn_t zero_mfn = _mfn(virt_to_mfn(zero_page));
+ mfn_t zero_mfns[ARRAY_SIZE(v->arch.pv.gdt_frames)];
unsigned int i;
ASSERT(v == current || !vcpu_cpu_dirty(v));
v->arch.pv.gdt_ents = 0;
- for ( i = 0; i < FIRST_RESERVED_GDT_PAGE; i++ )
+
+ for ( i = 0; i < ARRAY_SIZE(zero_mfns); i++ )
{
- mfn_t mfn = l1e_get_mfn(pl1e[i]);
+ zero_mfns[i] = zero_mfn;
- if ( (l1e_get_flags(pl1e[i]) & _PAGE_PRESENT) &&
- !mfn_eq(mfn, zero_mfn) )
- put_page_and_type(mfn_to_page(mfn));
+ /* MFN 0 can never pass get_page_and_type(), so 0 marks unused slots. */
+ if ( !v->arch.pv.gdt_frames[i] )
+ continue;
- l1e_write(&pl1e[i], zero_l1e);
+ put_page_and_type(mfn_to_page(_mfn(v->arch.pv.gdt_frames[i])));
v->arch.pv.gdt_frames[i] = 0;
}
+
+ /*
+ * Point every slot at the zero page, read-only: a descriptor fetch from
+ * the unused part of the GDT then finds a not-present descriptor rather
+ * than a missing mapping, so LAR/LSL/VERR/VERW on a selector beyond the
+ * guest's limit clear ZF as they do on native, instead of faulting.
+ */
+ populate_perdomain_mapping(v, GDT_VIRT_START(v), zero_mfns,
+ ARRAY_SIZE(zero_mfns), __PAGE_HYPERVISOR_RO);
}
int pv_set_gdt(struct vcpu *v, const unsigned long frames[],
unsigned int entries)
{
struct domain *d = v->domain;
- l1_pgentry_t *pl1e;
unsigned int i, nr_frames = DIV_ROUND_UP(entries, 512);
+ mfn_t mfns[ARRAY_SIZE(v->arch.pv.gdt_frames)];
ASSERT(v == current || !vcpu_cpu_dirty(v));
@@ -90,6 +99,8 @@ int pv_set_gdt(struct vcpu *v, const unsigned long frames[],
if ( !mfn_valid(mfn) ||
!get_page_and_type(mfn_to_page(mfn), d, PGT_seg_desc_page) )
goto fail;
+
+ mfns[i] = mfn;
}
/* Tear down the old GDT. */
@@ -97,12 +108,10 @@ int pv_set_gdt(struct vcpu *v, const unsigned long frames[],
/* Install the new GDT. */
v->arch.pv.gdt_ents = entries;
- pl1e = pv_gdt_ptes(v);
for ( i = 0; i < nr_frames; i++ )
- {
v->arch.pv.gdt_frames[i] = frames[i];
- l1e_write(&pl1e[i], l1e_from_pfn(frames[i], __PAGE_HYPERVISOR_RW));
- }
+ populate_perdomain_mapping(v, GDT_VIRT_START(v), mfns, nr_frames,
+ __PAGE_HYPERVISOR_RW);
return 0;
--
2.55.0
^ permalink raw reply related [flat|nested] 44+ messages in thread* Re: [PATCH v2 04/14] x86/pv: set/clear guest GDT mappings using populate_perdomain_mapping()
2026-09-02 9:43 ` [PATCH v2 04/14] x86/pv: set/clear guest GDT mappings using populate_perdomain_mapping() George Dunlap
@ 2026-09-07 12:50 ` Jan Beulich
2026-09-07 13:51 ` George Dunlap
0 siblings, 1 reply; 44+ messages in thread
From: Jan Beulich @ 2026-09-07 12:50 UTC (permalink / raw)
To: George Dunlap
Cc: Andrew Cooper, Roger Pau Monné, Alejandro Vallejo,
Teddy Astie, Anthony PERARD, Michal Orzel, Julien Grall,
Stefano Stabellini, George Dunlap, xen-devel
On 02.09.2026 11:43, George Dunlap wrote:
> From: Roger Pau Monné <roger.pau@citrix.com>
>
> Until the previous patch, update_xen_slot_in_full_gdt() used the
Please can we avoid "previous patch" (also again below) and alike in
commit messages?
> stashed pointer in d->arch.pv.gdt_ldt_l1tab to update the incoming
> vCPU's page tables with Xen's GDT; this was previously necessary
> because map_domain_page() couldn't be called in a context switch.
> Having a handy pointer to an always-mapped version of the GDT/LDT L1
> table, other sites which modify the table started using it for
> convenience, even if they weren't called from within a context switch.
> One example is pv_{set,destroy}_gdt().
>
> The previous patch switched the main user of the stashed reference to
> use populate_perdomain_mapping() instead. Continue that process by
> switching both pv_{set,destroy}_gdt() to it as well.
>
> pv_destroy_gdt() currently loops over the L1 entries directly,
> extracting the MFN from each, dropping the type and reference unless
> it was the zero page, and replacing the entry with a read-only mapping
> of the zero page. Rather than reading from the stashed L1, drop the
> references using v->arch.pv.gdt_frames[] instead, and install the
> zero-page mappings with a single populate_perdomain_mapping() call.
> This makes gdt_frames[] consistently the source of truth for MFNs.
>
> Note that we must maintain the invariant introduced in cf6d39f819
> ("x86/PV: properly populate descriptor tables"): pv_destroy_gdt() maps
> the zero page read-only in torn-down slots rather than unmapping them,
> so that LAR/LSL/VERR/VERW on a selector beyond the guest's limit clear
> ZF as on native rather than taking a #PF-converted #GP. (And since
> pv_set_gdt() tears down the old GDT before installing the new one,
> guests never see unmapped entries, only zero-page entries.)
>
> In the case of pv_set_gdt(), we have a slightly awkward situation with
> types. The ABI with the guest uses unsigned long[], but
> populate_perdomain_mapping() wants an array of mfn_t.
> v->arch.pv.gdt_frames being unsigned long means we can just copy from
> it across the guest ABI with no conversions. We could in theory
> convert it to mfn_t[] instead, and then pass v->arch.pv.gdt_frames
> into populate_perdomain_mapping(); but then we'd need to add a
> conversion on all the places where frames are copied out. We choose
> instead to copy frames into a temporary mfn_t array on the stack to
> pass into populate_perdomain_mapping().
>
> Signed-off-by: Roger Pau Monné <roger.pau@citrix.com>
> Assisted-by: Claude Code:claude-fable-5, Claude Code:claude-opus-4-8
> Signed-off-by: George Dunlap <gwd@xenproject.org>
With the adjustment above:
Reviewed-by: Jan Beulich <jbeulich@suse.com>
> --- a/xen/arch/x86/pv/descriptor-tables.c
> +++ b/xen/arch/x86/pv/descriptor-tables.c
> @@ -49,33 +49,42 @@ bool pv_destroy_ldt(struct vcpu *v)
>
> void pv_destroy_gdt(struct vcpu *v)
> {
> - l1_pgentry_t *pl1e = pv_gdt_ptes(v);
> - mfn_t zero_mfn = _mfn(virt_to_mfn(zero_page));
> - l1_pgentry_t zero_l1e = l1e_from_mfn(zero_mfn, __PAGE_HYPERVISOR_RO);
> + const mfn_t zero_mfn = _mfn(virt_to_mfn(zero_page));
> + mfn_t zero_mfns[ARRAY_SIZE(v->arch.pv.gdt_frames)];
I wonder if having an initializer here might not result in better code. The
array could likely be filled with REP STOSQ here, and ...
> unsigned int i;
>
> ASSERT(v == current || !vcpu_cpu_dirty(v));
>
> v->arch.pv.gdt_ents = 0;
> - for ( i = 0; i < FIRST_RESERVED_GDT_PAGE; i++ )
> +
> + for ( i = 0; i < ARRAY_SIZE(zero_mfns); i++ )
> {
> - mfn_t mfn = l1e_get_mfn(pl1e[i]);
> + zero_mfns[i] = zero_mfn;
... the live range of zero_mfn would reduce to about nothing.
Jan
^ permalink raw reply [flat|nested] 44+ messages in thread* Re: [PATCH v2 04/14] x86/pv: set/clear guest GDT mappings using populate_perdomain_mapping()
2026-09-07 12:50 ` Jan Beulich
@ 2026-09-07 13:51 ` George Dunlap
2026-09-07 14:57 ` Jan Beulich
0 siblings, 1 reply; 44+ messages in thread
From: George Dunlap @ 2026-09-07 13:51 UTC (permalink / raw)
To: Jan Beulich
Cc: Andrew Cooper, Roger Pau Monné, Alejandro Vallejo,
Teddy Astie, Anthony PERARD, Michal Orzel, Julien Grall,
Stefano Stabellini, xen-devel
On Mon, Sep 7, 2026 at 1:51 PM Jan Beulich <jbeulich@suse.com> wrote:
>
> On 02.09.2026 11:43, George Dunlap wrote:
> > From: Roger Pau Monné <roger.pau@citrix.com>
> >
> > Until the previous patch, update_xen_slot_in_full_gdt() used the
>
> Please can we avoid "previous patch" (also again below) and alike in
> commit messages?
Something like this then?
8<---
update_xen_slot_in_full_gdt() used to update the incoming vCPU's page
tables with Xen's GDT through the stashed pointer in
d->arch.pv.gdt_ldt_l1tab, because map_domain_page() couldn't be called
in a context switch. Having a handy pointer to an always-mapped
version of the GDT/LDT L1 table, other sites which modify the table
started using it for convenience, even though they aren't called from
within a context switch. One example is pv{set,destroy}gdt().
With update_xen_slot_in_full_gdt() now using
populate_perdomain_mapping() (see "x86/pv: use
populate_perdomain_mapping() to map the Xen GDT"), the stashed
reference has lost the user that justified it. Switch
pv{set,destroy}gdt() to populate_perdomain_mapping() as well; the LDT
paths are the remaining users, after which the stash can go.
--->8
> > Signed-off-by: Roger Pau Monné <roger.pau@citrix.com>
> > Assisted-by: Claude Code:claude-fable-5, Claude Code:claude-opus-4-8
> > Signed-off-by: George Dunlap <gwd@xenproject.org>
>
> With the adjustment above:
> Reviewed-by: Jan Beulich <jbeulich@suse.com>
Thanks!
> > --- a/xen/arch/x86/pv/descriptor-tables.c
> > +++ b/xen/arch/x86/pv/descriptor-tables.c
> > @@ -49,33 +49,42 @@ bool pv_destroy_ldt(struct vcpu *v)
> >
> > void pv_destroy_gdt(struct vcpu *v)
> > {
> > - l1_pgentry_t *pl1e = pv_gdt_ptes(v);
> > - mfn_t zero_mfn = _mfn(virt_to_mfn(zero_page));
> > - l1_pgentry_t zero_l1e = l1e_from_mfn(zero_mfn, __PAGE_HYPERVISOR_RO);
> > + const mfn_t zero_mfn = _mfn(virt_to_mfn(zero_page));
> > + mfn_t zero_mfns[ARRAY_SIZE(v->arch.pv.gdt_frames)];
>
> I wonder if having an initializer here might not result in better code. The
> array could likely be filled with REP STOSQ here, and ...
>
> > unsigned int i;
> >
> > ASSERT(v == current || !vcpu_cpu_dirty(v));
> >
> > v->arch.pv.gdt_ents = 0;
> > - for ( i = 0; i < FIRST_RESERVED_GDT_PAGE; i++ )
> > +
> > + for ( i = 0; i < ARRAY_SIZE(zero_mfns); i++ )
> > {
> > - mfn_t mfn = l1e_get_mfn(pl1e[i]);
> > + zero_mfns[i] = zero_mfn;
>
> ... the live range of zero_mfn would reduce to about nothing.
Tried it: GCC didn't use REP STOSQ [1], but seems worth it for clarity
and register efficiency anyway.
-George
[1] Fable experimented with a few different compiler flags, but
couldn't get a REP operation; it said "The reason is structural:
GCC's string-op machinery handles block clears and byte-splat fills
(memset means a repeated byte); a fill with a non-constant 8-byte
value is not a memset and never reaches that expander".
^ permalink raw reply [flat|nested] 44+ messages in thread* Re: [PATCH v2 04/14] x86/pv: set/clear guest GDT mappings using populate_perdomain_mapping()
2026-09-07 13:51 ` George Dunlap
@ 2026-09-07 14:57 ` Jan Beulich
0 siblings, 0 replies; 44+ messages in thread
From: Jan Beulich @ 2026-09-07 14:57 UTC (permalink / raw)
To: George Dunlap
Cc: Andrew Cooper, Roger Pau Monné, Alejandro Vallejo,
Teddy Astie, Anthony PERARD, Michal Orzel, Julien Grall,
Stefano Stabellini, xen-devel
On 07.09.2026 15:51, George Dunlap wrote:
> On Mon, Sep 7, 2026 at 1:51 PM Jan Beulich <jbeulich@suse.com> wrote:
>>
>> On 02.09.2026 11:43, George Dunlap wrote:
>>> From: Roger Pau Monné <roger.pau@citrix.com>
>>>
>>> Until the previous patch, update_xen_slot_in_full_gdt() used the
>>
>> Please can we avoid "previous patch" (also again below) and alike in
>> commit messages?
>
> Something like this then?
>
> 8<---
> update_xen_slot_in_full_gdt() used to update the incoming vCPU's page
> tables with Xen's GDT through the stashed pointer in
> d->arch.pv.gdt_ldt_l1tab, because map_domain_page() couldn't be called
> in a context switch. Having a handy pointer to an always-mapped
> version of the GDT/LDT L1 table, other sites which modify the table
> started using it for convenience, even though they aren't called from
> within a context switch. One example is pv{set,destroy}gdt().
>
> With update_xen_slot_in_full_gdt() now using
> populate_perdomain_mapping() (see "x86/pv: use
> populate_perdomain_mapping() to map the Xen GDT"), the stashed
> reference has lost the user that justified it. Switch
> pv{set,destroy}gdt() to populate_perdomain_mapping() as well; the LDT
> paths are the remaining users, after which the stash can go.
> --->8
Sgtm, thanks.
>>> --- a/xen/arch/x86/pv/descriptor-tables.c
>>> +++ b/xen/arch/x86/pv/descriptor-tables.c
>>> @@ -49,33 +49,42 @@ bool pv_destroy_ldt(struct vcpu *v)
>>>
>>> void pv_destroy_gdt(struct vcpu *v)
>>> {
>>> - l1_pgentry_t *pl1e = pv_gdt_ptes(v);
>>> - mfn_t zero_mfn = _mfn(virt_to_mfn(zero_page));
>>> - l1_pgentry_t zero_l1e = l1e_from_mfn(zero_mfn, __PAGE_HYPERVISOR_RO);
>>> + const mfn_t zero_mfn = _mfn(virt_to_mfn(zero_page));
>>> + mfn_t zero_mfns[ARRAY_SIZE(v->arch.pv.gdt_frames)];
>>
>> I wonder if having an initializer here might not result in better code. The
>> array could likely be filled with REP STOSQ here, and ...
>>
>>> unsigned int i;
>>>
>>> ASSERT(v == current || !vcpu_cpu_dirty(v));
>>>
>>> v->arch.pv.gdt_ents = 0;
>>> - for ( i = 0; i < FIRST_RESERVED_GDT_PAGE; i++ )
>>> +
>>> + for ( i = 0; i < ARRAY_SIZE(zero_mfns); i++ )
>>> {
>>> - mfn_t mfn = l1e_get_mfn(pl1e[i]);
>>> + zero_mfns[i] = zero_mfn;
>>
>> ... the live range of zero_mfn would reduce to about nothing.
>
> Tried it: GCC didn't use REP STOSQ [1], but seems worth it for clarity
> and register efficiency anyway.
Interesting - room for improvement there then.
Jan
> [1] Fable experimented with a few different compiler flags, but
> couldn't get a REP operation; it said "The reason is structural:
> GCC's string-op machinery handles block clears and byte-splat fills
> (memset means a repeated byte); a fill with a non-constant 8-byte
> value is not a memset and never reaches that expander".
^ permalink raw reply [flat|nested] 44+ messages in thread
* [PATCH v2 05/14] x86/pv: update guest LDT mappings using {populate,destroy}_perdomain_mapping()
2026-09-02 9:43 [PATCH v2 00/14] x86: Address Space Isolation, part 2: asi= option and per-vCPU page tables George Dunlap
` (3 preceding siblings ...)
2026-09-02 9:43 ` [PATCH v2 04/14] x86/pv: set/clear guest GDT mappings using populate_perdomain_mapping() George Dunlap
@ 2026-09-02 9:43 ` George Dunlap
2026-09-07 16:06 ` Jan Beulich
2026-09-02 9:43 ` [PATCH v2 06/14] x86/pv: remove stashing of GDT/LDT L1 page-tables George Dunlap
` (8 subsequent siblings)
13 siblings, 1 reply; 44+ messages in thread
From: George Dunlap @ 2026-09-02 9:43 UTC (permalink / raw)
To: xen-devel
Cc: Roger Pau Monné, Jan Beulich, Andrew Cooper,
Roger Pau Monné, Alejandro Vallejo, Teddy Astie,
Anthony PERARD, Michal Orzel, Julien Grall, Stefano Stabellini,
George Dunlap
From: Roger Pau Monné <roger.pau@citrix.com>
Until two patches ago, update_xen_slot_in_full_gdt() used the stashed
pointer in d->arch.pv.gdt_ldt_l1tab to update the incoming vCPU's page
tables with Xen's GDT; this was previously necessary because
map_domain_page() couldn't be called in a context switch. Having a
handy pointer to an always-mapped version of the GDT/LDT L1 table,
other sites which modify the table started using it for convenience,
even if they weren't called from within a context switch. These
include pv_map_ldt_shadow_page() and pv_destroy_ldt().
Continue the process of switching users of the stashed reference to use
populate_perdomain_mapping() instead.
pv_map_ldt_shadow_page() is, by definition, always modifying the
currently-running vCPU: it runs from the #PF handler for a descriptor
fetch on the guest's behalf, and running the guest implies its page
tables are loaded. So it could simply write the linear recursive
mappings directly. Go through populate_perdomain_mapping() anyway, to
keep a single writer for the per-domain area.
For pv_destroy_ldt(), use destroy_perdomain_mapping().
Previously, pv_destroy_ldt() used the L1 LDT entries themselves to
determine which MFNs to drop type and count references to. Rather
than reading from the stashed L1, keep the MFNs corresponding to L1
slots in an array in the vCPU structure, as we do in the GDT case.
(Note that unlike the GDT case, these are not part of a public ABI, so
can be mfn_t, avoiding a recast-and-copy.)
Note that mappings_dropped (the return value of pv_destroy_ldt()) now
reflects the *number of valid MFNs in this array*, not *the number of
non-empty L1 entries*. This introduces an invariant we must maintain:
pv_map_ldt_shadow_page() writes both the array entry and the mapping,
and pv_destroy_ldt() clears both, so the two stay in lockstep.
Also note that, unlike pv_destroy_gdt() from the previous patch,
pv_destroy_ldt() doesn't fill in the values with zero_l1e (see
61031e64d3), so there's no change here.
Signed-off-by: Roger Pau Monné <roger.pau@citrix.com>
Assisted-by: Claude Code:claude-fable-5, Claude Code:claude-opus-4-8
Signed-off-by: George Dunlap <gwd@xenproject.org>
---
Changes in v2:
- Reword the commit message
Changes since the previously posted version:
- Initialise ldt_frames[] ahead of the first fail-able initialisation step.
- Call destroy_perdomain_mapping() with its existing domain parameter;
the switch to a vCPU parameter moves to a future patch.
- Use populate_perdomain_mapping() in pv_map_ldt_shadow_page() rather
than open-coding the linear-map write; retitle accordingly.
- Comment the INVALID_MFN skip in pv_destroy_ldt(): the LDT is
demand-faulted, so its pages may be sparsely mapped (Alejandro's
question on v2).
- Rewrite commit message (including describing the ldt_frames[]
array's role directly, as Jan asked).
---
xen/arch/x86/include/asm/domain.h | 2 ++
xen/arch/x86/pv/descriptor-tables.c | 20 +++++++++++---------
xen/arch/x86/pv/domain.c | 4 ++++
xen/arch/x86/pv/mm.c | 16 ++++++++++++----
4 files changed, 29 insertions(+), 13 deletions(-)
diff --git a/xen/arch/x86/include/asm/domain.h b/xen/arch/x86/include/asm/domain.h
index 2d0a915410..61a9fe00f0 100644
--- a/xen/arch/x86/include/asm/domain.h
+++ b/xen/arch/x86/include/asm/domain.h
@@ -541,6 +541,8 @@ struct pv_vcpu
struct trap_info *trap_ctxt;
unsigned long gdt_frames[FIRST_RESERVED_GDT_PAGE];
+ /* Max LDT entries is 8192, so 8192 * 8 = 64KiB (16 pages). */
+ mfn_t ldt_frames[16];
unsigned long ldt_base;
unsigned int gdt_ents, ldt_ents;
diff --git a/xen/arch/x86/pv/descriptor-tables.c b/xen/arch/x86/pv/descriptor-tables.c
index 5dda5bffe3..261bf29c90 100644
--- a/xen/arch/x86/pv/descriptor-tables.c
+++ b/xen/arch/x86/pv/descriptor-tables.c
@@ -20,28 +20,30 @@
*/
bool pv_destroy_ldt(struct vcpu *v)
{
- l1_pgentry_t *pl1e;
+ const unsigned int nr_frames = ARRAY_SIZE(v->arch.pv.ldt_frames);
unsigned int i, mappings_dropped = 0;
- struct page_info *page;
ASSERT(!in_irq());
ASSERT(v == current || !vcpu_cpu_dirty(v));
- pl1e = pv_ldt_ptes(v);
+ destroy_perdomain_mapping(v->domain, LDT_VIRT_START(v), nr_frames);
- for ( i = 0; i < 16; i++ )
+ for ( i = 0; i < nr_frames; i++ )
{
- if ( !(l1e_get_flags(pl1e[i]) & _PAGE_PRESENT) )
- continue;
+ mfn_t mfn = v->arch.pv.ldt_frames[i];
+ struct page_info *page;
- page = l1e_get_page(pl1e[i]);
- l1e_write(&pl1e[i], l1e_empty());
- mappings_dropped++;
+ /* The LDT is demand-faulted, so its pages may be sparsely mapped. */
+ if ( mfn_eq(mfn, INVALID_MFN) )
+ continue;
+ v->arch.pv.ldt_frames[i] = INVALID_MFN;
+ page = mfn_to_page(mfn);
ASSERT_PAGE_IS_TYPE(page, PGT_seg_desc_page);
ASSERT_PAGE_IS_DOMAIN(page, v->domain);
put_page_and_type(page);
+ mappings_dropped++;
}
return mappings_dropped;
diff --git a/xen/arch/x86/pv/domain.c b/xen/arch/x86/pv/domain.c
index 0c42ae58aa..7ddab1949f 100644
--- a/xen/arch/x86/pv/domain.c
+++ b/xen/arch/x86/pv/domain.c
@@ -340,10 +340,14 @@ void pv_vcpu_destroy(struct vcpu *v)
int pv_vcpu_initialise(struct vcpu *v)
{
struct domain *d = v->domain;
+ unsigned int i;
int rc;
ASSERT(!is_idle_domain(d));
+ for ( i = 0; i < ARRAY_SIZE(v->arch.pv.ldt_frames); i++ )
+ v->arch.pv.ldt_frames[i] = INVALID_MFN;
+
rc = pv_create_gdt_ldt_l1tab(v);
if ( rc )
return rc;
diff --git a/xen/arch/x86/pv/mm.c b/xen/arch/x86/pv/mm.c
index 5378299b8c..da280d7757 100644
--- a/xen/arch/x86/pv/mm.c
+++ b/xen/arch/x86/pv/mm.c
@@ -53,7 +53,8 @@ bool pv_map_ldt_shadow_page(unsigned int offset)
struct vcpu *curr = current;
struct domain *currd = curr->domain;
struct page_info *page;
- l1_pgentry_t gl1e, *pl1e, nl1e;
+ l1_pgentry_t gl1e;
+ mfn_t mfn;
unsigned long linear = curr->arch.pv.ldt_base + offset;
BUG_ON(in_irq());
@@ -87,10 +88,17 @@ bool pv_map_ldt_shadow_page(unsigned int offset)
return false;
}
- pl1e = &pv_ldt_ptes(curr)[offset >> PAGE_SHIFT];
- nl1e = l1e_from_pfn(l1e_get_pfn(gl1e), __PAGE_HYPERVISOR_RW);
+ mfn = page_to_mfn(page);
+ curr->arch.pv.ldt_frames[offset >> PAGE_SHIFT] = mfn;
- l1e_write(pl1e, nl1e);
+ /*
+ * Running the guest implies its page-tables are loaded, so the linear
+ * mappings would do; go through the interface anyway to keep a single
+ * writer for the per-domain area.
+ */
+ populate_perdomain_mapping(curr,
+ LDT_VIRT_START(curr) + (offset & PAGE_MASK),
+ &mfn, 1, __PAGE_HYPERVISOR_RW);
return true;
}
--
2.55.0
^ permalink raw reply related [flat|nested] 44+ messages in thread* Re: [PATCH v2 05/14] x86/pv: update guest LDT mappings using {populate,destroy}_perdomain_mapping()
2026-09-02 9:43 ` [PATCH v2 05/14] x86/pv: update guest LDT mappings using {populate,destroy}_perdomain_mapping() George Dunlap
@ 2026-09-07 16:06 ` Jan Beulich
2026-09-09 19:29 ` George Dunlap
0 siblings, 1 reply; 44+ messages in thread
From: Jan Beulich @ 2026-09-07 16:06 UTC (permalink / raw)
To: George Dunlap
Cc: Andrew Cooper, Roger Pau Monné, Alejandro Vallejo,
Teddy Astie, Anthony PERARD, Michal Orzel, Julien Grall,
Stefano Stabellini, George Dunlap, xen-devel
On 02.09.2026 11:43, George Dunlap wrote:
> From: Roger Pau Monné <roger.pau@citrix.com>
>
> Until two patches ago, update_xen_slot_in_full_gdt() used the stashed
With wording at the start here and ...
> pointer in d->arch.pv.gdt_ldt_l1tab to update the incoming vCPU's page
> tables with Xen's GDT; this was previously necessary because
> map_domain_page() couldn't be called in a context switch. Having a
> handy pointer to an always-mapped version of the GDT/LDT L1 table,
> other sites which modify the table started using it for convenience,
> even if they weren't called from within a context switch. These
> include pv_map_ldt_shadow_page() and pv_destroy_ldt().
>
> Continue the process of switching users of the stashed reference to use
> populate_perdomain_mapping() instead.
>
> pv_map_ldt_shadow_page() is, by definition, always modifying the
> currently-running vCPU: it runs from the #PF handler for a descriptor
> fetch on the guest's behalf, and running the guest implies its page
> tables are loaded. So it could simply write the linear recursive
> mappings directly. Go through populate_perdomain_mapping() anyway, to
> keep a single writer for the per-domain area.
>
> For pv_destroy_ldt(), use destroy_perdomain_mapping().
>
> Previously, pv_destroy_ldt() used the L1 LDT entries themselves to
> determine which MFNs to drop type and count references to. Rather
> than reading from the stashed L1, keep the MFNs corresponding to L1
> slots in an array in the vCPU structure, as we do in the GDT case.
> (Note that unlike the GDT case, these are not part of a public ABI, so
> can be mfn_t, avoiding a recast-and-copy.)
>
> Note that mappings_dropped (the return value of pv_destroy_ldt()) now
> reflects the *number of valid MFNs in this array*, not *the number of
> non-empty L1 entries*. This introduces an invariant we must maintain:
> pv_map_ldt_shadow_page() writes both the array entry and the mapping,
> and pv_destroy_ldt() clears both, so the two stay in lockstep.
>
> Also note that, unlike pv_destroy_gdt() from the previous patch,
... here adjusted as per the comment on the earlier patch, ...
> pv_destroy_ldt() doesn't fill in the values with zero_l1e (see
> 61031e64d3), so there's no change here.
>
> Signed-off-by: Roger Pau Monné <roger.pau@citrix.com>
> Assisted-by: Claude Code:claude-fable-5, Claude Code:claude-opus-4-8
> Signed-off-by: George Dunlap <gwd@xenproject.org>
Reviewed-by: Jan Beulich <jbeulich@suse.com>
On the basis that ...
> --- a/xen/arch/x86/include/asm/domain.h
> +++ b/xen/arch/x86/include/asm/domain.h
> @@ -541,6 +541,8 @@ struct pv_vcpu
> struct trap_info *trap_ctxt;
>
> unsigned long gdt_frames[FIRST_RESERVED_GDT_PAGE];
> + /* Max LDT entries is 8192, so 8192 * 8 = 64KiB (16 pages). */
> + mfn_t ldt_frames[16];
> unsigned long ldt_base;
> unsigned int gdt_ents, ldt_ents;
... this not really insignificant size increase is okay-ish as long as
struct hvm_vcpu is about three times the size (i.e. is still more than
double the size after this change).
Jan
^ permalink raw reply [flat|nested] 44+ messages in thread* Re: [PATCH v2 05/14] x86/pv: update guest LDT mappings using {populate,destroy}_perdomain_mapping()
2026-09-07 16:06 ` Jan Beulich
@ 2026-09-09 19:29 ` George Dunlap
0 siblings, 0 replies; 44+ messages in thread
From: George Dunlap @ 2026-09-09 19:29 UTC (permalink / raw)
To: Jan Beulich
Cc: Andrew Cooper, Roger Pau Monné, Alejandro Vallejo,
Teddy Astie, Anthony PERARD, Michal Orzel, Julien Grall,
Stefano Stabellini, xen-devel
On Mon, Sep 7, 2026 at 5:06 PM Jan Beulich <jbeulich@suse.com> wrote:
>
> On 02.09.2026 11:43, George Dunlap wrote:
> > From: Roger Pau Monné <roger.pau@citrix.com>
> >
> > Until two patches ago, update_xen_slot_in_full_gdt() used the stashed
>
> With wording at the start here and ...
>
> > pointer in d->arch.pv.gdt_ldt_l1tab to update the incoming vCPU's page
> > tables with Xen's GDT; this was previously necessary because
> > map_domain_page() couldn't be called in a context switch. Having a
> > handy pointer to an always-mapped version of the GDT/LDT L1 table,
> > other sites which modify the table started using it for convenience,
> > even if they weren't called from within a context switch. These
> > include pv_map_ldt_shadow_page() and pv_destroy_ldt().
> >
> > Continue the process of switching users of the stashed reference to use
> > populate_perdomain_mapping() instead.
> >
> > pv_map_ldt_shadow_page() is, by definition, always modifying the
> > currently-running vCPU: it runs from the #PF handler for a descriptor
> > fetch on the guest's behalf, and running the guest implies its page
> > tables are loaded. So it could simply write the linear recursive
> > mappings directly. Go through populate_perdomain_mapping() anyway, to
> > keep a single writer for the per-domain area.
> >
> > For pv_destroy_ldt(), use destroy_perdomain_mapping().
> >
> > Previously, pv_destroy_ldt() used the L1 LDT entries themselves to
> > determine which MFNs to drop type and count references to. Rather
> > than reading from the stashed L1, keep the MFNs corresponding to L1
> > slots in an array in the vCPU structure, as we do in the GDT case.
> > (Note that unlike the GDT case, these are not part of a public ABI, so
> > can be mfn_t, avoiding a recast-and-copy.)
> >
> > Note that mappings_dropped (the return value of pv_destroy_ldt()) now
> > reflects the *number of valid MFNs in this array*, not *the number of
> > non-empty L1 entries*. This introduces an invariant we must maintain:
> > pv_map_ldt_shadow_page() writes both the array entry and the mapping,
> > and pv_destroy_ldt() clears both, so the two stay in lockstep.
> >
> > Also note that, unlike pv_destroy_gdt() from the previous patch,
>
> ... here adjusted as per the comment on the earlier patch, ...
>
> > pv_destroy_ldt() doesn't fill in the values with zero_l1e (see
> > 61031e64d3), so there's no change here.
> >
> > Signed-off-by: Roger Pau Monné <roger.pau@citrix.com>
> > Assisted-by: Claude Code:claude-fable-5, Claude Code:claude-opus-4-8
> > Signed-off-by: George Dunlap <gwd@xenproject.org>
>
> Reviewed-by: Jan Beulich <jbeulich@suse.com>
>
> On the basis that ...
>
> > --- a/xen/arch/x86/include/asm/domain.h
> > +++ b/xen/arch/x86/include/asm/domain.h
> > @@ -541,6 +541,8 @@ struct pv_vcpu
> > struct trap_info *trap_ctxt;
> >
> > unsigned long gdt_frames[FIRST_RESERVED_GDT_PAGE];
> > + /* Max LDT entries is 8192, so 8192 * 8 = 64KiB (16 pages). */
> > + mfn_t ldt_frames[16];
> > unsigned long ldt_base;
> > unsigned int gdt_ents, ldt_ents;
>
> ... this not really insignificant size increase is okay-ish as long as
> struct hvm_vcpu is about three times the size (i.e. is still more than
> double the size after this change).
FYI it looks like by the end of the whole series struct pv_vcpu has
net zero change: we add this array, but then move the mapcache
structure out into arch_vcpu. In turn arch_vcpu by the end is an
extra 256 bytes, but with some rearrangement, we can reduce padding
and increase struct vcpu by only 192 bytes.
-George
^ permalink raw reply [flat|nested] 44+ messages in thread
* [PATCH v2 06/14] x86/pv: remove stashing of GDT/LDT L1 page-tables
2026-09-02 9:43 [PATCH v2 00/14] x86: Address Space Isolation, part 2: asi= option and per-vCPU page tables George Dunlap
` (4 preceding siblings ...)
2026-09-02 9:43 ` [PATCH v2 05/14] x86/pv: update guest LDT mappings using {populate,destroy}_perdomain_mapping() George Dunlap
@ 2026-09-02 9:43 ` George Dunlap
2026-09-08 14:29 ` Jan Beulich
2026-09-02 9:43 ` [PATCH v2 07/14] x86/mm: simplify create_perdomain_mapping() interface George Dunlap
` (7 subsequent siblings)
13 siblings, 1 reply; 44+ messages in thread
From: George Dunlap @ 2026-09-02 9:43 UTC (permalink / raw)
To: xen-devel
Cc: Roger Pau Monné, Jan Beulich, Andrew Cooper,
Roger Pau Monné, Alejandro Vallejo, Teddy Astie,
Anthony PERARD, Michal Orzel, Julien Grall, Stefano Stabellini,
George Dunlap
From: Roger Pau Monné <roger.pau@citrix.com>
There are no remaining users of the stashed L1 page-tables in
pv_domain.gdt_ldt_l1tab. Remove it, and all helpers. This removes a
globally-mapped xenheap allocation, and sets the stage for per-vCPU
root page tables.
pv_create_gdt_ldt_l1tab() now passes NIL() rather than the stash
array. This will cause create_perdomain_mapping() to still eagerly
allocate the L1 tables covering the GDT/LDT range; but their addresses
are no longer handed back. Doing this is necessary because
populate_perdomain_mapping() only fills existing tables, and treats
missing structure as a bug.
Another side effect of passing NIL() rather than a pointer is that the
L1 tables move from the xenheap to the domheap. Residing in the
xenheap was only ever a requirement when the stashed pointer had to
stay usable; with that requirement dropped, we can relax the
allocation requirement as well.
Signed-off-by: Roger Pau Monné <roger.pau@citrix.com>
Assisted-by: Claude Code:claude-fable-5
Signed-off-by: George Dunlap <gwd@xenproject.org>
---
Changes in v2:
- With "x86/mm: allocate the per-domain page-tables from the xenheap"
dropped from the series, passing NIL() now does move the GDT/LDT L1
tables to the domheap (upstream's allocation for non-capture mode);
in v1 they stayed in the xenheap in all modes. Reword the commit
message accordingly.
Changes since the previously posted version:
- Note the implications of changing from pointer to NIL() in
pv_create_gdt_ldt_l1tab(). In v2 this also changed where new
GDT/LDT L1 tables were allocated from: upstream's capture mode
takes them from the xenheap (the stashed pointer has to stay
usable), the NIL() mode from the domheap. Here they come from the
xenheap in all modes ("x86/mm: allocate the per-domain page-tables
from the xenheap"), so the switch only stops the addresses being
handed back.
---
xen/arch/x86/include/asm/domain.h | 9 ---------
xen/arch/x86/pv/domain.c | 10 +---------
2 files changed, 1 insertion(+), 18 deletions(-)
diff --git a/xen/arch/x86/include/asm/domain.h b/xen/arch/x86/include/asm/domain.h
index 61a9fe00f0..5c7fad26a6 100644
--- a/xen/arch/x86/include/asm/domain.h
+++ b/xen/arch/x86/include/asm/domain.h
@@ -287,8 +287,6 @@ struct time_scale {
struct pv_domain
{
- l1_pgentry_t **gdt_ldt_l1tab;
-
atomic_t nr_l4_pages;
/* Is a 32-bit PV guest? */
@@ -524,13 +522,6 @@ struct arch_domain
#define has_pirq(d) (!!((d)->arch.emulation_flags & X86_EMU_USE_PIRQ))
#define has_vpci(d) (!!((d)->arch.emulation_flags & X86_EMU_VPCI))
-#define gdt_ldt_pt_idx(v) \
- ((v)->vcpu_id >> (PAGETABLE_ORDER - GDT_LDT_VCPU_SHIFT))
-#define pv_gdt_ptes(v) \
- ((v)->domain->arch.pv.gdt_ldt_l1tab[gdt_ldt_pt_idx(v)] + \
- (((v)->vcpu_id << GDT_LDT_VCPU_SHIFT) & (L1_PAGETABLE_ENTRIES - 1)))
-#define pv_ldt_ptes(v) (pv_gdt_ptes(v) + 16)
-
struct pv_vcpu
{
/* map_domain_page() mapping cache. */
diff --git a/xen/arch/x86/pv/domain.c b/xen/arch/x86/pv/domain.c
index 7ddab1949f..35d1761c9c 100644
--- a/xen/arch/x86/pv/domain.c
+++ b/xen/arch/x86/pv/domain.c
@@ -315,7 +315,7 @@ static int pv_create_gdt_ldt_l1tab(struct vcpu *v)
{
return create_perdomain_mapping(v->domain, GDT_VIRT_START(v),
1U << GDT_LDT_VCPU_SHIFT,
- v->domain->arch.pv.gdt_ldt_l1tab,
+ NIL(l1_pgentry_t *),
NULL);
}
@@ -389,8 +389,6 @@ void pv_domain_destroy(struct domain *d)
GDT_LDT_MBYTES << (20 - PAGE_SHIFT));
XFREE(d->arch.pv.cpuidmasks);
-
- FREE_XENHEAP_PAGE(d->arch.pv.gdt_ldt_l1tab);
}
void noreturn cf_check continue_pv_domain(void);
@@ -406,12 +404,6 @@ int pv_domain_initialise(struct domain *d)
pv_l1tf_domain_init(d);
- d->arch.pv.gdt_ldt_l1tab =
- alloc_xenheap_pages(0, MEMF_node(domain_to_node(d)));
- if ( !d->arch.pv.gdt_ldt_l1tab )
- goto fail;
- clear_page(d->arch.pv.gdt_ldt_l1tab);
-
if ( levelling_caps & ~LCAP_faulting &&
(d->arch.pv.cpuidmasks = xmemdup(&cpuidmask_defaults)) == NULL )
goto fail;
--
2.55.0
^ permalink raw reply related [flat|nested] 44+ messages in thread* Re: [PATCH v2 06/14] x86/pv: remove stashing of GDT/LDT L1 page-tables
2026-09-02 9:43 ` [PATCH v2 06/14] x86/pv: remove stashing of GDT/LDT L1 page-tables George Dunlap
@ 2026-09-08 14:29 ` Jan Beulich
0 siblings, 0 replies; 44+ messages in thread
From: Jan Beulich @ 2026-09-08 14:29 UTC (permalink / raw)
To: George Dunlap
Cc: Andrew Cooper, Roger Pau Monné, Alejandro Vallejo,
Teddy Astie, Anthony PERARD, Michal Orzel, Julien Grall,
Stefano Stabellini, George Dunlap, xen-devel
On 02.09.2026 11:43, George Dunlap wrote:
> From: Roger Pau Monné <roger.pau@citrix.com>
>
> There are no remaining users of the stashed L1 page-tables in
> pv_domain.gdt_ldt_l1tab. Remove it, and all helpers. This removes a
> globally-mapped xenheap allocation, and sets the stage for per-vCPU
> root page tables.
>
> pv_create_gdt_ldt_l1tab() now passes NIL() rather than the stash
> array. This will cause create_perdomain_mapping() to still eagerly
> allocate the L1 tables covering the GDT/LDT range; but their addresses
> are no longer handed back. Doing this is necessary because
> populate_perdomain_mapping() only fills existing tables, and treats
> missing structure as a bug.
>
> Another side effect of passing NIL() rather than a pointer is that the
> L1 tables move from the xenheap to the domheap. Residing in the
> xenheap was only ever a requirement when the stashed pointer had to
> stay usable; with that requirement dropped, we can relax the
> allocation requirement as well.
>
> Signed-off-by: Roger Pau Monné <roger.pau@citrix.com>
> Assisted-by: Claude Code:claude-fable-5
> Signed-off-by: George Dunlap <gwd@xenproject.org>
Reviewed-by: Jan Beulich <jbeulich@suse.com>
^ permalink raw reply [flat|nested] 44+ messages in thread
* [PATCH v2 07/14] x86/mm: simplify create_perdomain_mapping() interface
2026-09-02 9:43 [PATCH v2 00/14] x86: Address Space Isolation, part 2: asi= option and per-vCPU page tables George Dunlap
` (5 preceding siblings ...)
2026-09-02 9:43 ` [PATCH v2 06/14] x86/pv: remove stashing of GDT/LDT L1 page-tables George Dunlap
@ 2026-09-02 9:43 ` George Dunlap
2026-09-08 14:39 ` Jan Beulich
2026-09-02 9:43 ` [PATCH v2 08/14] x86/mm: purge unneeded destroy_perdomain_mapping() George Dunlap
` (6 subsequent siblings)
13 siblings, 1 reply; 44+ messages in thread
From: George Dunlap @ 2026-09-02 9:43 UTC (permalink / raw)
To: xen-devel
Cc: Roger Pau Monné, Jan Beulich, Andrew Cooper,
Roger Pau Monné, Alejandro Vallejo, Teddy Astie,
Anthony PERARD, Michal Orzel, Julien Grall, Stefano Stabellini,
George Dunlap
From: Roger Pau Monné <roger.pau@citrix.com>
create_perdomain_mapping()'s interface is richer than any caller needs.
The paging structure for the requested range is built to a depth
selected by the pl1tab and ppg arguments, each of which distinguishes
NULL from NIL() from a real pointer:
- nr == 0: only ensure the per-domain L3 exists; nothing else is
allocated, and the other arguments are ignored.
- pl1tab == a pointer: allocate the L1 tables covering the range from
the *xenheap*, and return their (stable, direct-map) addresses in
the array -- the mode that existed to build the GDT/LDT stash.
- pl1tab == NIL(): allocate the L1 tables from the domain heap, and
return nothing.
- pl1tab == NULL: do not plumb L1 tables for their own sake (they are
still allocated on demand if data-page population requires them).
- ppg == a pointer: allocate and install zeroed data pages across the
range, and return their struct page_info pointers in the array.
- ppg == NIL(): allocate and install the zeroed data pages, but hand
nothing back; the pages are reachable only through the mapping.
- ppg == NULL: do not allocate data pages.
- both NULL, nr > 0: stop after the slot's L2; do not plumb L1 tables
at all.
Very few of these modes have users now. The last user of the
pl1tab capture mode was removed when we removed the GDT/LDT stash.
The ppg capture mode never had any users. Nothing uses the both-NULL
L2-only mode with nr != 0. What remains is exactly one bit of
information: whether the caller wants the range populated with zeroed,
area-owned data pages, or merely plumbed down to the L1 tables, ready
for populate_perdomain_mapping() to install caller-owned pages.
Replace the two arguments with a boolean expressing that bit. With the
stashing mode gone the NIL()/IS_NIL() macros lose their last user, so
drop them as well; and document the resulting interface.
No caller changes behaviour: every existing call maps onto the boolean
exactly.
Signed-off-by: Roger Pau Monné <roger.pau@citrix.com>
Assisted-by: Claude Code:claude-fable-5
Signed-off-by: George Dunlap <gwd@xenproject.org>
---
Changes in v2:
- Describe, in the mode enumeration, which heap each pl1tab mode
allocates the L1 tables from (capture mode: xenheap; NIL(): the
domain heap). With "x86/mm: allocate the per-domain page-tables
from the xenheap" dropped from the series, upstream's heap split is
back in force at this point, and the previous patch's message
refers to it.
Changes since the previously posted version:
- Drop the now-unused NIL()/IS_NIL() macros as requested during review
- Describe the prior interface in the commit message and add a doc
comment for the simplified one.
---
xen/arch/x86/domain_page.c | 10 +++---
xen/arch/x86/hvm/hvm.c | 2 +-
xen/arch/x86/include/asm/mm.h | 6 +---
xen/arch/x86/mm.c | 67 ++++++++++++++++-------------------
xen/arch/x86/pv/domain.c | 4 +--
xen/arch/x86/x86_64/mm.c | 3 +-
6 files changed, 39 insertions(+), 53 deletions(-)
diff --git a/xen/arch/x86/domain_page.c b/xen/arch/x86/domain_page.c
index 1fc1580e62..b42cf1c8cf 100644
--- a/xen/arch/x86/domain_page.c
+++ b/xen/arch/x86/domain_page.c
@@ -279,8 +279,7 @@ int mapcache_domain_init(struct domain *d)
spin_lock_init(&dcache->lock);
return create_perdomain_mapping(d, (unsigned long)dcache->inuse,
- 2 * bitmap_pages + 1,
- NIL(l1_pgentry_t *), NULL);
+ 2 * bitmap_pages + 1, false);
}
int mapcache_vcpu_init(struct vcpu *v)
@@ -297,16 +296,15 @@ int mapcache_vcpu_init(struct vcpu *v)
if ( ents > dcache->entries )
{
/* Populate page tables. */
- int rc = create_perdomain_mapping(d, MAPCACHE_VIRT_START, ents,
- NIL(l1_pgentry_t *), NULL);
+ int rc = create_perdomain_mapping(d, MAPCACHE_VIRT_START, ents, false);
/* Populate bit maps. */
if ( !rc )
rc = create_perdomain_mapping(d, (unsigned long)dcache->inuse,
- nr, NULL, NIL(struct page_info *));
+ nr, true);
if ( !rc )
rc = create_perdomain_mapping(d, (unsigned long)dcache->garbage,
- nr, NULL, NIL(struct page_info *));
+ nr, true);
if ( rc )
return rc;
diff --git a/xen/arch/x86/hvm/hvm.c b/xen/arch/x86/hvm/hvm.c
index 9a4147b62e..a6e0818468 100644
--- a/xen/arch/x86/hvm/hvm.c
+++ b/xen/arch/x86/hvm/hvm.c
@@ -620,7 +620,7 @@ int hvm_domain_initialise(struct domain *d,
INIT_LIST_HEAD(&d->arch.hvm.mmcfg_regions);
INIT_LIST_HEAD(&d->arch.hvm.msix_tables);
- rc = create_perdomain_mapping(d, PERDOMAIN_VIRT_START, 0, NULL, NULL);
+ rc = create_perdomain_mapping(d, PERDOMAIN_VIRT_START, 0, false);
if ( rc )
goto fail;
diff --git a/xen/arch/x86/include/asm/mm.h b/xen/arch/x86/include/asm/mm.h
index 1888807394..30eaec9179 100644
--- a/xen/arch/x86/include/asm/mm.h
+++ b/xen/arch/x86/include/asm/mm.h
@@ -600,12 +600,8 @@ long arch_memory_op(unsigned long cmd, XEN_GUEST_HANDLE_PARAM(void) arg);
long subarch_memory_op(unsigned long cmd, XEN_GUEST_HANDLE_PARAM(void) arg);
int compat_arch_memory_op(unsigned long cmd, XEN_GUEST_HANDLE_PARAM(void) arg);
-#define NIL(type) ((type *)-sizeof(type))
-#define IS_NIL(ptr) (!((uintptr_t)(ptr) + sizeof(*(ptr))))
-
int create_perdomain_mapping(struct domain *d, unsigned long va,
- unsigned int nr, l1_pgentry_t **pl1tab,
- struct page_info **ppg);
+ unsigned int nr, bool populate);
void populate_perdomain_mapping(const struct vcpu *v, unsigned long va,
const mfn_t *mfn, unsigned int nr,
unsigned int flags);
diff --git a/xen/arch/x86/mm.c b/xen/arch/x86/mm.c
index 552559ecf1..48d1b427c5 100644
--- a/xen/arch/x86/mm.c
+++ b/xen/arch/x86/mm.c
@@ -6211,9 +6211,31 @@ static bool perdomain_l1e_needs_freeing(l1_pgentry_t l1e)
(_PAGE_PRESENT | _PAGE_AVAIL0);
}
+/*
+ * Ensure the paging structure for [va, va + nr * PAGE_SIZE) of d's
+ * per-domain area is in place, allocating whichever levels are missing:
+ * the (domain-wide) L3 root, the slot's L2, and all L1 tables covering
+ * the range. All allocations come from the domain heap. The range must
+ * lie within a single per-domain slot (one L3 entry), and already-present
+ * levels and entries are left untouched, so calls are idempotent over
+ * existing ranges.
+ *
+ * nr == 0: only ensure the per-domain L3 itself exists; populate is
+ * ignored. Used to set the area up before any sub-range is known.
+ *
+ * populate == false: stop once the L1 tables are in place. The range is
+ * then ready for caller-owned pages to be mapped and unmapped via
+ * populate_perdomain_mapping() / destroy_perdomain_mapping(), which only
+ * fill existing tables and treat missing structure as a bug.
+ *
+ * populate == true: additionally install a freshly allocated, zeroed page
+ * at every not-yet-present entry in the range. Such pages are marked
+ * _PAGE_AVAIL0, "owned by the per-domain area": teardown frees them (see
+ * perdomain_l1e_needs_freeing()), whereas caller-owned mappings are only
+ * ever unmapped.
+ */
int create_perdomain_mapping(struct domain *d, unsigned long va,
- unsigned int nr, l1_pgentry_t **pl1tab,
- struct page_info **ppg)
+ unsigned int nr, bool populate)
{
struct page_info *pg;
l3_pgentry_t *l3tab;
@@ -6262,55 +6284,32 @@ int create_perdomain_mapping(struct domain *d, unsigned long va,
unmap_domain_page(l3tab);
- if ( !pl1tab && !ppg )
- {
- unmap_domain_page(l2tab);
- return 0;
- }
-
for ( l1tab = NULL; !rc && nr--; )
{
l2_pgentry_t *pl2e = l2tab + l2_table_offset(va);
if ( !(l2e_get_flags(*pl2e) & _PAGE_PRESENT) )
{
- if ( pl1tab && !IS_NIL(pl1tab) )
- {
- l1tab = alloc_xenheap_pages(0, MEMF_node(domain_to_node(d)));
- if ( !l1tab )
- {
- rc = -ENOMEM;
- break;
- }
- ASSERT(!pl1tab[l2_table_offset(va)]);
- pl1tab[l2_table_offset(va)] = l1tab;
- pg = virt_to_page(l1tab);
- }
- else
+ pg = alloc_domheap_page(d, MEMF_no_owner);
+ if ( !pg )
{
- pg = alloc_domheap_page(d, MEMF_no_owner);
- if ( !pg )
- {
- rc = -ENOMEM;
- break;
- }
- l1tab = __map_domain_page(pg);
+ rc = -ENOMEM;
+ break;
}
+ l1tab = __map_domain_page(pg);
clear_page(l1tab);
*pl2e = l2e_from_page(pg, __PAGE_HYPERVISOR_RW);
}
else if ( !l1tab )
l1tab = map_l1t_from_l2e(*pl2e);
- if ( ppg &&
+ if ( populate &&
!(l1e_get_flags(l1tab[l1_table_offset(va)]) & _PAGE_PRESENT) )
{
pg = alloc_domheap_page(d, MEMF_no_owner);
if ( pg )
{
clear_domain_page(page_to_mfn(pg));
- if ( !IS_NIL(ppg) )
- *ppg++ = pg;
l1tab[l1_table_offset(va)] =
l1e_from_page(pg, __PAGE_HYPERVISOR_RW | _PAGE_AVAIL0);
l2e_add_flags(*pl2e, _PAGE_AVAIL0);
@@ -6322,7 +6321,6 @@ int create_perdomain_mapping(struct domain *d, unsigned long va,
va += PAGE_SIZE;
if ( rc || !nr || !l1_table_offset(va) )
{
- /* Note that this is a no-op for the alloc_xenheap_page() case. */
unmap_domain_page(l1tab);
l1tab = NULL;
}
@@ -6543,10 +6541,7 @@ void free_perdomain_mappings(struct domain *d)
unmap_domain_page(l1tab);
}
- if ( is_xen_heap_page(l1pg) )
- free_xenheap_page(page_to_virt(l1pg));
- else
- free_domheap_page(l1pg);
+ free_domheap_page(l1pg);
}
unmap_domain_page(l2tab);
diff --git a/xen/arch/x86/pv/domain.c b/xen/arch/x86/pv/domain.c
index 35d1761c9c..15a8238aff 100644
--- a/xen/arch/x86/pv/domain.c
+++ b/xen/arch/x86/pv/domain.c
@@ -314,9 +314,7 @@ int switch_compat(struct domain *d)
static int pv_create_gdt_ldt_l1tab(struct vcpu *v)
{
return create_perdomain_mapping(v->domain, GDT_VIRT_START(v),
- 1U << GDT_LDT_VCPU_SHIFT,
- NIL(l1_pgentry_t *),
- NULL);
+ 1U << GDT_LDT_VCPU_SHIFT, false);
}
static void pv_destroy_gdt_ldt_l1tab(struct vcpu *v)
diff --git a/xen/arch/x86/x86_64/mm.c b/xen/arch/x86/x86_64/mm.c
index 8eadab7933..ffeda06e08 100644
--- a/xen/arch/x86/x86_64/mm.c
+++ b/xen/arch/x86/x86_64/mm.c
@@ -733,8 +733,7 @@ void __init zap_low_mappings(void)
int setup_compat_arg_xlat(struct vcpu *v)
{
return create_perdomain_mapping(v->domain, ARG_XLAT_START(v),
- PFN_UP(COMPAT_ARG_XLAT_SIZE),
- NULL, NIL(struct page_info *));
+ PFN_UP(COMPAT_ARG_XLAT_SIZE), true);
}
void free_compat_arg_xlat(struct vcpu *v)
--
2.55.0
^ permalink raw reply related [flat|nested] 44+ messages in thread* Re: [PATCH v2 07/14] x86/mm: simplify create_perdomain_mapping() interface
2026-09-02 9:43 ` [PATCH v2 07/14] x86/mm: simplify create_perdomain_mapping() interface George Dunlap
@ 2026-09-08 14:39 ` Jan Beulich
0 siblings, 0 replies; 44+ messages in thread
From: Jan Beulich @ 2026-09-08 14:39 UTC (permalink / raw)
To: George Dunlap
Cc: Andrew Cooper, Roger Pau Monné, Alejandro Vallejo,
Teddy Astie, Anthony PERARD, Michal Orzel, Julien Grall,
Stefano Stabellini, George Dunlap, xen-devel
On 02.09.2026 11:43, George Dunlap wrote:
> From: Roger Pau Monné <roger.pau@citrix.com>
>
> create_perdomain_mapping()'s interface is richer than any caller needs.
> The paging structure for the requested range is built to a depth
> selected by the pl1tab and ppg arguments, each of which distinguishes
> NULL from NIL() from a real pointer:
>
> - nr == 0: only ensure the per-domain L3 exists; nothing else is
> allocated, and the other arguments are ignored.
> - pl1tab == a pointer: allocate the L1 tables covering the range from
> the *xenheap*, and return their (stable, direct-map) addresses in
> the array -- the mode that existed to build the GDT/LDT stash.
> - pl1tab == NIL(): allocate the L1 tables from the domain heap, and
> return nothing.
> - pl1tab == NULL: do not plumb L1 tables for their own sake (they are
> still allocated on demand if data-page population requires them).
> - ppg == a pointer: allocate and install zeroed data pages across the
> range, and return their struct page_info pointers in the array.
> - ppg == NIL(): allocate and install the zeroed data pages, but hand
> nothing back; the pages are reachable only through the mapping.
> - ppg == NULL: do not allocate data pages.
> - both NULL, nr > 0: stop after the slot's L2; do not plumb L1 tables
> at all.
>
> Very few of these modes have users now. The last user of the
> pl1tab capture mode was removed when we removed the GDT/LDT stash.
> The ppg capture mode never had any users. Nothing uses the both-NULL
> L2-only mode with nr != 0. What remains is exactly one bit of
> information: whether the caller wants the range populated with zeroed,
> area-owned data pages, or merely plumbed down to the L1 tables, ready
> for populate_perdomain_mapping() to install caller-owned pages.
>
> Replace the two arguments with a boolean expressing that bit. With the
> stashing mode gone the NIL()/IS_NIL() macros lose their last user, so
> drop them as well; and document the resulting interface.
>
> No caller changes behaviour: every existing call maps onto the boolean
> exactly.
>
> Signed-off-by: Roger Pau Monné <roger.pau@citrix.com>
> Assisted-by: Claude Code:claude-fable-5
> Signed-off-by: George Dunlap <gwd@xenproject.org>
Reviewed-by: Jan Beulich <jbeulich@suse.com>
^ permalink raw reply [flat|nested] 44+ messages in thread
* [PATCH v2 08/14] x86/mm: purge unneeded destroy_perdomain_mapping()
2026-09-02 9:43 [PATCH v2 00/14] x86: Address Space Isolation, part 2: asi= option and per-vCPU page tables George Dunlap
` (6 preceding siblings ...)
2026-09-02 9:43 ` [PATCH v2 07/14] x86/mm: simplify create_perdomain_mapping() interface George Dunlap
@ 2026-09-02 9:43 ` George Dunlap
2026-09-08 15:03 ` Jan Beulich
2026-09-02 9:43 ` [PATCH v2 09/14] x86/mm: prepare destroy_perdomain_mapping() for per-vCPU perdomain areas George Dunlap
` (5 subsequent siblings)
13 siblings, 1 reply; 44+ messages in thread
From: George Dunlap @ 2026-09-02 9:43 UTC (permalink / raw)
To: xen-devel
Cc: Roger Pau Monné, Jan Beulich, Andrew Cooper,
Roger Pau Monné, Alejandro Vallejo, Teddy Astie,
Anthony PERARD, Michal Orzel, Julien Grall, Stefano Stabellini,
Alejandro Vallejo, George Dunlap
From: Roger Pau Monné <roger.pau@citrix.com>
We want to change per-domain mappings to be per-vCPU mappings. In
preparation for that, we want to arrange that
destroy_perdomain_mapping() work either with a single perdomain area,
or with a per-vCPU perdomain area.
There are two calls made from domain-scoped contexts; both calls turn
out to be unnecessary:
- destroy_perdomain_mapping() is not logically the undo of
create_perdomain_mapping(), as the name and its use in
hvm_domain_initialise() suggest. create_ allocates a per-domain L3,
but destroy_ tears down mappings without freeing it; and since the
call here passes nr == 0, it tears down nothing at all. The
per-domain L3 page is actually freed by free_perdomain_mappings(),
which hvm_domain_initialise()'s caller, arch_domain_create(),
already invokes on its failure path.
- The call in pv_domain_destroy() is redundant: arch_domain_destroy()
unconditionally calls free_perdomain_mappings(), which tears down
the same entries and additionally frees the page-table structures.
Signed-off-by: Roger Pau Monné <roger.pau@citrix.com>
Reviewed-by: Alejandro Vallejo <alejandro.vallejo@cloud.com>
Assisted-by: Claude Code:claude-fable-5
Signed-off-by: George Dunlap <gwd@xenproject.org>
---
Changes in v2:
- Added to the series
Changes since the previously posted version:
- Reworked the commit message to make it more clear how it fits in
with the larger series. No functional change.
---
xen/arch/x86/hvm/hvm.c | 1 -
xen/arch/x86/pv/domain.c | 3 ---
2 files changed, 4 deletions(-)
diff --git a/xen/arch/x86/hvm/hvm.c b/xen/arch/x86/hvm/hvm.c
index a6e0818468..cd425c3342 100644
--- a/xen/arch/x86/hvm/hvm.c
+++ b/xen/arch/x86/hvm/hvm.c
@@ -730,7 +730,6 @@ int hvm_domain_initialise(struct domain *d,
XFREE(d->arch.hvm.irq);
fail0:
hvm_destroy_cacheattr_region_list(d);
- destroy_perdomain_mapping(d, PERDOMAIN_VIRT_START, 0);
fail:
hvm_domain_relinquish_resources(d);
XFREE(d->arch.hvm.io_handler);
diff --git a/xen/arch/x86/pv/domain.c b/xen/arch/x86/pv/domain.c
index 15a8238aff..b936ca9b26 100644
--- a/xen/arch/x86/pv/domain.c
+++ b/xen/arch/x86/pv/domain.c
@@ -383,9 +383,6 @@ void pv_domain_destroy(struct domain *d)
{
pv_l1tf_domain_destroy(d);
- destroy_perdomain_mapping(d, GDT_LDT_VIRT_START,
- GDT_LDT_MBYTES << (20 - PAGE_SHIFT));
-
XFREE(d->arch.pv.cpuidmasks);
}
--
2.55.0
^ permalink raw reply related [flat|nested] 44+ messages in thread* Re: [PATCH v2 08/14] x86/mm: purge unneeded destroy_perdomain_mapping()
2026-09-02 9:43 ` [PATCH v2 08/14] x86/mm: purge unneeded destroy_perdomain_mapping() George Dunlap
@ 2026-09-08 15:03 ` Jan Beulich
0 siblings, 0 replies; 44+ messages in thread
From: Jan Beulich @ 2026-09-08 15:03 UTC (permalink / raw)
To: George Dunlap
Cc: Andrew Cooper, Roger Pau Monné, Alejandro Vallejo,
Teddy Astie, Anthony PERARD, Michal Orzel, Julien Grall,
Stefano Stabellini, George Dunlap, xen-devel
On 02.09.2026 11:43, George Dunlap wrote:
> From: Roger Pau Monné <roger.pau@citrix.com>
>
> We want to change per-domain mappings to be per-vCPU mappings. In
> preparation for that, we want to arrange that
> destroy_perdomain_mapping() work either with a single perdomain area,
> or with a per-vCPU perdomain area.
>
> There are two calls made from domain-scoped contexts; both calls turn
> out to be unnecessary:
>
> - destroy_perdomain_mapping() is not logically the undo of
> create_perdomain_mapping(), as the name and its use in
> hvm_domain_initialise() suggest. create_ allocates a per-domain L3,
> but destroy_ tears down mappings without freeing it; and since the
> call here passes nr == 0, it tears down nothing at all. The
> per-domain L3 page is actually freed by free_perdomain_mappings(),
> which hvm_domain_initialise()'s caller, arch_domain_create(),
> already invokes on its failure path.
>
> - The call in pv_domain_destroy() is redundant: arch_domain_destroy()
> unconditionally calls free_perdomain_mappings(), which tears down
> the same entries and additionally frees the page-table structures.
pv_domain_destroy() has a 2nd call site (the error path of
pv_domain_initialise()), but the situation is the same there:
arch_domain_create()'s error path also calls free_perdomain_mappings().
> Signed-off-by: Roger Pau Monné <roger.pau@citrix.com>
> Reviewed-by: Alejandro Vallejo <alejandro.vallejo@cloud.com>
> Assisted-by: Claude Code:claude-fable-5
> Signed-off-by: George Dunlap <gwd@xenproject.org>
With the description amended:
Reviewed-by: Jan Beulich <jbeulich@suse.com>
Jan
^ permalink raw reply [flat|nested] 44+ messages in thread
* [PATCH v2 09/14] x86/mm: prepare destroy_perdomain_mapping() for per-vCPU perdomain areas
2026-09-02 9:43 [PATCH v2 00/14] x86: Address Space Isolation, part 2: asi= option and per-vCPU page tables George Dunlap
` (7 preceding siblings ...)
2026-09-02 9:43 ` [PATCH v2 08/14] x86/mm: purge unneeded destroy_perdomain_mapping() George Dunlap
@ 2026-09-02 9:43 ` George Dunlap
2026-09-08 15:36 ` Jan Beulich
2026-09-02 9:43 ` [PATCH v2 10/14] x86/domain_page: drop redundant create_perdomain_mapping() call George Dunlap
` (4 subsequent siblings)
13 siblings, 1 reply; 44+ messages in thread
From: George Dunlap @ 2026-09-02 9:43 UTC (permalink / raw)
To: xen-devel
Cc: Roger Pau Monné, Jan Beulich, Andrew Cooper,
Roger Pau Monné, Alejandro Vallejo, Teddy Astie,
Anthony PERARD, Michal Orzel, Julien Grall, Stefano Stabellini,
George Dunlap
From: Roger Pau Monné <roger.pau@citrix.com>
We want to change per-domain mappings to be per-vCPU mappings. In
preparation for that, we want to arrange that
destroy_perdomain_mapping() work either with a single perdomain area,
or with a per-vCPU perdomain area.
The remaining callers are already in a vCPU context, so we just need
to change the parameter from a domain pointer to a vCPU pointer.
Since we now have a specific vCPU in mind, we have the option of using
the linear page table mapping rather than map-and-walk. As in
populate_perdomain_mapping(), the linear page table fast path is keyed
off this_cpu(pgtable_vcpu) matching the target vCPU, which is the
conditional that implies "the linear mapping area points to v's
per-domain area".
Signed-off-by: Roger Pau Monné <roger.pau@citrix.com>
Assisted-by: Claude Code:claude-fable-5, Claude Code:claude-opus-4-8
Signed-off-by: George Dunlap <gwd@xenproject.org>
---
Changes in v2:
- Added to the series
Changes since the previously posted version:
- Key the fast path off pgtable_vcpu instead of current, matching
populate_perdomain_mapping(), and drop the sync_local_execstate()
call.
- Also convert the pv_destroy_ldt() call, added by the stash-removal
batch.
- Reword and retitle for clarity (was: "x86/mm: switch
destroy_perdomain_mapping() parameter from domain to vCPU").
---
xen/arch/x86/include/asm/mm.h | 2 +-
xen/arch/x86/mm.c | 23 ++++++++++++++++++++++-
xen/arch/x86/pv/descriptor-tables.c | 2 +-
xen/arch/x86/pv/domain.c | 3 +--
xen/arch/x86/x86_64/mm.c | 2 +-
5 files changed, 26 insertions(+), 6 deletions(-)
diff --git a/xen/arch/x86/include/asm/mm.h b/xen/arch/x86/include/asm/mm.h
index 30eaec9179..9a8fda782e 100644
--- a/xen/arch/x86/include/asm/mm.h
+++ b/xen/arch/x86/include/asm/mm.h
@@ -605,7 +605,7 @@ int create_perdomain_mapping(struct domain *d, unsigned long va,
void populate_perdomain_mapping(const struct vcpu *v, unsigned long va,
const mfn_t *mfn, unsigned int nr,
unsigned int flags);
-void destroy_perdomain_mapping(struct domain *d, unsigned long va,
+void destroy_perdomain_mapping(const struct vcpu *v, unsigned long va,
unsigned int nr);
void free_perdomain_mappings(struct domain *d);
diff --git a/xen/arch/x86/mm.c b/xen/arch/x86/mm.c
index 48d1b427c5..fc524ef0c3 100644
--- a/xen/arch/x86/mm.c
+++ b/xen/arch/x86/mm.c
@@ -6456,10 +6456,11 @@ void populate_perdomain_mapping(const struct vcpu *v, unsigned long va,
local_irq_restore(irq_flags);
}
-void destroy_perdomain_mapping(struct domain *d, unsigned long va,
+void destroy_perdomain_mapping(const struct vcpu *v, unsigned long va,
unsigned int nr)
{
const l3_pgentry_t *l3tab, *pl3e;
+ const struct domain *d = v->domain;
ASSERT(va >= PERDOMAIN_VIRT_START &&
va < PERDOMAIN_VIRT_SLOT(PERDOMAIN_SLOTS));
@@ -6468,6 +6469,26 @@ void destroy_perdomain_mapping(struct domain *d, unsigned long va,
if ( !d->arch.perdomain_l3_pg )
return;
+ if ( likely(this_cpu(pgtable_vcpu) == v) )
+ {
+ l1_pgentry_t *pl1e;
+
+ /*
+ * Fast path: v's page-tables are loaded on this pCPU, so the L1
+ * entries can be zapped using the recursive linear mappings.
+ */
+ pl1e = &__linear_l1_table[l1_linear_offset(va)];
+
+ for ( ; nr--; pl1e++ )
+ {
+ if ( perdomain_l1e_needs_freeing(*pl1e) )
+ free_domheap_page(l1e_get_page(*pl1e));
+ l1e_write(pl1e, l1e_empty());
+ }
+
+ return;
+ }
+
l3tab = __map_domain_page(d->arch.perdomain_l3_pg);
pl3e = l3tab + l3_table_offset(va);
diff --git a/xen/arch/x86/pv/descriptor-tables.c b/xen/arch/x86/pv/descriptor-tables.c
index 261bf29c90..0c1ea4ce3a 100644
--- a/xen/arch/x86/pv/descriptor-tables.c
+++ b/xen/arch/x86/pv/descriptor-tables.c
@@ -27,7 +27,7 @@ bool pv_destroy_ldt(struct vcpu *v)
ASSERT(v == current || !vcpu_cpu_dirty(v));
- destroy_perdomain_mapping(v->domain, LDT_VIRT_START(v), nr_frames);
+ destroy_perdomain_mapping(v, LDT_VIRT_START(v), nr_frames);
for ( i = 0; i < nr_frames; i++ )
{
diff --git a/xen/arch/x86/pv/domain.c b/xen/arch/x86/pv/domain.c
index b936ca9b26..40b834e1a4 100644
--- a/xen/arch/x86/pv/domain.c
+++ b/xen/arch/x86/pv/domain.c
@@ -319,8 +319,7 @@ static int pv_create_gdt_ldt_l1tab(struct vcpu *v)
static void pv_destroy_gdt_ldt_l1tab(struct vcpu *v)
{
- destroy_perdomain_mapping(v->domain, GDT_VIRT_START(v),
- 1U << GDT_LDT_VCPU_SHIFT);
+ destroy_perdomain_mapping(v, GDT_VIRT_START(v), 1U << GDT_LDT_VCPU_SHIFT);
}
void pv_vcpu_destroy(struct vcpu *v)
diff --git a/xen/arch/x86/x86_64/mm.c b/xen/arch/x86/x86_64/mm.c
index ffeda06e08..aa74acec82 100644
--- a/xen/arch/x86/x86_64/mm.c
+++ b/xen/arch/x86/x86_64/mm.c
@@ -738,7 +738,7 @@ int setup_compat_arg_xlat(struct vcpu *v)
void free_compat_arg_xlat(struct vcpu *v)
{
- destroy_perdomain_mapping(v->domain, ARG_XLAT_START(v),
+ destroy_perdomain_mapping(v, ARG_XLAT_START(v),
PFN_UP(COMPAT_ARG_XLAT_SIZE));
}
--
2.55.0
^ permalink raw reply related [flat|nested] 44+ messages in thread* Re: [PATCH v2 09/14] x86/mm: prepare destroy_perdomain_mapping() for per-vCPU perdomain areas
2026-09-02 9:43 ` [PATCH v2 09/14] x86/mm: prepare destroy_perdomain_mapping() for per-vCPU perdomain areas George Dunlap
@ 2026-09-08 15:36 ` Jan Beulich
2026-09-10 11:38 ` George Dunlap
0 siblings, 1 reply; 44+ messages in thread
From: Jan Beulich @ 2026-09-08 15:36 UTC (permalink / raw)
To: George Dunlap
Cc: Andrew Cooper, Roger Pau Monné, Alejandro Vallejo,
Teddy Astie, Anthony PERARD, Michal Orzel, Julien Grall,
Stefano Stabellini, George Dunlap, xen-devel
On 02.09.2026 11:43, George Dunlap wrote:
> --- a/xen/arch/x86/mm.c
> +++ b/xen/arch/x86/mm.c
> @@ -6456,10 +6456,11 @@ void populate_perdomain_mapping(const struct vcpu *v, unsigned long va,
> local_irq_restore(irq_flags);
> }
>
> -void destroy_perdomain_mapping(struct domain *d, unsigned long va,
> +void destroy_perdomain_mapping(const struct vcpu *v, unsigned long va,
> unsigned int nr)
> {
> const l3_pgentry_t *l3tab, *pl3e;
> + const struct domain *d = v->domain;
>
> ASSERT(va >= PERDOMAIN_VIRT_START &&
> va < PERDOMAIN_VIRT_SLOT(PERDOMAIN_SLOTS));
> @@ -6468,6 +6469,26 @@ void destroy_perdomain_mapping(struct domain *d, unsigned long va,
> if ( !d->arch.perdomain_l3_pg )
> return;
>
> + if ( likely(this_cpu(pgtable_vcpu) == v) )
Right now this looks to be relevant only to some of the call sites of
pv_destroy_ldt(). I expect this is going to change down the road?
> + {
> + l1_pgentry_t *pl1e;
> +
> + /*
> + * Fast path: v's page-tables are loaded on this pCPU, so the L1
> + * entries can be zapped using the recursive linear mappings.
> + */
> + pl1e = &__linear_l1_table[l1_linear_offset(va)];
> +
> + for ( ; nr--; pl1e++ )
> + {
> + if ( perdomain_l1e_needs_freeing(*pl1e) )
> + free_domheap_page(l1e_get_page(*pl1e));
> + l1e_write(pl1e, l1e_empty());
> + }
> +
> + return;
> + }
> +
> l3tab = __map_domain_page(d->arch.perdomain_l3_pg);
> pl3e = l3tab + l3_table_offset(va);
The slow path is resilient against an area never having been mapped. This
fast path looks like it would crash on a missing L2 or L1 table. I.e. for
domain cleanup (and error handling during domain creation) we'd
implicitly rely on those never taking the fast path. May be worth making
explicit in the description.
> --- a/xen/arch/x86/pv/domain.c
> +++ b/xen/arch/x86/pv/domain.c
> @@ -319,8 +319,7 @@ static int pv_create_gdt_ldt_l1tab(struct vcpu *v)
>
> static void pv_destroy_gdt_ldt_l1tab(struct vcpu *v)
> {
> - destroy_perdomain_mapping(v->domain, GDT_VIRT_START(v),
> - 1U << GDT_LDT_VCPU_SHIFT);
> + destroy_perdomain_mapping(v, GDT_VIRT_START(v), 1U << GDT_LDT_VCPU_SHIFT);
> }
Considering what patch 08 does - is this explicit call really needed?
There are no pages to free here, unlike ...
> --- a/xen/arch/x86/x86_64/mm.c
> +++ b/xen/arch/x86/x86_64/mm.c
> @@ -738,7 +738,7 @@ int setup_compat_arg_xlat(struct vcpu *v)
>
> void free_compat_arg_xlat(struct vcpu *v)
> {
> - destroy_perdomain_mapping(v->domain, ARG_XLAT_START(v),
> + destroy_perdomain_mapping(v, ARG_XLAT_START(v),
> PFN_UP(COMPAT_ARG_XLAT_SIZE));
> }
... here. Then again even the freeing is taken care of by
free_perdomain_mappings(), so even for this one the question arises
whether it's actually needed.
Jan
^ permalink raw reply [flat|nested] 44+ messages in thread* Re: [PATCH v2 09/14] x86/mm: prepare destroy_perdomain_mapping() for per-vCPU perdomain areas
2026-09-08 15:36 ` Jan Beulich
@ 2026-09-10 11:38 ` George Dunlap
2026-09-10 11:54 ` Jan Beulich
0 siblings, 1 reply; 44+ messages in thread
From: George Dunlap @ 2026-09-10 11:38 UTC (permalink / raw)
To: Jan Beulich
Cc: Andrew Cooper, Roger Pau Monné, Alejandro Vallejo,
Teddy Astie, Anthony PERARD, Michal Orzel, Julien Grall,
Stefano Stabellini, xen-devel
On Tue, Sep 8, 2026 at 4:36 PM Jan Beulich <jbeulich@suse.com> wrote:
>
> On 02.09.2026 11:43, George Dunlap wrote:
> > --- a/xen/arch/x86/mm.c
> > +++ b/xen/arch/x86/mm.c
> > @@ -6456,10 +6456,11 @@ void populate_perdomain_mapping(const struct vcpu *v, unsigned long va,
> > local_irq_restore(irq_flags);
> > }
> >
> > -void destroy_perdomain_mapping(struct domain *d, unsigned long va,
> > +void destroy_perdomain_mapping(const struct vcpu *v, unsigned long va,
> > unsigned int nr)
> > {
> > const l3_pgentry_t *l3tab, *pl3e;
> > + const struct domain *d = v->domain;
> >
> > ASSERT(va >= PERDOMAIN_VIRT_START &&
> > va < PERDOMAIN_VIRT_SLOT(PERDOMAIN_SLOTS));
> > @@ -6468,6 +6469,26 @@ void destroy_perdomain_mapping(struct domain *d, unsigned long va,
> > if ( !d->arch.perdomain_l3_pg )
> > return;
> >
> > + if ( likely(this_cpu(pgtable_vcpu) == v) )
>
> Right now this looks to be relevant only to some of the call sites of
> pv_destroy_ldt(). I expect this is going to change down the road?
It doesn't look like it, actually; I suspect Roger just added it for
consistency's sake (all perdomain functions take a vcpu, so they can
all use the linear pagetable).
> > + {
> > + l1_pgentry_t *pl1e;
> > +
> > + /*
> > + * Fast path: v's page-tables are loaded on this pCPU, so the L1
> > + * entries can be zapped using the recursive linear mappings.
> > + */
> > + pl1e = &__linear_l1_table[l1_linear_offset(va)];
> > +
> > + for ( ; nr--; pl1e++ )
> > + {
> > + if ( perdomain_l1e_needs_freeing(*pl1e) )
> > + free_domheap_page(l1e_get_page(*pl1e));
> > + l1e_write(pl1e, l1e_empty());
> > + }
> > +
> > + return;
> > + }
> > +
> > l3tab = __map_domain_page(d->arch.perdomain_l3_pg);
> > pl3e = l3tab + l3_table_offset(va);
>
> The slow path is resilient against an area never having been mapped. This
> fast path looks like it would crash on a missing L2 or L1 table. I.e. for
> domain cleanup (and error handling during domain creation) we'd
> implicitly rely on those never taking the fast path. May be worth making
> explicit in the description.
If we keep the fast path, it's probably worth making it resilient
against empty paths, just to save ourselves time in the future.
On the other hand, I'm not terribly attached to the fast path; it can
only really be used for the guest-driven LDT teardown, which isn't a
hot path. I'm as happy to rip it out.
> > --- a/xen/arch/x86/pv/domain.c
> > +++ b/xen/arch/x86/pv/domain.c
> > @@ -319,8 +319,7 @@ static int pv_create_gdt_ldt_l1tab(struct vcpu *v)
> >
> > static void pv_destroy_gdt_ldt_l1tab(struct vcpu *v)
> > {
> > - destroy_perdomain_mapping(v->domain, GDT_VIRT_START(v),
> > - 1U << GDT_LDT_VCPU_SHIFT);
> > + destroy_perdomain_mapping(v, GDT_VIRT_START(v), 1U << GDT_LDT_VCPU_SHIFT);
> > }
>
> Considering what patch 08 does - is this explicit call really needed?
> There are no pages to free here, unlike ...
>
> > --- a/xen/arch/x86/x86_64/mm.c
> > +++ b/xen/arch/x86/x86_64/mm.c
> > @@ -738,7 +738,7 @@ int setup_compat_arg_xlat(struct vcpu *v)
> >
> > void free_compat_arg_xlat(struct vcpu *v)
> > {
> > - destroy_perdomain_mapping(v->domain, ARG_XLAT_START(v),
> > + destroy_perdomain_mapping(v, ARG_XLAT_START(v),
> > PFN_UP(COMPAT_ARG_XLAT_SIZE));
> > }
>
> ... here. Then again even the freeing is taken care of by
> free_perdomain_mappings(), so even for this one the question arises
> whether it's actually needed.
It looks like this one is needed for the "undo_and_fail" path of
xen/arch/x86/pv/domain.c:switch_compat(); in theory the vcpu could
continue as a 64-bit vcpu afterwards.
But yes, given that we're allowing "free" to subsume "destroy", we
could just make that the policy, and drop a bunch of other redundant
calls as well. Right now there's only the one, but by the end of the
series there are a handful more that we could refrain to add.
I'm inclined to drop the linear map use, and also drop
pv_destroy_gdt_l1tab(). Any thoughts?
-George
^ permalink raw reply [flat|nested] 44+ messages in thread* Re: [PATCH v2 09/14] x86/mm: prepare destroy_perdomain_mapping() for per-vCPU perdomain areas
2026-09-10 11:38 ` George Dunlap
@ 2026-09-10 11:54 ` Jan Beulich
0 siblings, 0 replies; 44+ messages in thread
From: Jan Beulich @ 2026-09-10 11:54 UTC (permalink / raw)
To: George Dunlap
Cc: Andrew Cooper, Roger Pau Monné, Alejandro Vallejo,
Teddy Astie, Anthony PERARD, Michal Orzel, Julien Grall,
Stefano Stabellini, xen-devel
On 10.09.2026 13:38, George Dunlap wrote:
> I'm inclined to drop the linear map use, and also drop
> pv_destroy_gdt_l1tab(). Any thoughts?
Doing so would follow the "start simple" principle you mentioned earlier.
So: Yes, perhaps best.
Jan
^ permalink raw reply [flat|nested] 44+ messages in thread
* [PATCH v2 10/14] x86/domain_page: drop redundant create_perdomain_mapping() call
2026-09-02 9:43 [PATCH v2 00/14] x86: Address Space Isolation, part 2: asi= option and per-vCPU page tables George Dunlap
` (8 preceding siblings ...)
2026-09-02 9:43 ` [PATCH v2 09/14] x86/mm: prepare destroy_perdomain_mapping() for per-vCPU perdomain areas George Dunlap
@ 2026-09-02 9:43 ` George Dunlap
2026-09-08 15:55 ` Jan Beulich
2026-09-02 9:43 ` [PATCH v2 11/14] x86/mm: prepare create_perdomain_mapping() for per-vCPU perdomain areas George Dunlap
` (3 subsequent siblings)
13 siblings, 1 reply; 44+ messages in thread
From: George Dunlap @ 2026-09-02 9:43 UTC (permalink / raw)
To: xen-devel
Cc: Roger Pau Monné, Jan Beulich, Andrew Cooper,
Roger Pau Monné, Alejandro Vallejo, Teddy Astie,
Anthony PERARD, Michal Orzel, Julien Grall, Stefano Stabellini,
George Dunlap
From: Roger Pau Monné <roger.pau@citrix.com>
We want to change per-domain mappings to be per-vCPU mappings. In
preparation for that, we want to arrange that
create_perdomain_mapping() work either with a single perdomain area,
or with a per-vCPU perdomain area.
One of the current calls in a domain context turns out to be
unnecessary:
mapcache_domain_init() pre-plumbs L1 tables over the whole
inuse/garbage bitmap range (sized for the full MAPCACHE_ENTRIES
capacity), without populating any data pages. The plumbing is
redundant: mapcache_vcpu_init() installs the bitmap pages the domain
will actually use with populate=true calls, which allocate any missing
page-table structure on demand -- and since d->max_vcpus is fixed
before any vCPU is created, the range those calls cover never grows.
The pre-plumbed tail beyond it backs virtual addresses that are never
populated at all.
Drop the call. With the only fallible operation gone,
mapcache_domain_init() becomes void, and arch_domain_create()'s error
handling for it goes away.
Signed-off-by: Roger Pau Monné <roger.pau@citrix.com>
Assisted-by: Claude Code:claude-fable-5, Claude Code:claude-opus-4-8
Signed-off-by: George Dunlap <gwd@xenproject.org>
---
Changes in v2:
- Added to the series (split out of the following patch).
Changes since the previously posted version:
- Split out of "x86/mm: switch {create,destroy}_perdomain_mapping()
domain parameter to vCPU", where the removal was folded into the
parameter switch without its own rationale.
---
xen/arch/x86/domain.c | 3 +--
xen/arch/x86/domain_page.c | 7 ++-----
xen/arch/x86/include/asm/domain.h | 2 +-
3 files changed, 4 insertions(+), 8 deletions(-)
diff --git a/xen/arch/x86/domain.c b/xen/arch/x86/domain.c
index d8af06e533..efa72cd2f1 100644
--- a/xen/arch/x86/domain.c
+++ b/xen/arch/x86/domain.c
@@ -908,8 +908,7 @@ int arch_domain_create(struct domain *d,
}
else if ( is_pv_domain(d) )
{
- if ( (rc = mapcache_domain_init(d)) != 0 )
- goto fail;
+ mapcache_domain_init(d);
if ( (rc = pv_domain_initialise(d)) != 0 )
goto fail;
diff --git a/xen/arch/x86/domain_page.c b/xen/arch/x86/domain_page.c
index b42cf1c8cf..449d4f2a7d 100644
--- a/xen/arch/x86/domain_page.c
+++ b/xen/arch/x86/domain_page.c
@@ -256,7 +256,7 @@ void unmap_domain_page_irqoff(const void *ptr)
do_unmap_domain_page(ptr, true);
}
-int mapcache_domain_init(struct domain *d)
+void mapcache_domain_init(struct domain *d)
{
struct mapcache_domain *dcache = &d->arch.pv.mapcache;
unsigned int bitmap_pages;
@@ -265,7 +265,7 @@ int mapcache_domain_init(struct domain *d)
#ifdef NDEBUG
if ( !mem_hotplug && max_page <= PFN_DOWN(__pa(HYPERVISOR_VIRT_END - 1)) )
- return 0;
+ return;
#endif
BUILD_BUG_ON(MAPCACHE_VIRT_END + PAGE_SIZE * (3 +
@@ -277,9 +277,6 @@ int mapcache_domain_init(struct domain *d)
(bitmap_pages + 1) * PAGE_SIZE / sizeof(long);
spin_lock_init(&dcache->lock);
-
- return create_perdomain_mapping(d, (unsigned long)dcache->inuse,
- 2 * bitmap_pages + 1, false);
}
int mapcache_vcpu_init(struct vcpu *v)
diff --git a/xen/arch/x86/include/asm/domain.h b/xen/arch/x86/include/asm/domain.h
index 5c7fad26a6..38df5c376e 100644
--- a/xen/arch/x86/include/asm/domain.h
+++ b/xen/arch/x86/include/asm/domain.h
@@ -89,7 +89,7 @@ struct mapcache_domain {
unsigned long *garbage;
};
-int mapcache_domain_init(struct domain *d);
+void mapcache_domain_init(struct domain *d);
int mapcache_vcpu_init(struct vcpu *v);
/* x86/64: toggle guest between kernel and user modes. */
--
2.55.0
^ permalink raw reply related [flat|nested] 44+ messages in thread* Re: [PATCH v2 10/14] x86/domain_page: drop redundant create_perdomain_mapping() call
2026-09-02 9:43 ` [PATCH v2 10/14] x86/domain_page: drop redundant create_perdomain_mapping() call George Dunlap
@ 2026-09-08 15:55 ` Jan Beulich
2026-09-10 11:52 ` George Dunlap
0 siblings, 1 reply; 44+ messages in thread
From: Jan Beulich @ 2026-09-08 15:55 UTC (permalink / raw)
To: George Dunlap
Cc: Andrew Cooper, Roger Pau Monné, Alejandro Vallejo,
Teddy Astie, Anthony PERARD, Michal Orzel, Julien Grall,
Stefano Stabellini, George Dunlap, xen-devel
On 02.09.2026 11:43, George Dunlap wrote:
> --- a/xen/arch/x86/domain_page.c
> +++ b/xen/arch/x86/domain_page.c
> @@ -256,7 +256,7 @@ void unmap_domain_page_irqoff(const void *ptr)
> do_unmap_domain_page(ptr, true);
> }
>
> -int mapcache_domain_init(struct domain *d)
> +void mapcache_domain_init(struct domain *d)
> {
> struct mapcache_domain *dcache = &d->arch.pv.mapcache;
> unsigned int bitmap_pages;
> @@ -265,7 +265,7 @@ int mapcache_domain_init(struct domain *d)
>
> #ifdef NDEBUG
> if ( !mem_hotplug && max_page <= PFN_DOWN(__pa(HYPERVISOR_VIRT_END - 1)) )
> - return 0;
> + return;
> #endif
>
> BUILD_BUG_ON(MAPCACHE_VIRT_END + PAGE_SIZE * (3 +
> @@ -277,9 +277,6 @@ int mapcache_domain_init(struct domain *d)
> (bitmap_pages + 1) * PAGE_SIZE / sizeof(long);
>
> spin_lock_init(&dcache->lock);
> -
> - return create_perdomain_mapping(d, (unsigned long)dcache->inuse,
> - 2 * bitmap_pages + 1, false);
> }
At this point rather than removing this, all of what is done ...
> int mapcache_vcpu_init(struct vcpu *v)
... in this function (per-domain-mapping-wise) would want moving into
mapcache_domain_init(). The present arrangement, aiui, is a leftover from
when d->max_vcpus could change post-domain-creation. Question is - would
that go against further ASI plans? (Likely the answer is "yes".)
Jan
^ permalink raw reply [flat|nested] 44+ messages in thread* Re: [PATCH v2 10/14] x86/domain_page: drop redundant create_perdomain_mapping() call
2026-09-08 15:55 ` Jan Beulich
@ 2026-09-10 11:52 ` George Dunlap
0 siblings, 0 replies; 44+ messages in thread
From: George Dunlap @ 2026-09-10 11:52 UTC (permalink / raw)
To: Jan Beulich
Cc: Andrew Cooper, Roger Pau Monné, Alejandro Vallejo,
Teddy Astie, Anthony PERARD, Michal Orzel, Julien Grall,
Stefano Stabellini, xen-devel
On Tue, Sep 8, 2026 at 4:55 PM Jan Beulich <jbeulich@suse.com> wrote:
>
> On 02.09.2026 11:43, George Dunlap wrote:
> > --- a/xen/arch/x86/domain_page.c
> > +++ b/xen/arch/x86/domain_page.c
> > @@ -256,7 +256,7 @@ void unmap_domain_page_irqoff(const void *ptr)
> > do_unmap_domain_page(ptr, true);
> > }
> >
> > -int mapcache_domain_init(struct domain *d)
> > +void mapcache_domain_init(struct domain *d)
> > {
> > struct mapcache_domain *dcache = &d->arch.pv.mapcache;
> > unsigned int bitmap_pages;
> > @@ -265,7 +265,7 @@ int mapcache_domain_init(struct domain *d)
> >
> > #ifdef NDEBUG
> > if ( !mem_hotplug && max_page <= PFN_DOWN(__pa(HYPERVISOR_VIRT_END - 1)) )
> > - return 0;
> > + return;
> > #endif
> >
> > BUILD_BUG_ON(MAPCACHE_VIRT_END + PAGE_SIZE * (3 +
> > @@ -277,9 +277,6 @@ int mapcache_domain_init(struct domain *d)
> > (bitmap_pages + 1) * PAGE_SIZE / sizeof(long);
> >
> > spin_lock_init(&dcache->lock);
> > -
> > - return create_perdomain_mapping(d, (unsigned long)dcache->inuse,
> > - 2 * bitmap_pages + 1, false);
> > }
>
> At this point rather than removing this, all of what is done ...
>
> > int mapcache_vcpu_init(struct vcpu *v)
>
> ... in this function (per-domain-mapping-wise) would want moving into
> mapcache_domain_init(). The present arrangement, aiui, is a leftover from
> when d->max_vcpus could change post-domain-creation. Question is - would
> that go against further ASI plans? (Likely the answer is "yes".)
Yes, because soon we'll be introducing per-vCPU mapcaches, which will
very much want their own initialization function.
I can add a line to this effect in v3.
-George
^ permalink raw reply [flat|nested] 44+ messages in thread
* [PATCH v2 11/14] x86/mm: prepare create_perdomain_mapping() for per-vCPU perdomain areas
2026-09-02 9:43 [PATCH v2 00/14] x86: Address Space Isolation, part 2: asi= option and per-vCPU page tables George Dunlap
` (9 preceding siblings ...)
2026-09-02 9:43 ` [PATCH v2 10/14] x86/domain_page: drop redundant create_perdomain_mapping() call George Dunlap
@ 2026-09-02 9:43 ` George Dunlap
2026-09-02 9:43 ` [PATCH v2 12/14] x86/spec-ctrl: introduce Address Space Isolation command line option George Dunlap
` (2 subsequent siblings)
13 siblings, 0 replies; 44+ messages in thread
From: George Dunlap @ 2026-09-02 9:43 UTC (permalink / raw)
To: xen-devel
Cc: Roger Pau Monné, Jan Beulich, Andrew Cooper,
Roger Pau Monné, Alejandro Vallejo, Teddy Astie,
Anthony PERARD, Michal Orzel, Julien Grall, Stefano Stabellini,
George Dunlap
From: Roger Pau Monné <roger.pau@citrix.com>
We want to change per-domain mappings to be per-vCPU mappings. In
preparation for that, we want to arrange that
create_perdomain_mapping() work either with a single perdomain area,
or with a per-vCPU perdomain area.
Most of the remaining callers are already in a vCPU context. This is
no accident: the perdomain area has always been laid out in per-vCPU
slices -- each vCPU has its own GDT/LDT window, its own
COMPAT_ARG_XLAT pages, its own window of mapcache entries -- and each
vCPU's slice is set up as that vCPU is created. For these callers, we
just need to change the parameter from a domain pointer to a vCPU
pointer. Once the perdomain area itself becomes per-vCPU, the same
calls will populate the owning vCPU's own area rather than slices of a
shared one.
One exception is the call in hvm_domain_initialise(). An HVM vCPU's
monitor table is created during vCPU initialisation, and
init_xen_l4_slots() stamps the perdomain slot into it at that point --
far earlier than for PV, where the Xen slots are written only once
guest page tables are built. hvm_domain_initialise() therefore had an
explicit create_perdomain_mapping() call just to make the perdomain
root exist ahead of that. Move it to arch_vcpu_create(), covering HVM
and PV alike. With a single shared area, the call allocates at most
once per domain; but once each vCPU has its own perdomain area, this
is the call that will allocate every vCPU's root -- PV included --
before any page tables referencing it are built. vCPU creation is
where the call must end up; move it there directly. For PV guests
nothing observable changes: the root was already being created during
vCPU creation as a side effect (by mapcache_vcpu_init(), or failing
that pv_create_gdt_ldt_l1tab()); it now merely becomes explicit.
Note that we cannot yet do a parallel movement of
free_perdomain_mappings(): the per-domain page-table hierarchy is
still a single domain-wide structure shared by all vCPUs, so tearing
it down from a per-vCPU path would pull the mappings out from under
sibling vCPUs (e.g. on a partially failed, retryable
XEN_DOMCTL_max_vcpus), and vCPU-create error paths can rely on domain
destruction to free a partially set up hierarchy. Teardown will move
to vCPU scope only once the structure itself becomes per-vCPU.
Signed-off-by: Roger Pau Monné <roger.pau@citrix.com>
Assisted-by: Claude Code:claude-fable-5, Claude Code:claude-opus-4-8
Signed-off-by: George Dunlap <gwd@xenproject.org>
---
Changes in v2:
- Added to the series
Changes since the previously posted version:
- Keep free_perdomain_mappings() (and hence perdomain teardown)
domain-scoped.
- Keep the idle domain without a perdomain area.
- Retitle (was: "x86/mm: switch {create,destroy}_perdomain_mapping()
domain parameter to vCPU"); destroy_perdomain_mapping() was switched
in a separate patch.
- Split the removal of mapcache_domain_init()'s redundant
create_perdomain_mapping() call into its own (preceding) patch.
---
xen/arch/x86/domain.c | 10 ++++++++++
xen/arch/x86/domain_page.c | 6 +++---
xen/arch/x86/hvm/hvm.c | 5 -----
xen/arch/x86/include/asm/mm.h | 2 +-
xen/arch/x86/mm.c | 17 +++++++++--------
xen/arch/x86/pv/domain.c | 2 +-
xen/arch/x86/x86_64/mm.c | 2 +-
7 files changed, 25 insertions(+), 19 deletions(-)
diff --git a/xen/arch/x86/domain.c b/xen/arch/x86/domain.c
index efa72cd2f1..1f75d44fe0 100644
--- a/xen/arch/x86/domain.c
+++ b/xen/arch/x86/domain.c
@@ -517,6 +517,16 @@ int arch_vcpu_create(struct vcpu *v)
if ( !is_idle_domain(d) )
{
+ /*
+ * Make sure the per-domain L3 exists ahead of any consumer (e.g.
+ * init_xen_l4_slots() for the HVM monitor tables): with
+ * create_perdomain_mapping() taking a vCPU this can no longer be
+ * done when creating the domain.
+ */
+ rc = create_perdomain_mapping(v, PERDOMAIN_VIRT_START, 0, false);
+ if ( rc )
+ return rc;
+
paging_vcpu_init(v);
if ( (rc = vcpu_init_fpu(v)) != 0 )
diff --git a/xen/arch/x86/domain_page.c b/xen/arch/x86/domain_page.c
index 449d4f2a7d..9b375e438a 100644
--- a/xen/arch/x86/domain_page.c
+++ b/xen/arch/x86/domain_page.c
@@ -293,14 +293,14 @@ int mapcache_vcpu_init(struct vcpu *v)
if ( ents > dcache->entries )
{
/* Populate page tables. */
- int rc = create_perdomain_mapping(d, MAPCACHE_VIRT_START, ents, false);
+ int rc = create_perdomain_mapping(v, MAPCACHE_VIRT_START, ents, false);
/* Populate bit maps. */
if ( !rc )
- rc = create_perdomain_mapping(d, (unsigned long)dcache->inuse,
+ rc = create_perdomain_mapping(v, (unsigned long)dcache->inuse,
nr, true);
if ( !rc )
- rc = create_perdomain_mapping(d, (unsigned long)dcache->garbage,
+ rc = create_perdomain_mapping(v, (unsigned long)dcache->garbage,
nr, true);
if ( rc )
diff --git a/xen/arch/x86/hvm/hvm.c b/xen/arch/x86/hvm/hvm.c
index cd425c3342..a41ae35374 100644
--- a/xen/arch/x86/hvm/hvm.c
+++ b/xen/arch/x86/hvm/hvm.c
@@ -620,10 +620,6 @@ int hvm_domain_initialise(struct domain *d,
INIT_LIST_HEAD(&d->arch.hvm.mmcfg_regions);
INIT_LIST_HEAD(&d->arch.hvm.msix_tables);
- rc = create_perdomain_mapping(d, PERDOMAIN_VIRT_START, 0, false);
- if ( rc )
- goto fail;
-
hvm_init_cacheattr_region_list(d);
rc = paging_enable(d, PG_refcounts|PG_translate|PG_external);
@@ -730,7 +726,6 @@ int hvm_domain_initialise(struct domain *d,
XFREE(d->arch.hvm.irq);
fail0:
hvm_destroy_cacheattr_region_list(d);
- fail:
hvm_domain_relinquish_resources(d);
XFREE(d->arch.hvm.io_handler);
XFREE(d->arch.hvm.pl_time);
diff --git a/xen/arch/x86/include/asm/mm.h b/xen/arch/x86/include/asm/mm.h
index 9a8fda782e..97924a639b 100644
--- a/xen/arch/x86/include/asm/mm.h
+++ b/xen/arch/x86/include/asm/mm.h
@@ -600,7 +600,7 @@ long arch_memory_op(unsigned long cmd, XEN_GUEST_HANDLE_PARAM(void) arg);
long subarch_memory_op(unsigned long cmd, XEN_GUEST_HANDLE_PARAM(void) arg);
int compat_arch_memory_op(unsigned long cmd, XEN_GUEST_HANDLE_PARAM(void) arg);
-int create_perdomain_mapping(struct domain *d, unsigned long va,
+int create_perdomain_mapping(struct vcpu *v, unsigned long va,
unsigned int nr, bool populate);
void populate_perdomain_mapping(const struct vcpu *v, unsigned long va,
const mfn_t *mfn, unsigned int nr,
diff --git a/xen/arch/x86/mm.c b/xen/arch/x86/mm.c
index fc524ef0c3..6dfd75475a 100644
--- a/xen/arch/x86/mm.c
+++ b/xen/arch/x86/mm.c
@@ -6212,13 +6212,13 @@ static bool perdomain_l1e_needs_freeing(l1_pgentry_t l1e)
}
/*
- * Ensure the paging structure for [va, va + nr * PAGE_SIZE) of d's
- * per-domain area is in place, allocating whichever levels are missing:
- * the (domain-wide) L3 root, the slot's L2, and all L1 tables covering
- * the range. All allocations come from the domain heap. The range must
- * lie within a single per-domain slot (one L3 entry), and already-present
- * levels and entries are left untouched, so calls are idempotent over
- * existing ranges.
+ * Ensure the paging structure for [va, va + nr * PAGE_SIZE) of the
+ * per-domain area of v's domain is in place, allocating whichever levels
+ * are missing: the (domain-wide) L3 root, the slot's L2, and all L1
+ * tables covering the range. All allocations come from the domain heap.
+ * The range must lie within a single per-domain slot (one L3 entry), and
+ * already-present levels and entries are left untouched, so calls are
+ * idempotent over existing ranges.
*
* nr == 0: only ensure the per-domain L3 itself exists; populate is
* ignored. Used to set the area up before any sub-range is known.
@@ -6234,9 +6234,10 @@ static bool perdomain_l1e_needs_freeing(l1_pgentry_t l1e)
* perdomain_l1e_needs_freeing()), whereas caller-owned mappings are only
* ever unmapped.
*/
-int create_perdomain_mapping(struct domain *d, unsigned long va,
+int create_perdomain_mapping(struct vcpu *v, unsigned long va,
unsigned int nr, bool populate)
{
+ struct domain *d = v->domain;
struct page_info *pg;
l3_pgentry_t *l3tab;
l2_pgentry_t *l2tab;
diff --git a/xen/arch/x86/pv/domain.c b/xen/arch/x86/pv/domain.c
index 40b834e1a4..50f2d1284a 100644
--- a/xen/arch/x86/pv/domain.c
+++ b/xen/arch/x86/pv/domain.c
@@ -313,7 +313,7 @@ int switch_compat(struct domain *d)
static int pv_create_gdt_ldt_l1tab(struct vcpu *v)
{
- return create_perdomain_mapping(v->domain, GDT_VIRT_START(v),
+ return create_perdomain_mapping(v, GDT_VIRT_START(v),
1U << GDT_LDT_VCPU_SHIFT, false);
}
diff --git a/xen/arch/x86/x86_64/mm.c b/xen/arch/x86/x86_64/mm.c
index aa74acec82..bc7e49f4de 100644
--- a/xen/arch/x86/x86_64/mm.c
+++ b/xen/arch/x86/x86_64/mm.c
@@ -732,7 +732,7 @@ void __init zap_low_mappings(void)
int setup_compat_arg_xlat(struct vcpu *v)
{
- return create_perdomain_mapping(v->domain, ARG_XLAT_START(v),
+ return create_perdomain_mapping(v, ARG_XLAT_START(v),
PFN_UP(COMPAT_ARG_XLAT_SIZE), true);
}
--
2.55.0
^ permalink raw reply related [flat|nested] 44+ messages in thread* [PATCH v2 12/14] x86/spec-ctrl: introduce Address Space Isolation command line option
2026-09-02 9:43 [PATCH v2 00/14] x86: Address Space Isolation, part 2: asi= option and per-vCPU page tables George Dunlap
` (10 preceding siblings ...)
2026-09-02 9:43 ` [PATCH v2 11/14] x86/mm: prepare create_perdomain_mapping() for per-vCPU perdomain areas George Dunlap
@ 2026-09-02 9:43 ` George Dunlap
2026-09-02 9:43 ` [PATCH v2 13/14] x86/pv: clear the XPTI root_pgt per-domain slot on context-switch out George Dunlap
2026-09-02 9:43 ` [PATCH v2 14/14] x86/mm: introduce per-vCPU L3 page-table George Dunlap
13 siblings, 0 replies; 44+ messages in thread
From: George Dunlap @ 2026-09-02 9:43 UTC (permalink / raw)
To: xen-devel
Cc: Roger Pau Monné, Jan Beulich, Andrew Cooper,
Roger Pau Monné, Alejandro Vallejo, Teddy Astie,
Anthony PERARD, Michal Orzel, Julien Grall, Stefano Stabellini,
George Dunlap
From: Roger Pau Monné <roger.pau@citrix.com>
Introduce the `asi=` command line option, and the opt_vcpu_pt_{pv,hwdom,hvm}
knobs plus the per-domain d->arch.vcpu_pt setting it controls. The option is
introduced ahead of the functionality it enables, so that the newly added
code can be keyed on it from the start; all knobs currently default to off,
and enabling any of them taints the boot with a "not functional, development
purposes only" warning.
XPTI and per-vCPU page-tables are mutually exclusive (they are different
answers to the same problem, and the entry paths can only be built for one of
them at a time), so an explicit XPTI request takes precedence over vCPU-PT,
per axis: xpti=dom0 clears the hardware domain vCPU-PT knob (when dom0 is
PV), and xpti=domu clears the PV domU one. When XPTI is left to default, it
is turned off for those domain kinds that use vCPU-PT instead.
The boot log gains "ASI features for ..." lines for Dom0, HVM and PV
domains, so hardware-domain-only configurations remain visible, and the XPTI
line is printed unconditionally: users expecting to assert the state of XPTI
should not need to derive it from the ASI configuration.
Further per-mechanism tokens arrive with their mechanisms in later
patches.
Signed-off-by: Roger Pau Monné <roger.pau@citrix.com>
Assisted-by: Claude Code:claude-fable-5, Claude Code:claude-opus-4-8
Signed-off-by: George Dunlap <gwd@xenproject.org>
---
Changes in v2:
- Added to the series
Changes since the previously posted version:
- Make the XPTI exclusion per-axis: an explicit xpti=dom0 previously
cleared only the PV domU vCPU-PT knob, leaving dom0 with both XPTI
and vCPU-PT enabled.
- Include the hardware domain in the development warning and the boot
log summary.
- Print the XPTI status line unconditionally.
- Make the opt_vcpu_pt_* knobs plain booleans preinitialised to false,
dropping the late -1 resolution.
- Documentation: mention possible protection against unmitigated
attacks, use {pv,hvm} notation in the synopsis, and state that
pv=/hvm= do not affect the hardware domain.
- Rewrite the commit message.
---
docs/misc/xen-command-line.pandoc | 24 ++++++
xen/arch/x86/include/asm/domain.h | 3 +
xen/arch/x86/include/asm/spec_ctrl.h | 2 +
xen/arch/x86/spec_ctrl.c | 107 ++++++++++++++++++++++++++-
4 files changed, 134 insertions(+), 2 deletions(-)
diff --git a/docs/misc/xen-command-line.pandoc b/docs/misc/xen-command-line.pandoc
index 1c711fa980..834f6c57f2 100644
--- a/docs/misc/xen-command-line.pandoc
+++ b/docs/misc/xen-command-line.pandoc
@@ -202,6 +202,30 @@ to appropriate auditing by Xen. Argo is disabled by default.
This option is disabled by default, to protect domains from a DoS by a
buggy or malicious other domain spamming the ring.
+### asi (x86)
+> `= List of [ <bool>, pv=<bool>, hvm=<bool>,
+> vcpu-pt=<bool> | vcpu-pt={pv,hvm}=<bool> ]`
+
+> Default: `false`
+
+Offers control over whether the hypervisor will engage in Address Space
+Isolation, by not having potentially sensitive information permanently mapped
+in the VMM page-tables. Using this option might avoid the need to apply
+mitigations for certain speculative related attacks, at the cost of mapping
+sensitive information on-demand. It might also offer some protection against
+unmitigated speculation-related attacks.
+
+* `pv=` and `hvm=` sub-options allow enabling for specific guest types; they
+ do not affect the hardware domain, which follows the whole-feature forms
+ (the plain boolean, or an un-suffixed `vcpu-pt=<bool>`).
+
+**WARNING: manual de-selection of enabled options will invalidate any
+protection offered by the feature. The fine grained options provided below
+are meant to be used for debugging purposes only.**
+
+* `vcpu-pt` ensures each vCPU uses a unique top-level page-table and sets up
+ a virtual address space region to map memory on a per-vCPU basis.
+
### asid (x86)
> `= <boolean>`
diff --git a/xen/arch/x86/include/asm/domain.h b/xen/arch/x86/include/asm/domain.h
index 38df5c376e..fdf7b205ea 100644
--- a/xen/arch/x86/include/asm/domain.h
+++ b/xen/arch/x86/include/asm/domain.h
@@ -472,6 +472,9 @@ struct arch_domain
/* Don't unconditionally inject #GP for unhandled MSRs. */
bool msr_relaxed;
+ /* Use a per-vCPU root pt, and switch per-domain slot to per-vCPU. */
+ bool vcpu_pt;
+
/* Emulated devices enabled bitmap. */
uint32_t emulation_flags;
} __cacheline_aligned;
diff --git a/xen/arch/x86/include/asm/spec_ctrl.h b/xen/arch/x86/include/asm/spec_ctrl.h
index 8f82533c41..770c24f6a0 100644
--- a/xen/arch/x86/include/asm/spec_ctrl.h
+++ b/xen/arch/x86/include/asm/spec_ctrl.h
@@ -87,6 +87,8 @@ extern uint8_t default_scf;
extern int8_t opt_xpti_hwdom, opt_xpti_domu;
+extern bool opt_vcpu_pt_pv, opt_vcpu_pt_hwdom, opt_vcpu_pt_hvm;
+
extern bool cpu_has_bug_l1tf;
extern int8_t opt_pv_l1tf_hwdom, opt_pv_l1tf_domu;
extern bool opt_bp_spec_reduce;
diff --git a/xen/arch/x86/spec_ctrl.c b/xen/arch/x86/spec_ctrl.c
index bc8538a56f..55a02212a9 100644
--- a/xen/arch/x86/spec_ctrl.c
+++ b/xen/arch/x86/spec_ctrl.c
@@ -86,6 +86,14 @@ bool __ro_after_init opt_bp_spec_reduce = true;
static bool __initdata opt_ibpb_alt;
+/*
+ * Use a per-vCPU root page-table and switch the per-domain slot to per-vCPU.
+ * Off by default until the feature is complete.
+ */
+bool __ro_after_init opt_vcpu_pt_hvm;
+bool __ro_after_init opt_vcpu_pt_hwdom;
+bool __ro_after_init opt_vcpu_pt_pv;
+
static int __init cf_check parse_spec_ctrl(const char *s)
{
const char *ss;
@@ -383,6 +391,18 @@ int8_t __ro_after_init opt_xpti_domu = -1;
static __init void xpti_init_default(void)
{
+ if ( !opt_dom0_pvh && opt_xpti_hwdom == 1 && opt_vcpu_pt_hwdom )
+ {
+ printk(XENLOG_ERR
+ "XPTI incompatible with per-vCPU page-tables, disabling Dom0 vCPU-PT\n");
+ opt_vcpu_pt_hwdom = false;
+ }
+ if ( opt_xpti_domu == 1 && opt_vcpu_pt_pv )
+ {
+ printk(XENLOG_ERR
+ "XPTI incompatible with per-vCPU page-tables, disabling PV DomU vCPU-PT\n");
+ opt_vcpu_pt_pv = false;
+ }
if ( (boot_cpu_data.vendor & (X86_VENDOR_AMD | X86_VENDOR_HYGON)) ||
cpu_has_rdcl_no )
{
@@ -394,9 +414,9 @@ static __init void xpti_init_default(void)
else
{
if ( opt_xpti_hwdom < 0 )
- opt_xpti_hwdom = 1;
+ opt_xpti_hwdom = !opt_vcpu_pt_hwdom;
if ( opt_xpti_domu < 0 )
- opt_xpti_domu = 1;
+ opt_xpti_domu = !opt_vcpu_pt_pv;
}
}
@@ -487,6 +507,66 @@ static int __init cf_check parse_pv_l1tf(const char *s)
}
custom_param("pv-l1tf", parse_pv_l1tf);
+static int __init cf_check parse_asi(const char *s)
+{
+ const char *ss;
+ int val, rc = 0;
+
+ /* Interpret 'asi' alone in its positive boolean form. */
+ if ( *s == '\0' )
+ opt_vcpu_pt_pv = opt_vcpu_pt_hwdom = opt_vcpu_pt_hvm = true;
+
+ do {
+ ss = strchr(s, ',');
+ if ( !ss )
+ ss = strchr(s, '\0');
+
+ val = parse_bool(s, ss);
+ switch ( val )
+ {
+ case 0:
+ case 1:
+ opt_vcpu_pt_pv = opt_vcpu_pt_hwdom = opt_vcpu_pt_hvm = val;
+ break;
+
+ default:
+ if ( (val = parse_boolean("pv", s, ss)) >= 0 )
+ opt_vcpu_pt_pv = val;
+ else if ( (val = parse_boolean("hvm", s, ss)) >= 0 )
+ opt_vcpu_pt_hvm = val;
+ else if ( (val = parse_boolean("vcpu-pt", s, ss)) != -1 )
+ {
+ switch ( val )
+ {
+ case 1:
+ case 0:
+ opt_vcpu_pt_pv = opt_vcpu_pt_hvm = opt_vcpu_pt_hwdom = val;
+ break;
+
+ case -2:
+ s += strlen("vcpu-pt=");
+ if ( (val = parse_boolean("pv", s, ss)) >= 0 )
+ opt_vcpu_pt_pv = val;
+ else if ( (val = parse_boolean("hvm", s, ss)) >= 0 )
+ opt_vcpu_pt_hvm = val;
+ else
+ default:
+ rc = -EINVAL;
+ break;
+ }
+ }
+ else if ( *s )
+ rc = -EINVAL;
+ break;
+ }
+
+ s = ss + 1;
+ } while ( *ss );
+
+ return rc;
+}
+custom_param("asi", parse_asi);
+
static void __init print_details(enum ind_thunk thunk)
{
unsigned int _7d0 = 0, _7d2 = 0, e8b = 0, e21a = 0, e21c = 0, max = 0, tmp;
@@ -680,6 +760,20 @@ static void __init print_details(enum ind_thunk thunk)
opt_pv_l1tf_hwdom ? "enabled" : "disabled",
opt_pv_l1tf_domu ? "enabled" : "disabled");
#endif
+
+ printk(" ASI features for Dom0:%s%s\n",
+ opt_vcpu_pt_hwdom ? "" : " None",
+ opt_vcpu_pt_hwdom ? " vCPU-PT" : "");
+#ifdef CONFIG_HVM
+ printk(" ASI features for HVM VMs:%s%s\n",
+ opt_vcpu_pt_hvm ? "" : " None",
+ opt_vcpu_pt_hvm ? " vCPU-PT" : "");
+#endif
+#ifdef CONFIG_PV
+ printk(" ASI features for PV VMs:%s%s\n",
+ opt_vcpu_pt_pv ? "" : " None",
+ opt_vcpu_pt_pv ? " vCPU-PT" : "");
+#endif
}
static bool __init check_smt_enabled(void)
@@ -1866,6 +1960,10 @@ void spec_ctrl_init_domain(struct domain *d)
if ( pv )
d->arch.pv.xpti = is_hardware_domain(d) ? opt_xpti_hwdom
: opt_xpti_domu;
+
+ d->arch.vcpu_pt = is_hardware_domain(d) ? opt_vcpu_pt_hwdom
+ : pv ? opt_vcpu_pt_pv
+ : opt_vcpu_pt_hvm;
}
void __init init_speculation_mitigations(void)
@@ -2158,6 +2256,11 @@ void __init init_speculation_mitigations(void)
hw_smt_enabled && default_xen_spec_ctrl )
setup_force_cpu_cap(X86_FEATURE_SC_MSR_IDLE);
+ if ( opt_vcpu_pt_pv || opt_vcpu_pt_hwdom || opt_vcpu_pt_hvm )
+ warning_add(
+ "Address Space Isolation is not functional, this option is\n"
+ "intended to be used only for development purposes.\n");
+
xpti_init_default();
l1tf_calculations();
--
2.55.0
^ permalink raw reply related [flat|nested] 44+ messages in thread* [PATCH v2 13/14] x86/pv: clear the XPTI root_pgt per-domain slot on context-switch out
2026-09-02 9:43 [PATCH v2 00/14] x86: Address Space Isolation, part 2: asi= option and per-vCPU page tables George Dunlap
` (11 preceding siblings ...)
2026-09-02 9:43 ` [PATCH v2 12/14] x86/spec-ctrl: introduce Address Space Isolation command line option George Dunlap
@ 2026-09-02 9:43 ` George Dunlap
2026-09-02 9:43 ` [PATCH v2 14/14] x86/mm: introduce per-vCPU L3 page-table George Dunlap
13 siblings, 0 replies; 44+ messages in thread
From: George Dunlap @ 2026-09-02 9:43 UTC (permalink / raw)
To: xen-devel
Cc: George Dunlap, Jan Beulich, Andrew Cooper, Roger Pau Monné,
Alejandro Vallejo, Teddy Astie, Anthony PERARD, Michal Orzel,
Julien Grall, Stefano Stabellini
From: George Dunlap <gwd@xenproject.org>
XPTI maintains a per-pCPU root page table (root_pgt): a restricted L4,
with Xen largely unmapped, that an XPTI domain's guest context actually
runs on. Its guest mappings are copied in on the way back to guest
context; its per-domain slot is written by paravirt_ctxt_switch_to(),
so that it follows whichever domain is scheduled onto the pCPU.
Nothing ever clears the slot, however. When the pCPU switches from an
XPTI PV vCPU to one that does not refresh the slot -- an HVM vCPU, or
the idle vCPU after a lazy state flush -- the last PV domain's
per-domain L3 remains referenced from root_pgt. The reference isn't
cleared on domain destruction, so could even point to an already freed
page.
In theory, that slot should never be walked in this state; but it's
just generally safer not to leave dangling references around.
Consider that cleanup_cpu_root_pgt() frees pagetables by walking
root_pgt on CPU offline. Currently it correctly leaves slot 260 alone;
but one could easily imagine a mistake in which slot 260 is walked
erroneously.
Clear the slot in paravirt_ctxt_switch_from(), making the maintenance
a pair: cleared on the way out, installed on the way in. The slot is
now populated only while the vCPU using it runs, and an erroneous walk
in any other state faults cleanly. This change also makes robust
behavior simpler when we add per-vCPU areas in a subsequent patch.
Assisted-by: Claude Code:claude-fable-5
Signed-off-by: George Dunlap <gwd@xenproject.org>
---
Changes in v2:
- New patch.
---
xen/arch/x86/domain.c | 16 ++++++++++++++++
1 file changed, 16 insertions(+)
diff --git a/xen/arch/x86/domain.c b/xen/arch/x86/domain.c
index 1f75d44fe0..79555e6964 100644
--- a/xen/arch/x86/domain.c
+++ b/xen/arch/x86/domain.c
@@ -2006,8 +2006,18 @@ static void save_segments(struct vcpu *v)
void cf_check paravirt_ctxt_switch_from(struct vcpu *v)
{
+ root_pgentry_t *root_pgt = this_cpu(root_pgt);
+
save_segments(v);
+ /*
+ * Clear the XPTI per-domain slot: it is installed on the way in by
+ * paravirt_ctxt_switch_to(), and must not linger while another vCPU
+ * runs.
+ */
+ if ( root_pgt )
+ root_pgt[root_table_offset(PERDOMAIN_VIRT_START)] = l4e_empty();
+
/*
* Disable debug breakpoints. We do this aggressively because if we switch
* to an HVM guest we may load DR0-DR3 with values that can cause #DE
@@ -2022,6 +2032,12 @@ void cf_check paravirt_ctxt_switch_to(struct vcpu *v)
{
root_pgentry_t *root_pgt = this_cpu(root_pgt);
+ /*
+ * If XPTI is active, install the incoming domain's per-domain area
+ * in the per-domain slot of the L4 we run on while in guest mode.
+ * The slot was cleared on the way out (see
+ * paravirt_ctxt_switch_from()).
+ */
if ( root_pgt )
root_pgt[root_table_offset(PERDOMAIN_VIRT_START)] =
l4e_from_page(v->domain->arch.perdomain_l3_pg,
--
2.55.0
^ permalink raw reply related [flat|nested] 44+ messages in thread* [PATCH v2 14/14] x86/mm: introduce per-vCPU L3 page-table
2026-09-02 9:43 [PATCH v2 00/14] x86: Address Space Isolation, part 2: asi= option and per-vCPU page tables George Dunlap
` (12 preceding siblings ...)
2026-09-02 9:43 ` [PATCH v2 13/14] x86/pv: clear the XPTI root_pgt per-domain slot on context-switch out George Dunlap
@ 2026-09-02 9:43 ` George Dunlap
13 siblings, 0 replies; 44+ messages in thread
From: George Dunlap @ 2026-09-02 9:43 UTC (permalink / raw)
To: xen-devel
Cc: Roger Pau Monné, Jan Beulich, Andrew Cooper,
Roger Pau Monné, Alejandro Vallejo, Teddy Astie,
Anthony PERARD, Michal Orzel, Julien Grall, Stefano Stabellini,
George Dunlap
From: Roger Pau Monné <roger.pau@citrix.com>
The per-domain area is currently a single domain-wide structure: one L3,
referenced from every root page-table associated with the domain, so
every mapping in it is visible to every vCPU of the domain. Its
contents are already laid out in per-vCPU slices (each vCPU's GDT/LDT
window, COMPAT_ARG_XLAT pages, and mapcache entries), but the visibility
is domain-wide. Meanwhile, much of the per-vCPU state Xen maintains --
the VMCB, the VMX MSR load/save areas, FPU/XSAVE state -- lives in
always-mapped memory, and so do the pCPU stacks.
Allow the "per-domain" area to be per-vCPU instead ("VCPU-PT"). This
will immediately isolate the existing per-vCPU mappings from other
running vCPUs on HVM domains (PV SMP for VCPU-PT requires further
work; see below). We will later build on this, adding per-vCPU mapped
areas (into which we can put vCPU state currently in the xenheap,
mentioned above); per-vCPU mapcaches (which will eventually allow us
to remove domheap pages from the direct map), and finally transient
mappings of the pCPU stack on which the vCPU is currently running.
Add pervcpu_l3_pg to the arch_vcpu struct, to correspond to the
perdomain_l3_pg in the domain struct. (We retain both so that we can
switch between per-domain and per-vCPU on a domain-by-domain basis.)
In {create,populate,destroy}_perdomain_mapping(), if d->arch.vcpu_pt,
use pervcpu_l3_pg as the per-domain L3 (allocating it for a vCPU if
it's NULL, just as we allocate for a domain in !vcpu_pt mode);
otherwise, use perdomain_l3_pg. Introduce a helper, perdomain_l3(),
to consistently choose the correct one.
Introduce free_pervcpu_mappings() to free this tree, called from
arch_vcpu_destroy() on normal teardown. Since the vcpu structure holds
the only reference to pervcpu_l3_pg, arch_vcpu_create() must also call
it on its error paths: nothing else records the allocation once the
vcpu struct is torn down. The domain-wide free_perdomain_mappings() is
unchanged and keeps covering non-vCPU-PT domains.
Modify init_xen_l4_slots() to take a vCPU, and use perdomain_l3() to
select the value to install in slot 260. Most callers have the
specific vCPU in hand (the HVM monitor tables, setup_compat_l4(), PV
shadow L4s). Note that this includes dom0_construct(), since at that
point we're actually building vCPU 0, so passing in d->vcpu[0] is
exactly what we want.
In promote_l4_table() we pass in d->vcpu[0]. This is correct without
vCPU-PT, where every vCPU selects the same domain-wide L3. With
vCPU-PT this is a temporary arrangement: the promoted L4 carries vCPU
0's L3 until the per-pCPU shadow L4 -- which supersedes promoted L4s
as what the CPU actually runs on -- arrives in the series (see
the SMP note below). Since slot 260 is now keyed off d->vcpu[0],
promote_l4_table() refuses (-EINVAL) a domain that has no vCPUs yet,
as can happen if a toolstack pins page tables before creating vCPUs.
paravirt_ctxt_switch_to() now installs the XPTI root_pgt per-domain
slot only when the domain has a domain-wide perdomain area to install.
XPTI and vCPU-PT are mutually exclusive (xpti_init_default() disables
vCPU-PT if both are explicitly requested, and each defaults off when
the other is on) -- so the vCPU-PT slot stays empty, as
paravirt_ctxt_switch_from() left it.
vCPU-PT is not currently implemented for shadow paging. L4 shadows
are currently per-domain objects shared by all vCPUs shadowing the
same guest root, just as non-ASI non-shadow PV L4s are. Enabling
vCPU-PT for PV shadow guests would require adding vCPU-PT
functionality along all the shadow paths, which is outside the scope
of the current work.
This is guarded on every path that can turn shadow on for a PV domain:
paging_domctl() refuses shadow/log-dirty ops (xl save/migrate) for
such domains; dom0=shadow is ignored with a warning when dom0 uses
vCPU-PT; and shadow_one_bit_enable() refuses the mode with
-EOPNOTSUPP.
The PV L1TF mitigation cannot be refused up front: it acts at runtime,
when a guest installs a not-present PTE whose address is unsafe, by
forcing the domain into shadow mode -- which is what a vCPU-PT domain
cannot currently have. No special handling is needed, though:
pv_l1tf_check_pte() refuses the PTE write and schedules the shadowing
tasklet as usual; the tasklet's shadow_one_bit_enable() call fails
with -EOPNOTSUPP like any other enable failure; and the tasklet's
existing error handling crashes the domain. That is the right
disposition -- the entry being installed is precisely what the
mitigation exists to catch, so continuing unmitigated is not an option
-- and it matches what a build without CONFIG_SHADOW_PAGING does for
the same write, with a log trail showing the mitigation was attempted
and could not be enabled. Hardware without the erratum is unaffected,
the mitigation being off there by default.
Note SMP vCPU-PT PV guests are not yet functional at this point in the
series: promoted guest L4s embed vCPU#0's L3 in slot 260 for all
vCPUs; the per-pCPU shadow L4 that gives each vCPU its own slot 260
arrives with the guest_root_pt and per-pCPU-L4 patches later in the
series. HVM vCPU-PT guests are fully functional, SMP included:
monitor tables are already per-vCPU, so every HVM vCPU's root carries
its own L3 from creation.
Signed-off-by: Roger Pau Monné <roger.pau@citrix.com>
Assisted-by: Claude Code:claude-fable-5, Claude Code:claude-opus-4-8
Signed-off-by: George Dunlap <gwd@xenproject.org>
---
Changes in v2:
- Added to the series
Changes since the previously posted version:
- Expand the commit message with the motivation and the design
rationale.
- pv-l1tf: let the mitigation's shadowing request fail in
shadow_one_bit_enable() and rely on the tasklet's existing failure
handling to crash the domain, rather than special-casing vCPU-PT in
pv_l1tf_check_pte() or forcing the mitigation off at boot (an
earlier revision did the latter, silently withdrawing a protection
that is on by default on affected hardware).
- Keep domain-wide freeing intact and introduce a vCPU-scoped
free_pervcpu_mappings() instead of re-scoping
free_perdomain_mappings(); fix the arch_vcpu_create() error-path
leaks of a partially built per-vCPU hierarchy.
- Exclude PV shadow for vCPU-PT domains on all enable paths:
refuse shadow/log-dirty paging_domctl() ops (gate
moved here from the later per-pCPU-L4 patch so hazard and gate land
together), ignore dom0=shadow with a warning, and refuse the mode
(-EOPNOTSUPP) in shadow_one_bit_enable().
- Add a perdomain_l3() helper for the root selection, rather than
open-coding the vcpu_pt choice (and testing both root pointers) at
each site.
- Guard promote_l4_table() against vCPU-less domains.
- Install the XPTI root_pgt per-domain slot only when the domain has
a perdomain L3; the posted version computed an L4E from the NULL
pointer for vCPU-PT domains. (A new preparatory patch pairs the
slot's maintenance with a switch-out clear.)
- Do not log the refusal for XEN_DOMCTL_SHADOW_OP_OFF. Turning paging
off is the de-facto "make sure it is off" interface: the save path
issues it unconditionally as best-effort cleanup and discards the
result, so every save of a PV domain otherwise printed a hypervisor
error for an operation nothing was asking to succeed.
---
xen/arch/x86/domain.c | 21 +++++---
xen/arch/x86/include/asm/domain.h | 12 +++++
xen/arch/x86/include/asm/mm.h | 3 +-
xen/arch/x86/mm.c | 87 +++++++++++++++++++++++--------
xen/arch/x86/mm/hap/hap.c | 2 +-
xen/arch/x86/mm/paging.c | 14 +++++
xen/arch/x86/mm/shadow/common.c | 11 ++++
xen/arch/x86/mm/shadow/hvm.c | 2 +-
xen/arch/x86/mm/shadow/multi.c | 2 +-
xen/arch/x86/pv/dom0_build.c | 8 ++-
xen/arch/x86/pv/domain.c | 2 +-
11 files changed, 126 insertions(+), 38 deletions(-)
diff --git a/xen/arch/x86/domain.c b/xen/arch/x86/domain.c
index 79555e6964..3e571272d8 100644
--- a/xen/arch/x86/domain.c
+++ b/xen/arch/x86/domain.c
@@ -513,7 +513,7 @@ int arch_vcpu_create(struct vcpu *v)
rc = mapcache_vcpu_init(v);
if ( rc )
- return rc;
+ goto fail_early;
if ( !is_idle_domain(d) )
{
@@ -525,12 +525,12 @@ int arch_vcpu_create(struct vcpu *v)
*/
rc = create_perdomain_mapping(v, PERDOMAIN_VIRT_START, 0, false);
if ( rc )
- return rc;
+ goto fail_early;
paging_vcpu_init(v);
if ( (rc = vcpu_init_fpu(v)) != 0 )
- return rc;
+ goto fail_early;
vmce_init_vcpu(v);
@@ -578,6 +578,8 @@ int arch_vcpu_create(struct vcpu *v)
vcpu_destroy_fpu(v);
xfree(v->arch.msrs);
v->arch.msrs = NULL;
+ fail_early:
+ free_pervcpu_mappings(v);
return rc;
}
@@ -598,6 +600,8 @@ void arch_vcpu_destroy(struct vcpu *v)
pv_vcpu_destroy(v);
else
ASSERT_UNREACHABLE();
+
+ free_pervcpu_mappings(v);
}
int arch_sanitise_domain_config(struct xen_domctl_createdomain *config)
@@ -2033,12 +2037,13 @@ void cf_check paravirt_ctxt_switch_to(struct vcpu *v)
root_pgentry_t *root_pgt = this_cpu(root_pgt);
/*
- * If XPTI is active, install the incoming domain's per-domain area
- * in the per-domain slot of the L4 we run on while in guest mode.
- * The slot was cleared on the way out (see
- * paravirt_ctxt_switch_from()).
+ * If XPTI is active and the domain has a domain-wide perdomain area,
+ * install it in the per-domain slot of the L4 we run on while in
+ * guest mode. vCPU-PT domains have none (they don't use the XPTI
+ * machinery); their slot stays as paravirt_ctxt_switch_from() left
+ * it: empty.
*/
- if ( root_pgt )
+ if ( root_pgt && v->domain->arch.perdomain_l3_pg )
root_pgt[root_table_offset(PERDOMAIN_VIRT_START)] =
l4e_from_page(v->domain->arch.perdomain_l3_pg,
__PAGE_HYPERVISOR_RW);
diff --git a/xen/arch/x86/include/asm/domain.h b/xen/arch/x86/include/asm/domain.h
index fdf7b205ea..5c90d7627f 100644
--- a/xen/arch/x86/include/asm/domain.h
+++ b/xen/arch/x86/include/asm/domain.h
@@ -330,6 +330,11 @@ struct monitor_write_data {
struct arch_domain
{
+ /*
+ * Domain-wide L3 page-table for the L4 per-domain slot, used when
+ * the domain does not use per-vCPU page-tables (!d->arch.vcpu_pt).
+ * NULL otherwise (see v->arch.pervcpu_l3_pg and perdomain_l3()).
+ */
struct page_info *perdomain_l3_pg;
/* I/O-port admin-specified access capabilities. */
@@ -678,6 +683,13 @@ struct arch_vcpu
struct vcpu_msrs *msrs;
+ /*
+ * Per-vCPU L3 page-table for the L4 per-domain slot, used when the
+ * domain uses per-vCPU page-tables (d->arch.vcpu_pt). NULL
+ * otherwise (see d->arch.perdomain_l3_pg and perdomain_l3()).
+ */
+ struct page_info *pervcpu_l3_pg;
+
struct {
bool next_interrupt_enabled;
} monitor;
diff --git a/xen/arch/x86/include/asm/mm.h b/xen/arch/x86/include/asm/mm.h
index 97924a639b..acb553df48 100644
--- a/xen/arch/x86/include/asm/mm.h
+++ b/xen/arch/x86/include/asm/mm.h
@@ -370,7 +370,7 @@ int devalidate_page(struct page_info *page, unsigned long type,
void init_xen_pae_l2_slots(l2_pgentry_t *l2t, const struct domain *d);
void init_xen_l4_slots(l4_pgentry_t *l4t, mfn_t l4mfn,
- const struct domain *d, mfn_t sl4mfn, bool ro_mpt);
+ const struct vcpu *v, mfn_t sl4mfn, bool ro_mpt);
bool fill_ro_mpt(mfn_t mfn);
void zap_ro_mpt(mfn_t mfn);
@@ -608,6 +608,7 @@ void populate_perdomain_mapping(const struct vcpu *v, unsigned long va,
void destroy_perdomain_mapping(const struct vcpu *v, unsigned long va,
unsigned int nr);
void free_perdomain_mappings(struct domain *d);
+void free_pervcpu_mappings(struct vcpu *v);
void __iomem *ioremap_wc(paddr_t pa, size_t len);
diff --git a/xen/arch/x86/mm.c b/xen/arch/x86/mm.c
index 6dfd75475a..48266b1d27 100644
--- a/xen/arch/x86/mm.c
+++ b/xen/arch/x86/mm.c
@@ -1638,20 +1638,32 @@ static int promote_l3_table(struct page_info *page)
}
#endif /* CONFIG_PV */
+/*
+ * The root of the per-domain area in use by @v: the vCPU's own L3 for a
+ * vCPU-PT domain, the domain-wide one otherwise.
+ */
+static struct page_info *perdomain_l3(const struct vcpu *v)
+{
+ const struct domain *d = v->domain;
+
+ return d->arch.vcpu_pt ? v->arch.pervcpu_l3_pg : d->arch.perdomain_l3_pg;
+}
+
/*
* Fill an L4 with Xen entries.
*
* This function must write all ROOT_PAGETABLE_PV_XEN_SLOTS, to clobber any
* values a guest may have left there from promote_l4_table().
*
- * l4t, l4mfn, and d are mandatory, but l4mfn doesn't need to be the mfn under
+ * l4t, l4mfn, and v are mandatory, but l4mfn doesn't need to be the mfn under
* *l4t. All other parameters are optional and will either fill or zero the
* appropriate slots. Pagetables not shared with guests will gain the
* extended directmap.
*/
void init_xen_l4_slots(l4_pgentry_t *l4t, mfn_t l4mfn,
- const struct domain *d, mfn_t sl4mfn, bool ro_mpt)
+ const struct vcpu *v, mfn_t sl4mfn, bool ro_mpt)
{
+ const struct domain *d = v->domain;
/*
* PV vcpus need a shortened directmap. HVM and Idle vcpus get the full
* directmap.
@@ -1679,7 +1691,7 @@ void init_xen_l4_slots(l4_pgentry_t *l4t, mfn_t l4mfn,
/* Slot 260: Per-domain mappings. */
l4t[l4_table_offset(PERDOMAIN_VIRT_START)] =
- l4e_from_page(d->arch.perdomain_l3_pg, __PAGE_HYPERVISOR_RW);
+ l4e_from_page(perdomain_l3(v), __PAGE_HYPERVISOR_RW);
/* Slot 4: Per-domain mappings mirror. */
BUILD_BUG_ON(IS_ENABLED(CONFIG_PV32) &&
@@ -1755,11 +1767,17 @@ static int promote_l4_table(struct page_info *page)
{
struct domain *d = page_get_owner(page);
mfn_t l4mfn = page_to_mfn(page);
- l4_pgentry_t *pl4e = map_domain_page(l4mfn);
+ l4_pgentry_t *pl4e;
unsigned int i;
int rc = 0;
unsigned int partial_flags = page->partial_flags;
+ /* init_xen_l4_slots() needs a vCPU to key the per-domain slot off. */
+ if ( unlikely(!d->vcpu || !d->vcpu[0]) )
+ return -EINVAL;
+
+ pl4e = map_domain_page(l4mfn);
+
for ( i = page->nr_validated_ptes; i < L4_PAGETABLE_ENTRIES;
i++, partial_flags = 0 )
{
@@ -1834,8 +1852,15 @@ static int promote_l4_table(struct page_info *page)
if ( !rc )
{
+ /*
+ * Use vCPU#0 unconditionally. When not running with ASI enabled the
+ * per-domain table is shared between all vCPUs, so it doesn't matter
+ * which vCPU gets passed to init_xen_l4_slots(). When running with
+ * ASI enabled this L4 will not be used, as a shadow per-vCPU L4 is
+ * used instead.
+ */
init_xen_l4_slots(pl4e, l4mfn,
- d, INVALID_MFN, VM_ASSIST(d, m2p_strict));
+ d->vcpu[0], INVALID_MFN, VM_ASSIST(d, m2p_strict));
atomic_inc(&d->arch.pv.nr_l4_pages);
}
unmap_domain_page(pl4e);
@@ -6238,7 +6263,7 @@ int create_perdomain_mapping(struct vcpu *v, unsigned long va,
unsigned int nr, bool populate)
{
struct domain *d = v->domain;
- struct page_info *pg;
+ struct page_info *pg, *l3_pg = perdomain_l3(v);
l3_pgentry_t *l3tab;
l2_pgentry_t *l2tab;
l1_pgentry_t *l1tab;
@@ -6247,14 +6272,17 @@ int create_perdomain_mapping(struct vcpu *v, unsigned long va,
ASSERT(va >= PERDOMAIN_VIRT_START &&
va < PERDOMAIN_VIRT_SLOT(PERDOMAIN_SLOTS));
- if ( !d->arch.perdomain_l3_pg )
+ if ( !l3_pg )
{
pg = alloc_domheap_page(d, MEMF_no_owner);
if ( !pg )
return -ENOMEM;
l3tab = __map_domain_page(pg);
clear_page(l3tab);
- d->arch.perdomain_l3_pg = pg;
+ if ( d->arch.vcpu_pt )
+ v->arch.pervcpu_l3_pg = pg;
+ else
+ d->arch.perdomain_l3_pg = pg;
if ( !nr )
{
unmap_domain_page(l3tab);
@@ -6264,7 +6292,7 @@ int create_perdomain_mapping(struct vcpu *v, unsigned long va,
else if ( !nr )
return 0;
else
- l3tab = __map_domain_page(d->arch.perdomain_l3_pg);
+ l3tab = __map_domain_page(l3_pg);
ASSERT(!l3_table_offset(va ^ (va + nr * PAGE_SIZE - 1)));
@@ -6359,7 +6387,7 @@ void populate_perdomain_mapping(const struct vcpu *v, unsigned long va,
l1_pgentry_t *l1tab = NULL, *pl1e;
const l3_pgentry_t *l3tab;
const l2_pgentry_t *l2tab;
- struct domain *d = v->domain;
+ struct page_info *l3_pg;
unsigned long irq_flags;
ASSERT(va >= PERDOMAIN_VIRT_START &&
@@ -6401,7 +6429,8 @@ void populate_perdomain_mapping(const struct vcpu *v, unsigned long va,
return;
}
- BUG_ON(!d->arch.perdomain_l3_pg);
+ l3_pg = perdomain_l3(v);
+ BUG_ON(!l3_pg);
/*
* Slow path: walk v's per-domain page-table pages. All mappings are
@@ -6413,7 +6442,7 @@ void populate_perdomain_mapping(const struct vcpu *v, unsigned long va,
*/
local_irq_save(irq_flags);
- l3tab = __map_domain_page_irqoff(d->arch.perdomain_l3_pg);
+ l3tab = __map_domain_page_irqoff(l3_pg);
/*
* Missing page-table structure is a hypervisor bug: there is no safe
@@ -6461,13 +6490,13 @@ void destroy_perdomain_mapping(const struct vcpu *v, unsigned long va,
unsigned int nr)
{
const l3_pgentry_t *l3tab, *pl3e;
- const struct domain *d = v->domain;
+ struct page_info *l3_pg = perdomain_l3(v);
ASSERT(va >= PERDOMAIN_VIRT_START &&
va < PERDOMAIN_VIRT_SLOT(PERDOMAIN_SLOTS));
ASSERT(!nr || !l3_table_offset(va ^ (va + nr * PAGE_SIZE - 1)));
- if ( !d->arch.perdomain_l3_pg )
+ if ( !l3_pg )
return;
if ( likely(this_cpu(pgtable_vcpu) == v) )
@@ -6490,7 +6519,7 @@ void destroy_perdomain_mapping(const struct vcpu *v, unsigned long va,
return;
}
- l3tab = __map_domain_page(d->arch.perdomain_l3_pg);
+ l3tab = __map_domain_page(l3_pg);
pl3e = l3tab + l3_table_offset(va);
if ( l3e_get_flags(*pl3e) & _PAGE_PRESENT )
@@ -6529,16 +6558,11 @@ void destroy_perdomain_mapping(const struct vcpu *v, unsigned long va,
unmap_domain_page(l3tab);
}
-void free_perdomain_mappings(struct domain *d)
+static void free_perdomain_l3(struct page_info *l3pg)
{
- l3_pgentry_t *l3tab;
+ l3_pgentry_t *l3tab = __map_domain_page(l3pg);
unsigned int i;
- if ( !d->arch.perdomain_l3_pg )
- return;
-
- l3tab = __map_domain_page(d->arch.perdomain_l3_pg);
-
for ( i = 0; i < PERDOMAIN_SLOTS; ++i)
if ( l3e_get_flags(l3tab[i]) & _PAGE_PRESENT )
{
@@ -6571,10 +6595,27 @@ void free_perdomain_mappings(struct domain *d)
}
unmap_domain_page(l3tab);
- free_domheap_page(d->arch.perdomain_l3_pg);
+ free_domheap_page(l3pg);
+}
+
+void free_perdomain_mappings(struct domain *d)
+{
+ if ( !d->arch.perdomain_l3_pg )
+ return;
+
+ free_perdomain_l3(d->arch.perdomain_l3_pg);
d->arch.perdomain_l3_pg = NULL;
}
+void free_pervcpu_mappings(struct vcpu *v)
+{
+ if ( !v->arch.pervcpu_l3_pg )
+ return;
+
+ free_perdomain_l3(v->arch.pervcpu_l3_pg);
+ v->arch.pervcpu_l3_pg = NULL;
+}
+
static void write_sss_token(unsigned long *ptr)
{
/*
diff --git a/xen/arch/x86/mm/hap/hap.c b/xen/arch/x86/mm/hap/hap.c
index 0ede4181a0..aba4f77df9 100644
--- a/xen/arch/x86/mm/hap/hap.c
+++ b/xen/arch/x86/mm/hap/hap.c
@@ -407,7 +407,7 @@ static mfn_t hap_make_monitor_table(struct vcpu *v)
m4mfn = page_to_mfn(pg);
l4e = map_domain_page(m4mfn);
- init_xen_l4_slots(l4e, m4mfn, d, INVALID_MFN, false);
+ init_xen_l4_slots(l4e, m4mfn, v, INVALID_MFN, false);
unmap_domain_page(l4e);
return m4mfn;
diff --git a/xen/arch/x86/mm/paging.c b/xen/arch/x86/mm/paging.c
index 14ab7defd8..ab68dfa415 100644
--- a/xen/arch/x86/mm/paging.c
+++ b/xen/arch/x86/mm/paging.c
@@ -675,6 +675,20 @@ int paging_domctl(struct domain *d, struct xen_domctl_shadow_op *sc,
return -EINVAL;
}
+ if ( is_pv_domain(d) && d->arch.vcpu_pt )
+ {
+ /*
+ * Turning paging off is the de-facto "make sure it is off"
+ * interface: the save path issues it unconditionally as
+ * best-effort cleanup and discards the result, so logging an
+ * error for it is noise on every save of a PV domain.
+ */
+ if ( sc->op != XEN_DOMCTL_SHADOW_OP_OFF )
+ gprintk(XENLOG_ERR,
+ "Paging not supported on PV domains with ASI\n");
+ return -EOPNOTSUPP;
+ }
+
if ( resuming
? (d->arch.paging.preempt.dom != current->domain ||
d->arch.paging.preempt.op != sc->op)
diff --git a/xen/arch/x86/mm/shadow/common.c b/xen/arch/x86/mm/shadow/common.c
index e30c6c49e1..b559db84b1 100644
--- a/xen/arch/x86/mm/shadow/common.c
+++ b/xen/arch/x86/mm/shadow/common.c
@@ -2361,6 +2361,17 @@ static int shadow_one_bit_enable(struct domain *d, u32 mode)
return -EINVAL;
}
+ /*
+ * PV shadows embed the (per-vCPU) per-domain slot in L4 shadows shared
+ * by all vCPUs of the domain, so shadow modes are unavailable to
+ * domains using per-vCPU page-tables. Toolstack requests are refused
+ * in paging_domctl(); the pv-l1tf tasklet can still request
+ * PG_SH_forced at runtime, and crashes the domain when this refusal
+ * reaches it.
+ */
+ if ( is_pv_domain(d) && d->arch.vcpu_pt )
+ return -EOPNOTSUPP;
+
mode |= PG_SH_enable;
if ( d->arch.paging.total_pages < sh_min_allocation(d) )
diff --git a/xen/arch/x86/mm/shadow/hvm.c b/xen/arch/x86/mm/shadow/hvm.c
index e6fb97c4b6..0b23326214 100644
--- a/xen/arch/x86/mm/shadow/hvm.c
+++ b/xen/arch/x86/mm/shadow/hvm.c
@@ -760,7 +760,7 @@ mfn_t sh_make_monitor_table(const struct vcpu *v, unsigned int shadow_levels)
* shadow-linear mapping will either be inserted below when creating
* lower level monitor tables, or later in sh_update_cr3().
*/
- init_xen_l4_slots(l4e, m4mfn, d, INVALID_MFN, false);
+ init_xen_l4_slots(l4e, m4mfn, v, INVALID_MFN, false);
if ( shadow_levels < 4 )
{
diff --git a/xen/arch/x86/mm/shadow/multi.c b/xen/arch/x86/mm/shadow/multi.c
index 1ae1091acd..5b7dbc12dc 100644
--- a/xen/arch/x86/mm/shadow/multi.c
+++ b/xen/arch/x86/mm/shadow/multi.c
@@ -974,7 +974,7 @@ sh_make_shadow(struct vcpu *v, mfn_t gmfn, u32 shadow_type)
BUILD_BUG_ON(sizeof(l4_pgentry_t) != sizeof(shadow_l4e_t));
- init_xen_l4_slots(l4t, gmfn, d, smfn, (!is_pv_32bit_domain(d) &&
+ init_xen_l4_slots(l4t, gmfn, v, smfn, (!is_pv_32bit_domain(d) &&
VM_ASSIST(d, m2p_strict)));
unmap_domain_page(l4t);
}
diff --git a/xen/arch/x86/pv/dom0_build.c b/xen/arch/x86/pv/dom0_build.c
index ddeb144b06..52139cffb3 100644
--- a/xen/arch/x86/pv/dom0_build.c
+++ b/xen/arch/x86/pv/dom0_build.c
@@ -726,7 +726,7 @@ static int __init dom0_construct(const struct boot_domain *bd)
l4start = l4tab = __va(mpt_alloc); mpt_alloc += PAGE_SIZE;
clear_page(l4tab);
init_xen_l4_slots(l4tab, _mfn(virt_to_mfn(l4start)),
- d, INVALID_MFN, true);
+ d->vcpu[0], INVALID_MFN, true);
v->arch.guest_table = pagetable_from_paddr(__pa(l4start));
}
else
@@ -1048,7 +1048,11 @@ static int __init dom0_construct(const struct boot_domain *bd)
}
/* Activate shadow mode, if requested. Reuse the pv_l1tf tasklet. */
- if ( opt_dom0_shadow )
+ if ( opt_dom0_shadow && d->arch.vcpu_pt )
+ /* Shadow paging is incompatible with per-vCPU page-tables (ASI). */
+ printk(XENLOG_WARNING
+ "Ignoring dom0=shadow: incompatible with per-vCPU page-tables\n");
+ else if ( opt_dom0_shadow )
{
printk("Switching dom0 to using shadow paging\n");
tasklet_schedule(&d->arch.paging.shadow.pv_l1tf_tasklet);
diff --git a/xen/arch/x86/pv/domain.c b/xen/arch/x86/pv/domain.c
index 50f2d1284a..b1d57083f2 100644
--- a/xen/arch/x86/pv/domain.c
+++ b/xen/arch/x86/pv/domain.c
@@ -127,7 +127,7 @@ static int setup_compat_l4(struct vcpu *v)
mfn = page_to_mfn(pg);
l4tab = map_domain_page(mfn);
clear_page(l4tab);
- init_xen_l4_slots(l4tab, mfn, v->domain, INVALID_MFN, false);
+ init_xen_l4_slots(l4tab, mfn, v, INVALID_MFN, false);
unmap_domain_page(l4tab);
/* This page needs to look like a pagetable so that it can be shadowed */
--
2.55.0
^ permalink raw reply related [flat|nested] 44+ messages in thread