All of lore.kernel.org
 help / color / mirror / Atom feed
From: Jan Beulich <jbeulich@suse.com>
To: George Dunlap <dunlapg@umich.edu>
Cc: "Roger Pau Monné" <roger@xenproject.org>,
	"Andrew Cooper" <andrew.cooper3@citrix.com>,
	"Alejandro Vallejo" <agarciav@amd.com>,
	"Teddy Astie" <teddy.astie@vates.tech>,
	"Anthony PERARD" <anthony.perard@vates.tech>,
	"Michal Orzel" <michal.orzel@amd.com>,
	"Julien Grall" <julien@xen.org>,
	"Stefano Stabellini" <sstabellini@kernel.org>,
	"George Dunlap" <gwd@xenproject.org>,
	xen-devel@lists.xenproject.org
Subject: Re: [PATCH v2 02/14] x86/mm: introduce populate_perdomain_mapping()
Date: Thu, 3 Sep 2026 17:57:13 +0200	[thread overview]
Message-ID: <5071d4da-d9b1-45cd-8e9b-778df417ee7b@suse.com> (raw)
In-Reply-To: <20260901-asi-part2-2-ecc269f268b7@xenproject.org>

On 02.09.2026 11:43, George Dunlap wrote:
> --- a/xen/arch/x86/mm.c
> +++ b/xen/arch/x86/mm.c
> @@ -6334,6 +6334,130 @@ int create_perdomain_mapping(struct domain *d, unsigned long va,
>      return rc;
>  }
>  
> +/*
> + * Map @nr pages, @mfn[0..nr-1], at consecutive pages from @va in v's view of
> + * the per-domain area, with page-table @flags.  The range must lie within a
> + * single per-domain slot, and must already have been plumbed down to the L1
> + * tables by create_perdomain_mapping(): missing structure is a bug.  A
> + * present entry not owned by the area (no _PAGE_AVAIL0) is silently
> + * replaced, as that is how callers update their mappings; a present
> + * area-owned entry is freed and replaced, which constrains the calling
> + * context (see the comment in the body).  No TLB flushing is done: the
> + * caller decides whether the old translations can still be cached
> + * anywhere.
> + *
> + * When v's page-tables are loaded on this pCPU the L1 entries are reached
> + * through the recursive linear mappings; otherwise the walk maps the
> + * per-domain page-table pages transiently with IRQs off, so it needs
> + * nothing from the current address space and is usable from any context --
> + * including the context switch, before the incoming vcpu's page-tables are
> + * loaded.
> + */
> +void populate_perdomain_mapping(const struct vcpu *v, unsigned long va,
> +                                const mfn_t *mfn, unsigned int nr,
> +                                unsigned int flags)
> +{
> +    l1_pgentry_t *l1tab = NULL, *pl1e;
> +    const l3_pgentry_t *l3tab;
> +    const l2_pgentry_t *l2tab;
> +    struct domain *d = v->domain;
> +    unsigned long irq_flags;
> +
> +    ASSERT(va >= PERDOMAIN_VIRT_START &&
> +           va < PERDOMAIN_VIRT_SLOT(PERDOMAIN_SLOTS));
> +    ASSERT(!nr || !l3_table_offset(va ^ (va + nr * PAGE_SIZE - 1)));
> +    /* Area-owned pages are installed by create_perdomain_mapping() only. */
> +    ASSERT(!(flags & _PAGE_AVAIL0));
> +
> +    if ( likely(this_cpu(pgtable_vcpu) == v) )
> +    {
> +        unsigned int i;
> +
> +        /*
> +         * Fast path: v's page-tables are loaded on this pCPU, so the L1
> +         * entries can be reached using the recursive linear mappings.
> +         */
> +        pl1e = &__linear_l1_table[l1_linear_offset(va)];

As mentioned elsewhere, I'm concerned of this (or really any) new use of
the linear page tables. (Which, ftaod, isn't an objection.)

> +        for ( i = 0; i < nr; i++, pl1e++ )
> +        {
> +            /*
> +             * An area-owned entry (installed by create_perdomain_mapping(),
> +             * marked _PAGE_AVAIL0) holds the only reference to its page, so
> +             * displacing it means freeing it.  Nothing in this series
> +             * replaces area-owned backing, hence the ASSERT_UNREACHABLE();
> +             * any future caller doing so must run where freeing is
> +             * permitted -- IRQs enabled, not in interrupt context (see
> +             * ASSERT_ALLOC_CONTEXT()) -- which the context switch path is
> +             * not.
> +             */
> +            if ( unlikely(perdomain_l1e_needs_freeing(*pl1e)) )
> +            {
> +                ASSERT_UNREACHABLE();
> +                free_domheap_page(l1e_get_page(*pl1e));
> +            }
> +            l1e_write(pl1e, l1e_from_mfn(mfn[i], flags));
> +        }
> +
> +        return;
> +    }
> +
> +    BUG_ON(!d->arch.perdomain_l3_pg);
> +
> +    /*
> +     * Slow path: walk v's per-domain page-table pages.  All mappings are
> +     * local to this function, so disabling interrupts for the duration of
> +     * the walk satisfies the map_domain_page_irqoff() contract.  This in
> +     * turn makes this function usable from the context switch path, where
> +     * a plain map_domain_page() could recurse into __context_switch() via
> +     * sync_local_execstate().
> +     */
> +    local_irq_save(irq_flags);
> +
> +    l3tab = __map_domain_page_irqoff(d->arch.perdomain_l3_pg);
> +
> +    /*
> +     * Missing page-table structure is a hypervisor bug: there is no safe
> +     * continuation, least of all from the context switch, where the next
> +     * descriptor fetch through an unmapped GDT slot would be fatal.
> +     */
> +    BUG_ON(!(l3e_get_flags(l3tab[l3_table_offset(va)]) & _PAGE_PRESENT));
> +
> +    l2tab = map_domain_page_irqoff(l3e_get_mfn(l3tab[l3_table_offset(va)]));

l3tab[] isn't used any further, so I think it wants unmapping right away. No
need to have undue pressure on the number of active mappings.

> +    for ( ; nr--; va += PAGE_SIZE, mfn++ )
> +    {
> +        if ( !l1tab || !l1_table_offset(va) )
> +        {
> +            const l2_pgentry_t *pl2e = l2tab + l2_table_offset(va);
> +
> +            BUG_ON(!(l2e_get_flags(*pl2e) & _PAGE_PRESENT));
> +
> +            unmap_domain_page_irqoff(l1tab);
> +            l1tab = map_domain_page_irqoff(l2e_get_mfn(*pl2e));
> +        }
> +
> +        pl1e = &l1tab[l1_table_offset(va)];
> +
> +        /*
> +         * As the fast path -- and the slow path holds IRQs off throughout,
> +         * so replacing area-owned backing here is never permitted.
> +         */

With this comment I think ...

> +        if ( unlikely(perdomain_l1e_needs_freeing(*pl1e)) )
> +        {
> +            ASSERT_UNREACHABLE();
> +            free_domheap_page(l1e_get_page(*pl1e));

... this call should be removed from here (I would have suggested to comment
it out, but Misra dislikes that iirc). Otherwise it would in principle be
reachable in release builds.

Maybe instead of ASSERT_UNREACHABLE() it should really be BUG() here.

Jan


  reply	other threads:[~2026-09-03 15:57 UTC|newest]

Thread overview: 44+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-09-02  9:43 [PATCH v2 00/14] x86: Address Space Isolation, part 2: asi= option and per-vCPU page tables George Dunlap
2026-09-02  9:43 ` [PATCH v2 01/14] x86/domain_page: introduce IRQs-off variants of {,un}map_domain_page() George Dunlap
2026-09-03 14:07   ` Jan Beulich
2026-09-03 19:56     ` George Dunlap
2026-09-02  9:43 ` [PATCH v2 02/14] x86/mm: introduce populate_perdomain_mapping() George Dunlap
2026-09-03 15:57   ` Jan Beulich [this message]
2026-09-03 21:27     ` George Dunlap
2026-09-04  5:58       ` Jan Beulich
2026-09-04  5:47   ` Jan Beulich
2026-09-02  9:43 ` [PATCH v2 03/14] x86/pv: use populate_perdomain_mapping() to map the Xen GDT George Dunlap
2026-09-03 16:11   ` Jan Beulich
2026-09-03 22:35     ` George Dunlap
2026-09-04  6:00       ` Jan Beulich
2026-09-04  6:54         ` Jürgen Groß
2026-09-04  8:06           ` George Dunlap
2026-09-04  8:29             ` Jan Beulich
2026-09-04  8:50               ` George Dunlap
2026-09-04 10:11                 ` Jan Beulich
2026-09-04 10:34                 ` Roger Pau Monné
2026-09-07 13:58                   ` George Dunlap
2026-09-02  9:43 ` [PATCH v2 04/14] x86/pv: set/clear guest GDT mappings using populate_perdomain_mapping() George Dunlap
2026-09-07 12:50   ` Jan Beulich
2026-09-07 13:51     ` George Dunlap
2026-09-07 14:57       ` Jan Beulich
2026-09-02  9:43 ` [PATCH v2 05/14] x86/pv: update guest LDT mappings using {populate,destroy}_perdomain_mapping() George Dunlap
2026-09-07 16:06   ` Jan Beulich
2026-09-09 19:29     ` George Dunlap
2026-09-02  9:43 ` [PATCH v2 06/14] x86/pv: remove stashing of GDT/LDT L1 page-tables George Dunlap
2026-09-08 14:29   ` Jan Beulich
2026-09-02  9:43 ` [PATCH v2 07/14] x86/mm: simplify create_perdomain_mapping() interface George Dunlap
2026-09-08 14:39   ` Jan Beulich
2026-09-02  9:43 ` [PATCH v2 08/14] x86/mm: purge unneeded destroy_perdomain_mapping() George Dunlap
2026-09-08 15:03   ` Jan Beulich
2026-09-02  9:43 ` [PATCH v2 09/14] x86/mm: prepare destroy_perdomain_mapping() for per-vCPU perdomain areas George Dunlap
2026-09-08 15:36   ` Jan Beulich
2026-09-10 11:38     ` George Dunlap
2026-09-10 11:54       ` Jan Beulich
2026-09-02  9:43 ` [PATCH v2 10/14] x86/domain_page: drop redundant create_perdomain_mapping() call George Dunlap
2026-09-08 15:55   ` Jan Beulich
2026-09-10 11:52     ` George Dunlap
2026-09-02  9:43 ` [PATCH v2 11/14] x86/mm: prepare create_perdomain_mapping() for per-vCPU perdomain areas George Dunlap
2026-09-02  9:43 ` [PATCH v2 12/14] x86/spec-ctrl: introduce Address Space Isolation command line option George Dunlap
2026-09-02  9:43 ` [PATCH v2 13/14] x86/pv: clear the XPTI root_pgt per-domain slot on context-switch out George Dunlap
2026-09-02  9:43 ` [PATCH v2 14/14] x86/mm: introduce per-vCPU L3 page-table George Dunlap

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=5071d4da-d9b1-45cd-8e9b-778df417ee7b@suse.com \
    --to=jbeulich@suse.com \
    --cc=agarciav@amd.com \
    --cc=andrew.cooper3@citrix.com \
    --cc=anthony.perard@vates.tech \
    --cc=dunlapg@umich.edu \
    --cc=gwd@xenproject.org \
    --cc=julien@xen.org \
    --cc=michal.orzel@amd.com \
    --cc=roger@xenproject.org \
    --cc=sstabellini@kernel.org \
    --cc=teddy.astie@vates.tech \
    --cc=xen-devel@lists.xenproject.org \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is an external index of several public inboxes,
see mirroring instructions on how to clone and mirror
all data and code used by this external index.