All of lore.kernel.org
 help / color / mirror / Atom feed
From: "Jürgen Groß" <jgross@suse.com>
To: Jan Beulich <jbeulich@suse.com>, George Dunlap <dunlapg@umich.edu>
Cc: "Andrew Cooper" <andrew.cooper3@citrix.com>,
	"Roger Pau Monné" <roger@xenproject.org>,
	"Alejandro Vallejo" <agarciav@amd.com>,
	"Teddy Astie" <teddy.astie@vates.tech>,
	"Anthony PERARD" <anthony.perard@vates.tech>,
	"Michal Orzel" <michal.orzel@amd.com>,
	"Julien Grall" <julien@xen.org>,
	"Stefano Stabellini" <sstabellini@kernel.org>,
	xen-devel@lists.xenproject.org
Subject: Re: [PATCH v2 03/14] x86/pv: use populate_perdomain_mapping() to map the Xen GDT
Date: Fri, 4 Sep 2026 08:54:11 +0200	[thread overview]
Message-ID: <050f44ec-47e5-4754-b73b-1173458736b0@suse.com> (raw)
In-Reply-To: <07d332a4-a5a3-44d4-95ce-fc67d86edacf@suse.com>


[-- Attachment #1.1.1: Type: text/plain, Size: 4300 bytes --]

On 04.09.26 08:00, Jan Beulich wrote:
> On 04.09.2026 00:35, George Dunlap wrote:
>> On Thu, Sep 3, 2026 at 5:11 PM Jan Beulich <jbeulich@suse.com> wrote:
>>>
>>> On 02.09.2026 11:43, George Dunlap wrote:
>>>> From: Roger Pau Monné <roger.pau@citrix.com>
>>>>
>>>> Currently, update_xen_slot_in_full_gdt() uses the stashed direct-map
>>>> pointer in d->arch.pv.gdt_ldt_l1tab to update the incoming vcpu's
>>>> page tables with Xen's GDT, by writing a stashed per-cpu copy of a
>>>> pre-baked L1 entry (either 64-bit or compat version).
>>>>
>>>> Switch this to using populate_perdomain_mapping(), which doesn't rely
>>>> on the stashed address of the l1 page in the direct map.  Rather than
>>>> also stashing a pre-baked value for the payload, compute the mfn from
>>>> the per-cpu GDT pointer at use: the conversion is a handful of cycles
>>>> on a path costing thousands, and computing at use removes the
>>>> parallel {,compat_}gdt_l1e bookkeeping along with its boot-ordering
>>>> constraint (the cached value could only be generated after Xen's
>>>> physical relocation, and had to be in place before the first context
>>>> switch; a use-time lookup is correct by construction).  The flags on
>>>> the final mapping are identical.
>>>>
>>>> Signed-off-by: Roger Pau Monné <roger.pau@citrix.com>
>>>> Assisted-by: Claude Code:claude-fable-5, Claude Code:claude-opus-4-8
>>>> Signed-off-by: George Dunlap <gwd@xenproject.org>
>>>> ---
>>>> Changes in v2:
>>>> - Drop the {,compat_}gdt_mfn caching entirely (suggested by Andrew
>>>>     Cooper): compute virt_to_mfn() from the per-cpu GDT pointer at use.
>>>>     The PDX lookup behind it measures ~5-10 cycles warm against a
>>>>     ~1,500-cycle context switch, and this removes the double
>>>>     bookkeeping and the after-relocation caching constraint.  The
>>>>     cached-MFN assertion goes with the cache: a use-time computation
>>>>     from a live pointer needs no staleness check.
>>>
>>> This looks to contradict what 564d261687c0 ("x86/ctxt-switch: Document
>>> and improve GDT handling") used as justification to put in place the
>>> caching. Also Cc-ing Jürgen, who also was involved there, for possible
>>> further insight.
>>>
>>> Functionally the change looks okay to me, but the above will need
>>> sorting, at the very least by specifically discussing why effectively
>>> undoing that earlier change is okay.
>>
>> So looking back at the thread, Jürgen measured a 14% improvement for
>> something that might be described as a microbenchmark before and after
>> the patch (a benchmark purposely trying to set up an unusual scenario
>> to maximize the effect of context switch overhead, not one to
>> represent a typical workflow).  But are the numbers really plausible?
>> Even at an implausible 100k switches/s across the box, saving 100
>> cycles per switch is about 0.04% of eight 3 GHz cores.
>>
>> At any rate, we're already adding several map/unmap operations, and
>> about to add several more.  Keeping the PTE caching would require
>> adding a separate path that can write just PTEs, which then will
>> potentially further complication future paths where we need to make
>> sure we handle both domain-wide perdomain areas and per-vcpu areas.
>> If it were easy I would already have been keeping it.
>>
>> I'd be inclined to say: Since we're going to be adding more
>> populate_perdomain_mapping() calls anyway, let's do it the simple
>> correct way first; and then explore the idea of stashing mfns of
>> frequently-mapped L1s (rather than having to walk L3 -> L2 -> L1); and
>> at that time look into stashing baked l1es to avoid conversions.
> 
> Perhaps; I'd like to have Jürgen's and/or Andrew's input here, though.

At that time I implemented core scheduling in Xen. I noticed that very
subtle changes in the context switch path could result in unexpected large
performance differences. As I had the performance test for my purpose
already set up, I used it for Andrew's patch (which was a result of my
context switch path performance findings) and really did measure the
impressive effect of it.

Note that you can't only count instructions, often cache effects and
branch predictions are dominating the performance.


Juergen

[-- Attachment #1.1.2: OpenPGP public key --]
[-- Type: application/pgp-keys, Size: 3743 bytes --]

[-- Attachment #2: OpenPGP digital signature --]
[-- Type: application/pgp-signature, Size: 495 bytes --]

  reply	other threads:[~2026-09-04  6:54 UTC|newest]

Thread overview: 44+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-09-02  9:43 [PATCH v2 00/14] x86: Address Space Isolation, part 2: asi= option and per-vCPU page tables George Dunlap
2026-09-02  9:43 ` [PATCH v2 01/14] x86/domain_page: introduce IRQs-off variants of {,un}map_domain_page() George Dunlap
2026-09-03 14:07   ` Jan Beulich
2026-09-03 19:56     ` George Dunlap
2026-09-02  9:43 ` [PATCH v2 02/14] x86/mm: introduce populate_perdomain_mapping() George Dunlap
2026-09-03 15:57   ` Jan Beulich
2026-09-03 21:27     ` George Dunlap
2026-09-04  5:58       ` Jan Beulich
2026-09-04  5:47   ` Jan Beulich
2026-09-02  9:43 ` [PATCH v2 03/14] x86/pv: use populate_perdomain_mapping() to map the Xen GDT George Dunlap
2026-09-03 16:11   ` Jan Beulich
2026-09-03 22:35     ` George Dunlap
2026-09-04  6:00       ` Jan Beulich
2026-09-04  6:54         ` Jürgen Groß [this message]
2026-09-04  8:06           ` George Dunlap
2026-09-04  8:29             ` Jan Beulich
2026-09-04  8:50               ` George Dunlap
2026-09-04 10:11                 ` Jan Beulich
2026-09-04 10:34                 ` Roger Pau Monné
2026-09-07 13:58                   ` George Dunlap
2026-09-02  9:43 ` [PATCH v2 04/14] x86/pv: set/clear guest GDT mappings using populate_perdomain_mapping() George Dunlap
2026-09-07 12:50   ` Jan Beulich
2026-09-07 13:51     ` George Dunlap
2026-09-07 14:57       ` Jan Beulich
2026-09-02  9:43 ` [PATCH v2 05/14] x86/pv: update guest LDT mappings using {populate,destroy}_perdomain_mapping() George Dunlap
2026-09-07 16:06   ` Jan Beulich
2026-09-09 19:29     ` George Dunlap
2026-09-02  9:43 ` [PATCH v2 06/14] x86/pv: remove stashing of GDT/LDT L1 page-tables George Dunlap
2026-09-08 14:29   ` Jan Beulich
2026-09-02  9:43 ` [PATCH v2 07/14] x86/mm: simplify create_perdomain_mapping() interface George Dunlap
2026-09-08 14:39   ` Jan Beulich
2026-09-02  9:43 ` [PATCH v2 08/14] x86/mm: purge unneeded destroy_perdomain_mapping() George Dunlap
2026-09-08 15:03   ` Jan Beulich
2026-09-02  9:43 ` [PATCH v2 09/14] x86/mm: prepare destroy_perdomain_mapping() for per-vCPU perdomain areas George Dunlap
2026-09-08 15:36   ` Jan Beulich
2026-09-10 11:38     ` George Dunlap
2026-09-10 11:54       ` Jan Beulich
2026-09-02  9:43 ` [PATCH v2 10/14] x86/domain_page: drop redundant create_perdomain_mapping() call George Dunlap
2026-09-08 15:55   ` Jan Beulich
2026-09-10 11:52     ` George Dunlap
2026-09-02  9:43 ` [PATCH v2 11/14] x86/mm: prepare create_perdomain_mapping() for per-vCPU perdomain areas George Dunlap
2026-09-02  9:43 ` [PATCH v2 12/14] x86/spec-ctrl: introduce Address Space Isolation command line option George Dunlap
2026-09-02  9:43 ` [PATCH v2 13/14] x86/pv: clear the XPTI root_pgt per-domain slot on context-switch out George Dunlap
2026-09-02  9:43 ` [PATCH v2 14/14] x86/mm: introduce per-vCPU L3 page-table George Dunlap

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=050f44ec-47e5-4754-b73b-1173458736b0@suse.com \
    --to=jgross@suse.com \
    --cc=agarciav@amd.com \
    --cc=andrew.cooper3@citrix.com \
    --cc=anthony.perard@vates.tech \
    --cc=dunlapg@umich.edu \
    --cc=jbeulich@suse.com \
    --cc=julien@xen.org \
    --cc=michal.orzel@amd.com \
    --cc=roger@xenproject.org \
    --cc=sstabellini@kernel.org \
    --cc=teddy.astie@vates.tech \
    --cc=xen-devel@lists.xenproject.org \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is an external index of several public inboxes,
see mirroring instructions on how to clone and mirror
all data and code used by this external index.