Kernel KVM virtualization development
 help / color / mirror / Atom feed
From: Sean Christopherson <seanjc@google.com>
To: Rick P Edgecombe <rick.p.edgecombe@intel.com>
Cc: Yilun Xu <yilun.xu@intel.com>,
	Elena Reshetova <elena.reshetova@intel.com>,
	 Binbin Wu <binbin.wu@intel.com>,
	Dave Hansen <dave.hansen@intel.com>,
	 Vishal Annapurve <vannapurve@google.com>,
	"kas@kernel.org" <kas@kernel.org>,
	 "pbonzini@redhat.com" <pbonzini@redhat.com>,
	Peter Fang <peter.fang@intel.com>,
	 "kvm@vger.kernel.org" <kvm@vger.kernel.org>
Subject: Re: TDG quote analysis
Date: Fri, 18 Sep 2026 06:13:20 -0700	[thread overview]
Message-ID: <aq048KmPN5WPQ4sQ@google.com> (raw)
In-Reply-To: <56a61983c03806120e607dfd1e1b7c9046c1f54b.camel@intel.com>

On Fri, Sep 18, 2026, Rick P Edgecombe wrote:
> On Thu, 2026-09-17 at 17:04 -0700, Sean Christopherson wrote:
> >   If supported for the specific request type, the host VMM can request an
> >   interrupted request to be aborted.
> > 
> > Though per the above, apparently whether or not that's going to be supported is
> > TBD?
> 
> So extension aborts are a whole 'nother topic. And "support" has a nuanced
> answer to boot. I'd be very glad to hear your thoughts on it. Should we leave it
> for a later topic or get into it now? I *think* we don't need it, but the more
> we get into the details here, the more it's a notable thing we haven't covered.

How is saying "I no longer want to complete this quote" at all complex?  I can
see how aborting a request to the S3M might be somewhat pointless, but inserting
the equivalent to signal_pending() in a long-running operation doesn't seem that
difficult.

> > Regardless of where we end up with quoting, whatever RFC is sent next needs to be
> > a million times better.  My expectations are that someone with core TDX knowledge
> > and a basic understanding of attestation would be able to get a full and complete
> > understanding of the design and tradeoffs from the RFC alone.  If I have to
> > rereference a spec for anything other than double check params and magic values,
> > the entire series is getting ignored.
> 
> I mean. Confidential compute in general is trying to reinvent decades of
> virtualization solutions with one hand tied behind it's back. There *are* a lot
> of design choices.
> 
> Besides a rushed RFC, I think another factor is that this is the first TDX
> feature we have done that you don't already have at least some exposure to. The
> level of background delta need is way higher than S-EPT management for example.
> 
> And I'm less and less sure we are going to be successful distilling things like

It seems like part of the problem is that you're trying to "distill" into abstract
concepts.  I don't want abstract concepts, I want a description of what the TDX
Module code will literally do, using verbiage and terminology that a KVM developer
will natively understand.

> this down in a way that doesn't leave you frustrated. But if you want details
> there is just going to be a lot of them. Hmm.

This is what I want.  Note, this description is apparently wildly wrong, because
I wrote it before reading about whatever NRX modules are.  But I'm leaving it
because it highlights my point about not wanting TDX jargon rephrased as abstract
concepts.

  Under the hood, TDH.GET.QUOTE uses what TDX calls a "virtual thread pool".
  That just means there's a pre-allocated pool of "thread" memory that can be
  used to service interruptible requests (it's not true interruption, it's
  basically voluntary preemption).  The number of threads in the pool is
  configured by software during ??? (TDH.QUOTE.INIT?), and directly controls
  how many in-flight TDH.GET.QUOTE operations there can be at any given time.

  Each thread consumes ??? KiB of memory, and <reasoning>, so we chose to set
  the pool size to 1, i.e. to only allow a single quote to be in-flight across
  the entire system.

  TDH.GET.QUOTE is "interruptible" because the bulk of the quote crypto is done
  by the TDX module, i.e. on the CPU that invokes TDH.GET.QUOTE.  We don't have
  numbers for TDH.GET.QUOTE yet (why not?), but the very rough ballpark is that
  it will takes 1-2ms.  Note, the S3M is used only during ??? (TDH.QUOTE.INIT?)
  to get the platform key used to generate quotes, i.e. generating a quote is
  fully synchronous relative to the core.

> > As for guest-driven quoting, to me there is a fairly straightforward solution.
> > Assuming using guest-donated memory is too complex for the "virtual thread":
> > 
> >  1. Drop "virtual thread pools" from quoting and support at most one in-flight
> >     quote per TD.  I doubt whatever memory is needed for a virtual thread is so
> >     insanely ridiculous that we can't burn that much memory per TD.  32KiB is a
> >     no-brainer.  Above that and we'd have to reconsider, but if quoting requires
> >     more than 32KiB of scratch space, I have to wonder what on earth it's doing.
> 
> I think the abstract descriptions are not helping, so I'll try some more
> details. Today the "thread" involves something like a vCPU in a TD. You can see
> some references in the TDX base spec like:
>    Currently, TDX Module extensions are implemented as guests running in SEAM non-
>    root mode, called NRX Modules. Technically, each NRX Module is similar to a TD.

So my above statement that "it's not true interruption, it's basically voluntary
preemption" is wrong?  Because if the quote crud is running in a TD, then events
will trigger VM-Exit, and it will be impossible to make forward progress without
SEAMRETing to the host.

If that's true, then that changes my understanding of this just a bit, because
it means the quote operation isn't manually saving state at a "good stopping point",
it's relying on the VM-Exit to save/restore at whenever it happened to be.  But
given that you say "This idea came up before actually", maybe chunking the quote
operation isn't a big lift?

But the original mail says this:

   -- Guest Interrupts --

   It is not nice to keep the guest from running for too long. The TDH.QUOTE.GET
   SEAMCALL monitors for host interrupts, but not guest ones. Similarly, the SGX
            ^^^^^^^^^^^^^^^^^^^^^^^^^^^^
   quoting enclave doesn't know what is happening in the TD. So, to help reduce
   guest latencies, the GHCI exposes a way to register for a notification
   (SetupEventNotifyInterrupt) when the quoting is finished. A quote TDVMCALL
   handler can start the quote operation on another host thread, and resume
   guest execution waiting for the quote to finish.

which strongly suggests active polling.  Do you see why I don't want abstract
descriptions?  Similar to changelogs, there's a balance between "too detailed"
and "too abstract", and right now all of this is firmly on the "too abstract"
side of the world.

> >  2. To avoid having to check for guest interrupts, for the "SW based" synchronous
> >     method, simply resume the guest after processing a fixed amount of state.  Yes,
> >     that will increase the best case latency for a single quote, but *best* case
> >     latency isn't a huge concern, and I doubt the overhead of a VMX roundtrip will
> >     significantly impact that.  It's the tail latencies that will be problematic,
> >     and a forcing the serialization into the guest (the aforementioned mutex) means
> >     the tail latencies will only be affected by host activity on *that* CPU, which
> >     is more or less the status quo.
> > 
> >     E.g. even assuming an absurd 20% overhead for the VMX round trips, having a
> >     quote take ~1.2 *every* time is would be a far, far better experience than a
> >     quote taking 1ms - 10ms to complete.
> 
> Dave will probably not get a chance to respond this week, but we should maybe
> discuss what acceptable guest latency actually is.

Yes.  But it's not necessarily about "acceptable" guest latency, I care more about
*predictable* guest latency.  E.g. making up numbers to illustrate the point,
achieving 50us latency for the happy case is meaningless if the P99 latency is 10ms.

  reply	other threads:[~2026-09-18 13:13 UTC|newest]

Thread overview: 33+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-09-16  0:48 TDG quote analysis Edgecombe, Rick P
2026-09-16 18:00 ` Sean Christopherson
2026-09-16 18:49   ` Edgecombe, Rick P
2026-09-16 19:27     ` Sean Christopherson
2026-09-16 22:31       ` Edgecombe, Rick P
2026-09-16 23:00         ` Peter Fang
2026-09-16 23:57           ` Sean Christopherson
2026-09-17  0:56             ` Edgecombe, Rick P
2026-09-17 13:37               ` Sean Christopherson
2026-09-17 18:12                 ` Edgecombe, Rick P
2026-09-17 19:57                   ` Sean Christopherson
2026-09-17 21:32                     ` Edgecombe, Rick P
2026-09-18  0:04                       ` Sean Christopherson
2026-09-18  2:39                         ` Edgecombe, Rick P
2026-09-18 13:13                           ` Sean Christopherson [this message]
2026-09-18 18:12                             ` Edgecombe, Rick P
2026-09-18 21:32                               ` Sean Christopherson
2026-09-21 23:00                                 ` Peter Fang
2026-09-21 23:06                                   ` Dave Hansen
2026-09-23  4:09                                     ` Vishal Annapurve
2026-09-24  0:03                                       ` Vishal Annapurve
2026-09-24  0:18                                         ` Sean Christopherson
2026-09-24  0:44                                           ` Edgecombe, Rick P
2026-09-24 16:28                                             ` Sean Christopherson
2026-09-21 23:04                                 ` Dave Hansen
2026-09-21 23:19                                   ` Sean Christopherson
2026-09-21 23:36                                     ` Dave Hansen
2026-09-21 23:46                                       ` Sean Christopherson
2026-09-22  0:03                                         ` Dave Hansen
2026-09-22  6:48                                           ` Reshetova, Elena
2026-09-22 17:07                                             ` Edgecombe, Rick P
2026-09-23  8:50                                               ` Reshetova, Elena
2026-09-23 22:50                                                 ` Peter Fang

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=aq048KmPN5WPQ4sQ@google.com \
    --to=seanjc@google.com \
    --cc=binbin.wu@intel.com \
    --cc=dave.hansen@intel.com \
    --cc=elena.reshetova@intel.com \
    --cc=kas@kernel.org \
    --cc=kvm@vger.kernel.org \
    --cc=pbonzini@redhat.com \
    --cc=peter.fang@intel.com \
    --cc=rick.p.edgecombe@intel.com \
    --cc=vannapurve@google.com \
    --cc=yilun.xu@intel.com \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox