Kernel KVM virtualization development
 help / color / mirror / Atom feed
From: Sean Christopherson <seanjc@google.com>
To: Rick P Edgecombe <rick.p.edgecombe@intel.com>
Cc: Yilun Xu <yilun.xu@intel.com>,
	Elena Reshetova <elena.reshetova@intel.com>,
	 Binbin Wu <binbin.wu@intel.com>,
	Dave Hansen <dave.hansen@intel.com>,
	 Vishal Annapurve <vannapurve@google.com>,
	"kas@kernel.org" <kas@kernel.org>,
	 "pbonzini@redhat.com" <pbonzini@redhat.com>,
	Peter Fang <peter.fang@intel.com>,
	 "kvm@vger.kernel.org" <kvm@vger.kernel.org>
Subject: Re: TDG quote analysis
Date: Thu, 17 Sep 2026 17:04:32 -0700	[thread overview]
Message-ID: <aqyAED5ZK48eEZtq@google.com> (raw)
In-Reply-To: <906d0cb6515b396f4e9d74f1d546c902ea3ce49a.camel@intel.com>

On Thu, Sep 17, 2026, Rick P Edgecombe wrote:
> On Thu, 2026-09-17 at 12:57 -0700, Sean Christopherson wrote:
> > > Let's take a step back here. There is an existing attestation flow that was
> > > designed around some limitations that are changing (specifically whether the
> > > quoter has knowledge of the TD). The current host side quote discussion (KVM
> > > ioctl) is basically a straight forward evolution of the existing design,
> > > even though the limitations are getting removed.
> > > 
> > > You asked whether we could do a re-design that makes more sense in the
> > > context of the lack of that old limitation. At that point we are faced with
> > > the age-old question: how far into the fuzzy future should we design around?
> > 
> > Now I'm trying to understand what the *current* plan is.
> 
> The last Linux design for a host-based TDX DICE API, meaning the thing we were
> prepping for the next DICE posting before getting into this TDG analysis:
>  - Leave the existing report in the guest as is.
>  - Utilize the existing quote GHCI call from the guest to pass the old report,
> which goes through KVM to userspace.
>  - Add a new VM scoped ioctl to KVM to call into the TDH.QUOTE.GET. This gets
> the quote and passes it back to userspace. Then userspace uses the existing GHCI
> mechanisms to notify the guest that it is ready.
> 
> > 
> > > The nearest term thing is a SW based flow where work happens on the CPU. But
> > > the exact amount of time is not know yet. Peter gave a ballpark.
> > 
> > Why are we even discussing this?  I am so confused.  I thought there were two
> > options: SGX and S3M.  Now all of a sudden there's a third "let's do insane
> > things in software in the TDX module" option!?!?
> 
> ???
> 
> So when I said:
>    "HW" is talking about the S3M thing. The quote operation could go directly to
>    the S3M to get the quote. (HW based) Or it could get an intermediate key and
>    generate quotes using CPU instructions. (SW based). Think like a crypto library
>    in the TDX module.
>    
>    ...
> 
>    The software based flow would be expected to first. It would involve the CPU
>    doing crypto stuff as above.
>    
>    Then a HW based flow where the crypto happens on the limited HW resource. This
>    is where full parallelization is not possible, because the CPU is not doing the
>    heavy work. You might want this one instead for security reasons. But the main
>    point of discussing it is that you could expect some quotes to take a long time
>    and support a limited number of parallel quotes.
>    
> Did you interpret SW based flow to be talking about SGX? Like the intermediate key
> goes from S3M to SGX?

From *before* this conversation.  Forget this converation, what does the mock
TDX Module used as the basis for the RFC[1] do?  Because the RFC says absolutely
*nothing*.  I kinda sorta have a picture now, but it required hunting down an
additional spec, and the documentnation still leaves me wanting.  Because I still
don't know what it actually does.  Based on everything you're saying, I *assume*
it's this "SW based" flow, but *nothing* actually says that.

Under "Interruptibility", the quoting doc linked says:

  TDH.QUOTE.GET is interruptible. If a pending interrupt is detected during
  operation, TDH.QUOTE.GET returns with a TDX_INTERRUPTED_RESUMABLE status in RAX.

Then punts me to:

  For the general explanation on how interruption and resumption is handled for
  all Quoting Service functions please consult [Intel TDX Module Base Spec] section
  “Request Interruption and Resumption”.

And finishes with the wonderful:

  Rest of details are TBD

Once I finally found the "Request Interruption and Resumption", I discovered that
the TDX Module now has "virtual thread pools", which IS NEVER MENTIONED IN THE
RFC.  Seriously.  No one thought that was worth mentioning?

OMG.  I see it now:

+               /* Don't bother specifying the quote id */
+               .rdx = QUOTE_ID_MASK & (u64)-1,

So IIUC, there's a thread pool, but somewhere in the muck of RFC patches the
kernel initializes the pool with a size of 1?  And then adds a global mutex to
serialize quotes across the entire system.  That's certainly a choice.

This snippet from the RFC[2] doesn't help, because it makes it seem like there
is no state:

  2. Host just refills the previous args (may have been modified by seamcall
     output) on retry. such as TDH_EXT_INIT, TDH_EXT_MEM_ADD,
     TDH_QUOTE_INIT, TDH_QUOTE_GET...

But that's presumably just because the kernel is using a single REQUEST_ID.

That spec continues with:

  If supported for the specific request type, the host VMM can request an
  interrupted request to be aborted.

Though per the above, apparently whether or not that's going to be supported is
TBD?

I am beyond frustrated at this point.  I guess shame on me for not pouring over
dense docs?  But I shouldn't have to.  The *ENTIRE* point of an RFC like the one
that started all this is to get feedback on the design, but that's really hard to
achieve if you don't actually provide *any* details about the design.

And there are MASSIVE design choices in here.  Like sizing the "thread" pool to 1,
and thus deliberately serializing quotes across all VMs.  That's going to end well
when someone spins up 10s or 100s of TDs and they all try to attest at once.
Not listening for fatal signals while spinning on TDH.GET.QUOTE will save us though!
At least with PREEMPT_LAZY being forced, not doing cond_resched() is "fine".

What's worse, some of these patches that say nothing useful in the changelogs
have several Reviewed-by tags, which means multiple people looked at this as
thought it was all good.

Regardless of where we end up with quoting, whatever RFC is sent next needs to be
a million times better.  My expectations are that someone with core TDX knowledge
and a basic understanding of attestation would be able to get a full and complete
understanding of the design and tradeoffs from the RFC alone.  If I have to
rereference a spec for anything other than double check params and magic values,
the entire series is getting ignored.

As for guest-driven quoting, to me there is a fairly straightforward solution.
Assuming using guest-donated memory is too complex for the "virtual thread":

 1. Drop "virtual thread pools" from quoting and support at most one in-flight
    quote per TD.  I doubt whatever memory is needed for a virtual thread is so
    insanely ridiculous that we can't burn that much memory per TD.  32KiB is a
    no-brainer.  Above that and we'd have to reconsider, but if quoting requires
    more than 32KiB of scratch space, I have to wonder what on earth it's doing.

    Then, because the resource is fixed, have the host provide it at TD creation.
    The TDX-Module keeps track of whether or not a quote is in-flight, and rejects
    attempts to start a new one.  The guest can simply guard quotes with a mutex.

    The host never has to worry about canceling a quote request, because it just
    needs to reclaim the TD as normal.  If desired, the TDX Module can provide a
    TDCALL to let the guest cancel a quote.

 2. To avoid having to check for guest interrupts, for the "SW based" synchronous
    method, simply resume the guest after processing a fixed amount of state.  Yes,
    that will increase the best case latency for a single quote, but *best* case
    latency isn't a huge concern, and I doubt the overhead of a VMX roundtrip will
    significantly impact that.  It's the tail latencies that will be problematic,
    and a forcing the serialization into the guest (the aforementioned mutex) means
    the tail latencies will only be affected by host activity on *that* CPU, which
    is more or less the status quo.

    E.g. even assuming an absurd 20% overhead for the VMX round trips, having a
    quote take ~1.2 *every* time is would be a far, far better experience than a
    quote taking 1ms - 10ms to complete.

 3. Host interrupts work as they do for normal TD operation.  The state is all
    there, attached to the vCPU.  It's on the host to run the vCPU according to
    its SLOs, so as not to starve/DoS the guest.

 4. If/when S3M accesses during quoting come along (it's not clear to me if the
    SW-based quoting needs to access the S3M during every quote, or just during
    setup), then provide the host with a mechanism to throttle quote requests.
    E.g. a simple counter that triggers an exit to the host when it hits zero
    would probably suffice.

 5. If/when "HW based" asynchronous quoting comes along, then if S3M can send
    IRQs on completion, have a host-side driver wired up to call into the TDX
    Module to tell it a quote is ready.  Presumably the TDX Module can keep track
    of where the quote came from.  If the S3M can't do IRQs, then just put the
    onus on the guest to poll to see when it's quote is ready?  This needs more
    details on the S3M, but it doesn't seem insurmountable.

    But it's also not obvious to me why Linux/KVM would ever want to support this
    mode.  AFAICT, it adds more complexity (especially when accounting for noisy
    neighbor issues) in order to provide a worse customer experience. 

#2 above is going to be a problem with the host-based, CPU-intensive "SW" quoting,
which IIUC is what is currently being proposed.  Unless the host provides an overhead
CPU to do quoting (which is likely a non-starter for Google at least), from the
guest's perspective, a vCPU will disappear and become non-responsive for however
long it takes the host to generate the quote.

The SetupEventNotifyInterrupt will help a little, but won't fully alleviate the
problem, because that still requires a pCPU to do the work.  E.g. the vCPU wouldn't
be completely non-responsive, but it would observe extremely high steal-time until
the quote is generated, which is gross.  Presumably this is how the current SGX-
based quoting behaves?  But if these quotes are getting more massive, i.e. slower,
then it's probably something that needs to be addressed.

And that's assuming the virtual thread pool is sized such that there can be a thread
per VM.  If there is any cross-VM resource sharing, then we're going to have noisy
neighbor problems and the above issue becomes an order of magnitude worse.

> > > A future thing is a flow where S3M is engaged for every quote. That would be
> > > expected to take longer, and have greater limits on concurrency. But the
> > > details are not sorted on what exactly the user will want, or how it would
> > > be implemented.
> > >
> > > I think you are maybe wondering whether S3M could be used such that the
> > > guest could wait on a quote while not needing any saved state area?
> >
> > I'm trying to figure out if *any* path is viable.
> >
> > Burning 1ms of CPU time to generate a quote in uninterruptible code is a non-
> > starter. Hell, 100us is a non-starter.
>
> What code is uninterruptible? The whole point of the extensions thing is to make
> the operations broadly interruptible.

Due to the lack of useful changelogs and being allergic to TDX specs, it wasn't
clear to me that TDH.GET.QUOTE was "fully" interruptible *and* restartable,
i.e. would guarantee forward progress.

> > Waiting 2s for a quote to come back from the S3M is a non-starter.
>
> Not sure where 2s is coming from. But do you mean waiting in the guest? Or where
> is waiting a non-starter?

If each S3M quote takes "tens of milliseconds", and there's one S3M per socket,
it doesn't take that many concurrent quote requests for one of the TDs to observe
a 2s+ latency to get its quote.  E.g. quote takes 50ms, boot 40 TDs, and voila,
that last TD going through boot gets hit with a 2s+ quote latency.

> > The numbers matter, and *none* of this is reviewable, even in RFC format,
> > without a crisp understanding of what latencies we are talking about.
>
> I agree not having a measurement of the initial DICE behavior is a detriment for
> the TDG discussion. I didn't think we needed it though, 

Ignore the TDG discussion, I'm saying they're needed for *any* discussion, i.e.
for reviewing the TDH implementation as well.

> if we leaned on the original locking/scheduling reasons to prefer a host
> based quote flow.
>
> Previously you asked for ballparks and got them. So what do you need exactly?

Guarantees around orders of magnitude.  There is a enormous difference between
a CPU-intensive operation taking 10us versus 2ms.  I.e. "milliseconds or less"
isn't a ballpark, it's a country, maaaaybe a state.  For CPU-based in particular,
I was expecting ballparks with +/- 10us of precision, not "somewhere between 0
and several scheduler ticks".

[1] https://lore.kernel.org/all/20260522034128.3144354-1-yilun.xu@linux.intel.com
[2] https://lore.kernel.org/all/akaKmEZnTY4FO2gY@yilunxu-OptiPlex-7050

  reply	other threads:[~2026-09-18  0:04 UTC|newest]

Thread overview: 33+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-09-16  0:48 TDG quote analysis Edgecombe, Rick P
2026-09-16 18:00 ` Sean Christopherson
2026-09-16 18:49   ` Edgecombe, Rick P
2026-09-16 19:27     ` Sean Christopherson
2026-09-16 22:31       ` Edgecombe, Rick P
2026-09-16 23:00         ` Peter Fang
2026-09-16 23:57           ` Sean Christopherson
2026-09-17  0:56             ` Edgecombe, Rick P
2026-09-17 13:37               ` Sean Christopherson
2026-09-17 18:12                 ` Edgecombe, Rick P
2026-09-17 19:57                   ` Sean Christopherson
2026-09-17 21:32                     ` Edgecombe, Rick P
2026-09-18  0:04                       ` Sean Christopherson [this message]
2026-09-18  2:39                         ` Edgecombe, Rick P
2026-09-18 13:13                           ` Sean Christopherson
2026-09-18 18:12                             ` Edgecombe, Rick P
2026-09-18 21:32                               ` Sean Christopherson
2026-09-21 23:00                                 ` Peter Fang
2026-09-21 23:06                                   ` Dave Hansen
2026-09-23  4:09                                     ` Vishal Annapurve
2026-09-24  0:03                                       ` Vishal Annapurve
2026-09-24  0:18                                         ` Sean Christopherson
2026-09-24  0:44                                           ` Edgecombe, Rick P
2026-09-24 16:28                                             ` Sean Christopherson
2026-09-21 23:04                                 ` Dave Hansen
2026-09-21 23:19                                   ` Sean Christopherson
2026-09-21 23:36                                     ` Dave Hansen
2026-09-21 23:46                                       ` Sean Christopherson
2026-09-22  0:03                                         ` Dave Hansen
2026-09-22  6:48                                           ` Reshetova, Elena
2026-09-22 17:07                                             ` Edgecombe, Rick P
2026-09-23  8:50                                               ` Reshetova, Elena
2026-09-23 22:50                                                 ` Peter Fang

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=aqyAED5ZK48eEZtq@google.com \
    --to=seanjc@google.com \
    --cc=binbin.wu@intel.com \
    --cc=dave.hansen@intel.com \
    --cc=elena.reshetova@intel.com \
    --cc=kas@kernel.org \
    --cc=kvm@vger.kernel.org \
    --cc=pbonzini@redhat.com \
    --cc=peter.fang@intel.com \
    --cc=rick.p.edgecombe@intel.com \
    --cc=vannapurve@google.com \
    --cc=yilun.xu@intel.com \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox