Kernel KVM virtualization development
 help / color / mirror / Atom feed
From: "Edgecombe, Rick P" <rick.p.edgecombe@intel.com>
To: "seanjc@google.com" <seanjc@google.com>
Cc: "Xu, Yilun" <yilun.xu@intel.com>,
	"Reshetova, Elena" <elena.reshetova@intel.com>,
	"Wu, Binbin" <binbin.wu@intel.com>,
	"Hansen, Dave" <dave.hansen@intel.com>,
	"Annapurve, Vishal" <vannapurve@google.com>,
	"kas@kernel.org" <kas@kernel.org>,
	"pbonzini@redhat.com" <pbonzini@redhat.com>,
	"Fang, Peter" <peter.fang@intel.com>,
	"kvm@vger.kernel.org" <kvm@vger.kernel.org>
Subject: Re: TDG quote analysis
Date: Fri, 18 Sep 2026 02:39:25 +0000	[thread overview]
Message-ID: <56a61983c03806120e607dfd1e1b7c9046c1f54b.camel@intel.com> (raw)
In-Reply-To: <aqyAED5ZK48eEZtq@google.com>

Thanks for the detailed response.

On Thu, 2026-09-17 at 17:04 -0700, Sean Christopherson wrote:
> On Thu, Sep 17, 2026, Rick P Edgecombe wrote:
> > On Thu, 2026-09-17 at 12:57 -0700, Sean Christopherson wrote:
> > > > Let's take a step back here. There is an existing attestation flow that was
> > > > designed around some limitations that are changing (specifically whether the
> > > > quoter has knowledge of the TD). The current host side quote discussion (KVM
> > > > ioctl) is basically a straight forward evolution of the existing design,
> > > > even though the limitations are getting removed.
> > > > 
> > > > You asked whether we could do a re-design that makes more sense in the
> > > > context of the lack of that old limitation. At that point we are faced with
> > > > the age-old question: how far into the fuzzy future should we design around?
> > > 
> > > Now I'm trying to understand what the *current* plan is.
> > 
> > The last Linux design for a host-based TDX DICE API, meaning the thing we were
> > prepping for the next DICE posting before getting into this TDG analysis:
> >  - Leave the existing report in the guest as is.
> >  - Utilize the existing quote GHCI call from the guest to pass the old report,
> > which goes through KVM to userspace.
> >  - Add a new VM scoped ioctl to KVM to call into the TDH.QUOTE.GET. This gets
> > the quote and passes it back to userspace. Then userspace uses the existing GHCI
> > mechanisms to notify the guest that it is ready.
> > 
> > > 
> > > > The nearest term thing is a SW based flow where work happens on the CPU. But
> > > > the exact amount of time is not know yet. Peter gave a ballpark.
> > > 
> > > Why are we even discussing this?  I am so confused.  I thought there were two
> > > options: SGX and S3M.  Now all of a sudden there's a third "let's do insane
> > > things in software in the TDX module" option!?!?
> > 
> > ???
> > 
> > So when I said:
> >    "HW" is talking about the S3M thing. The quote operation could go directly to
> >    the S3M to get the quote. (HW based) Or it could get an intermediate key and
> >    generate quotes using CPU instructions. (SW based). Think like a crypto library
> >    in the TDX module.
> >    
> >    ...
> > 
> >    The software based flow would be expected to first. It would involve the CPU
> >    doing crypto stuff as above.
> >    
> >    Then a HW based flow where the crypto happens on the limited HW resource. This
> >    is where full parallelization is not possible, because the CPU is not doing the
> >    heavy work. You might want this one instead for security reasons. But the main
> >    point of discussing it is that you could expect some quotes to take a long time
> >    and support a limited number of parallel quotes.
> >    
> > Did you interpret SW based flow to be talking about SGX? Like the intermediate key
> > goes from S3M to SGX?
> 
> From *before* this conversation.  Forget this converation, what does the mock
> TDX Module used as the basis for the RFC[1] do?
> 

Ah ok, well I can see how the RFC was confusing. But I am surprised the
discussion up the thread didn't cover a lot of these points.

>   Because the RFC says absolutely
> *nothing*.  I kinda sorta have a picture now, but it required hunting down an
> additional spec, and the documentnation still leaves me wanting.  Because I still
> don't know what it actually does.  Based on everything you're saying, I *assume*
> it's this "SW based" flow, but *nothing* actually says that.
> 
> Under "Interruptibility", the quoting doc linked says:
> 
>   TDH.QUOTE.GET is interruptible. If a pending interrupt is detected during
>   operation, TDH.QUOTE.GET returns with a TDX_INTERRUPTED_RESUMABLE status in RAX.
> 
> Then punts me to:
> 
>   For the general explanation on how interruption and resumption is handled for
>   all Quoting Service functions please consult [Intel TDX Module Base Spec] section
>   “Request Interruption and Resumption”.
> 
> And finishes with the wonderful:
> 
>   Rest of details are TBD
> 
> Once I finally found the "Request Interruption and Resumption", I discovered that
> the TDX Module now has "virtual thread pools", which IS NEVER MENTIONED IN THE
> RFC.  Seriously.  No one thought that was worth mentioning?
> 
> OMG.  I see it now:
> 
> +               /* Don't bother specifying the quote id */
> +               .rdx = QUOTE_ID_MASK & (u64)-1,
> 
> So IIUC, there's a thread pool, but somewhere in the muck of RFC patches the
> kernel initializes the pool with a size of 1?
> 

Yea it is like a thread pool. Described above in more abstract terms of a
resumable operation that can save state and can have the number of save state
areas configured.

>   And then adds a global mutex to
> serialize quotes across the entire system.  That's certainly a choice.

Passing the GHCI request straight to the seamcall within KVM was a bad choice.
But I think a single mutex is a possible start for a KVM ioctl(). Some follow-on
patches could allow for configuring more options.

> 
> This snippet from the RFC[2] doesn't help, because it makes it seem like there
> is no state:
> 
>   2. Host just refills the previous args (may have been modified by seamcall
>      output) on retry. such as TDH_EXT_INIT, TDH_EXT_MEM_ADD,
>      TDH_QUOTE_INIT, TDH_QUOTE_GET...
> 
> But that's presumably just because the kernel is using a single REQUEST_ID.
> 
> That spec continues with:
> 
>   If supported for the specific request type, the host VMM can request an
>   interrupted request to be aborted.
> 
> Though per the above, apparently whether or not that's going to be supported is
> TBD?

So extension aborts are a whole 'nother topic. And "support" has a nuanced
answer to boot. I'd be very glad to hear your thoughts on it. Should we leave it
for a later topic or get into it now? I *think* we don't need it, but the more
we get into the details here, the more it's a notable thing we haven't covered.

> 
> I am beyond frustrated at this point.  I guess shame on me for not pouring over
> dense docs?  But I shouldn't have to.
> 

Yea, I agree you shouldn't have to. But I think a lot of this stuff did get
covered in this thread. If the post-RFC state was confusion, we can only try to
rectify it later from that point.

>   The *ENTIRE* point of an RFC like the one
> that started all this is to get feedback on the design, but that's really hard to
> achieve if you don't actually provide *any* details about the design.
> 
> And there are MASSIVE design choices in here.  Like sizing the "thread" pool to 1,
> and thus deliberately serializing quotes across all VMs.  That's going to end well
> when someone spins up 10s or 100s of TDs and they all try to attest at once.
> Not listening for fatal signals while spinning on TDH.GET.QUOTE will save us though!
> At least with PREEMPT_LAZY being forced, not doing cond_resched() is "fine".
> 
> What's worse, some of these patches that say nothing useful in the changelogs
> have several Reviewed-by tags, which means multiple people looked at this as
> thought it was all good.
> 
> Regardless of where we end up with quoting, whatever RFC is sent next needs to be
> a million times better.  My expectations are that someone with core TDX knowledge
> and a basic understanding of attestation would be able to get a full and complete
> understanding of the design and tradeoffs from the RFC alone.  If I have to
> rereference a spec for anything other than double check params and magic values,
> the entire series is getting ignored.

I mean. Confidential compute in general is trying to reinvent decades of
virtualization solutions with one hand tied behind it's back. There *are* a lot
of design choices.

Besides a rushed RFC, I think another factor is that this is the first TDX
feature we have done that you don't already have at least some exposure to. The
level of background delta need is way higher than S-EPT management for example.

And I'm less and less sure we are going to be successful distilling things like
this down in a way that doesn't leave you frustrated. But if you want details
there is just going to be a lot of them. Hmm.

> 
> As for guest-driven quoting, to me there is a fairly straightforward solution.
> Assuming using guest-donated memory is too complex for the "virtual thread":
> 
>  1. Drop "virtual thread pools" from quoting and support at most one in-flight
>     quote per TD.  I doubt whatever memory is needed for a virtual thread is so
>     insanely ridiculous that we can't burn that much memory per TD.  32KiB is a
>     no-brainer.  Above that and we'd have to reconsider, but if quoting requires
>     more than 32KiB of scratch space, I have to wonder what on earth it's doing.

I think the abstract descriptions are not helping, so I'll try some more
details. Today the "thread" involves something like a vCPU in a TD. You can see
some references in the TDX base spec like:
   Currently, TDX Module extensions are implemented as guests running in SEAM non-
   root mode, called NRX Modules. Technically, each NRX Module is similar to a TD.
   However, logically and from the usage perspective NRX Modules are part of the
   TDX Module and are hidden from the host VMM. The TDX Module builds the NRX
   Modules as part of the TDX Module’s initialization sequence and calls them when
   their functionality is required for execution of some TDX Module interface
   functions.
   
Think like the service TD concept got wrapped in the TDX module. *But* the idea
is to be more like an abstract thing. Being a TD is just an implementation
detail and not necessarily a long term fixed thing. So based on that, I'd think
the memory needs could be optimized.

Let's track down the exact size. I've seen it but forgot the exact number.

> 
>     Then, because the resource is fixed, have the host provide it at TD creation.
>     The TDX-Module keeps track of whether or not a quote is in-flight, and rejects
>     attempts to start a new one.  The guest can simply guard quotes with a mutex.
> 
>     The host never has to worry about canceling a quote request, because it just
>     needs to reclaim the TD as normal.  If desired, the TDX Module can provide a
>     TDCALL to let the guest cancel a quote.

This idea came up before actually, even on the host side solution IIRC. Yea. The
tradeoff is just overhead per-TD.

> 
>  2. To avoid having to check for guest interrupts, for the "SW based" synchronous
>     method, simply resume the guest after processing a fixed amount of state.  Yes,
>     that will increase the best case latency for a single quote, but *best* case
>     latency isn't a huge concern, and I doubt the overhead of a VMX roundtrip will
>     significantly impact that.  It's the tail latencies that will be problematic,
>     and a forcing the serialization into the guest (the aforementioned mutex) means
>     the tail latencies will only be affected by host activity on *that* CPU, which
>     is more or less the status quo.
> 
>     E.g. even assuming an absurd 20% overhead for the VMX round trips, having a
>     quote take ~1.2 *every* time is would be a far, far better experience than a
>     quote taking 1ms - 10ms to complete.

Dave will probably not get a chance to respond this week, but we should maybe
discuss what acceptable guest latency actually is.

This thread left me thinking that we might not need to handle guest interrupts
immediately. But it would be good for the solution to be able to extend to
support them later if needed.

> 
>  3. Host interrupts work as they do for normal TD operation.  The state is all
>     there, attached to the vCPU.  It's on the host to run the vCPU according to
>     its SLOs, so as not to starve/DoS the guest.
> 
>  4. If/when S3M accesses during quoting come along (it's not clear to me if the
>     SW-based quoting needs to access the S3M during every quote, or just during
>     setup), then provide the host with a mechanism to throttle quote requests.
>     E.g. a simple counter that triggers an exit to the host when it hits zero
>     would probably suffice.
> 
>  5. If/when "HW based" asynchronous quoting comes along, then if S3M can send
>     IRQs on completion, have a host-side driver wired up to call into the TDX
>     Module to tell it a quote is ready.  Presumably the TDX Module can keep track
>     of where the quote came from.  If the S3M can't do IRQs, then just put the
>     onus on the guest to poll to see when it's quote is ready?  This needs more
>     details on the S3M, but it doesn't seem insurmountable.
> 
>     But it's also not obvious to me why Linux/KVM would ever want to support this
>     mode.  AFAICT, it adds more complexity (especially when accounting for noisy
>     neighbor issues) in order to provide a worse customer experience. 

I think it is like "HW security" kind of value to some people. Think like a TPMs
value vs an in memory protected key. Depends on your outlook and other factors.
But I'm fine (glad) to not figure out "HW based" attestation today. Just want a
re-design to be able to grow for stuff like that.

> 
> #2 above is going to be a problem with the host-based, CPU-intensive "SW" quoting,
> which IIUC is what is currently being proposed.  Unless the host provides an overhead
> CPU to do quoting (which is likely a non-starter for Google at least), from the
> guest's perspective, a vCPU will disappear and become non-responsive for however
> long it takes the host to generate the quote.
> 
> The SetupEventNotifyInterrupt will help a little, but won't fully alleviate the
> problem, because that still requires a pCPU to do the work.  E.g. the vCPU wouldn't
> be completely non-responsive, but it would observe extremely high steal-time until
> the quote is generated, which is gross.  Presumably this is how the current SGX-
> based quoting behaves?  But if these quotes are getting more massive, i.e. slower,
> then it's probably something that needs to be addressed.

Hmm, yea. I mean it would be the same for SGX today, but I could totally see it
being a factor. Maybe some TDX users on CC will chime in on this one. This could
be the kind of thing that hinges the design clearly.

The other option would be, ugh, the TDH.QUOTE.GET becomes vCPU scoped and exits
for pending guest interrupts. Then it could use the vCPU thread. But it is a
long way from the guest at that point.

> 
> And that's assuming the virtual thread pool is sized such that there can be a thread
> per VM.  If there is any cross-VM resource sharing, then we're going to have noisy
> neighbor problems and the above issue becomes an order of magnitude worse.
> 
> > > > A future thing is a flow where S3M is engaged for every quote. That would be
> > > > expected to take longer, and have greater limits on concurrency. But the
> > > > details are not sorted on what exactly the user will want, or how it would
> > > > be implemented.
> > > > 
> > > > I think you are maybe wondering whether S3M could be used such that the
> > > > guest could wait on a quote while not needing any saved state area?
> > > 
> > > I'm trying to figure out if *any* path is viable.
> > > 
> > > Burning 1ms of CPU time to generate a quote in uninterruptible code is a non-
> > > starter. Hell, 100us is a non-starter.
> > 
> > What code is uninterruptible? The whole point of the extensions thing is to make
> > the operations broadly interruptible.
> 
> Due to the lack of useful changelogs and being allergic to TDX specs, it wasn't
> clear to me that TDH.GET.QUOTE was "fully" interruptible *and* restartable,
> i.e. would guarantee forward progress.

I do think we covered the interruptible part pretty thoroughly in this thread.

> 
> > > Waiting 2s for a quote to come back from the S3M is a non-starter.
> > 
> > Not sure where 2s is coming from. But do you mean waiting in the guest? Or where
> > is waiting a non-starter?
> 
> If each S3M quote takes "tens of milliseconds", and there's one S3M per socket,
> it doesn't take that many concurrent quote requests for one of the TDs to observe
> a 2s+ latency to get its quote.  E.g. quote takes 50ms, boot 40 TDs, and voila,
> that last TD going through boot gets hit with a 2s+ quote latency.

Ah I see what you are getting at.

> 
> > > The numbers matter, and *none* of this is reviewable, even in RFC format,
> > > without a crisp understanding of what latencies we are talking about.
> > 
> > I agree not having a measurement of the initial DICE behavior is a detriment for
> > the TDG discussion. I didn't think we needed it though, 
> 
> Ignore the TDG discussion, I'm saying they're needed for *any* discussion, i.e.
> for reviewing the TDH implementation as well.
> 
> > if we leaned on the original locking/scheduling reasons to prefer a host
> > based quote flow.
> > 
> > Previously you asked for ballparks and got them. So what do you need exactly?
> 
> Guarantees around orders of magnitude.  There is a enormous difference between
> a CPU-intensive operation taking 10us versus 2ms.  I.e. "milliseconds or less"
> isn't a ballpark, it's a country, maaaaybe a state.  For CPU-based in particular,
> I was expecting ballparks with +/- 10us of precision, not "somewhere between 0
> and several scheduler ticks".
> 

Ok, I think that Peter was giving best guesses. We can try to find a worst case
commitment.

But it does seem like the PQC transition could end up with some changes that are
totally out of control. Like that hybrid crypto lwn article I linked. How much
does 10us precision on the first thing really buy us when discussing a re-design
that would hopefully live a longer time than SGX based?

  reply	other threads:[~2026-09-18  2:39 UTC|newest]

Thread overview: 33+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-09-16  0:48 TDG quote analysis Edgecombe, Rick P
2026-09-16 18:00 ` Sean Christopherson
2026-09-16 18:49   ` Edgecombe, Rick P
2026-09-16 19:27     ` Sean Christopherson
2026-09-16 22:31       ` Edgecombe, Rick P
2026-09-16 23:00         ` Peter Fang
2026-09-16 23:57           ` Sean Christopherson
2026-09-17  0:56             ` Edgecombe, Rick P
2026-09-17 13:37               ` Sean Christopherson
2026-09-17 18:12                 ` Edgecombe, Rick P
2026-09-17 19:57                   ` Sean Christopherson
2026-09-17 21:32                     ` Edgecombe, Rick P
2026-09-18  0:04                       ` Sean Christopherson
2026-09-18  2:39                         ` Edgecombe, Rick P [this message]
2026-09-18 13:13                           ` Sean Christopherson
2026-09-18 18:12                             ` Edgecombe, Rick P
2026-09-18 21:32                               ` Sean Christopherson
2026-09-21 23:00                                 ` Peter Fang
2026-09-21 23:06                                   ` Dave Hansen
2026-09-23  4:09                                     ` Vishal Annapurve
2026-09-24  0:03                                       ` Vishal Annapurve
2026-09-24  0:18                                         ` Sean Christopherson
2026-09-24  0:44                                           ` Edgecombe, Rick P
2026-09-24 16:28                                             ` Sean Christopherson
2026-09-21 23:04                                 ` Dave Hansen
2026-09-21 23:19                                   ` Sean Christopherson
2026-09-21 23:36                                     ` Dave Hansen
2026-09-21 23:46                                       ` Sean Christopherson
2026-09-22  0:03                                         ` Dave Hansen
2026-09-22  6:48                                           ` Reshetova, Elena
2026-09-22 17:07                                             ` Edgecombe, Rick P
2026-09-23  8:50                                               ` Reshetova, Elena
2026-09-23 22:50                                                 ` Peter Fang

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=56a61983c03806120e607dfd1e1b7c9046c1f54b.camel@intel.com \
    --to=rick.p.edgecombe@intel.com \
    --cc=binbin.wu@intel.com \
    --cc=dave.hansen@intel.com \
    --cc=elena.reshetova@intel.com \
    --cc=kas@kernel.org \
    --cc=kvm@vger.kernel.org \
    --cc=pbonzini@redhat.com \
    --cc=peter.fang@intel.com \
    --cc=seanjc@google.com \
    --cc=vannapurve@google.com \
    --cc=yilun.xu@intel.com \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox