Kernel KVM virtualization development
 help / color / mirror / Atom feed
From: "Edgecombe, Rick P" <rick.p.edgecombe@intel.com>
To: "pbonzini@redhat.com" <pbonzini@redhat.com>,
	"Hansen, Dave" <dave.hansen@intel.com>,
	"seanjc@google.com" <seanjc@google.com>,
	"kas@kernel.org" <kas@kernel.org>
Cc: "Reshetova, Elena" <elena.reshetova@intel.com>,
	"Xu, Yilun" <yilun.xu@intel.com>,
	"kvm@vger.kernel.org" <kvm@vger.kernel.org>,
	"Annapurve, Vishal" <vannapurve@google.com>,
	"Fang, Peter" <peter.fang@intel.com>,
	"Wu, Binbin" <binbin.wu@intel.com>
Subject: TDG quote analysis
Date: Wed, 16 Sep 2026 00:48:02 +0000	[thread overview]
Message-ID: <24a751d16ba0dc37aa7774df5739f39734cf9aca.camel@intel.com> (raw)

Here is the overdue analysis of what a TDG based TDX quote operation could look
like, pulled together as writeup somewhat quickly. This was from Sean's request
originally. The writeup is based on a bunch of gathering by Peter, Elena and I.
Some TDX interrupt input from Binbin.

I included problem background, after the delay. If you remember it all, skip to
"Possible TDG call for quoting".

And of course, the proposals here are not remotely finalized or signed off on.
So take it like some ideas to discuss.

--

TDX splits and sometimes duplicates the functionality of virtualizing a guest
across the host VMM and the TDX module. The hard requirement from the TDX
module’s side is to ensure security for the guest, and from the VMM’s side the
hard requirement is to not interfere with the host’s other responsibilities.
Together they need to do both while keeping the guest running well.

One of the ways TDX avoids interfering with the hosts other responsibilities is
to have a limit for how long a SEAMCALL can be executing in the TDX module
without exiting back to the host to handle an interrupt. This way the host can
grab control back, handle important interrupts, etc. When a long running TDX
module SEAMCALL would exceed this time, it can be interrupted. At that point it
can either be restarted or resumed. 

TDX has grown some new internal machinery to support the resumable variant of
these long running SEAMCALLs. This machinery is called “TDX extensions”. One of
the first usages of this new machinery is a way to get a TD Quote. Today this is
done via a call into an SGX enclave. In the future the host can instead call
into one of these long running TDX module SEAMCALLs, called TDH.QUOTE.GET.

== Current Attestation Design == 

A quick refresher on why this is all so complicated. On initial TDX HW, the only
way to get the quote signed was in SGX. But the only thing that could know the
TD details that need to be in the quote was the TDX module. So the TD details
are gathered into a “report” and passed out of the TD all the way to host
userspace where SGX runs. Then userspace gets the quote out of SGX and passes it
back to the TD. There is one other piece of info that needs to be in the quote,
and that is the nonce. It comes from the remote party attesting the TD, so even
if SGX knew the TD details, the nonce still needs to get to the thing doing the
quote construction and signing. At the time this was so custom that the quote
format ended up a TDX specific thing. 

But now there is DICE, which is a cross industry format for doing attestation
stuff. TDX is making two changes at once. Supporting this new format and adding
quote generation via the TDX module. 

This DICE quote generation can take a long time. Exactly how long is not known,
but it is expected to vary and often exceed the time a SEAMCALL can be executing
in the TDX module by a considerable amount. At the same time it is expected to
not be easy to break apart into checkpoints such that it could be made a
resumable SEAMCALL the way the existing ones work. Hence this framework to make
it easier save and resume state for these long running TDX module SEAMCALLs.

== Quoting extension current arch == 

   -- Saving state while being interrupted and resuming -- 
   
   In order to resume an operation later, the TDX module needs a place to save
   state when interrupted. The number of save state areas can be configured at
   TDX module setup time, which translates roughly to the number of concurrent
   operations. But also, the quoting extension can involve a limited HW resource
   which is expected to in the future sometimes functionally limit quote
   operation to one at a time. 

   -- Locking -- 
   
   Since there will only be so many concurrent operations at runtime, to ensure
   fairness some synchronization is needed on the VMM side. This can be a simple
   mutex around the SEAMCALL.
   
   -- Host Interrupts -- 
   
   A host interrupt causes the extension SEAMCALL to return a
   TDX_INTERRUPTED_RESUMABLE error, and the interrupt fires on the host. This is
   pretty much exactly like the existing TDX_INTERRUPTED_RESUMABLE calls, except
   that the host has control of how many concurrent operations are possible. It
   has to manage the concurrency based on how it configured this. For example,
   if the host gets a TDX_INTERRUPTED_RESUMABLE return error and never returns
   to complete the operation, it will block any other operations from one of the
   concurrent state slots.
   
   -- Guest Interrupts -- 
   
   It is not nice to keep the guest from running for too long. The TDH.QUOTE.GET
   SEAMCALL monitors for host interrupts, but not guest ones. Similarly, the SGX
   quoting enclave doesn't know what is happening in the TD. So, to help reduce
   guest latencies, the GHCI exposes a way to register for a notification
   (SetupEventNotifyInterrupt) when the quoting is finished. A quote TDVMCALL
   handler can start the quote operation on another host thread, and resume
   guest execution waiting for the quote to finish.
   
== Possible TDG call for quoting == 

The question recently came up from Sean (actually not that recently at this
point): While this is all getting redone, why not just go directly to the TDX
module from the guest to get the quote? Skip all the host buffer passing. This
was considered difficult to do when previously considered. We investigated it
further. Here is what it would look like.

   -- Saving state while being interrupted and resuming -- 
   
   Whereas the host would be involved in ensuring fairness for the TDH call, for
   a TDG call, the TDX module would have to do this by itself. It turns out
   there is an existing (but unused) TDG call that has a similar problem to
   solve: TDG.MR.ASSIGNSVNS. It has a fixed rate at which the TD is allowed to
   call the TDG call, 200usec. If the TD calls it more frequently, the TDX
   Module exits to the host with a TDX_TDCALL_RATE_LIMIT.
   
   A resumable TDG quote call would have to solve a slightly different problem.
   A bad guest could refuse to resume the TDG call and hold the TDX module
   extension save state slot hostage.
   
   So instead, a TDG call could work by remotely canceling another TD’s paused
   operation if that TD didn't call back frequently enough. After a time-limit
   was reached on an extension save state slot usage, another TD’s TDG call
   could abort the stalled TD’s operation. There existing designs around doing
   this cancellation already in the TDX extension arch. It would need to be
   adapted to the guest side calls and incorporate the extra cancellation work
   into the time limit math. Some good timeouts would need to be picked.
   
   -- Locking -- 
   
   The TDX module would need to have some locks to manage the above fairness
   solution. It would have to do this in a way to ensure no TD ever got starved.
   This kind of global TDX locking is new for TDG calls.
   
   There would have to be some heuristic to handle the problem of a TD starving the
   other callers by repeatedly requesting quotes. Sean had previously tossed out
   the idea of having a TD get a limited number of quote operations. Maybe make it
   per some time period. A more complex fairness scheme would have the TDX module
   maintain a queue somehow. Another option still, could be handle it like
   TDX_TDCALL_RATE_LIMIT. Then the host could decide to throttle misbehaving TDs.
   
   A TDG quote call would otherwise expect to take similar locks from the guest
   side as are already taken by TDG.MR.REPORT. Long term, more details are
   expected in the quote, so a TDG direction would be signing up to potentially
   navigate more guest-host locking issues. Like we have in a few places.
   
   And like the other contention issue, some good limits would need to be picked.
   
   -- Host Interrupts -- 
   
   When a host interrupt happens, the TDX module needs to return to the host VMM
   to handle it. While executing a long running TDG call, the TDX module would
   need to do the same thing. But it also would have to remember what it was
   doing when the host attempts to re-enter the TD, and instead resume the guest
   side quote operation. If the VMM never re-entered the TD, the save state slot
   would have to be subject to reclaim similarly to if the guest never resumed
   it.
   
   -- Guest Interrupts -- 
   
   If a quote operation is long enough, the TDG SEAMCALL should probably support
   a resumable-like flow from the guest side too, where it also monitors for
   pending guest interrupts and returns from the TDG call to let the guest
   handle them. Then the host wouldn't need a thread and a completion guest
   notification mechanism. It basically transfers that complexity to the TDX
   module.
   
   But it also might transfer some of the control. The TDX module would have to
   embed some policy on how to decide when to inject guest interrupts. If the
   host and TDX module had different logic on when to actually re-enter the TD,
   that could be annoying.
   
== Summary ==

From what we found, I think it doesn't seem overly impossible for the module. But
the exact cancellation rules and timeout logic would need to be hashed out a bit
more. The simple mutex on the host side does still seem a fair amount simpler than
the TDX module's fairness options.

How do we feel about the roles and responsibility shifts? Any other problems to
solve? Or errors in these proposals? What kind of fairness heuristics are
acceptable for TDX users?

             reply	other threads:[~2026-09-16  0:48 UTC|newest]

Thread overview: 33+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-09-16  0:48 Edgecombe, Rick P [this message]
2026-09-16 18:00 ` TDG quote analysis Sean Christopherson
2026-09-16 18:49   ` Edgecombe, Rick P
2026-09-16 19:27     ` Sean Christopherson
2026-09-16 22:31       ` Edgecombe, Rick P
2026-09-16 23:00         ` Peter Fang
2026-09-16 23:57           ` Sean Christopherson
2026-09-17  0:56             ` Edgecombe, Rick P
2026-09-17 13:37               ` Sean Christopherson
2026-09-17 18:12                 ` Edgecombe, Rick P
2026-09-17 19:57                   ` Sean Christopherson
2026-09-17 21:32                     ` Edgecombe, Rick P
2026-09-18  0:04                       ` Sean Christopherson
2026-09-18  2:39                         ` Edgecombe, Rick P
2026-09-18 13:13                           ` Sean Christopherson
2026-09-18 18:12                             ` Edgecombe, Rick P
2026-09-18 21:32                               ` Sean Christopherson
2026-09-21 23:00                                 ` Peter Fang
2026-09-21 23:06                                   ` Dave Hansen
2026-09-23  4:09                                     ` Vishal Annapurve
2026-09-24  0:03                                       ` Vishal Annapurve
2026-09-24  0:18                                         ` Sean Christopherson
2026-09-24  0:44                                           ` Edgecombe, Rick P
2026-09-24 16:28                                             ` Sean Christopherson
2026-09-21 23:04                                 ` Dave Hansen
2026-09-21 23:19                                   ` Sean Christopherson
2026-09-21 23:36                                     ` Dave Hansen
2026-09-21 23:46                                       ` Sean Christopherson
2026-09-22  0:03                                         ` Dave Hansen
2026-09-22  6:48                                           ` Reshetova, Elena
2026-09-22 17:07                                             ` Edgecombe, Rick P
2026-09-23  8:50                                               ` Reshetova, Elena
2026-09-23 22:50                                                 ` Peter Fang

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=24a751d16ba0dc37aa7774df5739f39734cf9aca.camel@intel.com \
    --to=rick.p.edgecombe@intel.com \
    --cc=binbin.wu@intel.com \
    --cc=dave.hansen@intel.com \
    --cc=elena.reshetova@intel.com \
    --cc=kas@kernel.org \
    --cc=kvm@vger.kernel.org \
    --cc=pbonzini@redhat.com \
    --cc=peter.fang@intel.com \
    --cc=seanjc@google.com \
    --cc=vannapurve@google.com \
    --cc=yilun.xu@intel.com \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox