* TDG quote analysis
@ 2026-09-16 0:48 Edgecombe, Rick P
2026-09-16 18:00 ` Sean Christopherson
0 siblings, 1 reply; 33+ messages in thread
From: Edgecombe, Rick P @ 2026-09-16 0:48 UTC (permalink / raw)
To: pbonzini@redhat.com, Hansen, Dave, seanjc@google.com,
kas@kernel.org
Cc: Reshetova, Elena, Xu, Yilun, kvm@vger.kernel.org,
Annapurve, Vishal, Fang, Peter, Wu, Binbin
Here is the overdue analysis of what a TDG based TDX quote operation could look
like, pulled together as writeup somewhat quickly. This was from Sean's request
originally. The writeup is based on a bunch of gathering by Peter, Elena and I.
Some TDX interrupt input from Binbin.
I included problem background, after the delay. If you remember it all, skip to
"Possible TDG call for quoting".
And of course, the proposals here are not remotely finalized or signed off on.
So take it like some ideas to discuss.
--
TDX splits and sometimes duplicates the functionality of virtualizing a guest
across the host VMM and the TDX module. The hard requirement from the TDX
module’s side is to ensure security for the guest, and from the VMM’s side the
hard requirement is to not interfere with the host’s other responsibilities.
Together they need to do both while keeping the guest running well.
One of the ways TDX avoids interfering with the hosts other responsibilities is
to have a limit for how long a SEAMCALL can be executing in the TDX module
without exiting back to the host to handle an interrupt. This way the host can
grab control back, handle important interrupts, etc. When a long running TDX
module SEAMCALL would exceed this time, it can be interrupted. At that point it
can either be restarted or resumed.
TDX has grown some new internal machinery to support the resumable variant of
these long running SEAMCALLs. This machinery is called “TDX extensions”. One of
the first usages of this new machinery is a way to get a TD Quote. Today this is
done via a call into an SGX enclave. In the future the host can instead call
into one of these long running TDX module SEAMCALLs, called TDH.QUOTE.GET.
== Current Attestation Design ==
A quick refresher on why this is all so complicated. On initial TDX HW, the only
way to get the quote signed was in SGX. But the only thing that could know the
TD details that need to be in the quote was the TDX module. So the TD details
are gathered into a “report” and passed out of the TD all the way to host
userspace where SGX runs. Then userspace gets the quote out of SGX and passes it
back to the TD. There is one other piece of info that needs to be in the quote,
and that is the nonce. It comes from the remote party attesting the TD, so even
if SGX knew the TD details, the nonce still needs to get to the thing doing the
quote construction and signing. At the time this was so custom that the quote
format ended up a TDX specific thing.
But now there is DICE, which is a cross industry format for doing attestation
stuff. TDX is making two changes at once. Supporting this new format and adding
quote generation via the TDX module.
This DICE quote generation can take a long time. Exactly how long is not known,
but it is expected to vary and often exceed the time a SEAMCALL can be executing
in the TDX module by a considerable amount. At the same time it is expected to
not be easy to break apart into checkpoints such that it could be made a
resumable SEAMCALL the way the existing ones work. Hence this framework to make
it easier save and resume state for these long running TDX module SEAMCALLs.
== Quoting extension current arch ==
-- Saving state while being interrupted and resuming --
In order to resume an operation later, the TDX module needs a place to save
state when interrupted. The number of save state areas can be configured at
TDX module setup time, which translates roughly to the number of concurrent
operations. But also, the quoting extension can involve a limited HW resource
which is expected to in the future sometimes functionally limit quote
operation to one at a time.
-- Locking --
Since there will only be so many concurrent operations at runtime, to ensure
fairness some synchronization is needed on the VMM side. This can be a simple
mutex around the SEAMCALL.
-- Host Interrupts --
A host interrupt causes the extension SEAMCALL to return a
TDX_INTERRUPTED_RESUMABLE error, and the interrupt fires on the host. This is
pretty much exactly like the existing TDX_INTERRUPTED_RESUMABLE calls, except
that the host has control of how many concurrent operations are possible. It
has to manage the concurrency based on how it configured this. For example,
if the host gets a TDX_INTERRUPTED_RESUMABLE return error and never returns
to complete the operation, it will block any other operations from one of the
concurrent state slots.
-- Guest Interrupts --
It is not nice to keep the guest from running for too long. The TDH.QUOTE.GET
SEAMCALL monitors for host interrupts, but not guest ones. Similarly, the SGX
quoting enclave doesn't know what is happening in the TD. So, to help reduce
guest latencies, the GHCI exposes a way to register for a notification
(SetupEventNotifyInterrupt) when the quoting is finished. A quote TDVMCALL
handler can start the quote operation on another host thread, and resume
guest execution waiting for the quote to finish.
== Possible TDG call for quoting ==
The question recently came up from Sean (actually not that recently at this
point): While this is all getting redone, why not just go directly to the TDX
module from the guest to get the quote? Skip all the host buffer passing. This
was considered difficult to do when previously considered. We investigated it
further. Here is what it would look like.
-- Saving state while being interrupted and resuming --
Whereas the host would be involved in ensuring fairness for the TDH call, for
a TDG call, the TDX module would have to do this by itself. It turns out
there is an existing (but unused) TDG call that has a similar problem to
solve: TDG.MR.ASSIGNSVNS. It has a fixed rate at which the TD is allowed to
call the TDG call, 200usec. If the TD calls it more frequently, the TDX
Module exits to the host with a TDX_TDCALL_RATE_LIMIT.
A resumable TDG quote call would have to solve a slightly different problem.
A bad guest could refuse to resume the TDG call and hold the TDX module
extension save state slot hostage.
So instead, a TDG call could work by remotely canceling another TD’s paused
operation if that TD didn't call back frequently enough. After a time-limit
was reached on an extension save state slot usage, another TD’s TDG call
could abort the stalled TD’s operation. There existing designs around doing
this cancellation already in the TDX extension arch. It would need to be
adapted to the guest side calls and incorporate the extra cancellation work
into the time limit math. Some good timeouts would need to be picked.
-- Locking --
The TDX module would need to have some locks to manage the above fairness
solution. It would have to do this in a way to ensure no TD ever got starved.
This kind of global TDX locking is new for TDG calls.
There would have to be some heuristic to handle the problem of a TD starving the
other callers by repeatedly requesting quotes. Sean had previously tossed out
the idea of having a TD get a limited number of quote operations. Maybe make it
per some time period. A more complex fairness scheme would have the TDX module
maintain a queue somehow. Another option still, could be handle it like
TDX_TDCALL_RATE_LIMIT. Then the host could decide to throttle misbehaving TDs.
A TDG quote call would otherwise expect to take similar locks from the guest
side as are already taken by TDG.MR.REPORT. Long term, more details are
expected in the quote, so a TDG direction would be signing up to potentially
navigate more guest-host locking issues. Like we have in a few places.
And like the other contention issue, some good limits would need to be picked.
-- Host Interrupts --
When a host interrupt happens, the TDX module needs to return to the host VMM
to handle it. While executing a long running TDG call, the TDX module would
need to do the same thing. But it also would have to remember what it was
doing when the host attempts to re-enter the TD, and instead resume the guest
side quote operation. If the VMM never re-entered the TD, the save state slot
would have to be subject to reclaim similarly to if the guest never resumed
it.
-- Guest Interrupts --
If a quote operation is long enough, the TDG SEAMCALL should probably support
a resumable-like flow from the guest side too, where it also monitors for
pending guest interrupts and returns from the TDG call to let the guest
handle them. Then the host wouldn't need a thread and a completion guest
notification mechanism. It basically transfers that complexity to the TDX
module.
But it also might transfer some of the control. The TDX module would have to
embed some policy on how to decide when to inject guest interrupts. If the
host and TDX module had different logic on when to actually re-enter the TD,
that could be annoying.
== Summary ==
From what we found, I think it doesn't seem overly impossible for the module. But
the exact cancellation rules and timeout logic would need to be hashed out a bit
more. The simple mutex on the host side does still seem a fair amount simpler than
the TDX module's fairness options.
How do we feel about the roles and responsibility shifts? Any other problems to
solve? Or errors in these proposals? What kind of fairness heuristics are
acceptable for TDX users?
^ permalink raw reply [flat|nested] 33+ messages in thread
* Re: TDG quote analysis
2026-09-16 0:48 TDG quote analysis Edgecombe, Rick P
@ 2026-09-16 18:00 ` Sean Christopherson
2026-09-16 18:49 ` Edgecombe, Rick P
0 siblings, 1 reply; 33+ messages in thread
From: Sean Christopherson @ 2026-09-16 18:00 UTC (permalink / raw)
To: Rick P Edgecombe
Cc: pbonzini@redhat.com, Dave Hansen, kas@kernel.org, Elena Reshetova,
Yilun Xu, kvm@vger.kernel.org, Vishal Annapurve, Peter Fang,
Binbin Wu
On Wed, Sep 16, 2026, Rick P Edgecombe wrote:
> Here is the overdue analysis of what a TDG based TDX quote operation could look
> like, pulled together as writeup somewhat quickly. This was from Sean's request
> originally. The writeup is based on a bunch of gathering by Peter, Elena and I.
> Some TDX interrupt input from Binbin.
Thank you!
> This DICE quote generation can take a long time. Exactly how long is not known,
> but it is expected to vary and often exceed the time a SEAMCALL can be executing
> in the TDX module by a considerable amount.
What's the ballpark? Are we talking tens of nanoseconds, tens of microseconds,
tens of milliseconds?
Is the CPU spinning the entire time? Or is the quote degeneration I/O-like?
> It is not nice to keep the guest from running for too long. The TDH.QUOTE.GET
> SEAMCALL monitors for host interrupts, but not guest ones. Similarly, the SGX
> quoting enclave doesn't know what is happening in the TD. So, to help reduce
> guest latencies, the GHCI exposes a way to register for a notification
> (SetupEventNotifyInterrupt) when the quoting is finished. A quote TDVMCALL
> handler can start the quote operation on another host thread, and resume
> guest execution waiting for the quote to finish.
To me, this just *screams* for a virtio-like device to provide quotes to the guest,
especially if the slow part of quote generation is I/O-like. Even if it hogs a CPU,
an asynchronous virtio-like interface would be far easier to support. E.g. the
userspace device backend spawns a thread (affined to the same set of pCPUs as the
vCPU, i.e. to "tax" the guest instead of requiring dedicated "overhead" CPUs),
and kicks the guest when the quote is ready.
Ahh, that's more or less what SetupEventNotifyInterrupt is. I guess that's not
the end of the world, so long as the backend for SetupEventNotifyInterrupt is
handled entirely in host userspace.
> A resumable TDG quote call would have to solve a slightly different problem.
> A bad guest could refuse to resume the TDG call and hold the TDX module
> extension save state slot hostage.
The whole "what about guest interrupts?" thing makes me think the quote generation
is CPU-bound, i.e. not I/O-like.
Does the TDX Module *need* to provide a save slot? What would prevent TDX from
requiring the guest to provide storage for whatever in-flight data is needed?
That way there it doesn't matter if the guest never resumes the call, it can only
hurt itself.
> So instead, a TDG call could work by remotely canceling another TD’s paused
> operation if that TD didn't call back frequently enough. After a time-limit
> was reached on an extension save state slot usage, another TD’s TDG call
> could abort the stalled TD’s operation. There existing designs around doing
> this cancellation already in the TDX extension arch. It would need to be
> adapted to the guest side calls and incorporate the extra cancellation work
> into the time limit math. Some good timeouts would need to be picked.
...
> If a quote operation is long enough, the TDG SEAMCALL should probably support
> a resumable-like flow from the guest side too, where it also monitors for
> pending guest interrupts and returns from the TDG call to let the guest
> handle them.
>
> Then the host wouldn't need a thread and a completion guest notification
> mechanism. It basically transfers that complexity to the TDX module.
> But it also might transfer some of the control. The TDX module would have to
> embed some policy on how to decide when to inject guest interrupts. If the
> host and TDX module had different logic on when to actually re-enter the TD,
> that could be annoying.
But the TDX-Module already has some policy, no? In the sense that it decides
when to exit to the host because there's an interrupt pending.
> == Summary ==
>
> From what we found, I think it doesn't seem overly impossible for the module. But
> the exact cancellation rules and timeout logic would need to be hashed out a bit
> more. The simple mutex on the host side does still seem a fair amount simpler than
> the TDX module's fairness options.
>
> How do we feel about the roles and responsibility shifts? Any other problems to
> solve?
What about the whole "TD-specific data in the quote" thing? If getting a quote
(or report?) is punted to the host, I'd still like to keep it out of KVM.
^ permalink raw reply [flat|nested] 33+ messages in thread
* Re: TDG quote analysis
2026-09-16 18:00 ` Sean Christopherson
@ 2026-09-16 18:49 ` Edgecombe, Rick P
2026-09-16 19:27 ` Sean Christopherson
0 siblings, 1 reply; 33+ messages in thread
From: Edgecombe, Rick P @ 2026-09-16 18:49 UTC (permalink / raw)
To: seanjc@google.com
Cc: Xu, Yilun, Reshetova, Elena, Wu, Binbin, Hansen, Dave,
kas@kernel.org, Annapurve, Vishal, pbonzini@redhat.com,
Fang, Peter, kvm@vger.kernel.org
On Wed, 2026-09-16 at 11:00 -0700, Sean Christopherson wrote:
> On Wed, Sep 16, 2026, Rick P Edgecombe wrote:
> > This DICE quote generation can take a long time. Exactly how long is not
> > known, but it is expected to vary and often exceed the time a SEAMCALL can
> > be executing in the TDX module by a considerable amount.
>
> What's the ballpark? Are we talking tens of nanoseconds, tens of
> microseconds, tens of milliseconds?
I asked about this too. I have not seen any measurement yet for the real
implementation. AFAIU the PQC stuff is not settled enough at the industry level
to be sure in the long run. I have been under the impression of at least a ms or
2 for PQC signature based on general PQC signature benchmarks I've seen. That is
only my guess though.
If you are thinking the implementation could punt on how to handle the guest
interrupts for now, that seems reasonable. I'd think exiting to handle the host
interrupts could not be skipped though.
>
> Is the CPU spinning the entire time? Or is the quote degeneration I/O-like?
Today the CPU will mostly be working. In the future it could be waiting for the
shared HW. This would take much longer.
> >
>
> > A resumable TDG quote call would have to solve a slightly different
> > problem. A bad guest could refuse to resume the TDG call and hold the TDX
> > module extension save state slot hostage.
>
> The whole "what about guest interrupts?" thing makes me think the quote
> generation is CPU-bound, i.e. not I/O-like.
>
> Does the TDX Module *need* to provide a save slot? What would prevent TDX
> from requiring the guest to provide storage for whatever in-flight data is
> needed? That way there it doesn't matter if the guest never resumes the call,
> it can only hurt itself.
I was thinking about this too, but thought it sounded a bit like the NAKed SNP
secure AVIC pattern. At least the TDX module would need to handle getting asked
to zap the guest page that was given to be a saved state slot.
The host would probably need to be involved in unmapping the page from the S-EPT
while in use too, since it would need a TLB flush on each vCPU. Otherwise the
guest could see the intermediate memory of the operation. Or, hmm, I guess the
guest could coordinate entering the TDX module on each vCPU to get a flush. That
would be a new one.
>
> > If a quote operation is long enough, the TDG SEAMCALL should probably
> > support
> > a resumable-like flow from the guest side too, where it also monitors for
> > pending guest interrupts and returns from the TDG call to let the guest
> > handle them.
> >
> > Then the host wouldn't need a thread and a completion guest notification
> > mechanism. It basically transfers that complexity to the TDX module.
>
>
> > But it also might transfer some of the control. The TDX module would have
> > to embed some policy on how to decide when to inject guest interrupts. If
> > the host and TDX module had different logic on when to actually re-enter the
> > TD, that could be annoying.
>
> But the TDX-Module already has some policy, no? In the sense that it decides
> when to exit to the host because there's an interrupt pending.
Yes for exiting to the host. I was talking about re-entering the guest for a
guest pending event of some sort. Like if KVM decides to re-enter to handle some
event, but TDX module decides, nah we are going to work on the in-progress guest
call for a bit first.
>
> > == Summary ==
> >
> > From what we found, I think it doesn't seem overly impossible for the
> > module. But the exact cancellation rules and timeout logic would need to be
> > hashed out a bit more. The simple mutex on the host side does still seem a
> > fair amount simpler than the TDX module's fairness options.
> >
> > How do we feel about the roles and responsibility shifts? Any other problems
> > to solve?
>
> What about the whole "TD-specific data in the quote" thing? If getting a
> quote (or report?) is punted to the host, I'd still like to keep it out of
> KVM.
I checked another option for this: Write the nonce to the TDX module via some
new TDG call, then just ask for a TD scoped quote which fills in everything from
what the TDX module already knows. TD details and nonce. Apparently it could
work.
Then nothing gets passed out of the guest except for request for a quote. We'd
have to break compatibility with the SGX based quotes though. Is it worth it? Vs
just utilizing the upstream KVM functionality unchanged?
^ permalink raw reply [flat|nested] 33+ messages in thread
* Re: TDG quote analysis
2026-09-16 18:49 ` Edgecombe, Rick P
@ 2026-09-16 19:27 ` Sean Christopherson
2026-09-16 22:31 ` Edgecombe, Rick P
0 siblings, 1 reply; 33+ messages in thread
From: Sean Christopherson @ 2026-09-16 19:27 UTC (permalink / raw)
To: Rick P Edgecombe
Cc: Yilun Xu, Elena Reshetova, Binbin Wu, Dave Hansen, kas@kernel.org,
Vishal Annapurve, pbonzini@redhat.com, Peter Fang,
kvm@vger.kernel.org
On Wed, Sep 16, 2026, Rick P Edgecombe wrote:
> On Wed, 2026-09-16 at 11:00 -0700, Sean Christopherson wrote:
> > On Wed, Sep 16, 2026, Rick P Edgecombe wrote:
> > > This DICE quote generation can take a long time. Exactly how long is not
> > > known, but it is expected to vary and often exceed the time a SEAMCALL can
> > > be executing in the TDX module by a considerable amount.
> >
> > What's the ballpark? Are we talking tens of nanoseconds, tens of
> > microseconds, tens of milliseconds?
>
> I asked about this too. I have not seen any measurement yet for the real
> implementation. AFAIU the PQC stuff is not settled enough at the industry level
> to be sure in the long run. I have been under the impression of at least a ms or
> 2 for PQC signature based on general PQC signature benchmarks I've seen. That is
> only my guess though.
>
> If you are thinking the implementation could punt on how to handle the guest
> interrupts for now, that seems reasonable. I'd think exiting to handle the host
> interrupts could not be skipped though.
I'm just trying to understand the scope of the problem we're trying to solve.
E.g. I don't think any design will save us if generating a quote requires a CPU
to be spinning for multiple milliseconds.
> > Does the TDX Module *need* to provide a save slot? What would prevent TDX
> > from requiring the guest to provide storage for whatever in-flight data is
> > needed? That way there it doesn't matter if the guest never resumes the call,
> > it can only hurt itself.
>
> I was thinking about this too, but thought it sounded a bit like the NAKed SNP
> secure AVIC pattern. At least the TDX module would need to handle getting asked
> to zap the guest page that was given to be a saved state slot.
But doesn't the TDX Module already need to do this? It's writing guest memory,
no? So it needs to prevent the page from being freed while it's handling the
quote. Just put the onus on the guest to provide the same GPA when restarting
the TDG.
If the host yanks a page away from the guest, the guest is hosed no matter what.
If the guest frees an in-use page, e.g. converts it to SHARED, then that's 100%
a guest bug.
> The host would probably need to be involved in unmapping the page from the S-EPT
> while in use too, since it would need a TLB flush on each vCPU. Otherwise the
> guest could see the intermediate memory of the operation.
Do we care? I was and am assuming "no". Unless there is sensitive TDX-Module
data that needs to be saved, it's again on the guest not to read half-baked data.
> > What about the whole "TD-specific data in the quote" thing? If getting a
> > quote (or report?) is punted to the host, I'd still like to keep it out of
> > KVM.
>
> I checked another option for this: Write the nonce to the TDX module via some
> new TDG call, then just ask for a TD scoped quote which fills in everything from
> what the TDX module already knows. TD details and nonce. Apparently it could
> work.
>
> Then nothing gets passed out of the guest except for request for a quote. We'd
> have to break compatibility with the SGX based quotes though. Is it worth it? Vs
> just utilizing the upstream KVM functionality unchanged?
^ permalink raw reply [flat|nested] 33+ messages in thread
* Re: TDG quote analysis
2026-09-16 19:27 ` Sean Christopherson
@ 2026-09-16 22:31 ` Edgecombe, Rick P
2026-09-16 23:00 ` Peter Fang
0 siblings, 1 reply; 33+ messages in thread
From: Edgecombe, Rick P @ 2026-09-16 22:31 UTC (permalink / raw)
To: seanjc@google.com
Cc: Xu, Yilun, Reshetova, Elena, Wu, Binbin, Hansen, Dave,
kas@kernel.org, Annapurve, Vishal, pbonzini@redhat.com,
Fang, Peter, kvm@vger.kernel.org
On Wed, 2026-09-16 at 12:27 -0700, Sean Christopherson wrote:
> > If you are thinking the implementation could punt on how to handle the guest
> > interrupts for now, that seems reasonable. I'd think exiting to handle the
> > host interrupts could not be skipped though.
>
> I'm just trying to understand the scope of the problem we're trying to solve.
> E.g. I don't think any design will save us if generating a quote requires a
> CPU to be spinning for multiple milliseconds.
Elena, any updated estimates on CPU based quote generation and HW based?
>
> > > Does the TDX Module *need* to provide a save slot? What would prevent TDX
> > > from requiring the guest to provide storage for whatever in-flight data is
> > > needed? That way there it doesn't matter if the guest never resumes the
> > > call, it can only hurt itself.
> >
> > I was thinking about this too, but thought it sounded a bit like the NAKed
> > SNP secure AVIC pattern. At least the TDX module would need to handle
> > getting asked to zap the guest page that was given to be a saved state slot.
>
> But doesn't the TDX Module already need to do this? It's writing guest
> memory, no? So it needs to prevent the page from being freed while it's
> handling the quote. Just put the onus on the guest to provide the same GPA
> when restarting the TDG.
When you make a non-interruptible/non-restartable call, the TDX module doesn't
rush out if there is an interrupt pending. It lets the call fully complete, if
it is short. During that time, if it is operating on private memory it would
take internal locks to prevent the page from getting zapped out or swapped.
>
> If the host yanks a page away from the guest, the guest is hosed no matter
> what. If the guest frees an in-use page, e.g. converts it to SHARED, then
> that's 100% a guest bug.
Imagine the scenario of a vCPU executing the TDG quote operation. From another
CPU in the host, KVM zaps the page the TDG call is using for it's save state.
TDX module would need to know it can't yank the page that was actively being
used by the TDG quote operation. So I guess it would have to leave the GPA's S-
EPT entry locked. The VMM could use the kick scheme to break through.
But I would think it would need to keep the save state page from getting swapped
out during the entire quote operation actually. (i.e. between the resumes).
Because it is probably not safe to reset data on some execution context
operating with secret keys, etc. In that case, then some other more destructive
operation would be needed to let the guest still zap the S-EPT. Because the kick
scheme would not unlock the S-EPT entry in that case.
There might be some other solutions to avoid swapping out the per-operation
state.... Not sure.
>
> > The host would probably need to be involved in unmapping the page from the
> > S-EPT while in use too, since it would need a TLB flush on each vCPU.
> > Otherwise the guest could see the intermediate memory of the operation.
>
> Do we care? I was and am assuming "no". Unless there is sensitive TDX-Module
> data that needs to be saved, it's again on the guest not to read half-baked
> data.
It is basically asking whether the extension operation would use save state area
mapped with the guests keyid or the TDX modules keyid.
This quoting operation is working with secret keys, so I don't think we can let
the guest modify or see the memory. Think trying to do a secret operation while
the guest could interrupt it and look at the registers.
^ permalink raw reply [flat|nested] 33+ messages in thread
* Re: TDG quote analysis
2026-09-16 22:31 ` Edgecombe, Rick P
@ 2026-09-16 23:00 ` Peter Fang
2026-09-16 23:57 ` Sean Christopherson
0 siblings, 1 reply; 33+ messages in thread
From: Peter Fang @ 2026-09-16 23:00 UTC (permalink / raw)
To: Edgecombe, Rick P
Cc: seanjc@google.com, Xu, Yilun, Reshetova, Elena, Wu, Binbin,
Hansen, Dave, kas@kernel.org, Annapurve, Vishal,
pbonzini@redhat.com, kvm@vger.kernel.org
On Wed, Sep 16, 2026 at 03:31:50PM -0700, Edgecombe, Rick P wrote:
> On Wed, 2026-09-16 at 12:27 -0700, Sean Christopherson wrote:
> > > If you are thinking the implementation could punt on how to handle the guest
> > > interrupts for now, that seems reasonable. I'd think exiting to handle the
> > > host interrupts could not be skipped though.
> >
> > I'm just trying to understand the scope of the problem we're trying to solve.
> > E.g. I don't think any design will save us if generating a quote requires a
> > CPU to be spinning for multiple milliseconds.
>
> Elena, any updated estimates on CPU based quote generation and HW based?
The current estimate is 10s of milliseconds for HW based, and
milliseconds or less for CPU based.
>
> >
> > > > Does the TDX Module *need* to provide a save slot? What would prevent TDX
> > > > from requiring the guest to provide storage for whatever in-flight data is
> > > > needed? That way there it doesn't matter if the guest never resumes the
> > > > call, it can only hurt itself.
> > >
> > > I was thinking about this too, but thought it sounded a bit like the NAKed
> > > SNP secure AVIC pattern. At least the TDX module would need to handle
> > > getting asked to zap the guest page that was given to be a saved state slot.
> >
> > But doesn't the TDX Module already need to do this? It's writing guest
> > memory, no? So it needs to prevent the page from being freed while it's
> > handling the quote. Just put the onus on the guest to provide the same GPA
> > when restarting the TDG.
>
> When you make a non-interruptible/non-restartable call, the TDX module doesn't
> rush out if there is an interrupt pending. It lets the call fully complete, if
> it is short. During that time, if it is operating on private memory it would
> take internal locks to prevent the page from getting zapped out or swapped.
Given that quoting could take 10s of milliseconds, I think this TDG
might need to be interruptible/restartable, at least during some types
of quoting. So when this TDCALL returns to the guest to handle an
interrupt, the host could take a lock, and then the next TDCALL fails
with a BUSY?
>
^ permalink raw reply [flat|nested] 33+ messages in thread
* Re: TDG quote analysis
2026-09-16 23:00 ` Peter Fang
@ 2026-09-16 23:57 ` Sean Christopherson
2026-09-17 0:56 ` Edgecombe, Rick P
0 siblings, 1 reply; 33+ messages in thread
From: Sean Christopherson @ 2026-09-16 23:57 UTC (permalink / raw)
To: Peter Fang
Cc: Rick P Edgecombe, Yilun Xu, Elena Reshetova, Binbin Wu,
Dave Hansen, kas@kernel.org, Vishal Annapurve,
pbonzini@redhat.com, kvm@vger.kernel.org
On Wed, Sep 16, 2026, Peter Fang wrote:
> On Wed, Sep 16, 2026 at 03:31:50PM -0700, Edgecombe, Rick P wrote:
> > On Wed, 2026-09-16 at 12:27 -0700, Sean Christopherson wrote:
> > > > If you are thinking the implementation could punt on how to handle the guest
> > > > interrupts for now, that seems reasonable. I'd think exiting to handle the
> > > > host interrupts could not be skipped though.
> > >
> > > I'm just trying to understand the scope of the problem we're trying to solve.
> > > E.g. I don't think any design will save us if generating a quote requires a
> > > CPU to be spinning for multiple milliseconds.
> >
> > Elena, any updated estimates on CPU based quote generation and HW based?
>
> The current estimate is 10s of milliseconds for HW based, and
> milliseconds or less for CPU based.
I am beyond confused. What is "HW based" versus "CPU based"?
At this point, I don't care about the gory details, I just want to understand
the basics:
- Does generating a quote require the CPU to actively execute instructions?
- If not, how does software know when a quote is ready?
I tried peeking at TDX Module source code, but I can't find anything relevant
even in the latest TDX 2.0 drop.
^ permalink raw reply [flat|nested] 33+ messages in thread
* Re: TDG quote analysis
2026-09-16 23:57 ` Sean Christopherson
@ 2026-09-17 0:56 ` Edgecombe, Rick P
2026-09-17 13:37 ` Sean Christopherson
0 siblings, 1 reply; 33+ messages in thread
From: Edgecombe, Rick P @ 2026-09-17 0:56 UTC (permalink / raw)
To: seanjc@google.com, Fang, Peter
Cc: Xu, Yilun, Reshetova, Elena, Wu, Binbin, Hansen, Dave,
kas@kernel.org, Annapurve, Vishal, pbonzini@redhat.com,
kvm@vger.kernel.org
On Wed, 2026-09-16 at 16:57 -0700, Sean Christopherson wrote:
> > The current estimate is 10s of milliseconds for HW based, and
> > milliseconds or less for CPU based.
>
> I am beyond confused.
Sorry if this is another gap. Trying to find the right level to cover the
important points. Did we not talk about S3M?
> What is "HW based" versus "CPU based"?
"HW" is talking about the S3M thing. The quote operation could go directly to
the S3M to get the quote. (HW based) Or it could get an intermediate key and
generate quotes using CPU instructions. (SW based). Think like a crypto library
in the TDX module.
>
> At this point, I don't care about the gory details, I just want to understand
> the basics:
>
> - Does generating a quote require the CPU to actively execute instructions?
> - If not, how does software know when a quote is ready?
Oh man. There are actually a ton of attestation plans. I don't know which will
become real. So these two behaviors we are enumerating are actually an editorial
decision.
The software based flow would be expected to first. It would involve the CPU
doing crypto stuff as above.
Then a HW based flow where the crypto happens on the limited HW resource. This
is where full parallelization is not possible, because the CPU is not doing the
heavy work. You might want this one instead for security reasons. But the main
point of discussing it is that you could expect some quotes to take a long time
and support a limited number of parallel quotes.
But neither of these solutions are actually nailed down yet. How should a off-
cpu based flow work? We can discuss it. I'd think to have some interface that
doesn't require guest changes all the time. Focusing on something that just
supports long quotes seems the most robust.
>
> I tried peeking at TDX Module source code, but I can't find anything relevant
> even in the latest TDX 2.0 drop.
I have not seen it, don't think Peter either. The posted DICE seamcall quote
code was actually tested on a mocked TDX module. So DICE is early enabling kind
of stages.
It is not the usual years old TDX feature. This is new stuff. Still open and
unfinished, etc. In some ways it's a more complex problem then "cram this TDX
interface into KVM as best you can".
^ permalink raw reply [flat|nested] 33+ messages in thread
* Re: TDG quote analysis
2026-09-17 0:56 ` Edgecombe, Rick P
@ 2026-09-17 13:37 ` Sean Christopherson
2026-09-17 18:12 ` Edgecombe, Rick P
0 siblings, 1 reply; 33+ messages in thread
From: Sean Christopherson @ 2026-09-17 13:37 UTC (permalink / raw)
To: Rick P Edgecombe
Cc: Peter Fang, Yilun Xu, Elena Reshetova, Binbin Wu, Dave Hansen,
kas@kernel.org, Vishal Annapurve, pbonzini@redhat.com,
kvm@vger.kernel.org
On Thu, Sep 17, 2026, Rick P Edgecombe wrote:
> On Wed, 2026-09-16 at 16:57 -0700, Sean Christopherson wrote:
> > > The current estimate is 10s of milliseconds for HW based, and
> > > milliseconds or less for CPU based.
> >
> > I am beyond confused.
>
> Sorry if this is another gap. Trying to find the right level to cover the
> important points. Did we not talk about S3M?
We talked about S3M, it's the "CPU based" thing that confused me. I wasn't aware
that manually doing "everything" in software, on the CPU, was an option.
> > What is "HW based" versus "CPU based"?
>
> "HW" is talking about the S3M thing. The quote operation could go directly to
> the S3M to get the quote. (HW based) Or it could get an intermediate key and
Get an intermediate key from where? I assume this is a privileged key that can't
be exposed outside of the TDX-Module? I.e. we can't hand that key to the guest
to let it complete the quoting process, in an inherently interruptible environment?
> generate quotes using CPU instructions. (SW based). Think like a crypto library
> in the TDX module.
That's not a library...
> > At this point, I don't care about the gory details, I just want to understand
> > the basics:
> >
> > - Does generating a quote require the CPU to actively execute instructions?
> > - If not, how does software know when a quote is ready?
I would still like an answer to this question. Even if the end-to-end solution
isn't nailed down, I hope that the physical capabilities of the S3M are defined
enough to cover this.
> Oh man. There are actually a ton of attestation plans. I don't know which will
> become real. So these two behaviors we are enumerating are actually an editorial
> decision.
>
> The software based flow would be expected to first. It would involve the CPU
> doing crypto stuff as above.
>
> Then a HW based flow where the crypto happens on the limited HW resource. This
> is where full parallelization is not possible, because the CPU is not doing the
> heavy work. You might want this one instead for security reasons. But the main
> point of discussing it is that you could expect some quotes to take a long time
> and support a limited number of parallel quotes.
It's not just the raw time that matters, where/how that time is spent also matters
greatly. There is a *massive* difference between "fire off an operation and get a
notification" and "churn on crypto stuff for the entire time", especially when the
thing churning on crypto stuff isn't interruptible by default.
> But neither of these solutions are actually nailed down yet. How should a off-
> cpu based flow work? We can discuss it. I'd think to have some interface that
> doesn't require guest changes all the time. Focusing on something that just
> supports long quotes seems the most robust.
No, because they are wildly different beasts. This is basically like comparing
zswap and traditional swap; yes, they're both swap, but they have *very* different
characteristics that need to be accounted for at the system level. Now make the
zswap (de)compression code completely uninterruptible. The whole problem space
changes, because either the host has to be ok with a CPU "disappearing" for an
extended duration, or the interface needs to be reworked to make the swap sequence
restartable.
It sounds to me like y'all need to take a step back and nail down your customer
requirements, including what is tolerable latency from the guest perspective.
FWIW, if S3M quoting isn't I/O-like, i.e. isn't fire and get notified, then IMO
it's completely broken and likely unusable.
^ permalink raw reply [flat|nested] 33+ messages in thread
* Re: TDG quote analysis
2026-09-17 13:37 ` Sean Christopherson
@ 2026-09-17 18:12 ` Edgecombe, Rick P
2026-09-17 19:57 ` Sean Christopherson
0 siblings, 1 reply; 33+ messages in thread
From: Edgecombe, Rick P @ 2026-09-17 18:12 UTC (permalink / raw)
To: seanjc@google.com
Cc: Xu, Yilun, Reshetova, Elena, Wu, Binbin, Hansen, Dave,
Annapurve, Vishal, kas@kernel.org, pbonzini@redhat.com,
Fang, Peter, kvm@vger.kernel.org
On Thu, 2026-09-17 at 06:37 -0700, Sean Christopherson wrote:
> On Thu, Sep 17, 2026, Rick P Edgecombe wrote:
>
> >
> > "HW" is talking about the S3M thing. The quote operation could go directly to
> > the S3M to get the quote. (HW based) Or it could get an intermediate key and
>
> Get an intermediate key from where? I assume this is a privileged key that can't
> be exposed outside of the TDX-Module?
>
Yea that is my understanding.
> I.e. we can't hand that key to the guest
> to let it complete the quoting process, in an inherently interruptible environment?
>
> > generate quotes using CPU instructions. (SW based). Think like a crypto library
> > in the TDX module.
>
> That's not a library...
Not sure what you mean. The TDX module build depends on a crypto library:
https://github.com/intel/confidential-computing.tdx.tdx-module/blob/tdx_1.5/BUILD.md
I don't know what the DICE implementation will use. But I'd assume it will not
implement its own crypto primitives.
>
> > > At this point, I don't care about the gory details, I just want to understand
> > > the basics:
> > >
> > > - Does generating a quote require the CPU to actively execute instructions?
> > > - If not, how does software know when a quote is ready?
>
> I would still like an answer to this question. Even if the end-to-end solution
> isn't nailed down, I hope that the physical capabilities of the S3M are defined
> enough to cover this.
The S3M itself has public docs, but how TDX module would use it in the "HW"
flows is not settled.
>
> > Oh man. There are actually a ton of attestation plans. I don't know which will
> > become real. So these two behaviors we are enumerating are actually an editorial
> > decision.
> >
> > The software based flow would be expected to first. It would involve the CPU
> > doing crypto stuff as above.
> >
> > Then a HW based flow where the crypto happens on the limited HW resource. This
> > is where full parallelization is not possible, because the CPU is not doing the
> > heavy work. You might want this one instead for security reasons. But the main
> > point of discussing it is that you could expect some quotes to take a long time
> > and support a limited number of parallel quotes.
>
> It's not just the raw time that matters, where/how that time is spent also matters
> greatly. There is a *massive* difference between "fire off an operation and get a
> notification" and "churn on crypto stuff for the entire time", especially when the
> thing churning on crypto stuff isn't interruptible by default.
In general, S3M is described as a mailbox. So my understanding is that the CPU
is not churning while S3M works. More below.
>
> > But neither of these solutions are actually nailed down yet. How should a off-
> > cpu based flow work? We can discuss it. I'd think to have some interface that
> > doesn't require guest changes all the time. Focusing on something that just
> > supports long quotes seems the most robust.
>
> No, because they are wildly different beasts. This is basically like comparing
> zswap and traditional swap; yes, they're both swap, but they have *very* different
> characteristics that need to be accounted for at the system level. Now make the
> zswap (de)compression code completely uninterruptible. The whole problem space
> changes, because either the host has to be ok with a CPU "disappearing" for an
> extended duration, or the interface needs to be reworked to make the swap sequence
> restartable.
>
> It sounds to me like y'all need to take a step back and nail down your customer
> requirements, including what is tolerable latency from the guest perspective.
>
> FWIW, if S3M quoting isn't I/O-like, i.e. isn't fire and get notified, then IMO
> it's completely broken and likely unusable.
Let's take a step back here. There is an existing attestation flow that was
designed around some limitations that are changing (specifically whether the
quoter has knowledge of the TD). The current host side quote discussion (KVM
ioctl) is basically a straight forward evolution of the existing design, even
though the limitations are getting removed.
You asked whether we could do a re-design that makes more sense in the context
of the lack of that old limitation. At that point we are faced with the age-old
question: how far into the fuzzy future should we design around?
The nearest term thing is a SW based flow where work happens on the CPU. But the
exact amount of time is not know yet. Peter gave a ballpark.
A future thing is a flow where S3M is engaged for every quote. That would be
expected to take longer, and have greater limits on concurrency. But the details
are not sorted on what exactly the user will want, or how it would be
implemented.
I think you are maybe wondering whether S3M could be used such that the guest
could wait on a quote while not needing any saved state area? If you want to
know more about S3M we can round up some more info on it. There are some public
docs on it, but the intel link seems to be dead.
AFAICT, outside of TDX and just in the world in general, crypto is in a
transition phase. I read this interesting article on lwn awhile back about
"hybrid crypto" to add safety during the transition:
https://lwn.net/Articles/1048978/
Not trying to toss FUD here, but I'm imagining all the combinations of CPU or
S3M work that could possibly come up unpredictably.
If we want a redesign for TDX attestation, I think we shouldn't fall into the
trap of designing everything up front before taking the next near term step. We
should either stick with the current approach (guest report and host quote). Or
do the thing where you try to build a flexible ABI when you know an area could
have a lot of growth, while not stalling to figure out the full future.
So how far and wide should we hash out here? My vote would be to not work
through the details of a full S3M-every-quote design for the TDX module. Let's
just consider that in the future quotes could take a longer time and have
greater concurrency limits.
^ permalink raw reply [flat|nested] 33+ messages in thread
* Re: TDG quote analysis
2026-09-17 18:12 ` Edgecombe, Rick P
@ 2026-09-17 19:57 ` Sean Christopherson
2026-09-17 21:32 ` Edgecombe, Rick P
0 siblings, 1 reply; 33+ messages in thread
From: Sean Christopherson @ 2026-09-17 19:57 UTC (permalink / raw)
To: Rick P Edgecombe
Cc: Yilun Xu, Elena Reshetova, Binbin Wu, Dave Hansen,
Vishal Annapurve, kas@kernel.org, pbonzini@redhat.com, Peter Fang,
kvm@vger.kernel.org
On Thu, Sep 17, 2026, Rick P Edgecombe wrote:
> On Thu, 2026-09-17 at 06:37 -0700, Sean Christopherson wrote:
> > > The software based flow would be expected to first. It would involve the CPU
> > > doing crypto stuff as above.
> > >
> > > Then a HW based flow where the crypto happens on the limited HW resource. This
> > > is where full parallelization is not possible, because the CPU is not doing the
> > > heavy work. You might want this one instead for security reasons. But the main
> > > point of discussing it is that you could expect some quotes to take a long time
> > > and support a limited number of parallel quotes.
> >
> > It's not just the raw time that matters, where/how that time is spent also matters
> > greatly. There is a *massive* difference between "fire off an operation and get a
> > notification" and "churn on crypto stuff for the entire time", especially when the
> > thing churning on crypto stuff isn't interruptible by default.
>
> In general, S3M is described as a mailbox. So my understanding is that the CPU
> is not churning while S3M works. More below.
>
> >
> > > But neither of these solutions are actually nailed down yet. How should a off-
> > > cpu based flow work? We can discuss it. I'd think to have some interface that
> > > doesn't require guest changes all the time. Focusing on something that just
> > > supports long quotes seems the most robust.
> >
> > No, because they are wildly different beasts. This is basically like comparing
> > zswap and traditional swap; yes, they're both swap, but they have *very* different
> > characteristics that need to be accounted for at the system level. Now make the
> > zswap (de)compression code completely uninterruptible. The whole problem space
> > changes, because either the host has to be ok with a CPU "disappearing" for an
> > extended duration, or the interface needs to be reworked to make the swap sequence
> > restartable.
> >
> > It sounds to me like y'all need to take a step back and nail down your customer
> > requirements, including what is tolerable latency from the guest perspective.
> >
> > FWIW, if S3M quoting isn't I/O-like, i.e. isn't fire and get notified, then IMO
> > it's completely broken and likely unusable.
>
> Let's take a step back here. There is an existing attestation flow that was
> designed around some limitations that are changing (specifically whether the
> quoter has knowledge of the TD). The current host side quote discussion (KVM
> ioctl) is basically a straight forward evolution of the existing design, even
> though the limitations are getting removed.
>
> You asked whether we could do a re-design that makes more sense in the context
> of the lack of that old limitation. At that point we are faced with the age-old
> question: how far into the fuzzy future should we design around?
Now I'm trying to understand what the *current* plan is.
> The nearest term thing is a SW based flow where work happens on the CPU. But the
> exact amount of time is not know yet. Peter gave a ballpark.
Why are we even discussing this? I am so confused. I thought there were two
options: SGX and S3M. Now all of a sudden there's a third "let's do insane things
in software in the TDX module" option!?!?
> A future thing is a flow where S3M is engaged for every quote. That would be
> expected to take longer, and have greater limits on concurrency. But the details
> are not sorted on what exactly the user will want, or how it would be
> implemented.
>
> I think you are maybe wondering whether S3M could be used such that the guest
> could wait on a quote while not needing any saved state area?
I'm trying to figure out if *any* path is viable.
Burning 1ms of CPU time to generate a quote in uninterruptible code is a non-starter.
Hell, 100us is a non-starter.
Waiting 2s for a quote to come back from the S3M is a non-starter.
The numbers matter, and *none* of this is reviewable, even in RFC format, without
a crisp understanding of what latencies we are talking about.
> If you want to know more about S3M we can round up some more info on it.
> There are some public docs on it, but the intel link seems to be dead.
>
> AFAICT, outside of TDX and just in the world in general, crypto is in a
> transition phase. I read this interesting article on lwn awhile back about
> "hybrid crypto" to add safety during the transition:
> https://lwn.net/Articles/1048978/
> Not trying to toss FUD here, but I'm imagining all the combinations of CPU or
> S3M work that could possibly come up unpredictably.
>
> If we want a redesign for TDX attestation, I think we shouldn't fall into the
> trap of designing everything up front before taking the next near term step. We
> should either stick with the current approach (guest report and host quote).
Forget redesigning anything, I want to know if there's a path forward with *any*
design based on the limitations of hardware.
^ permalink raw reply [flat|nested] 33+ messages in thread
* Re: TDG quote analysis
2026-09-17 19:57 ` Sean Christopherson
@ 2026-09-17 21:32 ` Edgecombe, Rick P
2026-09-18 0:04 ` Sean Christopherson
0 siblings, 1 reply; 33+ messages in thread
From: Edgecombe, Rick P @ 2026-09-17 21:32 UTC (permalink / raw)
To: seanjc@google.com
Cc: Xu, Yilun, Reshetova, Elena, Wu, Binbin, Hansen, Dave,
Annapurve, Vishal, kas@kernel.org, pbonzini@redhat.com,
Fang, Peter, kvm@vger.kernel.org
On Thu, 2026-09-17 at 12:57 -0700, Sean Christopherson wrote:
> > Let's take a step back here. There is an existing attestation flow that was
> > designed around some limitations that are changing (specifically whether the
> > quoter has knowledge of the TD). The current host side quote discussion (KVM
> > ioctl) is basically a straight forward evolution of the existing design,
> > even though the limitations are getting removed.
> >
> > You asked whether we could do a re-design that makes more sense in the
> > context of the lack of that old limitation. At that point we are faced with
> > the age-old question: how far into the fuzzy future should we design around?
>
> Now I'm trying to understand what the *current* plan is.
The last Linux design for a host-based TDX DICE API, meaning the thing we were
prepping for the next DICE posting before getting into this TDG analysis:
- Leave the existing report in the guest as is.
- Utilize the existing quote GHCI call from the guest to pass the old report,
which goes through KVM to userspace.
- Add a new VM scoped ioctl to KVM to call into the TDH.QUOTE.GET. This gets
the quote and passes it back to userspace. Then userspace uses the existing GHCI
mechanisms to notify the guest that it is ready.
>
> > The nearest term thing is a SW based flow where work happens on the CPU. But
> > the exact amount of time is not know yet. Peter gave a ballpark.
>
> Why are we even discussing this? I am so confused. I thought there were two
> options: SGX and S3M. Now all of a sudden there's a third "let's do insane
> things in software in the TDX module" option!?!?
???
So when I said:
"HW" is talking about the S3M thing. The quote operation could go directly to
the S3M to get the quote. (HW based) Or it could get an intermediate key and
generate quotes using CPU instructions. (SW based). Think like a crypto library
in the TDX module.
...
The software based flow would be expected to first. It would involve the CPU
doing crypto stuff as above.
Then a HW based flow where the crypto happens on the limited HW resource. This
is where full parallelization is not possible, because the CPU is not doing the
heavy work. You might want this one instead for security reasons. But the main
point of discussing it is that you could expect some quotes to take a long time
and support a limited number of parallel quotes.
Did you interpret SW based flow to be talking about SGX? Like the intermediate key
goes from S3M to SGX? Hmm, I guess I could see it. But now maybe it makes sense why
there would be a crypto library in TDX? And why to have the whole extensions thing
for interrupting TDX module code? And why to care about the guest having visibility
into the save state area?
Here is a hopefully more concise description:
HW mode: Each quote request goes to S3M for signing
SW mode (faster): At init time, an intermediate key is generated using S3M and
is stored somehow in the TDX module. Each quote request uses that intermediate
key to do the signing via CPU work in the TDX module. The initial DICE support,
meaning the first thing that would be happening inside TDH.QUOTE.GET, will do
this using the whole "extensions" thing described in the beginning of this
thread. SGX is not involved.
Is it clear? As for if either is insane, I'm told there are users that want both
modes for various reasons. I have not spoken to them personally. I don't think
Peter has either? Do you want to hear from them?
>
> > A future thing is a flow where S3M is engaged for every quote. That would be
> > expected to take longer, and have greater limits on concurrency. But the
> > details are not sorted on what exactly the user will want, or how it would
> > be implemented.
> >
> > I think you are maybe wondering whether S3M could be used such that the
> > guest could wait on a quote while not needing any saved state area?
>
> I'm trying to figure out if *any* path is viable.
>
> Burning 1ms of CPU time to generate a quote in uninterruptible code is a non-
> starter. Hell, 100us is a non-starter.
What code is uninterruptible? The whole point of the extensions thing is to make
the operations broadly interruptible.
>
> Waiting 2s for a quote to come back from the S3M is a non-starter.
Not sure where 2s is coming from. But do you mean waiting in the guest? Or where
is waiting a non-starter?
>
> The numbers matter, and *none* of this is reviewable, even in RFC format,
> without a crisp understanding of what latencies we are talking about.
I agree not having a measurement of the initial DICE behavior is a detriment for
the TDG discussion. I didn't think we needed it though, if we leaned on the
original locking/scheduling reasons to prefer a host based quote flow.
Previously you asked for ballparks and got them. So what do you need exactly?
Exact measurements of an full TDX DICE quote implementation? How many future
possible modes of operation?
^ permalink raw reply [flat|nested] 33+ messages in thread
* Re: TDG quote analysis
2026-09-17 21:32 ` Edgecombe, Rick P
@ 2026-09-18 0:04 ` Sean Christopherson
2026-09-18 2:39 ` Edgecombe, Rick P
0 siblings, 1 reply; 33+ messages in thread
From: Sean Christopherson @ 2026-09-18 0:04 UTC (permalink / raw)
To: Rick P Edgecombe
Cc: Yilun Xu, Elena Reshetova, Binbin Wu, Dave Hansen,
Vishal Annapurve, kas@kernel.org, pbonzini@redhat.com, Peter Fang,
kvm@vger.kernel.org
On Thu, Sep 17, 2026, Rick P Edgecombe wrote:
> On Thu, 2026-09-17 at 12:57 -0700, Sean Christopherson wrote:
> > > Let's take a step back here. There is an existing attestation flow that was
> > > designed around some limitations that are changing (specifically whether the
> > > quoter has knowledge of the TD). The current host side quote discussion (KVM
> > > ioctl) is basically a straight forward evolution of the existing design,
> > > even though the limitations are getting removed.
> > >
> > > You asked whether we could do a re-design that makes more sense in the
> > > context of the lack of that old limitation. At that point we are faced with
> > > the age-old question: how far into the fuzzy future should we design around?
> >
> > Now I'm trying to understand what the *current* plan is.
>
> The last Linux design for a host-based TDX DICE API, meaning the thing we were
> prepping for the next DICE posting before getting into this TDG analysis:
> - Leave the existing report in the guest as is.
> - Utilize the existing quote GHCI call from the guest to pass the old report,
> which goes through KVM to userspace.
> - Add a new VM scoped ioctl to KVM to call into the TDH.QUOTE.GET. This gets
> the quote and passes it back to userspace. Then userspace uses the existing GHCI
> mechanisms to notify the guest that it is ready.
>
> >
> > > The nearest term thing is a SW based flow where work happens on the CPU. But
> > > the exact amount of time is not know yet. Peter gave a ballpark.
> >
> > Why are we even discussing this? I am so confused. I thought there were two
> > options: SGX and S3M. Now all of a sudden there's a third "let's do insane
> > things in software in the TDX module" option!?!?
>
> ???
>
> So when I said:
> "HW" is talking about the S3M thing. The quote operation could go directly to
> the S3M to get the quote. (HW based) Or it could get an intermediate key and
> generate quotes using CPU instructions. (SW based). Think like a crypto library
> in the TDX module.
>
> ...
>
> The software based flow would be expected to first. It would involve the CPU
> doing crypto stuff as above.
>
> Then a HW based flow where the crypto happens on the limited HW resource. This
> is where full parallelization is not possible, because the CPU is not doing the
> heavy work. You might want this one instead for security reasons. But the main
> point of discussing it is that you could expect some quotes to take a long time
> and support a limited number of parallel quotes.
>
> Did you interpret SW based flow to be talking about SGX? Like the intermediate key
> goes from S3M to SGX?
From *before* this conversation. Forget this converation, what does the mock
TDX Module used as the basis for the RFC[1] do? Because the RFC says absolutely
*nothing*. I kinda sorta have a picture now, but it required hunting down an
additional spec, and the documentnation still leaves me wanting. Because I still
don't know what it actually does. Based on everything you're saying, I *assume*
it's this "SW based" flow, but *nothing* actually says that.
Under "Interruptibility", the quoting doc linked says:
TDH.QUOTE.GET is interruptible. If a pending interrupt is detected during
operation, TDH.QUOTE.GET returns with a TDX_INTERRUPTED_RESUMABLE status in RAX.
Then punts me to:
For the general explanation on how interruption and resumption is handled for
all Quoting Service functions please consult [Intel TDX Module Base Spec] section
“Request Interruption and Resumption”.
And finishes with the wonderful:
Rest of details are TBD
Once I finally found the "Request Interruption and Resumption", I discovered that
the TDX Module now has "virtual thread pools", which IS NEVER MENTIONED IN THE
RFC. Seriously. No one thought that was worth mentioning?
OMG. I see it now:
+ /* Don't bother specifying the quote id */
+ .rdx = QUOTE_ID_MASK & (u64)-1,
So IIUC, there's a thread pool, but somewhere in the muck of RFC patches the
kernel initializes the pool with a size of 1? And then adds a global mutex to
serialize quotes across the entire system. That's certainly a choice.
This snippet from the RFC[2] doesn't help, because it makes it seem like there
is no state:
2. Host just refills the previous args (may have been modified by seamcall
output) on retry. such as TDH_EXT_INIT, TDH_EXT_MEM_ADD,
TDH_QUOTE_INIT, TDH_QUOTE_GET...
But that's presumably just because the kernel is using a single REQUEST_ID.
That spec continues with:
If supported for the specific request type, the host VMM can request an
interrupted request to be aborted.
Though per the above, apparently whether or not that's going to be supported is
TBD?
I am beyond frustrated at this point. I guess shame on me for not pouring over
dense docs? But I shouldn't have to. The *ENTIRE* point of an RFC like the one
that started all this is to get feedback on the design, but that's really hard to
achieve if you don't actually provide *any* details about the design.
And there are MASSIVE design choices in here. Like sizing the "thread" pool to 1,
and thus deliberately serializing quotes across all VMs. That's going to end well
when someone spins up 10s or 100s of TDs and they all try to attest at once.
Not listening for fatal signals while spinning on TDH.GET.QUOTE will save us though!
At least with PREEMPT_LAZY being forced, not doing cond_resched() is "fine".
What's worse, some of these patches that say nothing useful in the changelogs
have several Reviewed-by tags, which means multiple people looked at this as
thought it was all good.
Regardless of where we end up with quoting, whatever RFC is sent next needs to be
a million times better. My expectations are that someone with core TDX knowledge
and a basic understanding of attestation would be able to get a full and complete
understanding of the design and tradeoffs from the RFC alone. If I have to
rereference a spec for anything other than double check params and magic values,
the entire series is getting ignored.
As for guest-driven quoting, to me there is a fairly straightforward solution.
Assuming using guest-donated memory is too complex for the "virtual thread":
1. Drop "virtual thread pools" from quoting and support at most one in-flight
quote per TD. I doubt whatever memory is needed for a virtual thread is so
insanely ridiculous that we can't burn that much memory per TD. 32KiB is a
no-brainer. Above that and we'd have to reconsider, but if quoting requires
more than 32KiB of scratch space, I have to wonder what on earth it's doing.
Then, because the resource is fixed, have the host provide it at TD creation.
The TDX-Module keeps track of whether or not a quote is in-flight, and rejects
attempts to start a new one. The guest can simply guard quotes with a mutex.
The host never has to worry about canceling a quote request, because it just
needs to reclaim the TD as normal. If desired, the TDX Module can provide a
TDCALL to let the guest cancel a quote.
2. To avoid having to check for guest interrupts, for the "SW based" synchronous
method, simply resume the guest after processing a fixed amount of state. Yes,
that will increase the best case latency for a single quote, but *best* case
latency isn't a huge concern, and I doubt the overhead of a VMX roundtrip will
significantly impact that. It's the tail latencies that will be problematic,
and a forcing the serialization into the guest (the aforementioned mutex) means
the tail latencies will only be affected by host activity on *that* CPU, which
is more or less the status quo.
E.g. even assuming an absurd 20% overhead for the VMX round trips, having a
quote take ~1.2 *every* time is would be a far, far better experience than a
quote taking 1ms - 10ms to complete.
3. Host interrupts work as they do for normal TD operation. The state is all
there, attached to the vCPU. It's on the host to run the vCPU according to
its SLOs, so as not to starve/DoS the guest.
4. If/when S3M accesses during quoting come along (it's not clear to me if the
SW-based quoting needs to access the S3M during every quote, or just during
setup), then provide the host with a mechanism to throttle quote requests.
E.g. a simple counter that triggers an exit to the host when it hits zero
would probably suffice.
5. If/when "HW based" asynchronous quoting comes along, then if S3M can send
IRQs on completion, have a host-side driver wired up to call into the TDX
Module to tell it a quote is ready. Presumably the TDX Module can keep track
of where the quote came from. If the S3M can't do IRQs, then just put the
onus on the guest to poll to see when it's quote is ready? This needs more
details on the S3M, but it doesn't seem insurmountable.
But it's also not obvious to me why Linux/KVM would ever want to support this
mode. AFAICT, it adds more complexity (especially when accounting for noisy
neighbor issues) in order to provide a worse customer experience.
#2 above is going to be a problem with the host-based, CPU-intensive "SW" quoting,
which IIUC is what is currently being proposed. Unless the host provides an overhead
CPU to do quoting (which is likely a non-starter for Google at least), from the
guest's perspective, a vCPU will disappear and become non-responsive for however
long it takes the host to generate the quote.
The SetupEventNotifyInterrupt will help a little, but won't fully alleviate the
problem, because that still requires a pCPU to do the work. E.g. the vCPU wouldn't
be completely non-responsive, but it would observe extremely high steal-time until
the quote is generated, which is gross. Presumably this is how the current SGX-
based quoting behaves? But if these quotes are getting more massive, i.e. slower,
then it's probably something that needs to be addressed.
And that's assuming the virtual thread pool is sized such that there can be a thread
per VM. If there is any cross-VM resource sharing, then we're going to have noisy
neighbor problems and the above issue becomes an order of magnitude worse.
> > > A future thing is a flow where S3M is engaged for every quote. That would be
> > > expected to take longer, and have greater limits on concurrency. But the
> > > details are not sorted on what exactly the user will want, or how it would
> > > be implemented.
> > >
> > > I think you are maybe wondering whether S3M could be used such that the
> > > guest could wait on a quote while not needing any saved state area?
> >
> > I'm trying to figure out if *any* path is viable.
> >
> > Burning 1ms of CPU time to generate a quote in uninterruptible code is a non-
> > starter. Hell, 100us is a non-starter.
>
> What code is uninterruptible? The whole point of the extensions thing is to make
> the operations broadly interruptible.
Due to the lack of useful changelogs and being allergic to TDX specs, it wasn't
clear to me that TDH.GET.QUOTE was "fully" interruptible *and* restartable,
i.e. would guarantee forward progress.
> > Waiting 2s for a quote to come back from the S3M is a non-starter.
>
> Not sure where 2s is coming from. But do you mean waiting in the guest? Or where
> is waiting a non-starter?
If each S3M quote takes "tens of milliseconds", and there's one S3M per socket,
it doesn't take that many concurrent quote requests for one of the TDs to observe
a 2s+ latency to get its quote. E.g. quote takes 50ms, boot 40 TDs, and voila,
that last TD going through boot gets hit with a 2s+ quote latency.
> > The numbers matter, and *none* of this is reviewable, even in RFC format,
> > without a crisp understanding of what latencies we are talking about.
>
> I agree not having a measurement of the initial DICE behavior is a detriment for
> the TDG discussion. I didn't think we needed it though,
Ignore the TDG discussion, I'm saying they're needed for *any* discussion, i.e.
for reviewing the TDH implementation as well.
> if we leaned on the original locking/scheduling reasons to prefer a host
> based quote flow.
>
> Previously you asked for ballparks and got them. So what do you need exactly?
Guarantees around orders of magnitude. There is a enormous difference between
a CPU-intensive operation taking 10us versus 2ms. I.e. "milliseconds or less"
isn't a ballpark, it's a country, maaaaybe a state. For CPU-based in particular,
I was expecting ballparks with +/- 10us of precision, not "somewhere between 0
and several scheduler ticks".
[1] https://lore.kernel.org/all/20260522034128.3144354-1-yilun.xu@linux.intel.com
[2] https://lore.kernel.org/all/akaKmEZnTY4FO2gY@yilunxu-OptiPlex-7050
^ permalink raw reply [flat|nested] 33+ messages in thread
* Re: TDG quote analysis
2026-09-18 0:04 ` Sean Christopherson
@ 2026-09-18 2:39 ` Edgecombe, Rick P
2026-09-18 13:13 ` Sean Christopherson
0 siblings, 1 reply; 33+ messages in thread
From: Edgecombe, Rick P @ 2026-09-18 2:39 UTC (permalink / raw)
To: seanjc@google.com
Cc: Xu, Yilun, Reshetova, Elena, Wu, Binbin, Hansen, Dave,
Annapurve, Vishal, kas@kernel.org, pbonzini@redhat.com,
Fang, Peter, kvm@vger.kernel.org
Thanks for the detailed response.
On Thu, 2026-09-17 at 17:04 -0700, Sean Christopherson wrote:
> On Thu, Sep 17, 2026, Rick P Edgecombe wrote:
> > On Thu, 2026-09-17 at 12:57 -0700, Sean Christopherson wrote:
> > > > Let's take a step back here. There is an existing attestation flow that was
> > > > designed around some limitations that are changing (specifically whether the
> > > > quoter has knowledge of the TD). The current host side quote discussion (KVM
> > > > ioctl) is basically a straight forward evolution of the existing design,
> > > > even though the limitations are getting removed.
> > > >
> > > > You asked whether we could do a re-design that makes more sense in the
> > > > context of the lack of that old limitation. At that point we are faced with
> > > > the age-old question: how far into the fuzzy future should we design around?
> > >
> > > Now I'm trying to understand what the *current* plan is.
> >
> > The last Linux design for a host-based TDX DICE API, meaning the thing we were
> > prepping for the next DICE posting before getting into this TDG analysis:
> > - Leave the existing report in the guest as is.
> > - Utilize the existing quote GHCI call from the guest to pass the old report,
> > which goes through KVM to userspace.
> > - Add a new VM scoped ioctl to KVM to call into the TDH.QUOTE.GET. This gets
> > the quote and passes it back to userspace. Then userspace uses the existing GHCI
> > mechanisms to notify the guest that it is ready.
> >
> > >
> > > > The nearest term thing is a SW based flow where work happens on the CPU. But
> > > > the exact amount of time is not know yet. Peter gave a ballpark.
> > >
> > > Why are we even discussing this? I am so confused. I thought there were two
> > > options: SGX and S3M. Now all of a sudden there's a third "let's do insane
> > > things in software in the TDX module" option!?!?
> >
> > ???
> >
> > So when I said:
> > "HW" is talking about the S3M thing. The quote operation could go directly to
> > the S3M to get the quote. (HW based) Or it could get an intermediate key and
> > generate quotes using CPU instructions. (SW based). Think like a crypto library
> > in the TDX module.
> >
> > ...
> >
> > The software based flow would be expected to first. It would involve the CPU
> > doing crypto stuff as above.
> >
> > Then a HW based flow where the crypto happens on the limited HW resource. This
> > is where full parallelization is not possible, because the CPU is not doing the
> > heavy work. You might want this one instead for security reasons. But the main
> > point of discussing it is that you could expect some quotes to take a long time
> > and support a limited number of parallel quotes.
> >
> > Did you interpret SW based flow to be talking about SGX? Like the intermediate key
> > goes from S3M to SGX?
>
> From *before* this conversation. Forget this converation, what does the mock
> TDX Module used as the basis for the RFC[1] do?
>
Ah ok, well I can see how the RFC was confusing. But I am surprised the
discussion up the thread didn't cover a lot of these points.
> Because the RFC says absolutely
> *nothing*. I kinda sorta have a picture now, but it required hunting down an
> additional spec, and the documentnation still leaves me wanting. Because I still
> don't know what it actually does. Based on everything you're saying, I *assume*
> it's this "SW based" flow, but *nothing* actually says that.
>
> Under "Interruptibility", the quoting doc linked says:
>
> TDH.QUOTE.GET is interruptible. If a pending interrupt is detected during
> operation, TDH.QUOTE.GET returns with a TDX_INTERRUPTED_RESUMABLE status in RAX.
>
> Then punts me to:
>
> For the general explanation on how interruption and resumption is handled for
> all Quoting Service functions please consult [Intel TDX Module Base Spec] section
> “Request Interruption and Resumption”.
>
> And finishes with the wonderful:
>
> Rest of details are TBD
>
> Once I finally found the "Request Interruption and Resumption", I discovered that
> the TDX Module now has "virtual thread pools", which IS NEVER MENTIONED IN THE
> RFC. Seriously. No one thought that was worth mentioning?
>
> OMG. I see it now:
>
> + /* Don't bother specifying the quote id */
> + .rdx = QUOTE_ID_MASK & (u64)-1,
>
> So IIUC, there's a thread pool, but somewhere in the muck of RFC patches the
> kernel initializes the pool with a size of 1?
>
Yea it is like a thread pool. Described above in more abstract terms of a
resumable operation that can save state and can have the number of save state
areas configured.
> And then adds a global mutex to
> serialize quotes across the entire system. That's certainly a choice.
Passing the GHCI request straight to the seamcall within KVM was a bad choice.
But I think a single mutex is a possible start for a KVM ioctl(). Some follow-on
patches could allow for configuring more options.
>
> This snippet from the RFC[2] doesn't help, because it makes it seem like there
> is no state:
>
> 2. Host just refills the previous args (may have been modified by seamcall
> output) on retry. such as TDH_EXT_INIT, TDH_EXT_MEM_ADD,
> TDH_QUOTE_INIT, TDH_QUOTE_GET...
>
> But that's presumably just because the kernel is using a single REQUEST_ID.
>
> That spec continues with:
>
> If supported for the specific request type, the host VMM can request an
> interrupted request to be aborted.
>
> Though per the above, apparently whether or not that's going to be supported is
> TBD?
So extension aborts are a whole 'nother topic. And "support" has a nuanced
answer to boot. I'd be very glad to hear your thoughts on it. Should we leave it
for a later topic or get into it now? I *think* we don't need it, but the more
we get into the details here, the more it's a notable thing we haven't covered.
>
> I am beyond frustrated at this point. I guess shame on me for not pouring over
> dense docs? But I shouldn't have to.
>
Yea, I agree you shouldn't have to. But I think a lot of this stuff did get
covered in this thread. If the post-RFC state was confusion, we can only try to
rectify it later from that point.
> The *ENTIRE* point of an RFC like the one
> that started all this is to get feedback on the design, but that's really hard to
> achieve if you don't actually provide *any* details about the design.
>
> And there are MASSIVE design choices in here. Like sizing the "thread" pool to 1,
> and thus deliberately serializing quotes across all VMs. That's going to end well
> when someone spins up 10s or 100s of TDs and they all try to attest at once.
> Not listening for fatal signals while spinning on TDH.GET.QUOTE will save us though!
> At least with PREEMPT_LAZY being forced, not doing cond_resched() is "fine".
>
> What's worse, some of these patches that say nothing useful in the changelogs
> have several Reviewed-by tags, which means multiple people looked at this as
> thought it was all good.
>
> Regardless of where we end up with quoting, whatever RFC is sent next needs to be
> a million times better. My expectations are that someone with core TDX knowledge
> and a basic understanding of attestation would be able to get a full and complete
> understanding of the design and tradeoffs from the RFC alone. If I have to
> rereference a spec for anything other than double check params and magic values,
> the entire series is getting ignored.
I mean. Confidential compute in general is trying to reinvent decades of
virtualization solutions with one hand tied behind it's back. There *are* a lot
of design choices.
Besides a rushed RFC, I think another factor is that this is the first TDX
feature we have done that you don't already have at least some exposure to. The
level of background delta need is way higher than S-EPT management for example.
And I'm less and less sure we are going to be successful distilling things like
this down in a way that doesn't leave you frustrated. But if you want details
there is just going to be a lot of them. Hmm.
>
> As for guest-driven quoting, to me there is a fairly straightforward solution.
> Assuming using guest-donated memory is too complex for the "virtual thread":
>
> 1. Drop "virtual thread pools" from quoting and support at most one in-flight
> quote per TD. I doubt whatever memory is needed for a virtual thread is so
> insanely ridiculous that we can't burn that much memory per TD. 32KiB is a
> no-brainer. Above that and we'd have to reconsider, but if quoting requires
> more than 32KiB of scratch space, I have to wonder what on earth it's doing.
I think the abstract descriptions are not helping, so I'll try some more
details. Today the "thread" involves something like a vCPU in a TD. You can see
some references in the TDX base spec like:
Currently, TDX Module extensions are implemented as guests running in SEAM non-
root mode, called NRX Modules. Technically, each NRX Module is similar to a TD.
However, logically and from the usage perspective NRX Modules are part of the
TDX Module and are hidden from the host VMM. The TDX Module builds the NRX
Modules as part of the TDX Module’s initialization sequence and calls them when
their functionality is required for execution of some TDX Module interface
functions.
Think like the service TD concept got wrapped in the TDX module. *But* the idea
is to be more like an abstract thing. Being a TD is just an implementation
detail and not necessarily a long term fixed thing. So based on that, I'd think
the memory needs could be optimized.
Let's track down the exact size. I've seen it but forgot the exact number.
>
> Then, because the resource is fixed, have the host provide it at TD creation.
> The TDX-Module keeps track of whether or not a quote is in-flight, and rejects
> attempts to start a new one. The guest can simply guard quotes with a mutex.
>
> The host never has to worry about canceling a quote request, because it just
> needs to reclaim the TD as normal. If desired, the TDX Module can provide a
> TDCALL to let the guest cancel a quote.
This idea came up before actually, even on the host side solution IIRC. Yea. The
tradeoff is just overhead per-TD.
>
> 2. To avoid having to check for guest interrupts, for the "SW based" synchronous
> method, simply resume the guest after processing a fixed amount of state. Yes,
> that will increase the best case latency for a single quote, but *best* case
> latency isn't a huge concern, and I doubt the overhead of a VMX roundtrip will
> significantly impact that. It's the tail latencies that will be problematic,
> and a forcing the serialization into the guest (the aforementioned mutex) means
> the tail latencies will only be affected by host activity on *that* CPU, which
> is more or less the status quo.
>
> E.g. even assuming an absurd 20% overhead for the VMX round trips, having a
> quote take ~1.2 *every* time is would be a far, far better experience than a
> quote taking 1ms - 10ms to complete.
Dave will probably not get a chance to respond this week, but we should maybe
discuss what acceptable guest latency actually is.
This thread left me thinking that we might not need to handle guest interrupts
immediately. But it would be good for the solution to be able to extend to
support them later if needed.
>
> 3. Host interrupts work as they do for normal TD operation. The state is all
> there, attached to the vCPU. It's on the host to run the vCPU according to
> its SLOs, so as not to starve/DoS the guest.
>
> 4. If/when S3M accesses during quoting come along (it's not clear to me if the
> SW-based quoting needs to access the S3M during every quote, or just during
> setup), then provide the host with a mechanism to throttle quote requests.
> E.g. a simple counter that triggers an exit to the host when it hits zero
> would probably suffice.
>
> 5. If/when "HW based" asynchronous quoting comes along, then if S3M can send
> IRQs on completion, have a host-side driver wired up to call into the TDX
> Module to tell it a quote is ready. Presumably the TDX Module can keep track
> of where the quote came from. If the S3M can't do IRQs, then just put the
> onus on the guest to poll to see when it's quote is ready? This needs more
> details on the S3M, but it doesn't seem insurmountable.
>
> But it's also not obvious to me why Linux/KVM would ever want to support this
> mode. AFAICT, it adds more complexity (especially when accounting for noisy
> neighbor issues) in order to provide a worse customer experience.
I think it is like "HW security" kind of value to some people. Think like a TPMs
value vs an in memory protected key. Depends on your outlook and other factors.
But I'm fine (glad) to not figure out "HW based" attestation today. Just want a
re-design to be able to grow for stuff like that.
>
> #2 above is going to be a problem with the host-based, CPU-intensive "SW" quoting,
> which IIUC is what is currently being proposed. Unless the host provides an overhead
> CPU to do quoting (which is likely a non-starter for Google at least), from the
> guest's perspective, a vCPU will disappear and become non-responsive for however
> long it takes the host to generate the quote.
>
> The SetupEventNotifyInterrupt will help a little, but won't fully alleviate the
> problem, because that still requires a pCPU to do the work. E.g. the vCPU wouldn't
> be completely non-responsive, but it would observe extremely high steal-time until
> the quote is generated, which is gross. Presumably this is how the current SGX-
> based quoting behaves? But if these quotes are getting more massive, i.e. slower,
> then it's probably something that needs to be addressed.
Hmm, yea. I mean it would be the same for SGX today, but I could totally see it
being a factor. Maybe some TDX users on CC will chime in on this one. This could
be the kind of thing that hinges the design clearly.
The other option would be, ugh, the TDH.QUOTE.GET becomes vCPU scoped and exits
for pending guest interrupts. Then it could use the vCPU thread. But it is a
long way from the guest at that point.
>
> And that's assuming the virtual thread pool is sized such that there can be a thread
> per VM. If there is any cross-VM resource sharing, then we're going to have noisy
> neighbor problems and the above issue becomes an order of magnitude worse.
>
> > > > A future thing is a flow where S3M is engaged for every quote. That would be
> > > > expected to take longer, and have greater limits on concurrency. But the
> > > > details are not sorted on what exactly the user will want, or how it would
> > > > be implemented.
> > > >
> > > > I think you are maybe wondering whether S3M could be used such that the
> > > > guest could wait on a quote while not needing any saved state area?
> > >
> > > I'm trying to figure out if *any* path is viable.
> > >
> > > Burning 1ms of CPU time to generate a quote in uninterruptible code is a non-
> > > starter. Hell, 100us is a non-starter.
> >
> > What code is uninterruptible? The whole point of the extensions thing is to make
> > the operations broadly interruptible.
>
> Due to the lack of useful changelogs and being allergic to TDX specs, it wasn't
> clear to me that TDH.GET.QUOTE was "fully" interruptible *and* restartable,
> i.e. would guarantee forward progress.
I do think we covered the interruptible part pretty thoroughly in this thread.
>
> > > Waiting 2s for a quote to come back from the S3M is a non-starter.
> >
> > Not sure where 2s is coming from. But do you mean waiting in the guest? Or where
> > is waiting a non-starter?
>
> If each S3M quote takes "tens of milliseconds", and there's one S3M per socket,
> it doesn't take that many concurrent quote requests for one of the TDs to observe
> a 2s+ latency to get its quote. E.g. quote takes 50ms, boot 40 TDs, and voila,
> that last TD going through boot gets hit with a 2s+ quote latency.
Ah I see what you are getting at.
>
> > > The numbers matter, and *none* of this is reviewable, even in RFC format,
> > > without a crisp understanding of what latencies we are talking about.
> >
> > I agree not having a measurement of the initial DICE behavior is a detriment for
> > the TDG discussion. I didn't think we needed it though,
>
> Ignore the TDG discussion, I'm saying they're needed for *any* discussion, i.e.
> for reviewing the TDH implementation as well.
>
> > if we leaned on the original locking/scheduling reasons to prefer a host
> > based quote flow.
> >
> > Previously you asked for ballparks and got them. So what do you need exactly?
>
> Guarantees around orders of magnitude. There is a enormous difference between
> a CPU-intensive operation taking 10us versus 2ms. I.e. "milliseconds or less"
> isn't a ballpark, it's a country, maaaaybe a state. For CPU-based in particular,
> I was expecting ballparks with +/- 10us of precision, not "somewhere between 0
> and several scheduler ticks".
>
Ok, I think that Peter was giving best guesses. We can try to find a worst case
commitment.
But it does seem like the PQC transition could end up with some changes that are
totally out of control. Like that hybrid crypto lwn article I linked. How much
does 10us precision on the first thing really buy us when discussing a re-design
that would hopefully live a longer time than SGX based?
^ permalink raw reply [flat|nested] 33+ messages in thread
* Re: TDG quote analysis
2026-09-18 2:39 ` Edgecombe, Rick P
@ 2026-09-18 13:13 ` Sean Christopherson
2026-09-18 18:12 ` Edgecombe, Rick P
0 siblings, 1 reply; 33+ messages in thread
From: Sean Christopherson @ 2026-09-18 13:13 UTC (permalink / raw)
To: Rick P Edgecombe
Cc: Yilun Xu, Elena Reshetova, Binbin Wu, Dave Hansen,
Vishal Annapurve, kas@kernel.org, pbonzini@redhat.com, Peter Fang,
kvm@vger.kernel.org
On Fri, Sep 18, 2026, Rick P Edgecombe wrote:
> On Thu, 2026-09-17 at 17:04 -0700, Sean Christopherson wrote:
> > If supported for the specific request type, the host VMM can request an
> > interrupted request to be aborted.
> >
> > Though per the above, apparently whether or not that's going to be supported is
> > TBD?
>
> So extension aborts are a whole 'nother topic. And "support" has a nuanced
> answer to boot. I'd be very glad to hear your thoughts on it. Should we leave it
> for a later topic or get into it now? I *think* we don't need it, but the more
> we get into the details here, the more it's a notable thing we haven't covered.
How is saying "I no longer want to complete this quote" at all complex? I can
see how aborting a request to the S3M might be somewhat pointless, but inserting
the equivalent to signal_pending() in a long-running operation doesn't seem that
difficult.
> > Regardless of where we end up with quoting, whatever RFC is sent next needs to be
> > a million times better. My expectations are that someone with core TDX knowledge
> > and a basic understanding of attestation would be able to get a full and complete
> > understanding of the design and tradeoffs from the RFC alone. If I have to
> > rereference a spec for anything other than double check params and magic values,
> > the entire series is getting ignored.
>
> I mean. Confidential compute in general is trying to reinvent decades of
> virtualization solutions with one hand tied behind it's back. There *are* a lot
> of design choices.
>
> Besides a rushed RFC, I think another factor is that this is the first TDX
> feature we have done that you don't already have at least some exposure to. The
> level of background delta need is way higher than S-EPT management for example.
>
> And I'm less and less sure we are going to be successful distilling things like
It seems like part of the problem is that you're trying to "distill" into abstract
concepts. I don't want abstract concepts, I want a description of what the TDX
Module code will literally do, using verbiage and terminology that a KVM developer
will natively understand.
> this down in a way that doesn't leave you frustrated. But if you want details
> there is just going to be a lot of them. Hmm.
This is what I want. Note, this description is apparently wildly wrong, because
I wrote it before reading about whatever NRX modules are. But I'm leaving it
because it highlights my point about not wanting TDX jargon rephrased as abstract
concepts.
Under the hood, TDH.GET.QUOTE uses what TDX calls a "virtual thread pool".
That just means there's a pre-allocated pool of "thread" memory that can be
used to service interruptible requests (it's not true interruption, it's
basically voluntary preemption). The number of threads in the pool is
configured by software during ??? (TDH.QUOTE.INIT?), and directly controls
how many in-flight TDH.GET.QUOTE operations there can be at any given time.
Each thread consumes ??? KiB of memory, and <reasoning>, so we chose to set
the pool size to 1, i.e. to only allow a single quote to be in-flight across
the entire system.
TDH.GET.QUOTE is "interruptible" because the bulk of the quote crypto is done
by the TDX module, i.e. on the CPU that invokes TDH.GET.QUOTE. We don't have
numbers for TDH.GET.QUOTE yet (why not?), but the very rough ballpark is that
it will takes 1-2ms. Note, the S3M is used only during ??? (TDH.QUOTE.INIT?)
to get the platform key used to generate quotes, i.e. generating a quote is
fully synchronous relative to the core.
> > As for guest-driven quoting, to me there is a fairly straightforward solution.
> > Assuming using guest-donated memory is too complex for the "virtual thread":
> >
> > 1. Drop "virtual thread pools" from quoting and support at most one in-flight
> > quote per TD. I doubt whatever memory is needed for a virtual thread is so
> > insanely ridiculous that we can't burn that much memory per TD. 32KiB is a
> > no-brainer. Above that and we'd have to reconsider, but if quoting requires
> > more than 32KiB of scratch space, I have to wonder what on earth it's doing.
>
> I think the abstract descriptions are not helping, so I'll try some more
> details. Today the "thread" involves something like a vCPU in a TD. You can see
> some references in the TDX base spec like:
> Currently, TDX Module extensions are implemented as guests running in SEAM non-
> root mode, called NRX Modules. Technically, each NRX Module is similar to a TD.
So my above statement that "it's not true interruption, it's basically voluntary
preemption" is wrong? Because if the quote crud is running in a TD, then events
will trigger VM-Exit, and it will be impossible to make forward progress without
SEAMRETing to the host.
If that's true, then that changes my understanding of this just a bit, because
it means the quote operation isn't manually saving state at a "good stopping point",
it's relying on the VM-Exit to save/restore at whenever it happened to be. But
given that you say "This idea came up before actually", maybe chunking the quote
operation isn't a big lift?
But the original mail says this:
-- Guest Interrupts --
It is not nice to keep the guest from running for too long. The TDH.QUOTE.GET
SEAMCALL monitors for host interrupts, but not guest ones. Similarly, the SGX
^^^^^^^^^^^^^^^^^^^^^^^^^^^^
quoting enclave doesn't know what is happening in the TD. So, to help reduce
guest latencies, the GHCI exposes a way to register for a notification
(SetupEventNotifyInterrupt) when the quoting is finished. A quote TDVMCALL
handler can start the quote operation on another host thread, and resume
guest execution waiting for the quote to finish.
which strongly suggests active polling. Do you see why I don't want abstract
descriptions? Similar to changelogs, there's a balance between "too detailed"
and "too abstract", and right now all of this is firmly on the "too abstract"
side of the world.
> > 2. To avoid having to check for guest interrupts, for the "SW based" synchronous
> > method, simply resume the guest after processing a fixed amount of state. Yes,
> > that will increase the best case latency for a single quote, but *best* case
> > latency isn't a huge concern, and I doubt the overhead of a VMX roundtrip will
> > significantly impact that. It's the tail latencies that will be problematic,
> > and a forcing the serialization into the guest (the aforementioned mutex) means
> > the tail latencies will only be affected by host activity on *that* CPU, which
> > is more or less the status quo.
> >
> > E.g. even assuming an absurd 20% overhead for the VMX round trips, having a
> > quote take ~1.2 *every* time is would be a far, far better experience than a
> > quote taking 1ms - 10ms to complete.
>
> Dave will probably not get a chance to respond this week, but we should maybe
> discuss what acceptable guest latency actually is.
Yes. But it's not necessarily about "acceptable" guest latency, I care more about
*predictable* guest latency. E.g. making up numbers to illustrate the point,
achieving 50us latency for the happy case is meaningless if the P99 latency is 10ms.
^ permalink raw reply [flat|nested] 33+ messages in thread
* Re: TDG quote analysis
2026-09-18 13:13 ` Sean Christopherson
@ 2026-09-18 18:12 ` Edgecombe, Rick P
2026-09-18 21:32 ` Sean Christopherson
0 siblings, 1 reply; 33+ messages in thread
From: Edgecombe, Rick P @ 2026-09-18 18:12 UTC (permalink / raw)
To: seanjc@google.com
Cc: Xu, Yilun, Reshetova, Elena, Wu, Binbin, Hansen, Dave,
Annapurve, Vishal, kas@kernel.org, pbonzini@redhat.com,
Fang, Peter, kvm@vger.kernel.org
On Fri, 2026-09-18 at 06:13 -0700, Sean Christopherson wrote:
> On Fri, Sep 18, 2026, Rick P Edgecombe wrote:
> > On Thu, 2026-09-17 at 17:04 -0700, Sean Christopherson wrote:
> > > If supported for the specific request type, the host VMM can request an
> > > interrupted request to be aborted.
> > >
> > > Though per the above, apparently whether or not that's going to be
> > > supported is
> > > TBD?
> >
> > So extension aborts are a whole 'nother topic. And "support" has a nuanced
> > answer to boot. I'd be very glad to hear your thoughts on it. Should we
> > leave it for a later topic or get into it now? I *think* we don't need it,
> > but the more we get into the details here, the more it's a notable thing we
> > haven't covered.
>
> How is saying "I no longer want to complete this quote" at all complex? I can
> see how aborting a request to the S3M might be somewhat pointless, but
> inserting the equivalent to signal_pending() in a long-running operation
> doesn't seem that difficult.
As an interface it is not complex. But as an implementation, there are tradeoffs
and various options.
> > >
> > > [...]
> > >
> >
> > I think the abstract descriptions are not helping, so I'll try some more
> > details. Today the "thread" involves something like a vCPU in a TD. You can
> > see
> > some references in the TDX base spec like:
> > Currently, TDX Module extensions are implemented as guests running in
> > SEAM non- root mode, called NRX Modules. Technically, each NRX Module is
> > similar to a TD.
>
> So my above statement that "it's not true interruption, it's basically
> voluntary preemption" is wrong? Because if the quote crud is running in a TD,
> then events will trigger VM-Exit, and it will be impossible to make forward
> progress without SEAMRETing to the host.
>
> If that's true, then that changes my understanding of this just a bit, because
> it means the quote operation isn't manually saving state at a "good stopping
> point", it's relying on the VM-Exit to save/restore at whenever it happened to
> be.
Yea, I'd think it is best not to mess with crypto implementations to add
anything like checkpoints. That is what extensions bring to the table. The other
resumable seamcalls basically work like checkpoints. But it is not suitable for
all operations.
> But given that you say "This idea came up before actually", maybe chunking
> the quote operation isn't a big lift?
The idea that came up was to have a saved state are per-TD. But this only
eliminated concurrency issues in the SW based flow. In the future a HW based
flow would be S3M limited. Which is why I keep saying a redesigned solution
should be robust to concurrency limitations.
>
> But the original mail says this:
>
> -- Guest Interrupts --
>
> It is not nice to keep the guest from running for too long. The
> TDH.QUOTE.GET
> SEAMCALL monitors for host interrupts, but not guest ones. Similarly, the
> SGX
> ^^^^^^^^^^^^^^^^^^^^^^^^^^^^
> quoting enclave doesn't know what is happening in the TD. So, to help
> reduce guest latencies, the GHCI exposes a way to register for a notification
> (SetupEventNotifyInterrupt) when the quoting is finished. A quote TDVMCALL
> handler can start the quote operation on another host thread, and resume
> guest execution waiting for the quote to finish.
>
> which strongly suggests active polling. Do you see why I don't want abstract
> descriptions? Similar to changelogs, there's a balance between "too detailed"
> and "too abstract", and right now all of this is firmly on the "too abstract"
> side of the world.
>
For a host interrupt, I expected basically the same flow as a normal TD. For the
guest flow, I don't know the exact vmx-level solution for it. I threw out
polling and heard there could be a cleaner option. Not sure what it was.
> >
> >
> > Dave will probably not get a chance to respond this week, but we should
> > maybe discuss what acceptable guest latency actually is.
>
> Yes. But it's not necessarily about "acceptable" guest latency, I care more
> about *predictable* guest latency. E.g. making up numbers to illustrate the
> point, achieving 50us latency for the happy case is meaningless if the P99
> latency is 10ms.
The SEAMCALLs today have max latency and they get measured through some testing
process. In practice, some exceed it based on the reasoning that they are only
run once or a few times (setup, etc). So for the host side that is how it got
discussed. It was in terms of worst case.
But now I'm wondering what your specific interest is. It seems you have more
interest in non-KVM in-the-loop guest side latency than you did on host side.
Which is fine. But it makes me want to double check: You want an average and
worse case latency for the first DICE attestation "SW" flow?
Or you want an average and worst case promise for the future flows?
Or you want an average and worse case based on whatever guest interrupt
detection scheme is used for TDG extension calls?
The last one seems the most relevant to me. And the point would be to determine
the tradeoffs between host thread based quote call and a TDG call. But of course
none of this is built yet such that it could be measured. This was the "here is
how it could look in general" discussion. So I think it would take some TDX
module POC work.
^ permalink raw reply [flat|nested] 33+ messages in thread
* Re: TDG quote analysis
2026-09-18 18:12 ` Edgecombe, Rick P
@ 2026-09-18 21:32 ` Sean Christopherson
2026-09-21 23:00 ` Peter Fang
2026-09-21 23:04 ` Dave Hansen
0 siblings, 2 replies; 33+ messages in thread
From: Sean Christopherson @ 2026-09-18 21:32 UTC (permalink / raw)
To: Rick P Edgecombe
Cc: Yilun Xu, Elena Reshetova, Binbin Wu, Dave Hansen,
Vishal Annapurve, kas@kernel.org, pbonzini@redhat.com, Peter Fang,
kvm@vger.kernel.org
On Fri, Sep 18, 2026, Rick P Edgecombe wrote:
> On Fri, 2026-09-18 at 06:13 -0700, Sean Christopherson wrote:
> > But given that you say "This idea came up before actually", maybe chunking
> > the quote operation isn't a big lift?
>
> The idea that came up was to have a saved state are per-TD. But this only
> eliminated concurrency issues in the SW based flow. In the future a HW based
> flow would be S3M limited. Which is why I keep saying a redesigned solution
> should be robust to concurrency limitations.
IMO, that's the complete wrong way to look at this. HW-based is inherently
serialized, it does NOT have concurrency, period. SW-based is has no inherent
concurrency restrictions beyond existing CPU contention.
Trying to find a perfect one-size-fits-all solution is impossible, because the
problems with each are so very different. The problem you want to solve for
SW-based is how to support preemption. The problem you want to solve for HW-based
is how to hide the fact that someone in the future might think following AMD's
lead and putting a precious resource into a tiny microprosser on a slow bus is a
fantastic idea.
By creating these system-wide thread pools, TDX has effectively created a bizarre
M:N scheduling problem, *and* introduced a completely avoidable noisy-neighbor
problem.
Assuming my understanding is (finally) correct, and the SW-based flows do all the
work on the CPU, then the only way making the SW-based flows asynchronous adds
value is if the host is willing to set aside CPU cores for such chores. And for
the use cases where TDX makes sense, AFAIK no CSP *wants* to do that. Not to
mention the RFC doesn't even support that, because KVM doesn't resume the vCPU
until the quote is ready.
> But now I'm wondering what your specific interest is. It seems you have more
> interest in non-KVM in-the-loop guest side latency than you did on host side.
> Which is fine. But it makes me want to double check: You want an average and
> worse case latency for the first DICE attestation "SW" flow?
I care about future me not getting pulled into a customer issue because a vCPU
got waylaid by a system-wide mutex for multiple seconds.
^ permalink raw reply [flat|nested] 33+ messages in thread
* Re: TDG quote analysis
2026-09-18 21:32 ` Sean Christopherson
@ 2026-09-21 23:00 ` Peter Fang
2026-09-21 23:06 ` Dave Hansen
2026-09-21 23:04 ` Dave Hansen
1 sibling, 1 reply; 33+ messages in thread
From: Peter Fang @ 2026-09-21 23:00 UTC (permalink / raw)
To: Sean Christopherson
Cc: Rick P Edgecombe, Yilun Xu, Elena Reshetova, Binbin Wu,
Dave Hansen, Vishal Annapurve, kas@kernel.org,
pbonzini@redhat.com, kvm@vger.kernel.org
On Fri, Sep 18, 2026 at 02:32:23PM -0700, Sean Christopherson wrote:
> On Fri, Sep 18, 2026, Rick P Edgecombe wrote:
> > On Fri, 2026-09-18 at 06:13 -0700, Sean Christopherson wrote:
> > > But given that you say "This idea came up before actually", maybe chunking
> > > the quote operation isn't a big lift?
> >
> > The idea that came up was to have a saved state are per-TD. But this only
> > eliminated concurrency issues in the SW based flow. In the future a HW based
> > flow would be S3M limited. Which is why I keep saying a redesigned solution
> > should be robust to concurrency limitations.
>
> IMO, that's the complete wrong way to look at this. HW-based is inherently
> serialized, it does NOT have concurrency, period. SW-based is has no inherent
> concurrency restrictions beyond existing CPU contention.
>
> Trying to find a perfect one-size-fits-all solution is impossible, because the
> problems with each are so very different. The problem you want to solve for
> SW-based is how to support preemption. The problem you want to solve for HW-based
^ more discussion below
> is how to hide the fact that someone in the future might think following AMD's
> lead and putting a precious resource into a tiny microprosser on a slow bus is a
> fantastic idea.
As shocking as it sounds, there are actually cloud providers out there
that are wanting this, or at the very least wanting to see how it goes
in real deployment. I think memory side channel attacks are a real
concern for some of them. But yeah this is not great from a performance
perspective.
>
> By creating these system-wide thread pools, TDX has effectively created a bizarre
> M:N scheduling problem, *and* introduced a completely avoidable noisy-neighbor
> problem.
Hmm... Then how about M:M? Then the problem can be reduced to 1:1. I
think it would take some TDX module investigative work to say for sure
what this would look like. But just throwing it out there as an idea?
>
> Assuming my understanding is (finally) correct, and the SW-based flows do all the
> work on the CPU, then the only way making the SW-based flows asynchronous adds
> value is if the host is willing to set aside CPU cores for such chores. And for
> the use cases where TDX makes sense, AFAIK no CSP *wants* to do that. Not to
> mention the RFC doesn't even support that, because KVM doesn't resume the vCPU
> until the quote is ready.
There was also the fairness/starvation problem at play. In RFC if the
vCPU thread waited on the lock (say using down_timeout(), or even some
kind of fancier waitqueue) and gave it up to reenter TD to handle a
guest interrupt, it would have to retry the lock again later. And then
there would be preemption-retry livelock risks. Moving this lock into
the TDX module could still create the same problem, because basically
the guest would be re-calling the TDCALL after preemption instead. So I
think there would need to be some sort of additional smarts to handle
both preemption and fairness at the same time.
>
> > But now I'm wondering what your specific interest is. It seems you have more
> > interest in non-KVM in-the-loop guest side latency than you did on host side.
> > Which is fine. But it makes me want to double check: You want an average and
> > worse case latency for the first DICE attestation "SW" flow?
>
> I care about future me not getting pulled into a customer issue because a vCPU
> got waylaid by a system-wide mutex for multiple seconds.
I think quoting is typically a once-per-TD-boot type of thing. Or it is
not usually expected to happen very often in a TD. Can you share what
kind of performance expectations you have for this? E.g. is seeing a
high steal time during TD boot a non-starter?
Thanks,
Peter
^ permalink raw reply [flat|nested] 33+ messages in thread
* Re: TDG quote analysis
2026-09-18 21:32 ` Sean Christopherson
2026-09-21 23:00 ` Peter Fang
@ 2026-09-21 23:04 ` Dave Hansen
2026-09-21 23:19 ` Sean Christopherson
1 sibling, 1 reply; 33+ messages in thread
From: Dave Hansen @ 2026-09-21 23:04 UTC (permalink / raw)
To: Sean Christopherson, Rick P Edgecombe
Cc: Yilun Xu, Elena Reshetova, Binbin Wu, Vishal Annapurve,
kas@kernel.org, pbonzini@redhat.com, Peter Fang,
kvm@vger.kernel.org
On 9/18/26 14:32, Sean Christopherson wrote:
> By creating these system-wide thread pools, TDX has effectively created a bizarre
> M:N scheduling problem, *and* introduced a completely avoidable noisy-neighbor
> problem.
To be honest, I just want to chop all of the complexity out of this.
I want to assume that the TDX module can do one quote at a time. If that
quote takes too long, software kills it and lets the next guy do a quote.
I don't even want to know that the TDX module *can* have 1 or 10 or 100
quoting threads. I just want one which I can make sane rules around
because 1 is *FINE*.
I also want to ignore that anyone dreamed up (or copied from AMD) the
S3M thing. My *hope* is that once we have sane 1-thread "takes too long"
semantics, we can let folks who crave S3M quoting just pop up that
constant by one or three orders of magnitude and go on with their
masochistic ways.
^ permalink raw reply [flat|nested] 33+ messages in thread
* Re: TDG quote analysis
2026-09-21 23:00 ` Peter Fang
@ 2026-09-21 23:06 ` Dave Hansen
2026-09-23 4:09 ` Vishal Annapurve
0 siblings, 1 reply; 33+ messages in thread
From: Dave Hansen @ 2026-09-21 23:06 UTC (permalink / raw)
To: Peter Fang, Sean Christopherson
Cc: Rick P Edgecombe, Yilun Xu, Elena Reshetova, Binbin Wu,
Vishal Annapurve, kas@kernel.org, pbonzini@redhat.com,
kvm@vger.kernel.org
On 9/21/26 16:00, Peter Fang wrote:
>> is how to hide the fact that someone in the future might think following AMD's
>> lead and putting a precious resource into a tiny microprosser on a slow bus is a
>> fantastic idea.
> As shocking as it sounds, there are actually cloud providers out there
> that are wanting this, or at the very least wanting to see how it goes
> in real deployment. I think memory side channel attacks are a real
> concern for some of them. But yeah this is not great from a performance
> perspective.
I'll believe it when I see it (on this mailing list).
If folks want complexity and low performance in the kernel for
side-channel defense, then I'm your guy. I've made half a career out of
it. All that I ask is that they come and ask for it in the open. Here,
please.
^ permalink raw reply [flat|nested] 33+ messages in thread
* Re: TDG quote analysis
2026-09-21 23:04 ` Dave Hansen
@ 2026-09-21 23:19 ` Sean Christopherson
2026-09-21 23:36 ` Dave Hansen
0 siblings, 1 reply; 33+ messages in thread
From: Sean Christopherson @ 2026-09-21 23:19 UTC (permalink / raw)
To: Dave Hansen
Cc: Rick P Edgecombe, Yilun Xu, Elena Reshetova, Binbin Wu,
Vishal Annapurve, kas@kernel.org, pbonzini@redhat.com, Peter Fang,
kvm@vger.kernel.org
On Mon, Sep 21, 2026, Dave Hansen wrote:
> On 9/18/26 14:32, Sean Christopherson wrote:
> > By creating these system-wide thread pools, TDX has effectively created a bizarre
> > M:N scheduling problem, *and* introduced a completely avoidable noisy-neighbor
> > problem.
>
> To be honest, I just want to chop all of the complexity out of this.
>
> I want to assume that the TDX module can do one quote at a time. If that
> quote takes too long, software kills it and lets the next guy do a quote.
>
> I don't even want to know that the TDX module *can* have 1 or 10 or 100
> quoting threads. I just want one which I can make sane rules around
> because 1 is *FINE*.
I'm not remotely convinced one per system is fine. I'm all for one per VM, but
I'm pretty sure the SEAMCALL alone will have a measurable performance impact,
especially when factoring in that it will require a VMCS shootdown to do VMCLEAR
on the previous pCPU. Mix in a few (tens of?) thousand more cycles, and any use
case that's trying to boot a decent number of TDX guests in parallel will be sad.
^ permalink raw reply [flat|nested] 33+ messages in thread
* Re: TDG quote analysis
2026-09-21 23:19 ` Sean Christopherson
@ 2026-09-21 23:36 ` Dave Hansen
2026-09-21 23:46 ` Sean Christopherson
0 siblings, 1 reply; 33+ messages in thread
From: Dave Hansen @ 2026-09-21 23:36 UTC (permalink / raw)
To: Sean Christopherson
Cc: Rick P Edgecombe, Yilun Xu, Elena Reshetova, Binbin Wu,
Vishal Annapurve, kas@kernel.org, pbonzini@redhat.com, Peter Fang,
kvm@vger.kernel.org
On 9/21/26 16:19, Sean Christopherson wrote:
> On Mon, Sep 21, 2026, Dave Hansen wrote:
>> On 9/18/26 14:32, Sean Christopherson wrote:
>>> By creating these system-wide thread pools, TDX has effectively created a bizarre
>>> M:N scheduling problem, *and* introduced a completely avoidable noisy-neighbor
>>> problem.
>> To be honest, I just want to chop all of the complexity out of this.
>>
>> I want to assume that the TDX module can do one quote at a time. If that
>> quote takes too long, software kills it and lets the next guy do a quote.
>>
>> I don't even want to know that the TDX module *can* have 1 or 10 or 100
>> quoting threads. I just want one which I can make sane rules around
>> because 1 is *FINE*.
> I'm not remotely convinced one per system is fine. I'm all for one per VM, but
> I'm pretty sure the SEAMCALL alone will have a measurable performance impact,
> especially when factoring in that it will require a VMCS shootdown to do VMCLEAR
> on the previous pCPU. Mix in a few (tens of?) thousand more cycles, and any use
> case that's trying to boot a decent number of TDX guests in parallel will be sad.
Sad, but way less sad than the folks who decided to offload their
quoting to an Arduino attached over a fancy serial port. ;)
I do think we can _start_ with one. One needs no ABI and no policy bike
shedding. If my crystal ball is right, we can stop there. If yours is
right, then we can go build the stuff we need to make it scale.
Oh, and folks can _theoretically_ do one quoting instance per VM. But I
think the overhead per quoting instance is in 100MB ballpark. So a bit
rotund to be per-TD in practice.
^ permalink raw reply [flat|nested] 33+ messages in thread
* Re: TDG quote analysis
2026-09-21 23:36 ` Dave Hansen
@ 2026-09-21 23:46 ` Sean Christopherson
2026-09-22 0:03 ` Dave Hansen
0 siblings, 1 reply; 33+ messages in thread
From: Sean Christopherson @ 2026-09-21 23:46 UTC (permalink / raw)
To: Dave Hansen
Cc: Rick P Edgecombe, Yilun Xu, Elena Reshetova, Binbin Wu,
Vishal Annapurve, kas@kernel.org, pbonzini@redhat.com, Peter Fang,
kvm@vger.kernel.org
On Mon, Sep 21, 2026, Dave Hansen wrote:
> On 9/21/26 16:19, Sean Christopherson wrote:
> > On Mon, Sep 21, 2026, Dave Hansen wrote:
> >> On 9/18/26 14:32, Sean Christopherson wrote:
> >>> By creating these system-wide thread pools, TDX has effectively created a bizarre
> >>> M:N scheduling problem, *and* introduced a completely avoidable noisy-neighbor
> >>> problem.
> >> To be honest, I just want to chop all of the complexity out of this.
> >>
> >> I want to assume that the TDX module can do one quote at a time. If that
> >> quote takes too long, software kills it and lets the next guy do a quote.
> >>
> >> I don't even want to know that the TDX module *can* have 1 or 10 or 100
> >> quoting threads. I just want one which I can make sane rules around
> >> because 1 is *FINE*.
> > I'm not remotely convinced one per system is fine. I'm all for one per VM, but
> > I'm pretty sure the SEAMCALL alone will have a measurable performance impact,
> > especially when factoring in that it will require a VMCS shootdown to do VMCLEAR
> > on the previous pCPU. Mix in a few (tens of?) thousand more cycles, and any use
> > case that's trying to boot a decent number of TDX guests in parallel will be sad.
>
> Sad, but way less sad than the folks who decided to offload their
> quoting to an Arduino attached over a fancy serial port. ;)
>
> I do think we can _start_ with one. One needs no ABI and no policy bike
> shedding. If my crystal ball is right, we can stop there. If yours is
> right, then we can go build the stuff we need to make it scale.
My preference *was* to not have to find out, because those types of scaling issues
tend to lead to a mad scramble and angry customers. But...
> Oh, and folks can _theoretically_ do one quoting instance per VM. But I
> think the overhead per quoting instance is in 100MB ballpark.
MB with an M!?!?! What!?!? How? I was assuming my 32KiB threshold was fairly
conservative; what is the TDX Module doing that it needs 100MiB? That's insane.
That's literally 3x the footprint of some full-blown VMMs.
If it really is 100MiB per instance, that definitely changes things. I'm much
more willing to roll the dice and hope quoting is never a bottleneck in practice
if it means saving 100MiB per VM. Yeesh.
> So a bit rotund to be per-TD in practice.
LOL, yeah, a "bit".
^ permalink raw reply [flat|nested] 33+ messages in thread
* Re: TDG quote analysis
2026-09-21 23:46 ` Sean Christopherson
@ 2026-09-22 0:03 ` Dave Hansen
2026-09-22 6:48 ` Reshetova, Elena
0 siblings, 1 reply; 33+ messages in thread
From: Dave Hansen @ 2026-09-22 0:03 UTC (permalink / raw)
To: Sean Christopherson
Cc: Rick P Edgecombe, Yilun Xu, Elena Reshetova, Binbin Wu,
Vishal Annapurve, kas@kernel.org, pbonzini@redhat.com, Peter Fang,
kvm@vger.kernel.org
On 9/21/26 16:46, Sean Christopherson wrote:
>> Oh, and folks can _theoretically_ do one quoting instance per VM. But I
>> think the overhead per quoting instance is in 100MB ballpark.
> MB with an M!?!?! What!?!? How? I was assuming my 32KiB threshold was fairly
> conservative; what is the TDX Module doing that it needs 100MiB? That's insane.
> That's literally 3x the footprint of some full-blown VMMs.
First of all, I hope I'm getting my wires horribly crossed here.
Wouldn't be the first time I threw a couple of 0's at the end of
something accidentally. It's why I'm not a pharmacist.
Elena, am I remembering this horribly wrong? I think the interesting
numbers here are probably what the smallest possible extension TD (aka.
NRX TD) are, and then how much each quoting TD adds on top of that. Any
chance you could fill in my horrible memory?
^ permalink raw reply [flat|nested] 33+ messages in thread
* RE: TDG quote analysis
2026-09-22 0:03 ` Dave Hansen
@ 2026-09-22 6:48 ` Reshetova, Elena
2026-09-22 17:07 ` Edgecombe, Rick P
0 siblings, 1 reply; 33+ messages in thread
From: Reshetova, Elena @ 2026-09-22 6:48 UTC (permalink / raw)
To: Hansen, Dave, Sean Christopherson
Cc: Edgecombe, Rick P, Xu, Yilun, Wu, Binbin, Annapurve, Vishal,
kas@kernel.org, pbonzini@redhat.com, Fang, Peter,
kvm@vger.kernel.org
> -----Original Message-----
> From: Hansen, Dave <dave.hansen@intel.com>
> Sent: Tuesday, September 22, 2026 3:03 AM
> To: Sean Christopherson <seanjc@google.com>
> Cc: Edgecombe, Rick P <rick.p.edgecombe@intel.com>; Xu, Yilun
> <yilun.xu@intel.com>; Reshetova, Elena <elena.reshetova@intel.com>; Wu,
> Binbin <binbin.wu@intel.com>; Annapurve, Vishal <vannapurve@google.com>;
> kas@kernel.org; pbonzini@redhat.com; Fang, Peter <peter.fang@intel.com>;
> kvm@vger.kernel.org
> Subject: Re: TDG quote analysis
>
> On 9/21/26 16:46, Sean Christopherson wrote:
> >> Oh, and folks can _theoretically_ do one quoting instance per VM. But I
> >> think the overhead per quoting instance is in 100MB ballpark.
> > MB with an M!?!?! What!?!? How? I was assuming my 32KiB threshold was
> fairly
> > conservative; what is the TDX Module doing that it needs 100MiB? That's
> insane.
> > That's literally 3x the footprint of some full-blown VMMs.
>
> First of all, I hope I'm getting my wires horribly crossed here.
> Wouldn't be the first time I threw a couple of 0's at the end of
> something accidentally. It's why I'm not a pharmacist.
>
> Elena, am I remembering this horribly wrong? I think the interesting
> numbers here are probably what the smallest possible extension TD (aka.
> NRX TD) are, and then how much each quoting TD adds on top of that. Any
> chance you could fill in my horrible memory?
The calculations for the memory overhead work like the following.
Let's assume that maximum number of TDs on a given platform equals
LP count on that platform. For the reference on GNR it is 1376.
So, TDX module needs to be able to create 1376 instances of Quoting
Service to have 1:1 mapping.
Each TDX Module extension (Quoting Service in this case) needs
memory to store its session and VCPU state. These must be pre-allocated
in order to guarantee availability at any moment.
Session state size (S) depends on concrete TDX Module extension, some
have more information to store (especially the ones that run multi-stage
communication authenticated key exchange protocols like SPDM), some
less. Quoting Service is luckily the one that needs very little, but it does
consume a single 4KB page (smallest granularity of memory allocation inside TDX
module). Again, for the reference, TPA (SPDM connection management
extension) consumes much more - 128 KB at the moment and it might grow
as a result of introduction of postquantum support.
So, the formula to calculate how much *total memory* is required to store
session state for all instances of a given TDX Module extension is:
#Max_TD_count * (S*2+2)
Plugging numbers for Quoting service above, gives us
1376 * (1*2 + 2) = 22.5MB total
Next, we need to add numbers for supporting separate VCPU state per each
Quoting Service instance. Math is the following:
#Max_TD_count * 6 (this assumes non-partitioned TDs)
Plugging numbers for Quoting service, gives us
1376 * 6 = 33.8MB total
There is also additional impact inside SEAMRR, but it is small in terms of 30KB
total for the above LP/TD count.
So, the total memory cost specifically for Quoting service to support
1:1 mapping given GNR max LP count would be:
22.5 + 33.8 = 56.3MB total
Hope this clarifies how to calculate memory impact.
^ permalink raw reply [flat|nested] 33+ messages in thread
* Re: TDG quote analysis
2026-09-22 6:48 ` Reshetova, Elena
@ 2026-09-22 17:07 ` Edgecombe, Rick P
2026-09-23 8:50 ` Reshetova, Elena
0 siblings, 1 reply; 33+ messages in thread
From: Edgecombe, Rick P @ 2026-09-22 17:07 UTC (permalink / raw)
To: Reshetova, Elena, Hansen, Dave, seanjc@google.com
Cc: Xu, Yilun, pbonzini@redhat.com, Annapurve, Vishal, kas@kernel.org,
Wu, Binbin, Fang, Peter, kvm@vger.kernel.org
On Tue, 2026-09-22 at 06:48 +0000, Reshetova, Elena wrote:
> Hope this clarifies how to calculate memory impact.
Sean is suggesting to have some save state area per-TD such that each TD can get
a quote without interfering with the other. I think this is saying inside TDX
module you need a "vCPU" and a "session", at which point you can execute
independently from the other quotes. Together that would take:
2 + 2 + 6 = 8 pages = 40KB?
(BTW not following the "* 2 + 2" bit, how does that square with "Quoting Service
is luckily the one that needs very little, but it does consume a single 4KB
page"?)
And then the quoting TD itself could be, I'd guess at least in the MB range?
So the two options:
One operation at a time per-system: A few MBs + 40KB + contention handling
One operation at a time per-TD: A few MBs + 40KB * NUM_TDs
Outside of the S3M based attestation type schemes, the per-TD solution seems
simpler to me. No one needs to think about contention. (except for inter-guest
contention) And it eliminates ever needing to configure contention handling
parameters from the host.
Then we leave contention handling for another day? In a way, the save state per-
TD solution is in the spirit of punting, because it makes the whole thing
simpler today and will require more work for an S3M each-each quote thing.
But that "simpler" is only about contention. We still have some opens on what
TDG extension calls need to handle. Dave, something else we were discussing up
the thread is whether a TDG quote call needs to have some
INTERRUBTIBLE_RESUMABLE behavior from the guest. So that the guest can take a
break from the long running guest call in order to handle *guest* interrupts. Do
we have the same concern there that we do in the host?
^ permalink raw reply [flat|nested] 33+ messages in thread
* Re: TDG quote analysis
2026-09-21 23:06 ` Dave Hansen
@ 2026-09-23 4:09 ` Vishal Annapurve
2026-09-24 0:03 ` Vishal Annapurve
0 siblings, 1 reply; 33+ messages in thread
From: Vishal Annapurve @ 2026-09-23 4:09 UTC (permalink / raw)
To: Dave Hansen
Cc: Peter Fang, Sean Christopherson, Rick P Edgecombe, Yilun Xu,
Elena Reshetova, Binbin Wu, kas@kernel.org, pbonzini@redhat.com,
kvm@vger.kernel.org
On Mon, Sep 21, 2026 at 4:06 PM Dave Hansen <dave.hansen@intel.com> wrote:
>
> On 9/21/26 16:00, Peter Fang wrote:
> >> is how to hide the fact that someone in the future might think following AMD's
> >> lead and putting a precious resource into a tiny microprosser on a slow bus is a
> >> fantastic idea.
> > As shocking as it sounds, there are actually cloud providers out there
> > that are wanting this, or at the very least wanting to see how it goes
> > in real deployment. I think memory side channel attacks are a real
> > concern for some of them. But yeah this is not great from a performance
> > perspective.
>
> I'll believe it when I see it (on this mailing list).
>
> If folks want complexity and low performance in the kernel for
> side-channel defense, then I'm your guy. I've made half a career out of
> it. All that I ask is that they come and ask for it in the open. Here,
> please.
Google is tracking use cases for deploying S3M-based attestation for
both TDX guest workloads and host workloads.
Based on my reading of the discussions so far, as long as guest VCPUs
are not blocked while quote signing occurs, we should be able to
devise a reasonable scheme to support S3M-based attestation for TDX
guest workloads without introducing "undue" complexity (I agree with
Sean's complaints on this front [1], [2]). Quote signing doesn't have
to be "fast" because it is not supposed to be a frequent operation in
the guest lifecycle, provided the guest can make forward progress
while quote signing occurs (similar to existing SGX based
attestation). Have we concluded that such a scheme is unachievable?
[1] https://lore.kernel.org/kvm/ak7PHTGEuxAZCvXw@google.com/
[2] https://lore.kernel.org/kvm/aq2t5wJGLtsltMza@google.com/
^ permalink raw reply [flat|nested] 33+ messages in thread
* RE: TDG quote analysis
2026-09-22 17:07 ` Edgecombe, Rick P
@ 2026-09-23 8:50 ` Reshetova, Elena
2026-09-23 22:50 ` Peter Fang
0 siblings, 1 reply; 33+ messages in thread
From: Reshetova, Elena @ 2026-09-23 8:50 UTC (permalink / raw)
To: Edgecombe, Rick P, Hansen, Dave, seanjc@google.com
Cc: Xu, Yilun, pbonzini@redhat.com, Annapurve, Vishal, kas@kernel.org,
Wu, Binbin, Fang, Peter, kvm@vger.kernel.org
> On Tue, 2026-09-22 at 06:48 +0000, Reshetova, Elena wrote:
> > Hope this clarifies how to calculate memory impact.
>
> Sean is suggesting to have some save state area per-TD such that each TD can
> get
> a quote without interfering with the other. I think this is saying inside TDX
> module you need a "vCPU" and a "session", at which point you can execute
> independently from the other quotes. Together that would take:
>
> 2 + 2 + 6 = 8 pages = 40KB?
Correct for the current Quoting Service.
>
> (BTW not following the "* 2 + 2" bit, how does that square with "Quoting
> Service
> is luckily the one that needs very little, but it does consume a single 4KB
> page"?)
Multiplication by 2 currently comes from internal implementation inside TDX
Module where it keeps two copies of session state: active and saved and quickly
flips one to another when TDX Module extension (Quoting service in this case)
finishes working on a session state. It is generic optimization mainly targeted
for TDX Module extensions that keep bigger state to avoid copying this state.
A fixed addition by 2 comes from the 2 pages that are required to store paging
structures to map active and saved session state into Quoting Service extension
when it executes.
>
> And then the quoting TD itself could be, I'd guess at least in the MB range?
Yes. Again, the size will increase with introduction of PQC, but still in this range.
^ permalink raw reply [flat|nested] 33+ messages in thread
* Re: TDG quote analysis
2026-09-23 8:50 ` Reshetova, Elena
@ 2026-09-23 22:50 ` Peter Fang
0 siblings, 0 replies; 33+ messages in thread
From: Peter Fang @ 2026-09-23 22:50 UTC (permalink / raw)
To: Reshetova, Elena
Cc: Edgecombe, Rick P, Hansen, Dave, seanjc@google.com, Xu, Yilun,
pbonzini@redhat.com, Annapurve, Vishal, kas@kernel.org,
Wu, Binbin, kvm@vger.kernel.org
On Wed, Sep 23, 2026 at 01:50:17AM -0700, Reshetova, Elena wrote:
>
> > On Tue, 2026-09-22 at 06:48 +0000, Reshetova, Elena wrote:
> > > Hope this clarifies how to calculate memory impact.
> >
> > Sean is suggesting to have some save state area per-TD such that each TD can
> > get
> > a quote without interfering with the other. I think this is saying inside TDX
> > module you need a "vCPU" and a "session", at which point you can execute
> > independently from the other quotes. Together that would take:
> >
A recap of the PUCK meeting discussion today:
1. The TDG quoting idea would make things simpler for the host, but at
the cost of chunking up and disturbing crypto primitives. So in the
end we all agreed to keep the original TDH direction.
2. Use concurrent TDH quoting threads (number of physical cores).
3. Punt on the S3M-based quoting for now.
4. Use the existing mechanisms (polling or async notification) to handle
quoting completion.
Please feel free to correct or add to this. Thanks all for the feedback!
^ permalink raw reply [flat|nested] 33+ messages in thread
* Re: TDG quote analysis
2026-09-23 4:09 ` Vishal Annapurve
@ 2026-09-24 0:03 ` Vishal Annapurve
2026-09-24 0:18 ` Sean Christopherson
0 siblings, 1 reply; 33+ messages in thread
From: Vishal Annapurve @ 2026-09-24 0:03 UTC (permalink / raw)
To: Dave Hansen
Cc: Peter Fang, Sean Christopherson, Rick P Edgecombe, Yilun Xu,
Elena Reshetova, Binbin Wu, kas@kernel.org, pbonzini@redhat.com,
kvm@vger.kernel.org
On Tue, Sep 22, 2026 at 9:09 PM Vishal Annapurve <vannapurve@google.com> wrote:
>
> On Mon, Sep 21, 2026 at 4:06 PM Dave Hansen <dave.hansen@intel.com> wrote:
> >
> > On 9/21/26 16:00, Peter Fang wrote:
> > >> is how to hide the fact that someone in the future might think following AMD's
> > >> lead and putting a precious resource into a tiny microprosser on a slow bus is a
> > >> fantastic idea.
> > > As shocking as it sounds, there are actually cloud providers out there
> > > that are wanting this, or at the very least wanting to see how it goes
> > > in real deployment. I think memory side channel attacks are a real
> > > concern for some of them. But yeah this is not great from a performance
> > > perspective.
> >
> > I'll believe it when I see it (on this mailing list).
> >
> > If folks want complexity and low performance in the kernel for
> > side-channel defense, then I'm your guy. I've made half a career out of
> > it. All that I ask is that they come and ask for it in the open. Here,
> > please.
>
> Google is tracking use cases for deploying S3M-based attestation for
> both TDX guest workloads and host workloads.
>
To clarify this statement further and address the discussion from the
PUCK meeting: Google is tracking use cases for deploying S3M based
signing of attestation quotes (or S3M-based quoting as referred to in
this thread) for TDX guest workloads.
^ permalink raw reply [flat|nested] 33+ messages in thread
* Re: TDG quote analysis
2026-09-24 0:03 ` Vishal Annapurve
@ 2026-09-24 0:18 ` Sean Christopherson
2026-09-24 0:44 ` Edgecombe, Rick P
0 siblings, 1 reply; 33+ messages in thread
From: Sean Christopherson @ 2026-09-24 0:18 UTC (permalink / raw)
To: Vishal Annapurve
Cc: Dave Hansen, Peter Fang, Rick P Edgecombe, Yilun Xu,
Elena Reshetova, Binbin Wu, kas@kernel.org, pbonzini@redhat.com,
kvm@vger.kernel.org
On Wed, Sep 23, 2026, Vishal Annapurve wrote:
> On Tue, Sep 22, 2026 at 9:09 PM Vishal Annapurve <vannapurve@google.com> wrote:
> >
> > On Mon, Sep 21, 2026 at 4:06 PM Dave Hansen <dave.hansen@intel.com> wrote:
> > >
> > > On 9/21/26 16:00, Peter Fang wrote:
> > > >> is how to hide the fact that someone in the future might think following AMD's
> > > >> lead and putting a precious resource into a tiny microprosser on a slow bus is a
> > > >> fantastic idea.
> > > > As shocking as it sounds, there are actually cloud providers out there
> > > > that are wanting this, or at the very least wanting to see how it goes
> > > > in real deployment. I think memory side channel attacks are a real
> > > > concern for some of them. But yeah this is not great from a performance
> > > > perspective.
> > >
> > > I'll believe it when I see it (on this mailing list).
> > >
> > > If folks want complexity and low performance in the kernel for
> > > side-channel defense, then I'm your guy. I've made half a career out of
> > > it. All that I ask is that they come and ask for it in the open. Here,
> > > please.
> >
> > Google is tracking use cases for deploying S3M-based attestation for
> > both TDX guest workloads and host workloads.
> >
>
> To clarify this statement further and address the discussion from the
> PUCK meeting: Google is tracking use cases for deploying S3M based
> signing of attestation quotes (or S3M-based quoting as referred to in
> this thread) for TDX guest workloads.
Good timing, I was just chatting with Erdem about this. The TL;DR of "why" is
that anything that's done on-CPU is more vulnerable to side channel attacks, so
hyper-paranoid use cases that *also* know the tradeoffs involved may want to
use much slower, but more secure (in theory) S3M signing.
From an upstream perspective, I think we don't care? Or rather, we *shouldn't*
have to care beyond letting the admin flip a TDX Module switch. I'm a-ok letting
end users opt-in to ultra-paranoid mode, with a pile of disclaimers around the
shortcomings and tradeoffs involved. I.e. I don't really care how the quote is
generated, I just don't want to carry any kernel/KVM code to make it less painful
(and it should absolutely be opt-in).
^ permalink raw reply [flat|nested] 33+ messages in thread
* Re: TDG quote analysis
2026-09-24 0:18 ` Sean Christopherson
@ 2026-09-24 0:44 ` Edgecombe, Rick P
2026-09-24 16:28 ` Sean Christopherson
0 siblings, 1 reply; 33+ messages in thread
From: Edgecombe, Rick P @ 2026-09-24 0:44 UTC (permalink / raw)
To: Annapurve, Vishal, seanjc@google.com
Cc: Xu, Yilun, Reshetova, Elena, Wu, Binbin, Hansen, Dave,
kas@kernel.org, pbonzini@redhat.com, Fang, Peter,
kvm@vger.kernel.org
On Wed, 2026-09-23 at 17:18 -0700, Sean Christopherson wrote:
> From an upstream perspective, I think we don't care? Or rather, we *shouldn't*
> have to care beyond letting the admin flip a TDX Module switch. I'm a-ok letting
> end users opt-in to ultra-paranoid mode, with a pile of disclaimers around the
> shortcomings and tradeoffs involved. I.e. I don't really care how the quote is
> generated, I just don't want to carry any kernel/KVM code to make it less painful
> (and it should absolutely be opt-in).
I think the direction is to basically punt on Linux S3M design for the initial
upstream implementation. But based on today, I think we would have several
options. This is good thing to do for now I think, because the TDX arch side of
S3M based attestation is still fuzzy.
Thanks all for all the attention in helping close out a direction!
^ permalink raw reply [flat|nested] 33+ messages in thread
* Re: TDG quote analysis
2026-09-24 0:44 ` Edgecombe, Rick P
@ 2026-09-24 16:28 ` Sean Christopherson
0 siblings, 0 replies; 33+ messages in thread
From: Sean Christopherson @ 2026-09-24 16:28 UTC (permalink / raw)
To: Rick P Edgecombe
Cc: Vishal Annapurve, Yilun Xu, Elena Reshetova, Binbin Wu,
Dave Hansen, kas@kernel.org, pbonzini@redhat.com, Peter Fang,
kvm@vger.kernel.org
On Thu, Sep 24, 2026, Rick P Edgecombe wrote:
> On Wed, 2026-09-23 at 17:18 -0700, Sean Christopherson wrote:
> > From an upstream perspective, I think we don't care? Or rather, we *shouldn't*
> > have to care beyond letting the admin flip a TDX Module switch. I'm a-ok letting
> > end users opt-in to ultra-paranoid mode, with a pile of disclaimers around the
> > shortcomings and tradeoffs involved. I.e. I don't really care how the quote is
> > generated, I just don't want to carry any kernel/KVM code to make it less painful
> > (and it should absolutely be opt-in).
>
> I think the direction is to basically punt on Linux S3M design for the initial
> upstream implementation. But based on today, I think we would have several
> options. This is good thing to do for now I think, because the TDX arch side of
> S3M based attestation is still fuzzy.
Ya, and what I'm saying is that one of the goals of the TDX Module should be to
provide a common interface for signing quotes, regardless of what is doing the
actual signing. I.e. beyond flipping a switch at boot time, we shouldn't need
to modify the host or guest kernels. Host userspace and maybe guest userspace
will need to be aware of which signing implementation is being used, because it
will significantly impact VM scheduling and behavior, but the kernel really
shouldn't have to care.
^ permalink raw reply [flat|nested] 33+ messages in thread
end of thread, other threads:[~2026-09-24 16:28 UTC | newest]
Thread overview: 33+ messages (download: mbox.gz follow: Atom feed
-- links below jump to the message on this page --
2026-09-16 0:48 TDG quote analysis Edgecombe, Rick P
2026-09-16 18:00 ` Sean Christopherson
2026-09-16 18:49 ` Edgecombe, Rick P
2026-09-16 19:27 ` Sean Christopherson
2026-09-16 22:31 ` Edgecombe, Rick P
2026-09-16 23:00 ` Peter Fang
2026-09-16 23:57 ` Sean Christopherson
2026-09-17 0:56 ` Edgecombe, Rick P
2026-09-17 13:37 ` Sean Christopherson
2026-09-17 18:12 ` Edgecombe, Rick P
2026-09-17 19:57 ` Sean Christopherson
2026-09-17 21:32 ` Edgecombe, Rick P
2026-09-18 0:04 ` Sean Christopherson
2026-09-18 2:39 ` Edgecombe, Rick P
2026-09-18 13:13 ` Sean Christopherson
2026-09-18 18:12 ` Edgecombe, Rick P
2026-09-18 21:32 ` Sean Christopherson
2026-09-21 23:00 ` Peter Fang
2026-09-21 23:06 ` Dave Hansen
2026-09-23 4:09 ` Vishal Annapurve
2026-09-24 0:03 ` Vishal Annapurve
2026-09-24 0:18 ` Sean Christopherson
2026-09-24 0:44 ` Edgecombe, Rick P
2026-09-24 16:28 ` Sean Christopherson
2026-09-21 23:04 ` Dave Hansen
2026-09-21 23:19 ` Sean Christopherson
2026-09-21 23:36 ` Dave Hansen
2026-09-21 23:46 ` Sean Christopherson
2026-09-22 0:03 ` Dave Hansen
2026-09-22 6:48 ` Reshetova, Elena
2026-09-22 17:07 ` Edgecombe, Rick P
2026-09-23 8:50 ` Reshetova, Elena
2026-09-23 22:50 ` Peter Fang
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox