From: Borislav Petkov <bp@alien8.de>
To: Zack Rusin <zack.rusin@broadcom.com>
Cc: Kiryl Shutsemau <kas@kernel.org>,
x86@kernel.org, Dennis Zhou <dennis@kernel.org>,
Tejun Heo <tj@kernel.org>, Arnd Bergmann <arnd@arndb.de>,
Rick Edgecombe <rick.p.edgecombe@intel.com>,
Tom Lendacky <thomas.lendacky@amd.com>,
Wei Liu <wei.liu@kernel.org>, Dexuan Cui <decui@microsoft.com>,
Paolo Bonzini <pbonzini@redhat.com>,
Vitaly Kuznetsov <vkuznets@redhat.com>,
Ajay Kaher <ajay.kaher@broadcom.com>,
Alexey Makhalov <alexey.makhalov@broadcom.com>,
Thomas Gleixner <tglx@kernel.org>, Ingo Molnar <mingo@redhat.com>,
Dave Hansen <dave.hansen@linux.intel.com>,
"H. Peter Anvin" <hpa@zytor.com>,
virtualization@lists.linux.dev,
bcm-kernel-feedback-list@broadcom.com,
linux-kernel@vger.kernel.org, Christoph Lameter <cl@gentwo.org>,
Andrew Morton <akpm@linux-foundation.org>,
Bo Gan <bo.gan@broadcom.com>,
linux-mm@kvack.org, linux-arch@vger.kernel.org,
linux-coco@lists.linux.dev, kvm@vger.kernel.org,
Jonathan Corbet <corbet@lwn.net>,
"K. Y. Srinivasan" <kys@microsoft.com>,
Haiyang Zhang <haiyangz@microsoft.com>,
Long Li <longli@microsoft.com>, Andy Lutomirski <luto@kernel.org>,
Peter Zijlstra <peterz@infradead.org>,
linux-doc@vger.kernel.org, linux-hyperv@vger.kernel.org,
Nathan Chancellor <nathan@kernel.org>,
Kees Cook <kees@kernel.org>, Ashish Kalra <ashish.kalra@amd.com>
Subject: Re: [PATCH v2 0/6] x86/percpu: Share decrypted storage before guest setup
Date: Thu, 1 Oct 2026 16:59:50 -0700 [thread overview]
Message-ID: <20261001235950.GHar7z9v3zfk8odoTh@fat_crate.local> (raw)
In-Reply-To: <CABQX2QNmza+=i1FXgDtFS3e3O5oXxPZppRbRrL54P1LmrB2gBw@mail.gmail.com>
On Tue, Sep 29, 2026 at 01:25:09PM -0400, Zack Rusin wrote:
> Partly. So steal time tells Linux how much time a virtual CPU spent
> ready to run but waiting for the hypervisor to schedule it.
>
> For example, during a 100 ms interval, the vCPU might execute for 70
> ms and wait for a physical CPU for 30 ms. Reporting those 30 ms helps
> kernel:
> - account for CPU usage accurately: avoid charging applications for
> time when the hypervisor wasn't running their vCPU.
> - make fairer scheduling decisions: base task execution accounting on
> the CPU time tasks actually received.
> - expose host contention: the st field in top and counters in
> /proc/stat help explain why a VM is slow even though its applications
> don't appear to consume all available CPU time.
Aaaha, IOW, that's the "st" column here:
$ vmstat
procs -----------memory---------- ---swap-- -----io---- -system-- -------cpu-------
r b swpd free buff cache si so bi bo in cs us sy id wa st gu
1 0 0 1714852 9748 85300 0 0 3793 51 1445 0 0 1 99 0 0 0
In any case, you could keep that helpful explanation in yout 0th message. :)
> If options are binary then "no" :) It's not for monitoring tools, it's
> for the kernel's own steal-time accounting above, which currently
> doesn't work in confidential VMware guests. So the issue is that in a
> confidential guest (AMD SEV SNP, Intel TDX) memory is private by
> default, so the hypervisor can't read or write it. Any page the
> hypervisor has to write must first be converted to shared by the
> guest. Linux gives us the address of the steal-time buffer without
> converting it. The hypervisor then tries to write into a private page,
> which doesn't work. Currently ESXi deliberately powers off the VM when
> that happens. That only affects VMs with the steal clock enabled,
> which ESXi leaves off by default except for Photon guests. I'll
> probably change ESXi to just disable steal time when a guest gives us
> a private page, but that only avoids the power-off; steal time still
> won't work in these guests without this series.
>
> KVM has the same need for three of its per-CPU buffers (steal time,
> async page faults and PV EOI). It already converts them, but only on
> AMD and in its own loop. Kiryl asked on v1 for one common place that
> converts all such per CPU buffers early, on both AMD and Intel,
Right, why early?
I mean, I am still trying to see the justification for this diffstat
14 files changed, 247 insertions(+), 51 deletions(-)
and whether it is really worth it.
> instead of each hypervisor driver doing it. That's patches 2-6.
>
> Patch 1 fixes an old layout bug in uniprocessor kernels, where these
> buffers can share a page with unrelated data. The series also fixes
> SEV and SEV-SNP guests on KVM running kernels built with CONFIG_SMP=n,
> which currently hang at boot (we reproduced the hang).
That should tell you how much we care about UP. We would even love to make SMP
the default.
> Fair enough. I'll rework the cover letter and the commit messages so
> each one starts with the problem. The series originally wasn't really
> touching x86 core parts and I haven't updated it for a larger crowd.
> Would you like an explanation of steal time, like the above, in the
> cover as well?
Yes please.
Also, we have some blurb about how to write those:
https://docs.kernel.org/process/maintainer-tip.html#patch-subject
and
https://docs.kernel.org/process/submitting-patches.html
In talking to Peter about it, we were wondering whether this can be made
simpler. Like do not touch perCPU but do a normal page for each CPU's steal
time gunk and thus do not split the large page and then that early
enc/decrypting of memory I don't like either.
Perhaps we should start with the simplest approach first.
Thx.
--
Regards/Gruss,
Boris.
https://people.kernel.org/tglx/notes-about-netiquette
next prev parent reply other threads:[~2026-10-02 0:00 UTC|newest]
Thread overview: 14+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-09-29 4:02 [PATCH v2 0/6] x86/percpu: Share decrypted storage before guest setup Zack Rusin
2026-09-29 4:02 ` [PATCH v2 1/6] percpu: Page-align decrypted data in UP kernels Zack Rusin
2026-09-29 4:02 ` [PATCH v2 2/6] percpu: Bound decrypted storage for all x86 encrypted guests Zack Rusin
2026-09-29 7:54 ` Peter Zijlstra
2026-09-29 17:12 ` Zack Rusin
2026-09-30 8:54 ` Peter Zijlstra
2026-09-29 4:02 ` [PATCH v2 3/6] x86/percpu: Require embedded allocation in " Zack Rusin
2026-09-29 4:02 ` [PATCH v2 4/6] x86/mm: Provide common early memory decryption Zack Rusin
2026-09-29 4:02 ` [PATCH v2 5/6] x86/tdx: Support early sharing of kernel data Zack Rusin
2026-09-29 4:02 ` [PATCH v2 6/6] x86/percpu: Share decrypted storage before guest CPU setup Zack Rusin
2026-09-29 5:39 ` [PATCH v2 0/6] x86/percpu: Share decrypted storage before guest setup Borislav Petkov
2026-09-29 17:25 ` Zack Rusin
2026-10-01 23:59 ` Borislav Petkov [this message]
2026-10-02 6:06 ` Zack Rusin
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=20261001235950.GHar7z9v3zfk8odoTh@fat_crate.local \
--to=bp@alien8.de \
--cc=ajay.kaher@broadcom.com \
--cc=akpm@linux-foundation.org \
--cc=alexey.makhalov@broadcom.com \
--cc=arnd@arndb.de \
--cc=ashish.kalra@amd.com \
--cc=bcm-kernel-feedback-list@broadcom.com \
--cc=bo.gan@broadcom.com \
--cc=cl@gentwo.org \
--cc=corbet@lwn.net \
--cc=dave.hansen@linux.intel.com \
--cc=decui@microsoft.com \
--cc=dennis@kernel.org \
--cc=haiyangz@microsoft.com \
--cc=hpa@zytor.com \
--cc=kas@kernel.org \
--cc=kees@kernel.org \
--cc=kvm@vger.kernel.org \
--cc=kys@microsoft.com \
--cc=linux-arch@vger.kernel.org \
--cc=linux-coco@lists.linux.dev \
--cc=linux-doc@vger.kernel.org \
--cc=linux-hyperv@vger.kernel.org \
--cc=linux-kernel@vger.kernel.org \
--cc=linux-mm@kvack.org \
--cc=longli@microsoft.com \
--cc=luto@kernel.org \
--cc=mingo@redhat.com \
--cc=nathan@kernel.org \
--cc=pbonzini@redhat.com \
--cc=peterz@infradead.org \
--cc=rick.p.edgecombe@intel.com \
--cc=tglx@kernel.org \
--cc=thomas.lendacky@amd.com \
--cc=tj@kernel.org \
--cc=virtualization@lists.linux.dev \
--cc=vkuznets@redhat.com \
--cc=wei.liu@kernel.org \
--cc=x86@kernel.org \
--cc=zack.rusin@broadcom.com \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox